Research · measured 11 August 2026
What line shopping is worth.
Bookmakers do not all price the same match the same way, and the difference is not decoration — it is most of the margin. Across 27,323 completed matches, taking the best surveyed closing price instead of the market average removed 5.77 points of house margin. This page is the measurement, the method, and what it does not say.
Market average
5.62%
±0.02
Pinnacle
2.91%
±0.01
Best available
−0.14%
±0.02
Mean overround on the 1X2 close, 27,323 matches, 81,969 outcomes. Intervals are ±2 standard errors.
1
The question, and why it is the one we can answer
A bookmaker prices a football match so that the three outcomes add up to more than certainty. That excess — the overround, or the vig — is what the market charges you for taking a position, and it is deducted whether you are right or wrong. Every other question in this business sits downstream of it: a model that is genuinely better than the market still loses if the margin it pays exceeds the advantage it found.
Clauseground has asked the upstream question repeatedly and answered it honestly. Five separate attempts to make the model predict football results better than a strong market measured no, and the most careful of them measured worse than the priors already being served. Nothing on this site claims a proven historical edge, and the public performance framework publishes the current sample with its interval rather than a headline.
The margin question is different, because it is not about prediction at all. It asks what the market costs at each of the prices actually on offer at the moment a match kicks off. That has an answer, it does not depend on our model being right, and it is reproducible by anyone holding the same closing prices.
2
How it was measured
The corpus is historical football matches carrying a full set of closing 1X2 prices: the Pinnacle close, the average close across the surveyed books, and the best close across the surveyed books, for all three outcomes. Matches missing any of those nine numbers are not scored, which is why the sample is 27,323 matches — 81,969 individual outcomes — rather than the full corpus table.
Fair means the de-vigged Pinnacle close, computed with the power method — the one the backtest and calibration paths use, not the proportional method the live board runs. Pinnacle is the yardstick because it is the sharpest series in this corpus: its close scores a better Brier than its own open (0.1877 against 0.1885). Beyond Pinnacle the corpus carries the market average and the best-of-book close, not individual books, so no ranking against other single books is made here. So “this price beats fair” means “this price beats the sharpest estimate this corpus holds at kickoff”, which is a demanding test rather than a flattering one.
Three rows were excluded because their recorded maximum sat below their own recorded average, which is impossible for a maximum and is a defect in the upstream source rather than in parsing. They are dropped rather than averaged in, because a negative shopping gap is precisely the quantity this measurement exists to detect — silently keeping three impossible rows would bias the headline downward and hide the reason.
Every figure carries ±2 standard errors, computed with Welford’s algorithm so a mean over tens of thousands of rows reports an honest width rather than a comfortable one. Reproduce the whole thing with npx tsx scripts/price-dispersion.ts.
3
Result: the margin you pay depends on where you take the price
The same three outcomes, priced three ways. At the market average the books held 5.62% (±0.02). At Pinnacle, a book that competes on price and makes its money on volume, they held 2.91% (±0.01) — roughly half. Assembling each outcome from whichever surveyed book priced it highest gives −0.14% (±0.02).
That last number is negative, and the sign is the entire finding. A set of prices collected from the top of every book no longer sums to a book: the implied probabilities of home, draw and away add to slightly less than certainty. The distance between the first row and the last is 5.77 points of margin — larger than the whole margin Pinnacle charges, and far larger than any model disagreement this product has been able to demonstrate.
Per individual outcome rather than per match, the best surveyed price was 6.83% (±0.03) higher than the market average for the same selection. On a price of 3.00 that is the difference between 3.00 and 3.20 — not a rounding error, and it applies to every position taken rather than to a chosen few.
4
Result: how many outcomes clear fair, and where
The margin figures describe the cost of doing business. The second measurement asks the product question directly: of the 81,969 outcomes scored, how many were priced above fair — that is, above the de-vigged sharpest estimate available at kickoff?
- At the best surveyed price: 34,141 outcomes, 41.7% of the set.
- At the market average: 795 outcomes, 1.0%.
- At Pinnacle itself: 3 outcomes, 0.0% — approximately zero by construction, because fair is derived from the Pinnacle close and a number cannot meaningfully beat itself.
Shopping multiplies the count by 42.9×. That third line is published on purpose: without it, the first two look like a discovery about bookmakers, when they are partly a statement about the yardstick. Read together they say something narrower and more useful — measured against the sharpest single price in the market, almost nothing clears at the average, and a large minority clears at the top of the book.
This is a count of prices, not a count of winners. It says how often some book was offering more than the sharpest available estimate of an outcome’s probability. It does not say those outcomes went on to happen more often than priced, and this page deliberately does not report what they returned.
5
What this does not say
The best price here is the best across all surveyed books, and that survey includes books that limit winning accounts, books unavailable in any given jurisdiction, and quotes that were never simultaneously obtainable. Every best-price figure on this page is an upper bound on what a real bettor captures, never an achieved return.
Three further limits, stated plainly. First, this is a measurement of the market, not of Clauseground: nothing above depends on our model, and nothing above is evidence about what our own published selections have done. The forward record and its interval are on the performance page, and this site does not currently claim a proven historical edge.
Second, a gap between books is not a forecast of where a price is going. We have measured price movement separately and found the opening line about as sharp as the close, so nothing here should be read as an argument for timing a market.
Third, the figures are closing prices from a historical corpus. The board you can see today is a live one, with its own coverage gaps, its own timestamps and its own stale quotes. What generalises from this measurement is the shape — books disagree, and the disagreement is worth more than the margin the sharpest book charges — not any particular number on any particular fixture.
6
The same measurement, per competition
The gap is not uniform. Competitions that draw the deepest, most heavily traded markets tend to show tighter agreement between books; competitions priced by fewer desks show wider. Below is the per-leg gap between the best surveyed closing price and the market average, with the number of matches behind each figure.
| Competition | Best vs average | Matches |
|---|---|---|
| Liga MX | 8.13% | 4,437 |
| Major League Soccer | 7.40% | 5,798 |
| Eredivisie | 7.24% | 1,904 |
| Bundesliga | 6.60% | 1,985 |
| Ligue 1 | 6.52% | 2,184 |
| Serie A | 6.48% | 2,478 |
| La Liga | 6.29% | 2,468 |
| Premier League | 6.21% | 2,490 |
| EFL Championship | 5.45% | 3,579 |
The UEFA Champions League is covered by the live board and is absent from this table. The historical corpus has no Champions League source at all, so there is no measured gap for it — not a small one, not an estimated one. A figure interpolated from the domestic leagues would look exactly like the rows above and would be invented, so the competition simply has no row.
7
What follows from it, for this product
Clauseground surveys venues, records each price with the book it came from and the time it was seen, and shows the best one it collected next to a margin-free market probability and a separate model view. The measurement above is the argument for that design: the largest reliably measurable quantity in this market is not our disagreement with the price, it is the disagreement between the books quoting it.
That is also why the product is careful about what it puts next to a price. A best price with no venue is not actionable; a best price with no timestamp is a claim about a moment that has passed. Both are shown on every fixture, and the methodology states what “best” means: the most favourable usable price in the latest collected set, not a promise that every venue on earth was searched.
None of this makes a bet a good idea, and none of it is advice. It makes one part of the decision measurable, which is the whole of what this product claims to do.