Skip to content

Methodology · updated 10 September 2026

What each Clauseground number means.

This is the current implementation, not an idealized future version. It explains where the data comes from, how each probability is produced, and where the process can fail.

Three labels to keep separate

Market
A no-vig probability derived from a complete designated reference book.
Model
The independent statistical football-model output.
Combined
A fixed, market-specific weighted combination of the market and model views.

1

Data categories and coverage

Clauseground ingests fixture details, results, team ratings, venue context, sportsbook odds, supported exchange or prediction-market quotes, and selected live-match data. Odds are the category with real depth: tens of millions of timestamped price snapshots across dozens of venues, collected on a roughly 20-minute cycle.

Contextual enrichment is much thinner than that, and the page should say so. Injury and player-availability data is not currently available for any active competition — the upstream provider is not configured for the leagues on the board, so the availability input is inert rather than best-effort. Weather is fetched from a public forecast service and applies only to the minority of fixtures that carry a linked venue record with coordinates. Confirmed lineups are not ingested at all.

Coverage is data-dependent. The public Markets page is the source of truth for competitions and fixtures currently available, and each match page names the books behind every price. The product does not promise one universal update interval or complete venue coverage.

2

Odds normalization and no-vig probability

Decimal odds are converted to raw implied probability using 1 ÷ odds. For a complete market, those probabilities usually sum to more than 100%. Clauseground currently removes that margin with the multiplicative method, dividing each raw probability by the total.

The current implementation checks designated reference books in configured order and uses the first one with every required outcome. It does not yet build a liquidity-weighted consensus across every venue. Public copy therefore describes this output as the reference market view, not a universal market consensus.

3

Football model

The core is an independent Poisson scoring model with a Dixon-Coles correction for common low-scoring outcomes. Expected goals are based on team attack and defence strengths, the competition scoring baseline, and home or neutral-venue context.

The engine can apply bounded modifiers for rest days, venue altitude and venue type. It also has modifiers for weather and for injury availability, but those two are currently doing nothing for almost every fixture: the availability feed is not configured for any active competition, and weather is only fetched where a fixture is linked to a venue record. In practice the probabilities you see are driven by team strength, competition baseline and home advantage.

Missing context falls back to neutral treatment. Absence of an adjustment must not be read as confirmation that no real-world effect exists — most of the time it means the input was unavailable, not that it was checked and found irrelevant.

4

Simulation and market outputs

The default run uses 50,000 Monte Carlo simulations per fixture. It stores the run count, engine version, inputs, timestamp, and resulting probabilities. The same scoring distribution supports 1X2, totals, both-teams-to-score, and selected knockout or player markets where the required inputs exist.

More simulations reduce sampling noise; they do not remove modelling error. A precisely calculated probability can still be wrong because its assumptions or inputs are wrong.

5

Market–model combination

Clauseground keeps the reference-market and model probabilities visible, then can calculate a combined estimate. The current weights are fixed by market type and lean toward the market on major match markets. They are not dynamically changed by lineup confirmation, time to kickoff, or a hidden AI judgment.

Combined estimates are useful decision inputs, but they are not described as no-vig market probabilities. If only one source exists, the output may fall back to that source and should be read with its label.

6

Difference, expected value, and best price

The model–market difference is model probability minus no-vig reference-market probability. It describes disagreement only. For a specific venue price, Clauseground compares the combined estimate with the price-implied probability and calculates expected return per unit staked.

The best available price means the most favorable usable price in the latest collected set after data-hygiene rules. It is not a claim that every venue was searched or that the displayed price is still available. Check the venue before acting.

7

Price dispersion: what the best price is worth

The claim in §6 — that the best available price is materially better than the average one — is measurable, and it has been measured. The corpus is 27,323 completed matches carrying a full set of closing 1X2 prices: the Pinnacle close, the average close across the surveyed books, and the best close across those books, for all three outcomes. That is 81,969 individual priced outcomes. Matches missing any of those nine numbers are not scored.

Fair is the de-vigged Pinnacle close computed with the power method — the one the backtest and calibration paths use, not the proportional method §2 describes and the live board runs. Three source rows whose recorded maximum sat below their own recorded average were excluded: that is impossible for a maximum, and a negative shopping gap is exactly the quantity being measured, so keeping them would bias the result toward the null and hide why. Every figure carries ±2 standard errors.

Mean overround on the close, by the price taken:

  • Market average: 5.62% (±0.02)
  • Pinnacle: 2.91% (±0.01)
  • Best available: −0.14% (±0.02)

The distance between the first and the last is 5.77 points of margin. Per individual selection, the best surveyed price sat 6.83% (±0.03) above the market average. Of the 81,969 outcomes, 34,141 (41.7%) were priced above fair at the best surveyed price against 795 (1.0%) at the market average — a factor of 42.9×. At Pinnacle itself the count is 3, which is approximately zero by construction, because fair is derived from that price.

The best price here is the best across all surveyed books, and that survey includes books that limit winning accounts, books unavailable in any given jurisdiction, and quotes that were never simultaneously obtainable. Every best-price figure on this page is an upper bound on what a real bettor captures, never an achieved return.

Two boundaries on how this may be read. It is a property of the market and not a result of Clauseground’s modelling: nothing in it depends on our probabilities being right, and it is not a record of returns. Clauseground does not currently claim a proven historical edge — see §10 and the public performance framework for that. And the corpus covers 9 domestic competitions; it holds no UEFA Champions League matches, so no per-competition figure is published for that competition. The full write-up, including the per-competition table, is at what line shopping is worth. Reproduce it with npx tsx scripts/price-dispersion.ts.

8

Stake sizing, and why none is shown

Kelly sizing converts an estimated advantage and a price into a bankroll fraction. Clauseground applies a conservative fractional-Kelly approach and additional caps because small probability errors can create large sizing errors, especially on longshots.

No stake size is displayed anywhere in the product. The calculation is used internally to order the price-comparison list, and that is all. It cannot know your finances, limits, legal status, or tolerance for loss, so publishing a number would imply a recommendation the product is not in a position to make. What you see instead is the price, the estimate it is being compared against, how many books are pricing it, and how it has moved. How much to stake — including nothing — is your decision.

9

Freshness and provenance

Odds snapshots and model runs are timestamped. Public surfaces show the latest available relevant update, but upstream delays, failed polling, event-status lag, and moved prices can still make a number stale. Clauseground is not presented as a guaranteed real-time feed.

Each probability should retain a market, model, or combined label. Product text that lacks a source or timestamp should not be treated as a current numeric claim.

10

Closing line value

Closing line value compares the price we published with the de-vigged closing price of the sharp reference market for the same selection — same match, market, line and outcome. Positive means the published price was better than the closing fair line. CLV = closing fair probability − implied probability of the price taken (1 ÷ the best collected decimal odds at publication), in probability points.

The closing reference is the last quote before kickoff from the first designated sharp book (Pinnacle, then Betfair Exchange (EU), then William Hill) with a complete and coherent set of prices for that market and line, de-vigged with the multiplicative method — the same rule and the same code the board uses for its no-vig market view. It is the designated-reference rule of §2 applied to the last quote before kickoff rather than the latest one: a book whose prices do not sum to a coherent market is skipped and the next book in the order is tried, and the book that answered is stored with the grade, so the mix of reference books behind a figure is reported rather than implied.

A close counts only if it was captured within 12 hours of kickoff. A prediction with no usable close is reported as such and excluded from the figure, never counted as zero. The reference book, its closing timestamp and a method identifier are stored on every graded prediction, and the closing reference is archived before raw prices are pruned, so a grade can be reproduced later and a change of definition is detectable rather than silent.

Predictions enter the record at their first evaluation — typically days before kickoff, when only a few venues quote the market — so the comparison is between an early published price and a close that is usually days later. The median publication lead is shown beside the figure on the performance page.

It measures the price, not the outcome. Positive closing line value is not a claim of profit, and a record that beats the close can still lose. The figure is informative only while the publication rule does not itself select on beating the reference book: a selection published because its price already beat that book’s line at publication would show positive closing line value largely by construction. Should the rule ever change that way, the figure will be reported in two parts — the margin at publication and the market’s movement to the close — rather than as one number.

When the definition or the lookup changes, every affected prediction is re-graded in one audited pass, and each pass is listed with its date, row count and reason on the performance page. No prediction, price or difference is ever rewritten.

11

Calibration and performance

The system includes Brier-score, log-loss, closing-line-value (§10), and return calculations. Historical development and retro-simulation records can help test the machinery, but they are not clean forward evidence that the model beats the market.

Headline performance uses only predictions recorded before kickoff under a stated inclusion policy. A prediction whose match is later moved by more than 72 hours after the kickoff it was published under is void — counted and disclosed, scored in no figure, and never compared with the re-dated market’s close — which is the usual bookmaker rule for a postponement. The scoreboard publishes automatically once at least 50 predictions have been scored, and every headline figure carries a 95% interval; the current status and figures are on the public performance page.

12

What the AI analyst does — and does not do

The analyst retrieves structured Clauseground data, calls the deterministic tools, compares outputs, and explains them in natural language. It can help identify uncertainty, ask whether a price has moved, and translate technical concepts.

It does not create probabilities by intuition, guarantee winners, know every piece of team news, place bets, or replace your judgment. If a numeric answer cannot be traced to a product tool or stored record, it should not be presented as fact.

13

Known limitations

  • The current market view can come from one designated complete reference book rather than a multi-book consensus.
  • Team ratings and contextual modifiers can lag real changes in team strength.
  • Confirmed lineups are not ingested, so a probability never reflects team news.
  • Injury and player-availability enrichment is inert for every active competition — it is configured but receiving no data.
  • Weather enrichment only reaches fixtures with a linked venue record, which is a minority of the board.
  • Venue prices can move between collection and display, and a “best price” is the best in our latest collected set, not the best in the world.
  • Price movement is measured on a single reference book across two timestamps; where that book has no comparable earlier quote, no movement is claimed.
  • Model-market disagreement is not proof of an exploitable advantage.
  • The price-dispersion figures in §7 are an upper bound: the best price is taken across every surveyed book, some of which limit winning accounts or are unavailable in a given jurisdiction, and no two quotes were necessarily obtainable at the same instant.
  • That measurement covers 9 domestic competitions and no UEFA Champions League matches, so it is not published per-competition for that competition.
  • The clean-forward record’s return and closing-line-value intervals still span zero, so no performance claim is made from it.

Inspect the current market board.

Use the labels and timestamps on each match, then decide whether the evidence is strong enough to act — or whether the disciplined answer is to pass.

Decision support, not betting advice. No guaranteed outcomes. 18+/21+ where applicable. Gamble responsibly.