🤖 AI benchmark: hit-rate of 7 models

Prematch ライブ (in-play)

Seven external AI models (Hermes contour) independently analyze the same 試合 — predicting the 結果 (1X2), total (オーバー/アンダー), both チーム to スコア (BTTS) and the exact スコア. Here we honestly compare their 予想 against the real result after the final whistle and combine everything into a single accuracy rating. An informational and analytical snapshot, not betting advice.

⚠️ Data is still accumulating — counting starts from 09.07.2026, so all models are compared on the same events (early test 予想 are excluded). The sample is still small and not representative. Right now the snapshot holds 2928 試合(es), 9548 settled AI 予想 (サッカー, テニス, バスケットボール, バレーボール, ホッケー, 卓球). The figures below are N, not «a percentage you can trust»: the more 試合 are played out, the more reliable the snapshot becomes. We show it transparently from day one, not only once the sample becomes «convenient».

Leaderboard

Model N (settled) 1X2 ダブルチャンス (1X) Exact スコア Composite accuracy
ChatGPT
263 65.0%(171/263) 74.6%(47/63) 17.2%(35/203) 44.2%(206/466)
Claude
2417 60.7%(1467/2417) 77.3%(683/884) 22.4%(479/2138) 42.7%(1946/4555)
Kimi
1109 60.8%(674/1109) 73.8%(253/343) 20.3%(196/964) 42.0%(870/2073)
Google AI
3062 59.7%(1828/3060) 74.4%(804/1080) 20.3%(612/3020) 40.1%(2440/6080)
GLM 5.2
1285 59.5%(765/1285) 74.4%(302/406) 19.5%(246/1261) 39.7%(1011/2546)
DeepSeek
175 60.6%(106/175) 56.1%(23/41) 18.5%(32/173) 39.7%(138/348)
Qwen
1237 59.3%(733/1237) 75.0%(267/356) 16.8%(196/1166) 38.7%(929/2403)

grey — sample <5, not representative; «—» — the モデル has not made a settled 予想 yet.

ダブルチャンス (1X) — the same ピック counts as a 勝利 if the chosen side won or the 試合 drew. Of the 1X2 losses in football/hockey: 739 draws, 794 underdog (total settled 1X2 ピック in these sports: 3173, double chance combined 75.0% (2379/3173)). The models almost always take the favorite and don't bet on a 引き分け — double chance shows how many bets are eaten specifically by draws.

Composite accuracy — the share of correct 予想 横断 all 表示 マーケット together: (sum of correct ピック) ÷ (sum of all settled ピック) 横断 the マーケット 1X2 + Exact スコア. Each マーケット-ピック weighs equally. This is hit-rate, not profitability — for money/ROI by モデル see /ai-agent. «Exact スコア» — the full final スコア was guessed correctly (H and A matched); for tennis, sets are compared; 予想 with no recognized スコア do not count toward the denominator. ダブルチャンス (1X): a ピック counts as a 勝利 if the chosen side won OR it was a 引き分け — it accounts for frequent draws that «eat» bets on the favorite. This metric is informational and is not included in composite accuracy. On «全スポーツ» we don't show total and BTTS — they're tied to the sport and make no sense in a mixed pool.

Composite モデル rating · all マーケット

Bar height = the モデル's composite accuracy 横断 all applicable マーケット on the current sample. Sorted from best to worst.

44.2% (206/466)
GPT 5.5
42.7% (1946/4555)
Opus 4.8
42.0% (870/2073)
Kimi 2.6
40.1% (2440/6080)
Gemini 3.5 Flash
39.7% (1011/2546)
GLM 5.2
39.7% (138/348)
DeepSeek V4 Pro
38.7% (929/2403)
Qwen 3.7 Plus

Bars are AI models by version; grey/dimmed — sample <5, not representative. The snapshot is informational, not betting advice.

Accuracy by マーケット

Where each モデル is strong: one mini-bar per applicable マーケット, with the percentage and (hits/sample).

ChatGPT Composite 44.2%
1X2
65.0% (171/263)
Exact スコア
17.2% (35/203)
Claude Composite 42.7%
1X2
60.7% (1467/2417)
Exact スコア
22.4% (479/2138)
Kimi Composite 42.0%
1X2
60.8% (674/1109)
Exact スコア
20.3% (196/964)
Google AI Composite 40.1%
1X2
59.7% (1828/3060)
Exact スコア
20.3% (612/3020)
GLM 5.2 Composite 39.7%
1X2
59.5% (765/1285)
Exact スコア
19.5% (246/1261)
DeepSeek Composite 39.7%
1X2
60.6% (106/175)
Exact スコア
18.5% (32/173)
Qwen Composite 38.7%
1X2
59.3% (733/1237)
Exact スコア
16.8% (196/1166)

The モデル's favorite by 1X2 = the max of P1/X/P2 in its probabilities; for sports without a 引き分け (tennis, volleyball, etc.) the «X» option doesn't participate. grey — sample <5, not representative. Not betting advice.