The leaderboard built on verifiable scores.
12 established benchmarks, ~50 questions each, sampled with a fixed seed and scored entirely by code — regex extraction, exact match, and sandboxed test execution. Every number on this page is reproducible.
- 5
- Models tested
- 12
- Benchmarks
- 600
- Questions sampled
- 42
- Sampling seed
- 🏆 50.1%
- Top · Deepseek v4 Pro Max
- Jul 28, 2026
- Last run (UTC)
RANKINGS
Overall Leaderboard
5 models · seed 42 · n≈50 per benchmark
| # | Model | Overall | 💻 Coding | 🔬 Science | 📐 Math | 📚 Knowledge | 📋 Instruction |
|---|---|---|---|---|---|---|---|
| 🥇 | Deepseek v4 Pro Max 🧠 max DeepSeek · deepseek/deepseek-v4-pro | 50.1% | 31.5% | 53.0% | 60.0% | 59.0% | 47.0% |
| Per-benchmark details · 12/12 completed Click headers to sort BigCodeBench-Hard 22%
(11/50 ±11.2pp)
HumanEval+ 98%
(49/50 ±5.1pp)
MBPP+ 6%
(3/50 ±7.1pp)
GPQA Diamond 60%
(30/50 ±13.1pp)
SciBench 46%
(23/50 ±13.3pp)
AIME 2024/2025 36%
(18/50 ±12.9pp)
MATH-500 84%
(42/50 ±10.1pp)
MMLU-Pro 78%
(39/50 ±11.2pp)
IFEval 84%
(42/50 ±10.1pp)
SciCode 0%
(0/50 ±3.6pp)
SuperGPQA 40%
(20/50 ±13.1pp)
Tau-Bench (Retail) 10%
(5/50 ±8.5pp)
| |||||||
| 🥈 | Deepseek v4 Flash Max 🧠 max DeepSeek · deepseek/deepseek-v4-flash | 47.8% | 30.0% | 52.0% | 55.0% | 53.0% | 49.0% |
| Per-benchmark details · 12/12 completed Click headers to sort BigCodeBench-Hard 22%
(11/50 ±11.2pp)
HumanEval+ 88%
(44/50 ±9.1pp)
MBPP+ 10%
(5/50 ±8.5pp)
GPQA Diamond 56%
(28/50 ±13.3pp)
SciBench 48%
(24/50 ±13.3pp)
AIME 2024/2025 28%
(14/50 ±12.1pp)
MATH-500 82%
(41/50 ±10.5pp)
MMLU-Pro 78%
(39/50 ±11.2pp)
IFEval 86%
(43/50 ±9.6pp)
SciCode 0%
(0/50 ±3.6pp)
SuperGPQA 28%
(14/50 ±12.1pp)
Tau-Bench (Retail) 12%
(6/50 ±9.1pp)
| |||||||
| 🥉 | Deepseek v4 Flash DeepSeek · deepseek/deepseek-v4-flash | 46.1% | 28.5% | 46.0% | 55.0% | 55.0% | 46.0% |
| Per-benchmark details · 12/12 completed Click headers to sort BigCodeBench-Hard 16%
(8/50 ±10.1pp)
HumanEval+ 88%
(44/50 ±9.1pp)
MBPP+ 10%
(5/50 ±8.5pp)
GPQA Diamond 52%
(26/50 ±13.3pp)
SciBench 40%
(20/50 ±13.1pp)
AIME 2024/2025 26%
(13/50 ±11.8pp)
MATH-500 84%
(42/50 ±10.1pp)
MMLU-Pro 78%
(39/50 ±11.2pp)
IFEval 84%
(42/50 ±10.1pp)
SciCode 0%
(0/50 ±3.6pp)
SuperGPQA 32%
(16/50 ±12.5pp)
Tau-Bench (Retail) 8%
(4/50 ±7.8pp)
| |||||||
| 4 | Gemma 4 31B 🧠 high Google · gemini/gemma-4-31b-it | 20.6% | 0.0% | 42.0% | 20.0% | — | — |
| Per-benchmark details · 4/12 completed Click headers to sort BigCodeBench-Hard 0%
(0/47 ±3.8pp)
HumanEval+ — MBPP+ — GPQA Diamond 62%
(29/47 ±13.4pp)
SciBench 22%
(10/45 ±11.9pp)
AIME 2024/2025 20%
(1/5 ±29.4pp)
MATH-500 — MMLU-Pro — IFEval — SciCode — SuperGPQA — Tau-Bench (Retail) — | |||||||
| 5 | Gemma 4 26B a4b 🧠 high Google · gemini/gemma-4-26B-A4B-it | — | — | — | — | — | — |
| Per-benchmark details · 0/12 completed Click headers to sort BigCodeBench-Hard —
(0/0 ±0pp)
HumanEval+ — MBPP+ — GPQA Diamond —
(0/0 ±0pp)
SciBench — AIME 2024/2025 — MATH-500 — MMLU-Pro — IFEval — SciCode — SuperGPQA — Tau-Bench (Retail) — | |||||||
Click a row for per-benchmark detail · click headers to sort · ±pp = 95% Wilson confidence interval Showing 5 of 5 models
UNDER THE HOOD
Analysis
hover any chart for exact values
Quality vs. Cost
Overall score against estimated API cost for a full run. Bubble size = output speed (tokens/sec).
Category Radar
Each model's profile across the five categories. Larger area = stronger all-round.
Category Breakdown
Head-to-head grouped comparison across categories.
Benchmark Heatmap
Per-benchmark scores for every model. Hover for exact counts and Wilson confidence intervals.
Token Mix
Input, thinking, and output tokens consumed per full run.
Thinking vs. Score
Do more thinking tokens buy higher scores?
THE GAUNTLET
Benchmark Catalog
12 benchmarks · fixed-seed sampling
| Benchmark | Category | Full dataset | Sampled | Verification | Source |
|---|---|---|---|---|---|
| BigCodeBench-Hard | coding | 148 | 50 | Python unittest execution (explicit opt-in required) | bigcode/bigcodebench-hard (v0.1.4) ↗ |
| HumanEval+ | coding | 164 | 50 | Python test execution (explicit opt-in required) | evalplus/humanevalplus ↗ |
| MBPP+ | coding | 378 | 50 | Python test execution (explicit opt-in required) | evalplus/mbppplus ↗ |
| GPQA Diamond | science | 198 | 50 | Multiple choice (4 options) | nichenshun/gpqa_diamond (community mirror of Idavidrein/gpqa) ↗ |
| SciBench | science | 692 | 50 | Numerical / Formula exact match | xw27/scibench ↗ |
| AIME 2024/2025 | math | 90 | 50 | Integer exact match (000-999) | AI-MO/aimo-validation-aime ↗ |
| MATH-500 | math | 500 | 50 | Exact match / \boxed{} extraction | HuggingFaceH4/MATH-500 ↗ |
| MMLU-Pro | knowledge | 12,032 | 50 | Multiple choice (10 options) | TIGER-Lab/MMLU-Pro ↗ |
| IFEval | instruction | 541 | 50 | 25 programmatic verifiers (strict) | google/IFEval ↗ |
| SciCode | coding | 65 | 50 | Python code execution & unit test assertions | SciCode1/SciCode ↗ |
| SuperGPQA | knowledge | 26,529 (7,050 hard) | 50 | Multiple choice (up to 10 options) | m-a-p/SuperGPQA ↗ |
| Tau-Bench (Retail) | instruction | 82 | 50 | Agentic tool-call function & argument matching | amityco/tau-bench-retail-train-next-action ↗ |
Hover a row for the benchmark description and citation.
EFFICIENCY
Tokens & Performance
| Model | Input | Output | Thinking | Total | Mix | TPS ⓘ | Avg time | Est. cost ⓘ |
|---|---|---|---|---|---|---|---|---|
| Deepseek v4 Pro Max | 273,638 | 85,211 | 970,381 | 1,329,230 | 49.9 | 38.3s | $1.04 | |
| Deepseek v4 Flash Max | 273,638 | 71,636 | 1,012,225 | 1,357,499 | 93.0 | 21.1s | $0.3410 | |
| Deepseek v4 Flash | 273,638 | 79,896 | 995,022 | 1,348,556 | 92.9 | 21.1s | $0.3344 | |
| Gemma 4 31B | 30,457 | 25,685 | 0 | 511,427 | 7.4 | 98.9s | — | |
| Gemma 4 26B a4b | 0 | 0 | 0 | 0 | — | — | — |
TPS = output tokens/second (cloud APIs only) · cost via LiteLLM tables · totals per full benchmark run
TRUST THE NUMBERS
Methodology
🎲 Deterministic sampling
- ~50 questions are sampled from each benchmark's full dataset using a fixed seed (42) — the exact same questions on every run.
- Samples of n≈50 carry 95% Wilson confidence intervals of roughly ±7–14pp. Small ranking gaps are noise, not signal.
- Sampling and scoring strictness follow schema v2 (v0.2.0+); older runs are not directly comparable.
✅ Programmatic scoring only
- No LLM-as-judge anywhere. Answers are verified by code: letter extraction for multiple choice, boxed/numeric comparison for math, 25 strict verifiers for IFEval, function + argument matching for Tau-Bench.
- Code benchmarks require explicit opt-in and run in a 3-layer sandbox: AST scan → hardened subprocess (no keys, no network, temp dir) → Windows Job Object confinement.
📊 Score computation
- Category score = average of its benchmark scores across Coding, Science, Math, Knowledge, and Instruction.
- Overall score = equal-weight average of completed category scores. Provider outages are excluded, not scored as zero, so one bad API day doesn't sink a model.
⚙️ Inference settings
- temperature 0 · max_tokens 4096 · 300s timeout per request.
- Transient errors retry with exponential backoff until a good response arrives; permanent errors (context length, content filter) are never retried.