The leaderboard built on verifiable scores.
12 established benchmarks, ~50 questions each, sampled with a fixed seed and scored entirely by code - regex extraction, exact match, and sandboxed test execution. Every number on this page is reproducible.
- 10
- Models tested
- 12
- Benchmarks
- 600
- Questions sampled
- 42
- Sampling seed
- 🏆 73.9%
- Top · Gemini 3.7 Flash
- Sep 3, 2026
- Last run (UTC)
RANKINGS
Overall Leaderboard
10 models · seed 42 · n≈50 per benchmark
| # | Model | Overall | 💻 Coding | 🔬 Science | 📐 Math | 📚 Knowledge | 📋 Instruction |
|---|---|---|---|---|---|---|---|
| 🥇 | Gemini 3.7 Flash ⚡ no thinking Google · gemini/gemini-3.7-flash | 73.9% | 48.5% | 84.0% | 97.0% | 82.0% | 58.0% |
| Per-benchmark details · 12/12 completed Click headers to sort BigCodeBench-Hard 6%
(3/50 ±7.1pp)
HumanEval+ 100%
(50/50 ±3.6pp)
MBPP+ 88%
(44/50 ±9.1pp)
GPQA Diamond 100%
(50/50 ±3.6pp)
SciBench 68%
(34/50 ±12.5pp)
AIME 2024/2025 98%
(49/50 ±5.1pp)
MATH-500 96%
(48/50 ±6.2pp)
MMLU-Pro 90%
(45/50 ±8.5pp)
IFEval 92%
(46/50 ±7.8pp)
SciCode 0%
(0/50 ±3.6pp)
SuperGPQA 74%
(37/50 ±11.8pp)
Tau-Bench (Retail) 24%
(12/50 ±11.6pp)
| |||||||
| 🥈 | Gemini 3.5 Flash Lite ⚡ no thinking Google · gemini/gemini-3.5-flash-lite | 67.6% | 46.0% | 77.0% | 88.0% | 74.0% | 53.0% |
| Per-benchmark details · 12/12 completed Click headers to sort BigCodeBench-Hard 4%
(2/50 ±6.2pp)
HumanEval+ 100%
(50/50 ±3.6pp)
MBPP+ 80%
(40/50 ±10.9pp)
GPQA Diamond 88%
(44/50 ±9.1pp)
SciBench 66%
(33/50 ±12.7pp)
AIME 2024/2025 80%
(40/50 ±10.9pp)
MATH-500 96%
(48/50 ±6.2pp)
MMLU-Pro 90%
(45/50 ±8.5pp)
IFEval 92%
(46/50 ±7.8pp)
SciCode 0%
(0/50 ±3.6pp)
SuperGPQA 58%
(29/50 ±13.2pp)
Tau-Bench (Retail) 14%
(7/50 ±9.6pp)
| |||||||
| 🥉 | Mimo v2.5 ⚡ no thinking Mimo · mimo-v2.5 | 62.0% | 62.0% | - | - | - | - |
| Per-benchmark details · 3/12 completed Click headers to sort BigCodeBench-Hard 12%
(6/50 ±9.1pp)
HumanEval+ 98%
(49/50 ±5.1pp)
MBPP+ 76%
(38/50 ±11.6pp)
GPQA Diamond - SciBench - AIME 2024/2025 - MATH-500 - MMLU-Pro - IFEval - SciCode - SuperGPQA - Tau-Bench (Retail) - | |||||||
| 4 | DeepSeek v4 Flash 🧠 max DeepSeek · bai/deepseek-v4-flash | 56.8% | 42.0% | 52.0% | 77.0% | 68.0% | 45.0% |
| Per-benchmark details · 12/12 completed Click headers to sort BigCodeBench-Hard 2%
(1/50 ±5.1pp)
HumanEval+ 86%
(43/50 ±9.6pp)
MBPP+ 80%
(40/50 ±10.9pp)
GPQA Diamond 68%
(34/50 ±12.5pp)
SciBench 36%
(18/50 ±12.9pp)
AIME 2024/2025 78%
(39/50 ±11.2pp)
MATH-500 76%
(38/50 ±11.6pp)
MMLU-Pro 78%
(39/50 ±11.2pp)
IFEval 78%
(39/50 ±11.2pp)
SciCode 0%
(0/50 ±3.6pp)
SuperGPQA 58%
(29/50 ±13.2pp)
Tau-Bench (Retail) 12%
(6/50 ±9.1pp)
| |||||||
| 5 | Qwen 3.8 Flash 🧠 max Qwen · bai/qwen3.8-flash | 55.3% | 40.5% | 58.0% | 93.0% | 41.0% | 44.0% |
| Per-benchmark details · 12/12 completed Click headers to sort BigCodeBench-Hard 2%
(1/50 ±5.1pp)
HumanEval+ 90%
(45/50 ±8.5pp)
MBPP+ 70%
(35/50 ±12.3pp)
GPQA Diamond 64%
(32/50 ±12.9pp)
SciBench 52%
(26/50 ±13.3pp)
AIME 2024/2025 94%
(47/50 ±7.1pp)
MATH-500 92%
(46/50 ±7.8pp)
MMLU-Pro 82%
(41/50 ±10.5pp)
IFEval 88%
(44/50 ±9.1pp)
SciCode 0%
(0/50 ±3.6pp)
SuperGPQA 0%
(0/50 ±3.6pp)
Tau-Bench (Retail) 0%
(0/50 ±3.6pp)
| |||||||
| 6 | DeepSeek v4 Flash Vision Exp 🧠 max DeepSeek · bai/deepseek-v4-flash-vision-exp | 54.8% | 37.0% | 70.0% | 61.0% | 63.0% | 43.0% |
| Per-benchmark details · 12/12 completed Click headers to sort BigCodeBench-Hard 4%
(2/50 ±6.2pp)
HumanEval+ 68%
(34/50 ±12.5pp)
MBPP+ 76%
(38/50 ±11.6pp)
GPQA Diamond 92%
(46/50 ±7.8pp)
SciBench 48%
(24/50 ±13.3pp)
AIME 2024/2025 48%
(24/50 ±13.3pp)
MATH-500 74%
(37/50 ±11.8pp)
MMLU-Pro 64%
(32/50 ±12.9pp)
IFEval 76%
(38/50 ±11.6pp)
SciCode 0%
(0/50 ±3.6pp)
SuperGPQA 62%
(31/50 ±13pp)
Tau-Bench (Retail) 10%
(5/50 ±8.5pp)
| |||||||
| 7 | HY3 🧠 max Hunyuan · hy3 | 44.9% | 51.3% | 69.0% | 88.0% | 16.0% | 0.0% |
| Per-benchmark details · 9/12 completed Click headers to sort BigCodeBench-Hard 2%
(1/50 ±5.1pp)
HumanEval+ 72%
(36/50 ±12.1pp)
MBPP+ 80%
(40/50 ±10.9pp)
GPQA Diamond 86%
(43/50 ±9.6pp)
SciBench 52%
(26/50 ±13.3pp)
AIME 2024/2025 82%
(41/50 ±10.5pp)
MATH-500 94%
(47/50 ±7.1pp)
MMLU-Pro 16%
(8/50 ±10.1pp)
IFEval 0%
(0/50 ±3.6pp)
SciCode - SuperGPQA - Tau-Bench (Retail) - | |||||||
| - | Laguna S 2.1 Laguna · laguna-s-2.1 | - | - | - | - | - | - |
| Per-benchmark details · 0/12 completed Click headers to sort BigCodeBench-Hard - HumanEval+ - MBPP+ - GPQA Diamond - SciBench - AIME 2024/2025 - MATH-500 - MMLU-Pro - IFEval - SciCode - SuperGPQA - Tau-Bench (Retail) - | |||||||
| - | MiniMax M3 MiniMax · nvidia/minimaxai/minimax-m3 | - | - | - | - | - | - |
| Per-benchmark details · 0/12 completed Click headers to sort BigCodeBench-Hard - HumanEval+ - MBPP+ - GPQA Diamond - SciBench - AIME 2024/2025 - MATH-500 - MMLU-Pro - IFEval - SciCode - SuperGPQA - Tau-Bench (Retail) - | |||||||
| - | Qwen 3.8 Max Free Qwen · qwen/qwen3.8-max-free | - | - | - | - | - | - |
| Per-benchmark details · 0/12 completed Click headers to sort BigCodeBench-Hard - HumanEval+ - MBPP+ - GPQA Diamond - SciBench - AIME 2024/2025 - MATH-500 - MMLU-Pro - IFEval - SciCode - SuperGPQA - Tau-Bench (Retail) - | |||||||
Click a row for per-benchmark detail · click headers to sort · ±pp = 95% Wilson confidence interval Showing 10 of 10 models
DOMAIN SPECIALISTS
Category Arenas
Compare models strictly against peers within specialized task domains.
Best Coding AI Leader
Evaluated exclusively across BigCodeBench-Hard, HumanEval+, MBPP+, SciCode.
Coding Domain Ranking
Aggregated pass rate across coding benchmarks (hover bar for sub-test breakdown)
INDIVIDUAL BENCHMARK DRILL-DOWN
Benchmark Results & Curves
Inspect granular accuracy, question counts, and Wilson confidence intervals for any specific benchmark.
BigCodeBench-Hard
The hardest 148 practical Python programming tasks from BigCodeBench requiring deep integration of complex real-world libraries (pandas, numpy, scipy, etc.).
BigCodeBench-Hard · Accuracy Breakdown
Model score (%) and exact pass/total fraction (hover for Wilson confidence interval)
UNDER THE HOOD
Deep Analysis
hover any chart for exact values
Quality vs. Output Speed
Artificial Analysis style comparison: Overall Quality (%) against output throughput (tokens/second). Top-right is optimal.
Quality vs. Cost
Overall score against estimated API cost for a full run. Bubble size = output speed (tokens/sec).
Category Radar
Each model's profile across the five categories. Larger area = stronger all-round.
Category Breakdown
Head-to-head grouped comparison across categories.
Benchmark Heatmap
Per-benchmark scores for every model. Hover for exact counts and Wilson confidence intervals.
Token Mix
Input, thinking, and output tokens consumed per full run.
Thinking vs. Score
Do more thinking tokens buy higher scores?
THE GAUNTLET
Benchmark Catalog
12 benchmarks · fixed-seed sampling
| Benchmark | Category | Full dataset | Sampled | Verification | Source |
|---|---|---|---|---|---|
| BigCodeBench-Hard | coding | 148 | 50 | Python unittest execution (explicit opt-in required) | bigcode/bigcodebench-hard (v0.1.4) ↗ |
| HumanEval+ | coding | 164 | 50 | Python test execution (explicit opt-in required) | evalplus/humanevalplus ↗ |
| MBPP+ | coding | 378 | 50 | Python test execution (explicit opt-in required) | evalplus/mbppplus ↗ |
| GPQA Diamond | science | 198 | 50 | Multiple choice (4 options) | nichenshun/gpqa_diamond (community mirror of Idavidrein/gpqa) ↗ |
| SciBench | science | 692 | 50 | Numerical / Formula exact match | xw27/scibench ↗ |
| AIME 2024/2025 | math | 90 | 50 | Integer exact match (000-999) | AI-MO/aimo-validation-aime ↗ |
| MATH-500 | math | 500 | 50 | Exact match / \boxed{} extraction | HuggingFaceH4/MATH-500 ↗ |
| MMLU-Pro | knowledge | 12,032 | 50 | Multiple choice (10 options) | TIGER-Lab/MMLU-Pro ↗ |
| IFEval | instruction | 541 | 50 | 25 programmatic verifiers (strict) | google/IFEval ↗ |
| SciCode | coding | 65 | 50 | Python code execution & unit test assertions | SciCode1/SciCode ↗ |
| SuperGPQA | knowledge | 26,529 (7,050 hard) | 50 | Multiple choice (up to 10 options) | m-a-p/SuperGPQA ↗ |
| Tau-Bench (Retail) | instruction | 82 | 50 | Agentic tool-call function & argument matching | amityco/tau-bench-retail-train-next-action ↗ |
Hover a row for the benchmark description and citation.
EFFICIENCY
Tokens & Performance
| Model | Input | Output | Thinking | Total | Mix | TPS ⓘ | Avg time | Est. cost ⓘ |
|---|---|---|---|---|---|---|---|---|
| Gemini 3.7 Flash | 314,999 | 204,193 | 0 | 519,192 | 22.7 | 14.5s | - | |
| Gemini 3.5 Flash Lite | 314,999 | 202,689 | 0 | 517,688 | 78.4 | 6.2s | - | |
| Mimo v2.5 | 57,723 | 234,435 | 0 | 292,158 | 61.4 | 23.8s | - | |
| DeepSeek v4 Flash | 243,326 | 116,314 | 1,500,068 | 1,859,708 | 89.1 | 43.7s | - | |
| Qwen 3.8 Flash | 98,983 | 99,178 | 2,038,234 | 2,236,395 | 93.4 | 69.9s | - | |
| DeepSeek v4 Flash Vision Exp | 241,244 | 69,356 | 1,663,159 | 1,973,759 | 85.0 | 61.6s | - | |
| HY3 | 53,307 | 54,796 | 1,897,742 | 2,005,845 | 103.7 | 72.3s | - | |
| Laguna S 2.1 | 0 | 0 | 0 | 0 | - | - | - | |
| MiniMax M3 | 0 | 0 | 0 | 0 | - | - | - | |
| Qwen 3.8 Max Free | 0 | 0 | 0 | 0 | - | - | - |
TPS = output tokens/second (cloud APIs only) · cost via LiteLLM tables · totals per full benchmark run
TRUST THE NUMBERS
Methodology
🎲 Deterministic sampling
- ~50 questions are sampled from each benchmark's full dataset using a fixed seed (42) - the exact same questions on every run.
- Samples of n≈50 carry 95% Wilson confidence intervals of roughly ±7–14pp. Small ranking gaps are noise, not signal.
- Sampling and scoring strictness follow schema v2 (v0.2.0+); older runs are not directly comparable.
✅ Programmatic scoring only
- No LLM-as-judge anywhere. Answers are verified by code: letter extraction for multiple choice, boxed/numeric comparison for math, 25 strict verifiers for IFEval, function + argument matching for Tau-Bench.
- Code benchmarks require explicit opt-in and run in a 3-layer sandbox: AST scan → hardened subprocess (no keys, no network, temp dir) → Windows Job Object confinement.
📊 Score computation
- Category score = average of its benchmark scores across Coding, Science, Math, Knowledge, and Instruction.
- Overall score = equal-weight average of completed category scores. Provider outages are excluded, not scored as zero, so one bad API day doesn't sink a model.
⚙️ Inference settings
- temperature 0 · max_tokens 16384 · 300s timeout per request.
- All runs request max reasoning. Models badged ⚡ no thinking returned zero thinking tokens (the endpoint exposed no reasoning output) - their scores reflect answers alone.
- Transient errors retry with backoff until a good response arrives (same failure 3x in a row stops the question); permanent errors (context length, content filter) are never retried.