PROGRAMMATIC LLM EVALUATION · NO LLM-AS-JUDGE

The leaderboard built on verifiable scores.

12 established benchmarks, ~50 questions each, sampled with a fixed seed and scored entirely by code - regex extraction, exact match, and sandboxed test execution. Every number on this page is reproducible.

10
Models tested
12
Benchmarks
600
Questions sampled
42
Sampling seed
🏆 73.9%
Top · Gemini 3.7 Flash
Sep 3, 2026
Last run (UTC)

RANKINGS

Overall Leaderboard

10 models · seed 42 · n≈50 per benchmark

# Model Overall 💻 Coding 🔬 Science 📐 Math 📚 Knowledge 📋 Instruction
🥇 Gemini 3.7 Flash ⚡ no thinking Google · gemini/gemini-3.7-flash 73.9% 48.5% 84.0% 97.0% 82.0% 58.0%

Per-benchmark details · 12/12 completed

Click headers to sort
BigCodeBench-Hard 6% (3/50 ±7.1pp) HumanEval+ 100% (50/50 ±3.6pp) MBPP+ 88% (44/50 ±9.1pp) GPQA Diamond 100% (50/50 ±3.6pp) SciBench 68% (34/50 ±12.5pp) AIME 2024/2025 98% (49/50 ±5.1pp) MATH-500 96% (48/50 ±6.2pp) MMLU-Pro 90% (45/50 ±8.5pp) IFEval 92% (46/50 ±7.8pp) SciCode 0% (0/50 ±3.6pp) SuperGPQA 74% (37/50 ±11.8pp) Tau-Bench (Retail) 24% (12/50 ±11.6pp)
🥈 Gemini 3.5 Flash Lite ⚡ no thinking Google · gemini/gemini-3.5-flash-lite 67.6% 46.0% 77.0% 88.0% 74.0% 53.0%

Per-benchmark details · 12/12 completed

Click headers to sort
BigCodeBench-Hard 4% (2/50 ±6.2pp) HumanEval+ 100% (50/50 ±3.6pp) MBPP+ 80% (40/50 ±10.9pp) GPQA Diamond 88% (44/50 ±9.1pp) SciBench 66% (33/50 ±12.7pp) AIME 2024/2025 80% (40/50 ±10.9pp) MATH-500 96% (48/50 ±6.2pp) MMLU-Pro 90% (45/50 ±8.5pp) IFEval 92% (46/50 ±7.8pp) SciCode 0% (0/50 ±3.6pp) SuperGPQA 58% (29/50 ±13.2pp) Tau-Bench (Retail) 14% (7/50 ±9.6pp)
🥉 Mimo v2.5 ⚡ no thinking Mimo · mimo-v2.5 62.0% 62.0% - - - -

Per-benchmark details · 3/12 completed

Click headers to sort
BigCodeBench-Hard 12% (6/50 ±9.1pp) HumanEval+ 98% (49/50 ±5.1pp) MBPP+ 76% (38/50 ±11.6pp) GPQA Diamond - SciBench - AIME 2024/2025 - MATH-500 - MMLU-Pro - IFEval - SciCode - SuperGPQA - Tau-Bench (Retail) -
4 DeepSeek v4 Flash 🧠 max DeepSeek · bai/deepseek-v4-flash 56.8% 42.0% 52.0% 77.0% 68.0% 45.0%

Per-benchmark details · 12/12 completed

Click headers to sort
BigCodeBench-Hard 2% (1/50 ±5.1pp) HumanEval+ 86% (43/50 ±9.6pp) MBPP+ 80% (40/50 ±10.9pp) GPQA Diamond 68% (34/50 ±12.5pp) SciBench 36% (18/50 ±12.9pp) AIME 2024/2025 78% (39/50 ±11.2pp) MATH-500 76% (38/50 ±11.6pp) MMLU-Pro 78% (39/50 ±11.2pp) IFEval 78% (39/50 ±11.2pp) SciCode 0% (0/50 ±3.6pp) SuperGPQA 58% (29/50 ±13.2pp) Tau-Bench (Retail) 12% (6/50 ±9.1pp)
5 Qwen 3.8 Flash 🧠 max Qwen · bai/qwen3.8-flash 55.3% 40.5% 58.0% 93.0% 41.0% 44.0%

Per-benchmark details · 12/12 completed

Click headers to sort
BigCodeBench-Hard 2% (1/50 ±5.1pp) HumanEval+ 90% (45/50 ±8.5pp) MBPP+ 70% (35/50 ±12.3pp) GPQA Diamond 64% (32/50 ±12.9pp) SciBench 52% (26/50 ±13.3pp) AIME 2024/2025 94% (47/50 ±7.1pp) MATH-500 92% (46/50 ±7.8pp) MMLU-Pro 82% (41/50 ±10.5pp) IFEval 88% (44/50 ±9.1pp) SciCode 0% (0/50 ±3.6pp) SuperGPQA 0% (0/50 ±3.6pp) Tau-Bench (Retail) 0% (0/50 ±3.6pp)
6 DeepSeek v4 Flash Vision Exp 🧠 max DeepSeek · bai/deepseek-v4-flash-vision-exp 54.8% 37.0% 70.0% 61.0% 63.0% 43.0%

Per-benchmark details · 12/12 completed

Click headers to sort
BigCodeBench-Hard 4% (2/50 ±6.2pp) HumanEval+ 68% (34/50 ±12.5pp) MBPP+ 76% (38/50 ±11.6pp) GPQA Diamond 92% (46/50 ±7.8pp) SciBench 48% (24/50 ±13.3pp) AIME 2024/2025 48% (24/50 ±13.3pp) MATH-500 74% (37/50 ±11.8pp) MMLU-Pro 64% (32/50 ±12.9pp) IFEval 76% (38/50 ±11.6pp) SciCode 0% (0/50 ±3.6pp) SuperGPQA 62% (31/50 ±13pp) Tau-Bench (Retail) 10% (5/50 ±8.5pp)
7 HY3 🧠 max Hunyuan · hy3 44.9% 51.3% 69.0% 88.0% 16.0% 0.0%

Per-benchmark details · 9/12 completed

Click headers to sort
BigCodeBench-Hard 2% (1/50 ±5.1pp) HumanEval+ 72% (36/50 ±12.1pp) MBPP+ 80% (40/50 ±10.9pp) GPQA Diamond 86% (43/50 ±9.6pp) SciBench 52% (26/50 ±13.3pp) AIME 2024/2025 82% (41/50 ±10.5pp) MATH-500 94% (47/50 ±7.1pp) MMLU-Pro 16% (8/50 ±10.1pp) IFEval 0% (0/50 ±3.6pp) SciCode - SuperGPQA - Tau-Bench (Retail) -
- Laguna S 2.1 Laguna · laguna-s-2.1 - - - - - -

Per-benchmark details · 0/12 completed

Click headers to sort
BigCodeBench-Hard - HumanEval+ - MBPP+ - GPQA Diamond - SciBench - AIME 2024/2025 - MATH-500 - MMLU-Pro - IFEval - SciCode - SuperGPQA - Tau-Bench (Retail) -
- MiniMax M3 MiniMax · nvidia/minimaxai/minimax-m3 - - - - - -

Per-benchmark details · 0/12 completed

Click headers to sort
BigCodeBench-Hard - HumanEval+ - MBPP+ - GPQA Diamond - SciBench - AIME 2024/2025 - MATH-500 - MMLU-Pro - IFEval - SciCode - SuperGPQA - Tau-Bench (Retail) -
- Qwen 3.8 Max Free Qwen · qwen/qwen3.8-max-free - - - - - -

Per-benchmark details · 0/12 completed

Click headers to sort
BigCodeBench-Hard - HumanEval+ - MBPP+ - GPQA Diamond - SciBench - AIME 2024/2025 - MATH-500 - MMLU-Pro - IFEval - SciCode - SuperGPQA - Tau-Bench (Retail) -

Click a row for per-benchmark detail · click headers to sort · ±pp = 95% Wilson confidence interval Showing 10 of 10 models

DOMAIN SPECIALISTS

Category Arenas

Compare models strictly against peers within specialized task domains.

💻 Coding Domain 4 benchmarks

Best Coding AI Leader

Evaluated exclusively across BigCodeBench-Hard, HumanEval+, MBPP+, SciCode.

Domain Champion 👑
Mimo v2.5
Mimo · 61.4 tok/s
62.0% aggregate category pass rate
Constituent Benchmarks
BigCodeBench-Hard 50 questions
HumanEval+ 50 questions
MBPP+ 50 questions
SciCode 50 questions

Coding Domain Ranking

Aggregated pass rate across coding benchmarks (hover bar for sub-test breakdown)

higher is better

INDIVIDUAL BENCHMARK DRILL-DOWN

Benchmark Results & Curves

Inspect granular accuracy, question counts, and Wilson confidence intervals for any specific benchmark.

CODING 50 samples (seed 42)

BigCodeBench-Hard

The hardest 148 practical Python programming tasks from BigCodeBench requiring deep integration of complex real-world libraries (pandas, numpy, scipy, etc.).

Full Dataset Size: 148 items
Verification: Python unittest execution (explicit opt-in required)
Paper / Source: Zhuo et al. 2024
Error bounds: 95% Wilson Score Interval

BigCodeBench-Hard · Accuracy Breakdown

Model score (%) and exact pass/total fraction (hover for Wilson confidence interval)

n≈50

UNDER THE HOOD

Deep Analysis

hover any chart for exact values

Quality vs. Output Speed

Artificial Analysis style comparison: Overall Quality (%) against output throughput (tokens/second). Top-right is optimal.

Quality vs. Cost

Overall score against estimated API cost for a full run. Bubble size = output speed (tokens/sec).

Category Radar

Each model's profile across the five categories. Larger area = stronger all-round.

Category Breakdown

Head-to-head grouped comparison across categories.

Benchmark Heatmap

Per-benchmark scores for every model. Hover for exact counts and Wilson confidence intervals.

Token Mix

Input, thinking, and output tokens consumed per full run.

Thinking vs. Score

Do more thinking tokens buy higher scores?

THE GAUNTLET

Benchmark Catalog

12 benchmarks · fixed-seed sampling

Benchmark Category Full dataset Sampled Verification Source
BigCodeBench-Hard coding 148 50 Python unittest execution (explicit opt-in required) bigcode/bigcodebench-hard (v0.1.4) ↗
HumanEval+ coding 164 50 Python test execution (explicit opt-in required) evalplus/humanevalplus ↗
MBPP+ coding 378 50 Python test execution (explicit opt-in required) evalplus/mbppplus ↗
GPQA Diamond science 198 50 Multiple choice (4 options) nichenshun/gpqa_diamond (community mirror of Idavidrein/gpqa) ↗
SciBench science 692 50 Numerical / Formula exact match xw27/scibench ↗
AIME 2024/2025 math 90 50 Integer exact match (000-999) AI-MO/aimo-validation-aime ↗
MATH-500 math 500 50 Exact match / \boxed{} extraction HuggingFaceH4/MATH-500 ↗
MMLU-Pro knowledge 12,032 50 Multiple choice (10 options) TIGER-Lab/MMLU-Pro ↗
IFEval instruction 541 50 25 programmatic verifiers (strict) google/IFEval ↗
SciCode coding 65 50 Python code execution & unit test assertions SciCode1/SciCode ↗
SuperGPQA knowledge 26,529 (7,050 hard) 50 Multiple choice (up to 10 options) m-a-p/SuperGPQA ↗
Tau-Bench (Retail) instruction 82 50 Agentic tool-call function & argument matching amityco/tau-bench-retail-train-next-action ↗

Hover a row for the benchmark description and citation.

EFFICIENCY

Tokens & Performance

Input Thinking Output
Model Input Output Thinking Total Mix TPS ⓘ Avg time Est. cost ⓘ
Gemini 3.7 Flash 314,999 204,193 0 519,192 22.7 14.5s -
Gemini 3.5 Flash Lite 314,999 202,689 0 517,688 78.4 6.2s -
Mimo v2.5 57,723 234,435 0 292,158 61.4 23.8s -
DeepSeek v4 Flash 243,326 116,314 1,500,068 1,859,708 89.1 43.7s -
Qwen 3.8 Flash 98,983 99,178 2,038,234 2,236,395 93.4 69.9s -
DeepSeek v4 Flash Vision Exp 241,244 69,356 1,663,159 1,973,759 85.0 61.6s -
HY3 53,307 54,796 1,897,742 2,005,845 103.7 72.3s -
Laguna S 2.1 0 0 0 0 - - -
MiniMax M3 0 0 0 0 - - -
Qwen 3.8 Max Free 0 0 0 0 - - -

TPS = output tokens/second (cloud APIs only) · cost via LiteLLM tables · totals per full benchmark run

TRUST THE NUMBERS

Methodology

full methodology in the README →
🎲 Deterministic sampling
  • ~50 questions are sampled from each benchmark's full dataset using a fixed seed (42) - the exact same questions on every run.
  • Samples of n≈50 carry 95% Wilson confidence intervals of roughly ±7–14pp. Small ranking gaps are noise, not signal.
  • Sampling and scoring strictness follow schema v2 (v0.2.0+); older runs are not directly comparable.
✅ Programmatic scoring only
  • No LLM-as-judge anywhere. Answers are verified by code: letter extraction for multiple choice, boxed/numeric comparison for math, 25 strict verifiers for IFEval, function + argument matching for Tau-Bench.
  • Code benchmarks require explicit opt-in and run in a 3-layer sandbox: AST scan → hardened subprocess (no keys, no network, temp dir) → Windows Job Object confinement.
📊 Score computation
  • Category score = average of its benchmark scores across Coding, Science, Math, Knowledge, and Instruction.
  • Overall score = equal-weight average of completed category scores. Provider outages are excluded, not scored as zero, so one bad API day doesn't sink a model.
⚙️ Inference settings
  • temperature 0 · max_tokens 16384 · 300s timeout per request.
  • All runs request max reasoning. Models badged ⚡ no thinking returned zero thinking tokens (the endpoint exposed no reasoning output) - their scores reflect answers alone.
  • Transient errors retry with backoff until a good response arrives (same failure 3x in a row stops the question); permanent errors (context length, content filter) are never retried.