PROGRAMMATIC LLM EVALUATION · NO LLM-AS-JUDGE

The leaderboard built on verifiable scores.

12 established benchmarks, ~50 questions each, sampled with a fixed seed and scored entirely by code — regex extraction, exact match, and sandboxed test execution. Every number on this page is reproducible.

5
Models tested
12
Benchmarks
600
Questions sampled
42
Sampling seed
🏆 50.1%
Top · Deepseek v4 Pro Max
Jul 28, 2026
Last run (UTC)

RANKINGS

Overall Leaderboard

5 models · seed 42 · n≈50 per benchmark

# Model Overall 💻 Coding 🔬 Science 📐 Math 📚 Knowledge 📋 Instruction
🥇 Deepseek v4 Pro Max 🧠 max DeepSeek · deepseek/deepseek-v4-pro 50.1% 31.5% 53.0% 60.0% 59.0% 47.0%

Per-benchmark details · 12/12 completed

Click headers to sort
BigCodeBench-Hard 22% (11/50 ±11.2pp) HumanEval+ 98% (49/50 ±5.1pp) MBPP+ 6% (3/50 ±7.1pp) GPQA Diamond 60% (30/50 ±13.1pp) SciBench 46% (23/50 ±13.3pp) AIME 2024/2025 36% (18/50 ±12.9pp) MATH-500 84% (42/50 ±10.1pp) MMLU-Pro 78% (39/50 ±11.2pp) IFEval 84% (42/50 ±10.1pp) SciCode 0% (0/50 ±3.6pp) SuperGPQA 40% (20/50 ±13.1pp) Tau-Bench (Retail) 10% (5/50 ±8.5pp)
🥈 Deepseek v4 Flash Max 🧠 max DeepSeek · deepseek/deepseek-v4-flash 47.8% 30.0% 52.0% 55.0% 53.0% 49.0%

Per-benchmark details · 12/12 completed

Click headers to sort
BigCodeBench-Hard 22% (11/50 ±11.2pp) HumanEval+ 88% (44/50 ±9.1pp) MBPP+ 10% (5/50 ±8.5pp) GPQA Diamond 56% (28/50 ±13.3pp) SciBench 48% (24/50 ±13.3pp) AIME 2024/2025 28% (14/50 ±12.1pp) MATH-500 82% (41/50 ±10.5pp) MMLU-Pro 78% (39/50 ±11.2pp) IFEval 86% (43/50 ±9.6pp) SciCode 0% (0/50 ±3.6pp) SuperGPQA 28% (14/50 ±12.1pp) Tau-Bench (Retail) 12% (6/50 ±9.1pp)
🥉 Deepseek v4 Flash DeepSeek · deepseek/deepseek-v4-flash 46.1% 28.5% 46.0% 55.0% 55.0% 46.0%

Per-benchmark details · 12/12 completed

Click headers to sort
BigCodeBench-Hard 16% (8/50 ±10.1pp) HumanEval+ 88% (44/50 ±9.1pp) MBPP+ 10% (5/50 ±8.5pp) GPQA Diamond 52% (26/50 ±13.3pp) SciBench 40% (20/50 ±13.1pp) AIME 2024/2025 26% (13/50 ±11.8pp) MATH-500 84% (42/50 ±10.1pp) MMLU-Pro 78% (39/50 ±11.2pp) IFEval 84% (42/50 ±10.1pp) SciCode 0% (0/50 ±3.6pp) SuperGPQA 32% (16/50 ±12.5pp) Tau-Bench (Retail) 8% (4/50 ±7.8pp)
4 Gemma 4 31B 🧠 high Google · gemini/gemma-4-31b-it 20.6% 0.0% 42.0% 20.0%

Per-benchmark details · 4/12 completed

Click headers to sort
BigCodeBench-Hard 0% (0/47 ±3.8pp) HumanEval+ MBPP+ GPQA Diamond 62% (29/47 ±13.4pp) SciBench 22% (10/45 ±11.9pp) AIME 2024/2025 20% (1/5 ±29.4pp) MATH-500 MMLU-Pro IFEval SciCode SuperGPQA Tau-Bench (Retail)
5 Gemma 4 26B a4b 🧠 high Google · gemini/gemma-4-26B-A4B-it

Per-benchmark details · 0/12 completed

Click headers to sort
BigCodeBench-Hard (0/0 ±0pp) HumanEval+ MBPP+ GPQA Diamond (0/0 ±0pp) SciBench AIME 2024/2025 MATH-500 MMLU-Pro IFEval SciCode SuperGPQA Tau-Bench (Retail)

Click a row for per-benchmark detail · click headers to sort · ±pp = 95% Wilson confidence interval Showing 5 of 5 models

UNDER THE HOOD

Analysis

hover any chart for exact values

Quality vs. Cost

Overall score against estimated API cost for a full run. Bubble size = output speed (tokens/sec).

Category Radar

Each model's profile across the five categories. Larger area = stronger all-round.

Category Breakdown

Head-to-head grouped comparison across categories.

Benchmark Heatmap

Per-benchmark scores for every model. Hover for exact counts and Wilson confidence intervals.

Token Mix

Input, thinking, and output tokens consumed per full run.

Thinking vs. Score

Do more thinking tokens buy higher scores?

THE GAUNTLET

Benchmark Catalog

12 benchmarks · fixed-seed sampling

Benchmark Category Full dataset Sampled Verification Source
BigCodeBench-Hard coding 148 50 Python unittest execution (explicit opt-in required) bigcode/bigcodebench-hard (v0.1.4) ↗
HumanEval+ coding 164 50 Python test execution (explicit opt-in required) evalplus/humanevalplus ↗
MBPP+ coding 378 50 Python test execution (explicit opt-in required) evalplus/mbppplus ↗
GPQA Diamond science 198 50 Multiple choice (4 options) nichenshun/gpqa_diamond (community mirror of Idavidrein/gpqa) ↗
SciBench science 692 50 Numerical / Formula exact match xw27/scibench ↗
AIME 2024/2025 math 90 50 Integer exact match (000-999) AI-MO/aimo-validation-aime ↗
MATH-500 math 500 50 Exact match / \boxed{} extraction HuggingFaceH4/MATH-500 ↗
MMLU-Pro knowledge 12,032 50 Multiple choice (10 options) TIGER-Lab/MMLU-Pro ↗
IFEval instruction 541 50 25 programmatic verifiers (strict) google/IFEval ↗
SciCode coding 65 50 Python code execution & unit test assertions SciCode1/SciCode ↗
SuperGPQA knowledge 26,529 (7,050 hard) 50 Multiple choice (up to 10 options) m-a-p/SuperGPQA ↗
Tau-Bench (Retail) instruction 82 50 Agentic tool-call function & argument matching amityco/tau-bench-retail-train-next-action ↗

Hover a row for the benchmark description and citation.

EFFICIENCY

Tokens & Performance

Input Thinking Output
Model Input Output Thinking Total Mix TPS ⓘ Avg time Est. cost ⓘ
Deepseek v4 Pro Max 273,638 85,211 970,381 1,329,230 49.9 38.3s $1.04
Deepseek v4 Flash Max 273,638 71,636 1,012,225 1,357,499 93.0 21.1s $0.3410
Deepseek v4 Flash 273,638 79,896 995,022 1,348,556 92.9 21.1s $0.3344
Gemma 4 31B 30,457 25,685 0 511,427 7.4 98.9s
Gemma 4 26B a4b 0 0 0 0

TPS = output tokens/second (cloud APIs only) · cost via LiteLLM tables · totals per full benchmark run

TRUST THE NUMBERS

Methodology

full methodology in the README →
🎲 Deterministic sampling
  • ~50 questions are sampled from each benchmark's full dataset using a fixed seed (42) — the exact same questions on every run.
  • Samples of n≈50 carry 95% Wilson confidence intervals of roughly ±7–14pp. Small ranking gaps are noise, not signal.
  • Sampling and scoring strictness follow schema v2 (v0.2.0+); older runs are not directly comparable.
Programmatic scoring only
  • No LLM-as-judge anywhere. Answers are verified by code: letter extraction for multiple choice, boxed/numeric comparison for math, 25 strict verifiers for IFEval, function + argument matching for Tau-Bench.
  • Code benchmarks require explicit opt-in and run in a 3-layer sandbox: AST scan → hardened subprocess (no keys, no network, temp dir) → Windows Job Object confinement.
📊 Score computation
  • Category score = average of its benchmark scores across Coding, Science, Math, Knowledge, and Instruction.
  • Overall score = equal-weight average of completed category scores. Provider outages are excluded, not scored as zero, so one bad API day doesn't sink a model.
⚙️ Inference settings
  • temperature 0 · max_tokens 4096 · 300s timeout per request.
  • Transient errors retry with exponential backoff until a good response arrives; permanent errors (context length, content filter) are never retried.