Clean leaderboard

The dedicated results view for one combined ranking across every fully accepted mean@N protocol. Filters preserve the official rank.

Results 0.1.0Data through Aug 17, 2026

Primary results

Combined Clean ranking

Official competition rank across every accepted protocol. Filters narrow the view; they never recompute rank.

Download CSV

Showing the first 25 of 243 configurations while the full leaderboard loads.

Combined GPQA Diamond Clean ranking
Rank
Rank movement

Compares this configuration's standard competition rank on Clean 189 with its paired rank on Original 198.

Improved
Fell
Unchanged

Ranks use exact scores. Filters never recompute them.

ModelClean scoreClean changeBenchmark source
Benchmark source classes

How each score entered the combined ranking.

  • Strict

    Complete public evaluation evidence validated under the current contract.

  • Compatibility

    Complete legacy evidence translated through a versioned compatibility validator.

  • Derived uniform

    A globally complete repeat subset reconstructed consistently for every question.

  • Targeted recovery

    Original attempts plus only the exact failed source slots recovered under a declared route.

  • Complete native primary

    A complete independent all-198 run on one declared primary route.

  • Native transport composite

    A complete independent source run plus documented exact-route transport recovery.

Full admission method →
1↑3 GPT 5.5 Pro pre-releaseOpenAI · mean@7 · xhighDerived uniform Clean score97.4% Clean change+3.4 ppfrom 93.9%
Benchmark sourceDerived uniform
1↑3 Gemini 3.6 FlashGoogle · mean@1 · highNative transport composite Clean score97.4% Clean change+3.4 ppfrom 93.9%
Benchmark sourceNative transport composite
3 GPT 5.5 pre-releaseOpenAI · mean@8 · xhighStrict Clean score97.2% Clean change+3.2 ppfrom 94.0%
Benchmark sourceStrict
4↓3 Gemini 3.1 Pro PreviewGoogle · mean@8 · defaultTargeted recovery Clean score97.0% Clean change+2.5 ppfrom 94.5%
Benchmark sourceTargeted recovery
5↑1 GPT 5.6 SolOpenAI · mean@1 · maxComplete native primary Clean score96.8% Clean change+3.4 ppfrom 93.4%
Benchmark sourceComplete native primary
6↑1 GPT 5.4OpenAI · mean@8 · xhighTargeted recovery Clean score96.7% Clean change+3.4 ppfrom 93.2%
Benchmark sourceTargeted recovery
7↓5 Gemini 3.1 Pro PreviewGoogle · mean@1 · highStrict Clean score96.3% Clean change+1.9 ppfrom 94.4%
Benchmark sourceStrict
8↑2 Gemini 3.5 FlashGoogle · mean@8 · highStrict Clean score96.1% Clean change+3.3 ppfrom 92.8%
Benchmark sourceStrict
9↑3 Grok 4.6xAI · mean@1 · highComplete native primary Clean score95.8% Clean change+3.3 ppfrom 92.4%
Benchmark sourceComplete native primary
9↓1 DeepSeek V4 FlashDeepSeek · mean@1 · maxComplete native primary Clean score95.8% Clean change+2.8 ppfrom 92.9%
Benchmark sourceComplete native primary
11 Gemini 3 Pro PreviewGoogle · mean@8 · defaultStrict Clean score95.5% Clean change+2.9 ppfrom 92.6%
Benchmark sourceStrict
12 DeepSeek V4 ProDeepSeek · mean@1 · maxComplete native primary Clean score95.2% Clean change+2.8 ppfrom 92.4%
Benchmark sourceComplete native primary
12↓4 Claude Opus 5Anthropic · mean@1 · noneStrict Clean score95.2% Clean change+2.3 ppfrom 92.9%
Benchmark sourceStrict
14↑2 GPT 5.2OpenAI · mean@8 · xhighStrict Clean score95.0% Clean change+3.6 ppfrom 91.4%
Benchmark sourceStrict
15↓3 GPT 5.6 LunaOpenAI · mean@1 · maxComplete native primary Clean score94.7% Clean change+2.3 ppfrom 92.4%
Benchmark sourceComplete native primary
15↑2 Qwen3.7 MaxQwen · mean@1 · maxStrict Clean score94.7% Clean change+3.8 ppfrom 90.9%
Benchmark sourceStrict
15 Kimi K3Moonshot AI · mean@1 · highStrict Clean score94.7% Clean change+2.8 ppfrom 91.9%
Benchmark sourceStrict
18↑4 GPT 5.5OpenAI · mean@8 · lowStrict Clean score93.8% Clean change+3.2 ppfrom 90.7%
Benchmark sourceStrict
19↓2 Kimi K2.6Moonshot AI · mean@6 · defaultDerived uniform Clean score93.8% Clean change+2.9 ppfrom 90.9%
Benchmark sourceDerived uniform
20↓3 GPT 5.6 TerraOpenAI · mean@1 · maxComplete native primary Clean score93.7% Clean change+2.7 ppfrom 90.9%
Benchmark sourceComplete native primary
20↓3 MiniMax M3MiniMax · mean@1 · defaultStrict Clean score93.7% Clean change+2.7 ppfrom 90.9%
Benchmark sourceStrict
20↑6 GPT 5.4OpenAI · mean@1 · highStrict Clean score93.7% Clean change+3.8 ppfrom 89.9%
Benchmark sourceStrict
20↑4 DeepSeek V4 FlashDeepSeek · mean@1 · highComplete native primary Clean score93.7% Clean change+3.2 ppfrom 90.4%
Benchmark sourceComplete native primary
20↑5 Claude Opus 4.7Anthropic · mean@8 · xhighStrict Clean score93.7% Clean change+3.5 ppfrom 90.2%
Benchmark sourceStrict
20↑6 GPT 5.6 SolOpenAI · mean@1 · lowStrict Clean score93.7% Clean change+3.8 ppfrom 89.9%
Benchmark sourceStrict

How ranking works

All accepted protocols share one standard competition rank ordered by the exact Clean score. Equal exact rational scores share rank, and the next rank skips the tied places. Displayed one-decimal values may coincide; they never create a ranking tie.

How to read uncertainty

The observed rank remains primary. Each expanded record shows its plausible rank interval from one synchronized Record-ID bootstrap shared across configurations.