Cost–score frontier

Compare exact Clean score with the Cost of one normalized 189-question pass.
Cheaper configurations appear farther right; the black line is the global Pareto reference.

Results 0.1.0Data through Aug 17, 2026

Display
227 of 243 plotted16 Cost unavailable5 DeepSeek configurations use the selected pricing period

GPQA Diamond Clean · Cost

Global Cost–score frontier

Filters highlight configurations without changing the global reference. Only the DeepSeek pricing-period selector recomputes the frontier.

  • Color = manufacturer
  • Diamond = reasoning setting
  • Circle = none, minimal, or default
  • Black line = global Pareto reference

Loading verified Cost data…

Clean scores use exact numerators and denominators. Cost uses a logarithmic scale with cheaper configurations farther right.
Cost for 189 questions, one attempt eachPrices dated 2026-08-31Results 0.1.0 · source 39e9c3f05fef

Accessible chart summary

Verified chart data is loading.

How Cost and the frontier are calculated

Each x-coordinate is the public API Cost for the configuration’s observed average tokens across 189 retained questions, normalized to one attempt per question. The y-coordinate is its exact Clean score. A point is on the global Pareto frontier when no configuration is both cheaper and at least as accurate, or equally priced and more accurate. Configurations with unavailable Cost are excluded rather than treated as zero.

DeepSeek publishes scheduled off-peak and peak prices. Selecting a period changes those five x-coordinates and recomputes the global frontier. Manufacturer, model, effort, protocol, focus, and near-ceiling controls only change presentation.