Efficiency is always shown beside reliability and token-variance — never ranked alone.
Illustrative / sample data — not real benchmark results.
When labelled ILLUSTRATIVE / sample, these rows are hand-crafted demo data to show how the ranking behaves — they are NOT real benchmark results.
Cross-model caveat
Raw tokens-per-answer is NOT directly comparable across providers: different models tokenize the same text differently, so the same answer can cost a different number of tokens on each provider. Read the cross-model ranking with that in mind — it compares estimates, not like-for-like token counts.
How this is ranked: Ranked by composite score = correct_per_1k_output * reliability (reliability in [0,1]), so efficiency is discounted by how often the model is actually right. Ties break by reliability (higher first), then cost (lower first). Efficiency, reliability and token-variance (CV) are shown as their own columns — never collapse them into the score.
Models ranked best-first by composite score (= efficiency × reliability)
#
provider
model
score
reliability ↑
efficiency ↑
tok CV ↓
cost (est.) ↓
1
anthropic
claude-sonnet-4-6 illustrative
40.54
1.00
40.54
0.05
$0.006400
2
openai
gpt-5.4-mini illustrative
38.85
0.87
44.83
0.18
$0.003300
3
openai
speedy-but-flaky-mini illustrative
26.56
0.47
56.91
0.54
$0.001300
4
anthropic
claude-opus-4-8 illustrative
15.31
1.00
15.31
0.10
$0.027600
Legend (↑ = higher is better, ↓ = lower is better)
reliability ↑ — pass-rate across repetitions, in [0, 1].
efficiency ↑ — correct answers per 1,000 output tokens (correct_per_1k_output).
tok CV ↓ — run-to-run output-token coefficient of variation (mean of per-task CVs); lower = more predictable cost.
cost (est.) ↓ — estimated total USD; lower is cheaper.