EffBench Leaderboard

Efficiency is always shown beside reliability and token-variance — never ranked alone.

How this is ranked: Ranked by composite score = correct_per_1k_output * reliability (reliability in [0,1]), so efficiency is discounted by how often the model is actually right. Ties break by reliability (higher first), then cost (lower first). Efficiency, reliability and token-variance (CV) are shown as their own columns — never collapse them into the score.
Models ranked best-first by composite score (= efficiency × reliability)
# provider model score reliability ↑ efficiency ↑ tok CV ↓ cost (est.) ↓
1anthropicclaude-sonnet-4-6 illustrative40.541.0040.540.05$0.006400
2openaigpt-5.4-mini illustrative38.850.8744.830.18$0.003300
3openaispeedy-but-flaky-mini illustrative26.560.4756.910.54$0.001300
4anthropicclaude-opus-4-8 illustrative15.311.0015.310.10$0.027600
Legend (↑ = higher is better, ↓ = lower is better)
reliability ↑ — pass-rate across repetitions, in [0, 1].
efficiency ↑ — correct answers per 1,000 output tokens (correct_per_1k_output).
tok CV ↓ — run-to-run output-token coefficient of variation (mean of per-task CVs); lower = more predictable cost.
cost (est.) ↓ — estimated total USD; lower is cheaper.