Cheapest LLM that's good enough.
We run over 150,000 large-language-model API calls per week and let the outputs be evaluated by a panel of leading models. Cost per token doesn't track quality — and when you say something like "give me the model that's reliably 95% as good as the best one but most cost-effective", the best-value model typically saves north of 90% of the cost.
This benchmark publishes our results so other builders can answer the same question for their own pipelines: which model(s) are good enough and the most cost-efficient for each category of tasks. The data also shows that the most expensive model — even when your budget can pay for it — is not always the best performer.
Have a look. Tell us what you think. We can also run tests for you — reach us at llmbench@kapualabs.com.
Snapshot: 2026-07-20
Right-sized per step
Best-value model per capability, at your quality bar. Greyed cells mean no model qualifies on that category yet.
1.0x = this model is the best-value good-enough option, no overpayment; 3.2x = you pay ~3× more on average for the same quality).
★ marks any model that is the best-value pick on at least one of the category's tasks (multiple stars per column are possible — they're the per-task cost winners).
Greyed cells: model doesn't qualify on any task in that category at the current bar.
Rows grouped by provider; click a model name for the full per-task breakdown.By capability
Financial Analysis & Trading Decisions
Read business/financial docs deeply enough to make forward-looking judgments about opportunity and risk.
| Model | Tasks | Avg $/1k runs | Overpay |
|---|---|---|---|
| MiniMax M3 ★ | 3/4 | $3.02 | 1.3x |
| Qwen 3.5 Flash ★ | 1/2 | $3.11 | best value |
| GPT-5.4 Nano ★ | 2/2 | $7.01 | 1.3x |
| GPT-5.6 Luna | 1/3 | $7.68 | 2.5x |
| Qwen 3.7 Plus | 3/3 | $9.63 | 2.7x |
Structured Data & Fact Extraction
Precise pattern recognition and field-level information retrieval, strict schema adherence.
| Model | Tasks | Avg $/1k runs | Overpay |
|---|---|---|---|
| DeepSeek V4 Flash ★ | 3/3 | $0.58 | best value |
| Qwen 3.5 Flash | 1/2 | $0.75 | 1.1x |
| Gemini 3.1 Flash Lite | 1/3 | $0.83 | 1.2x |
| NVIDIA Nemotron-3 Super 120B | 1/1 | $1.39 | 2x |
| MiniMax M3 | 2/3 | $3.36 | 7.6x |
Content Summarization & Synthesis
Compress long input into essential nuggets without losing material detail.
| Model | Tasks | Avg $/1k runs | Overpay |
|---|---|---|---|
| DeepSeek V4 Flash ★ | 1/2 | $0.71 | best value |
| Tencent Hy3 | 2/2 | $2.54 | 2.7x |
| MiniMax M3 ★ | 2/3 | $3.71 | 1.6x |
| NVIDIA Nemotron-3 Ultra 550B | 1/2 | $4.53 | 6.4x |
| GPT-5.4 Nano ★ | 2/3 | $8.00 | 1.8x |
Long-form Content Generation
Sustained compositional skill, voice consistency, coherent extended prose.
| Model | Tasks | Avg $/1k runs | Overpay |
|---|---|---|---|
| NVIDIA Nemotron-3 Nano 30B-A3B ★ | 1/3 | $0.28 | best value |
| NVIDIA Nemotron-3 Super 120B ★ | 2/4 | $1.08 | best value |
| Tencent Hy3 ★ | 3/5 | $1.44 | 2.2x |
| DeepSeek V4 Flash | 1/2 | $1.53 | 1.8x |
| MiniMax M3 ★ | 4/4 | $1.59 | 1.9x |
Social & Promotional Content
Conciseness, platform-native conventions, engagement under tight character limits.
| Model | Tasks | Avg $/1k runs | Overpay |
|---|---|---|---|
| NVIDIA Nemotron-3 Nano 30B-A3B | 1/2 | $0.30 | 1.1x |
| DeepSeek V4 Flash ★ | 2/4 | $0.42 | best value |
| GPT-5.4 Mini | 1/5 | $0.52 | 1.9x |
| NVIDIA Nemotron-3 Super 120B | 1/3 | $0.85 | 3.1x |
| Tencent Hy3 ★ | 2/4 | $1.30 | 2.7x |
Relevance, Classification & Matching
Semantic similarity judgment: does this thing belong in that bucket / match that target?
| Model | Tasks | Avg $/1k runs | Overpay |
|---|---|---|---|
| NVIDIA Nemotron-3 Nano 30B-A3B ★ | 1/2 | $0.04 | best value |
| GPT-5.4 Nano | 2/9 | $0.14 | 1.6x |
| Gemini 3.1 Flash Lite ★ | 4/10 | $0.20 | 1.4x |
| GPT-5.4 Mini | 3/9 | $0.31 | 2.5x |
| NVIDIA Nemotron-3 Super 120B | 2/4 | $0.34 | 3.5x |
Topic Organization & Clustering
Discover natural groupings without external schema, name and order them coherently.
| Model | Tasks | Avg $/1k runs | Overpay |
|---|---|---|---|
| Tencent Hy3 ★ | 3/3 | $2.05 | 1.6x |
| DeepSeek V4 Flash ★ | 2/3 | $2.23 | 1.3x |
| GPT-5.4 Mini ★ | 1/3 | $2.58 | best value |
| MiniMax M3 | 1/2 | $3.08 | 1.2x |
| Qwen 3.5 Flash | 2/3 | $3.21 | 2.5x |
Infrastructure & Utility
Mechanical competence at format conversion, metadata manipulation, prompt rewriting, translation; minimal domain expertise required.
| Model | Tasks | Avg $/1k runs | Overpay |
|---|---|---|---|
| Gemini 3.1 Flash Lite ★ | 1/7 | $0.56 | best value |
| Tencent Hy3 ★ | 3/7 | $1.12 | 2.4x |
| DeepSeek V4 Flash ★ | 4/5 | $1.44 | 1.4x |
| GPT-5.6 Luna ★ | 5/7 | $2.23 | 3.8x |
| Qwen 3.5 Flash | 4/9 | $2.24 | 3.8x |
Methodology, briefly
Quality scores come from LLM-judge verdicts on production workloads, on a 0–10 scale. Only MEDIUM-or-better cells qualify. The quality bar is relative to the best-performing model on each task — adjust the slider above to see your own pipeline.
What changed this week
What changed — 2026-07-10 → 2026-07-20
- Tasks: 52 · Models: 30 (+8) · Evaluations: 905 (+243)
New models: GPT-5.6 Luna, GPT-5.6 Sol, GPT-5.6 Terra, Meta Muse Spark 1.1, NVIDIA Nemotron-3 Nano 30B-A3B, NVIDIA Nemotron-3 Super 120B, NVIDIA Nemotron-3 Ultra 550B, Tencent Hy3
Per-model movement (every model, 90% bar) — best value = tasks where the model is the best-value good-enough option; qualifying = tasks it’s good-enough on:
| Model | Best Value | Qualifying |
|---|---|---|
| DeepSeek V4 Flash | 14 (-4) | 18 |
| MiniMax M3 | 11 (+5) | 25 (+1) |
| Tencent Hy3 (new) | 4 (new) | 16 (new) |
| Gemini 3.5 Flash | 2 (-5) | 38 (+2) |
| Meta Muse Spark 1.1 (new) | 2 (new) | 27 (new) |
| GPT-5.6 Luna (new) | 2 (new) | 25 (new) |
| DeepSeek V4 Pro | 2 (-2) | 22 (-1) |
| Qwen 3.6 Flash | 2 (-3) | 17 (+1) |
| GPT-5.4 Nano | 2 (+1) | 9 (+3) |
| Gemini 3.1 Flash Lite | 2 (+1) | 6 (+1) |
| NVIDIA Nemotron-3 Super 120B (new) | 2 (new) | 6 (new) |
| NVIDIA Nemotron-3 Nano 30B-A3B (new) | 2 (new) | 3 (new) |
| GPT-5.6 Terra (new) | 1 (new) | 26 (new) |
| Claude Sonnet 5 | 1 (+1) | 24 (+2) |
| Qwen 3.5 Flash | 1 (-5) | 14 |
| GPT-5.4 Mini | 1 (+1) | 6 (+1) |
| Gemini 3.1 Flash Image Preview | 1 | 1 |
| GPT-5.5 | 0 | 30 (+1) |
| GPT-5.6 Sol (new) | 0 (new) | 28 (new) |
| Grok 4.5 | 0 | 28 (+2) |
| Kimi K2.6 | 0 | 27 (+3) |
| Qwen 3.6 Plus | 0 | 24 (+1) |
| Qwen 3.7 Plus | 0 (-2) | 21 (+1) |
| Claude Sonnet 4.6 | 0 | 19 (+2) |
| Gemini 3.1 Pro Preview | 0 | 16 (+1) |
| NVIDIA Nemotron-3 Ultra 550B (new) | 0 (new) | 15 (new) |
| Claude Opus 4.8 | 0 | 12 (+4) |
| Claude Haiku 4.5 | 0 (-1) | 9 |
| Gemini 3 Pro Image Preview | 0 | 1 |
| GPT-image-2 | 0 | 1 |
- Default quality bar 95% → 90%