Cheapest LLM that's good enough.
We run over 150,000 large-language-model API calls per week and let the outputs be evaluated by a panel of leading models. Cost per token doesn't track quality — and when you say something like "give me the model that's reliably 95% as good as the best one but most cost-effective", the best-value model typically saves north of 90% of the cost.
This benchmark publishes our results so other builders can answer the same question for their own pipelines: which model(s) are good enough and the most cost-efficient for each category of tasks. The data also shows that the most expensive model — even when your budget can pay for it — is not always the best performer.
Have a look. Tell us what you think. We can also run tests for you — reach us at llmbench@kapualabs.com.
Snapshot: 2026-09-15
Right-sized per step
Best-value model per capability, at your quality bar. Greyed cells mean no model qualifies on that category yet.
1.0x = this model is the best-value good-enough option, no overpayment; 3.2x = you pay ~3× more on average for the same quality).
★ marks any model that is the best-value pick on at least one of the category's tasks (multiple stars per column are possible — they're the per-task cost winners).
Greyed cells: model doesn't qualify on any task in that category at the current bar.
Rows grouped by provider; click a model name for the full per-task breakdown.By capability
Financial Analysis & Trading Decisions
Read business/financial docs deeply enough to make forward-looking judgments about opportunity and risk.
| Model | Tasks | Avg $/1k runs | Overpay |
|---|---|---|---|
| GPT-5.6 Luna ★ | 2/3 | $2.41 | best value |
| MiniMax M3 ★ | 3/4 | $5.15 | 2.1x |
| GPT-5.4 Nano | 2/4 | $7.47 | 2.9x |
| GLM-5.3 Flash | 2/3 | $8.29 | 4.1x |
| Qwen 3.7 Plus | 3/3 | $11.34 | 6.4x |
Structured Data & Fact Extraction
Precise pattern recognition and field-level information retrieval, strict schema adherence.
| Model | Tasks | Avg $/1k runs | Overpay |
|---|---|---|---|
| GPT-5.6 Luna ★ | 2/3 | $1.99 | 1.3x |
| NVIDIA Nemotron-3 Super 120B ★ | 1/1 | $2.03 | best value |
| Tencent Hy3 | 1/1 | $5.75 | 4.7x |
| MiniMax M3 | 2/3 | $5.83 | 3.1x |
| Qwen 3.8 Flash | 1/1 | $6.59 | 5.3x |
Content Summarization & Synthesis
Compress long input into essential nuggets without losing material detail.
| Model | Tasks | Avg $/1k runs | Overpay |
|---|---|---|---|
| Gemini 3.1 Flash Lite ★ | 1/3 | $1.88 | best value |
| GPT-5.6 Luna ★ | 2/3 | $2.69 | 1.1x |
| GLM-5.3 Flash ★ | 2/2 | $3.60 | 1.1x |
| Tencent Hy3 | 1/2 | $3.74 | 1.8x |
| MiniMax M3 ★ | 2/3 | $3.97 | 1.1x |
Long-form Content Generation
Sustained compositional skill, voice consistency, coherent extended prose.
| Model | Tasks | Avg $/1k runs | Overpay |
|---|---|---|---|
| MiniMax M3 ★ | 3/3 | $1.45 | 1.6x |
| GPT-5.6 Luna ★ | 4/4 | $1.83 | best value |
| GLM-5.3 Flash | 1/1 | $3.67 | 4.1x |
| GPT-5.4 Nano | 1/2 | $4.14 | 4.6x |
| NVIDIA Nemotron-3 Ultra 550B | 2/2 | $5.37 | 6x |
Social & Promotional Content
Conciseness, platform-native conventions, engagement under tight character limits.
| Model | Tasks | Avg $/1k runs | Overpay |
|---|---|---|---|
| GPT-5.6 Luna ★ | 3/5 | $0.50 | best value |
| NVIDIA Nemotron 3.5 Lightning | 1/1 | $1.04 | 4.1x |
| Tencent Hy3 | 2/5 | $1.59 | 3.7x |
| NVIDIA Nemotron-3 Super 120B | 2/4 | $2.61 | 6.9x |
| GPT-5.6 Terra | 3/4 | $3.09 | 6.5x |
Relevance, Classification & Matching
Semantic similarity judgment: does this thing belong in that bucket / match that target?
| Model | Tasks | Avg $/1k runs | Overpay |
|---|---|---|---|
| GPT-5.4 Nano | 2/8 | $0.15 | 1.3x |
| NVIDIA Nemotron-3 Nano 30B-A3B | 2/4 | $0.20 | 1.3x |
| Gemini 3.1 Flash Lite ★ | 3/9 | $0.21 | 1.3x |
| Gemini 3.5 Flash Lite | 3/7 | $0.23 | 1.5x |
| Tencent Hy3 ★ | 3/10 | $0.45 | 3.9x |
Topic Organization & Clustering
Discover natural groupings without external schema, name and order them coherently.
| Model | Tasks | Avg $/1k runs | Overpay |
|---|---|---|---|
| GLM-5.3 Flash | 1/1 | $0.98 | 3.9x |
| Tencent Hy3 ★ | 3/3 | $2.68 | 1.4x |
| MiniMax M3 ★ | 1/2 | $3.19 | best value |
| Gemini 3.8 Flash | 2/2 | $4.04 | 5.6x |
| GPT-5.6 Luna ★ | 2/3 | $4.15 | best value |
Infrastructure & Utility
Mechanical competence at format conversion, metadata manipulation, prompt rewriting, translation; minimal domain expertise required.
| Model | Tasks | Avg $/1k runs | Overpay |
|---|---|---|---|
| Gemini 3.1 Flash Lite ★ | 1/7 | $0.53 | best value |
| Gemini 3.5 Flash Lite | 1/4 | $0.66 | 1.9x |
| GPT-5.6 Luna ★ | 6/8 | $0.89 | best value |
| MiniMax M3 ★ | 3/5 | $2.54 | 2.5x |
| GLM-5.3 Flash | 4/4 | $2.80 | 6.5x |
Claim Verification & Review
Cross-check assertions against supplied evidence and apply a fixed rubric or checklist to issue a cited approve/reject verdict on accuracy, support, and wording.
No model qualifies in this category at the current bar.
Long-form Segmentation & Chunking
Divide long inputs into ordered, semantically coherent segments that respect structure and size constraints without compressing content.
No model qualifies in this category at the current bar.
Source-Grounded Dialogue Generation
Transform source material into natural multi-voice spoken dialogue that preserves the source's facts, caveats, and sequence while hitting a target duration and style.
No model qualifies in this category at the current bar.
Methodology, briefly
Quality scores come from LLM-judge verdicts on production workloads, on a 0–10 scale. Only MEDIUM-or-better cells qualify. The quality bar is relative to the best-performing model on each task — adjust the slider above to see your own pipeline.
What changed this week
What changed — 2026-09-05 → 2026-09-15
- Tasks: 52 · Models: 30 (-6) · Model–task pairs: 833 (-195) · Graded samples: 57,538 (-7,724)
Dropped models: Gemini 3.6 Flash, Gemini 3.7 Flash, Grok 4.5, Meta Muse Spark 1.1, Meta Muse Spark 1.2, Alibaba Qwen3.7-Flash
Per-model movement (every model, 90% bar) — best value = tasks where the model is the cheapest good-enough option, counted separately under each cost mode; qualifying = tasks it’s good-enough on (quality only, so it is the same under both modes). A model with no batch API prices batch at its sync rate, so the two best-value columns diverge wherever a provider batch discount decides the winner:
| Model | Best Value (Sync) | Best Value (Async/Batch) | Qualifying |
|---|---|---|---|
| GPT-5.6 Luna | 22 (+16) | 27 (+10) | 31 (+1) |
| MiniMax M3 | 11 (+4) | 8 (+4) | 20 |
| GLM-5.3 Flash | 3 (-9) | 3 (-4) | 19 (+1) |
| Gemini 3.1 Flash Lite | 3 (+1) | 3 (+1) | 5 |
| Tencent Hy3 | 3 (-2) | 2 (-2) | 13 (-2) |
| Thinking Machines Inkling Small | 2 (-1) | 1 (-1) | 26 |
| Claude Sonnet 5 | 1 (+1) | 2 (+1) | 25 (-1) |
| Gemini 3.5 Flash | 1 | 1 | 35 |
| Thinking Machines Inkling | 1 | 1 | 27 |
| DeepSeek V4 Flash | 1 (-3) | 1 (-1) | 18 |
| Claude Opus 5 | 1 (+1) | 1 (+1) | 17 (+1) |
| NVIDIA Nemotron-3 Super 120B | 1 | 1 | 8 |
| Qwen 3.7 Plus | 1 | 0 | 21 |
| GPT-5.4 Nano | 0 | 1 | 9 |
| NVIDIA Nemotron-3 Nano 30B-A3B | 1 | 0 (-1) | 2 |
| GPT-5.6 Terra | 0 | 0 | 30 |
| GPT-5.6 Sol | 0 | 0 | 29 |
| Moonshot Kimi K3 | 0 | 0 | 23 |
| DeepSeek V4 Pro | 0 | 0 | 22 |
| GLM-5.3 | 0 | 0 | 22 (+1) |
| Grok 4.6 | 0 | 0 | 22 |
| Meta Muse Spark 1.3 | 0 | 0 | 21 (+1) |
| Tencent Hy4 Preview | 0 | 0 | 18 |
| Gemini 3.8 Flash | 0 | 0 | 15 |
| NVIDIA Nemotron-3 Ultra 550B | 0 | 0 | 12 |
| Qwen 3.8 Max | 0 | 0 | 9 |
| Claude Haiku 4.5 | 0 | 0 | 8 |
| Gemini 3.5 Flash Lite | 0 | 0 | 5 (-1) |
| NVIDIA Nemotron 3.5 Lightning | 0 | 0 | 3 |
| Qwen 3.8 Flash | 0 | 0 | 3 |
| Alibaba Qwen3.7-Flash (dropped) | — (was 7) | — (was 7) | — (was 9) |
| Gemini 3.6 Flash (dropped) | — (was 0) | — (was 0) | — (was 18) |
| Gemini 3.7 Flash (dropped) | — (was 0) | — (was 0) | — (was 13) |
| Grok 4.5 (dropped) | — (was 0) | — (was 0) | — (was 37) |
| Meta Muse Spark 1.1 (dropped) | — (was 1) | — (was 1) | — (was 29) |
| Meta Muse Spark 1.2 (dropped) | — (was 0) | — (was 0) | — (was 21) |