Cheapest LLM that's good enough.

We run over 150,000 large-language-model API calls per week and let the outputs be evaluated by a panel of leading models. Cost per token doesn't track quality — and when you say something like "give me the model that's reliably 95% as good as the best one but most cost-effective", the best-value model typically saves north of 90% of the cost.

This benchmark publishes our results so other builders can answer the same question for their own pipelines: which model(s) are good enough and the most cost-efficient for each category of tasks. The data also shows that the most expensive model — even when your budget can pay for it — is not always the best performer.

Have a look. Tell us what you think. We can also run tests for you — reach us at llmbench@kapualabs.com.

Snapshot: 2026-07-20

Cost mode:

Right-sized per step

Best-value model per capability, at your quality bar. Greyed cells mean no model qualifies on that category yet.

ModelFinancial Analysis & Trading DecisionsStructured Data & Fact ExtractionContent Summarization & SynthesisLong-form Content GenerationSocial & Promotional ContentRelevance, Classification & MatchingTopic Organization & ClusteringInfrastructure & UtilityQualifies on
Qwen 3.6 Plus Alibaba Cloud (DashScope)3.1x52x16x15x22x20x4.7x15x24/52
Qwen 3.7 Plus Alibaba Cloud (DashScope)2.7x56x7.3x7.2x12x17x2.6x11x21/52
Qwen 3.6 Flash Alibaba Cloud (DashScope)3.1x 31x10x9.4x24x33x 4.3x15x17/52
Qwen 3.5 Flash Alibaba Cloud (DashScope)1.0x 1.1x7.4x9.2x2.5x3.8x14/52
Claude Sonnet 5 Anthropic7.7x16x39x13x18x24x8.7x17x 24/52
Claude Sonnet 4.6 Anthropic12x16x27x25x14x28x9.8x35x19/52
Claude Opus 4.8 Anthropic12x26x34x22x69x35x12/52
Claude Haiku 4.5 Anthropic7.9x8.1x15x11x3.1x9/52
DeepSeek V4 Pro DeepSeek1.1x11x5.4x4.5x 4.3x3.6x 2.1x5.8x22/52
DeepSeek V4 Flash DeepSeek1.0x 1.0x 1.8x1.0x 1.1x 1.3x 1.4x 18/52
Gemini 3.5 Flash Gemini5.6x45x7.6x 14x11x14x 10x17x38/52
Gemini 3.1 Pro Preview Gemini83x15x19x13x18x5.5x14x16/52
Gemini 3.1 Flash Lite Gemini1.2x1.4x 1.0x 6/52
Gemini 3 Pro Image Preview Gemini1.6x1/52
Gemini 3.1 Flash Image Preview Gemini1.0x 1/52
Meta Muse Spark 1.1 Meta7.6x19x20x24x32x 13x 28x27/52
MiniMax M3 MiniMax1.3x 7.6x1.6x 1.9x 1.7x 3x 1.2x2.4x 25/52
Kimi K2.6 Moonshot AI7.4x56x22x26x25x51x10x26x27/52
GPT-5.5 OpenAI22x183x51x71x14x29x16x62x30/52
GPT-5.6 Sol OpenAI14x84x31x21x13x19x20x20x28/52
GPT-5.6 Terra OpenAI4.5x4.5x8.9x7.5x4.8x9.3x 4.5x8.3x26/52
GPT-5.6 Luna OpenAI2.5x13x 5.6x3.7x2.5x5.1x1.8x3.8x 25/52
GPT-5.4 Nano OpenAI1.3x 19x1.8x 4.6x1.6x2.3x9/52
GPT-5.4 Mini OpenAI1.9x2.5x1.0x 7.2x6/52
GPT-image-2 OpenAI1.5x1/52
Tencent Hy3 OpenRouter2.7x2.2x 2.7x 3.9x1.6x 2.4x 16/52
NVIDIA Nemotron-3 Ultra 550B OpenRouter1.9x4.4x6.4x7.8x18x24x3.3x8.1x15/52
NVIDIA Nemotron-3 Super 120B OpenRouter2x1.0x 3.1x3.5x6/52
NVIDIA Nemotron-3 Nano 30B-A3B OpenRouter1.0x 1.1x1.0x 3/52
Grok 4.5 xAI6.6x73x22x19x29x33x7.4x19x28/52
Each cell shows how much you overpay vs the best-value good-enough model in that category: best value most expensive Cells show how much you overpay vs the best-value model clearing the bar (1.0x = this model is the best-value good-enough option, no overpayment; 3.2x = you pay ~3× more on average for the same quality). marks any model that is the best-value pick on at least one of the category's tasks (multiple stars per column are possible — they're the per-task cost winners). Greyed cells: model doesn't qualify on any task in that category at the current bar. Rows grouped by provider; click a model name for the full per-task breakdown.

By capability

Financial Analysis & Trading Decisions

Read business/financial docs deeply enough to make forward-looking judgments about opportunity and risk.

ModelTasksAvg $/1k runsOverpay
MiniMax M3 3/4$3.021.3x
Qwen 3.5 Flash 1/2$3.11best value
GPT-5.4 Nano 2/2$7.011.3x
GPT-5.6 Luna1/3$7.682.5x
Qwen 3.7 Plus3/3$9.632.7x

See all 20 qualifying models →

Structured Data & Fact Extraction

Precise pattern recognition and field-level information retrieval, strict schema adherence.

ModelTasksAvg $/1k runsOverpay
DeepSeek V4 Flash 3/3$0.58best value
Qwen 3.5 Flash1/2$0.751.1x
Gemini 3.1 Flash Lite1/3$0.831.2x
NVIDIA Nemotron-3 Super 120B1/1$1.392x
MiniMax M32/3$3.367.6x

See all 22 qualifying models →

Content Summarization & Synthesis

Compress long input into essential nuggets without losing material detail.

ModelTasksAvg $/1k runsOverpay
DeepSeek V4 Flash 1/2$0.71best value
Tencent Hy32/2$2.542.7x
MiniMax M3 2/3$3.711.6x
NVIDIA Nemotron-3 Ultra 550B1/2$4.536.4x
GPT-5.4 Nano 2/3$8.001.8x

See all 22 qualifying models →

Long-form Content Generation

Sustained compositional skill, voice consistency, coherent extended prose.

ModelTasksAvg $/1k runsOverpay
NVIDIA Nemotron-3 Nano 30B-A3B 1/3$0.28best value
NVIDIA Nemotron-3 Super 120B 2/4$1.08best value
Tencent Hy3 3/5$1.442.2x
DeepSeek V4 Flash1/2$1.531.8x
MiniMax M3 4/4$1.591.9x

See all 23 qualifying models →

Social & Promotional Content

Conciseness, platform-native conventions, engagement under tight character limits.

ModelTasksAvg $/1k runsOverpay
NVIDIA Nemotron-3 Nano 30B-A3B1/2$0.301.1x
DeepSeek V4 Flash 2/4$0.42best value
GPT-5.4 Mini1/5$0.521.9x
NVIDIA Nemotron-3 Super 120B1/3$0.853.1x
Tencent Hy3 2/4$1.302.7x

See all 24 qualifying models →

Relevance, Classification & Matching

Semantic similarity judgment: does this thing belong in that bucket / match that target?

ModelTasksAvg $/1k runsOverpay
NVIDIA Nemotron-3 Nano 30B-A3B 1/2$0.04best value
GPT-5.4 Nano2/9$0.141.6x
Gemini 3.1 Flash Lite 4/10$0.201.4x
GPT-5.4 Mini3/9$0.312.5x
NVIDIA Nemotron-3 Super 120B2/4$0.343.5x

See all 27 qualifying models →

Topic Organization & Clustering

Discover natural groupings without external schema, name and order them coherently.

ModelTasksAvg $/1k runsOverpay
Tencent Hy3 3/3$2.051.6x
DeepSeek V4 Flash 2/3$2.231.3x
GPT-5.4 Mini 1/3$2.58best value
MiniMax M31/2$3.081.2x
Qwen 3.5 Flash2/3$3.212.5x

See all 22 qualifying models →

Infrastructure & Utility

Mechanical competence at format conversion, metadata manipulation, prompt rewriting, translation; minimal domain expertise required.

ModelTasksAvg $/1k runsOverpay
Gemini 3.1 Flash Lite 1/7$0.56best value
Tencent Hy3 3/7$1.122.4x
DeepSeek V4 Flash 4/5$1.441.4x
GPT-5.6 Luna 5/7$2.233.8x
Qwen 3.5 Flash4/9$2.243.8x

See all 27 qualifying models →

Methodology, briefly

Quality scores come from LLM-judge verdicts on production workloads, on a 0–10 scale. Only MEDIUM-or-better cells qualify. The quality bar is relative to the best-performing model on each task — adjust the slider above to see your own pipeline.

Full methodology →

What changed this week

What changed — 2026-07-10 → 2026-07-20

  • Tasks: 52 · Models: 30 (+8) · Evaluations: 905 (+243)

New models: GPT-5.6 Luna, GPT-5.6 Sol, GPT-5.6 Terra, Meta Muse Spark 1.1, NVIDIA Nemotron-3 Nano 30B-A3B, NVIDIA Nemotron-3 Super 120B, NVIDIA Nemotron-3 Ultra 550B, Tencent Hy3

Per-model movement (every model, 90% bar) — best value = tasks where the model is the best-value good-enough option; qualifying = tasks it’s good-enough on:

ModelBest ValueQualifying
DeepSeek V4 Flash14 (-4)18
MiniMax M311 (+5)25 (+1)
Tencent Hy3 (new)4 (new)16 (new)
Gemini 3.5 Flash2 (-5)38 (+2)
Meta Muse Spark 1.1 (new)2 (new)27 (new)
GPT-5.6 Luna (new)2 (new)25 (new)
DeepSeek V4 Pro2 (-2)22 (-1)
Qwen 3.6 Flash2 (-3)17 (+1)
GPT-5.4 Nano2 (+1)9 (+3)
Gemini 3.1 Flash Lite2 (+1)6 (+1)
NVIDIA Nemotron-3 Super 120B (new)2 (new)6 (new)
NVIDIA Nemotron-3 Nano 30B-A3B (new)2 (new)3 (new)
GPT-5.6 Terra (new)1 (new)26 (new)
Claude Sonnet 51 (+1)24 (+2)
Qwen 3.5 Flash1 (-5)14
GPT-5.4 Mini1 (+1)6 (+1)
Gemini 3.1 Flash Image Preview11
GPT-5.5030 (+1)
GPT-5.6 Sol (new)0 (new)28 (new)
Grok 4.5028 (+2)
Kimi K2.6027 (+3)
Qwen 3.6 Plus024 (+1)
Qwen 3.7 Plus0 (-2)21 (+1)
Claude Sonnet 4.6019 (+2)
Gemini 3.1 Pro Preview016 (+1)
NVIDIA Nemotron-3 Ultra 550B (new)0 (new)15 (new)
Claude Opus 4.8012 (+4)
Claude Haiku 4.50 (-1)9
Gemini 3 Pro Image Preview01
GPT-image-201
  • Default quality bar 95% → 90%

Full changelog →