Cheapest LLM that's good enough.

We run over 150,000 large-language-model API calls per week and let the outputs be evaluated by a panel of leading models. Cost per token doesn't track quality — and when you say something like "give me the model that's reliably 95% as good as the best one but most cost-effective", the best-value model typically saves north of 90% of the cost.

This benchmark publishes our results so other builders can answer the same question for their own pipelines: which model(s) are good enough and the most cost-efficient for each category of tasks. The data also shows that the most expensive model — even when your budget can pay for it — is not always the best performer.

Have a look. Tell us what you think. We can also run tests for you — reach us at llmbench@kapualabs.com.

Snapshot: 2026-07-10

Cost mode:

Right-sized per step

Best-value model per capability, at your quality bar. Greyed cells mean no model qualifies on that category yet.

ModelFinancial Analysis & Trading DecisionsStructured Data & Fact ExtractionContent Summarization & SynthesisLong-form Content GenerationSocial & Promotional ContentRelevance, Classification & MatchingTopic Organization & ClusteringInfrastructure & UtilityQualifies on
Qwen 3.6 Plus Alibaba Cloud (DashScope)1.6x 14x3.5x6.6x5.2x 3.2x2.4x 14/52
Qwen 3.7 Plus Alibaba Cloud (DashScope)1.2x23x1.4x1.3x2.2x 1.0x 11/52
Qwen 3.6 Flash Alibaba Cloud (DashScope)1.0x 1.0x 1.2x 2.3x 7/52
Qwen 3.5 Flash Alibaba Cloud (DashScope)1.3x1.1x 1.0x 6/52
Claude Sonnet 5 Anthropic3.4x3.9x4x21x6.6x 7.5x12/52
Claude Sonnet 4.6 Anthropic2.2x5x5.7x14x3.4x4.8x9/52
Claude Haiku 4.5 Anthropic1.0x 1.3x2.9x1.1x4/52
Claude Opus 4.8 Anthropic7.6x23x3/52
DeepSeek V4 Pro DeepSeek2x 1.0x 3.1x3.3x1.2x1.0x 10/52
DeepSeek V4 Flash DeepSeek1.0x 1.0x 1.0x 1.0x 6/52
Gemini 3.5 Flash Gemini1.4x 6.6x1.4x 2.7x1.0x 2.8x 1.0x 3.2x 28/52
Gemini 3.1 Pro Preview Gemini14x3.5x1.0x 6.7x2.5x 8/52
Gemini 3.1 Flash Lite Gemini1.1x1.0x 4/52
Gemini 3 Pro Image Preview Gemini1.6x1/52
MiniMax M3 MiniMax1.0x 3.3x1.0x 1.1x 4.3x2x 1.0x 14/52
Kimi K2.6 Moonshot AI2.4x10x2.9x4.1x5.4x9.5x2.4x2.7x16/52
GPT-5.5 OpenAI13x22x30x44x12x11x15/52
GPT-5.4 Nano OpenAI2.4x1.0x 1.3x3/52
GPT-5.4 Mini OpenAI2.3x2.5x2/52
GPT-image-2 OpenAI1.0x 1/52
Grok 4.5 xAI4.4x4.6x5.3x7.9x3.8x15/52
Each cell shows how much you overpay vs the best-value good-enough model in that category: best value most expensive Cells show how much you overpay vs the best-value model clearing the bar (1.0x = this model is the best-value good-enough option, no overpayment; 3.2x = you pay ~3× more on average for the same quality). marks any model that is the best-value pick on at least one of the category's tasks (multiple stars per column are possible — they're the per-task cost winners). Greyed cells: model doesn't qualify on any task in that category at the current bar. Rows grouped by provider; click a model name for the full per-task breakdown.

By capability

Financial Analysis & Trading Decisions

Read business/financial docs deeply enough to make forward-looking judgments about opportunity and risk.

ModelTasksAvg $/1M inOverpay
Qwen 3.6 Flash 1/3$0.58best value
MiniMax M3 2/4$0.60best value
Qwen 3.7 Plus1/3$0.721.2x
Qwen 3.6 Plus 2/4$0.851.6x
Gemini 3.5 Flash 2/5$1.041.4x

See all 10 qualifying models →

Structured Data & Fact Extraction

Precise pattern recognition and field-level information retrieval, strict schema adherence.

ModelTasksAvg $/1M inOverpay
DeepSeek V4 Flash 1/3$0.47best value
Claude Haiku 4.5 1/2$0.91best value
DeepSeek V4 Pro 3/4$0.952x
GPT-5.4 Nano1/2$1.132.4x
MiniMax M32/4$1.563.3x

See all 11 qualifying models →

Content Summarization & Synthesis

Compress long input into essential nuggets without losing material detail.

ModelTasksAvg $/1M inOverpay
Qwen 3.6 Flash 1/2$0.41best value
GPT-5.4 Nano 1/3$0.50best value
Qwen 3.7 Plus1/3$0.571.4x
MiniMax M3 1/3$0.93best value
Gemini 3.5 Flash 2/4$1.371.4x

See all 9 qualifying models →

Long-form Content Generation

Sustained compositional skill, voice consistency, coherent extended prose.

ModelTasksAvg $/1M inOverpay
DeepSeek V4 Pro 2/3$0.81best value
Claude Haiku 4.51/3$1.301.3x
MiniMax M3 4/4$1.441.1x
Qwen 3.6 Flash 2/5$1.801.2x
Qwen 3.7 Plus3/5$2.281.3x

See all 13 qualifying models →

Social & Promotional Content

Conciseness, platform-native conventions, engagement under tight character limits.

ModelTasksAvg $/1M inOverpay
DeepSeek V4 Flash 1/4$0.15best value
Qwen 3.5 Flash1/4$0.201.3x
GPT-5.4 Mini1/5$0.362.3x
DeepSeek V4 Pro1/5$0.483.1x
MiniMax M31/4$0.664.3x

See all 11 qualifying models →

Relevance, Classification & Matching

Semantic similarity judgment: does this thing belong in that bucket / match that target?

ModelTasksAvg $/1M inOverpay
DeepSeek V4 Flash 3/10$0.17best value
GPT-5.4 Nano1/9$0.201.3x
Qwen 3.5 Flash 4/9$0.341.1x
Gemini 3.1 Flash Lite3/10$0.391.1x
GPT-5.4 Mini1/9$0.402.5x

See all 19 qualifying models →

Topic Organization & Clustering

Discover natural groupings without external schema, name and order them coherently.

ModelTasksAvg $/1M inOverpay
Qwen 3.5 Flash 1/3$0.15best value
Qwen 3.7 Plus 1/1$0.45best value
Claude Haiku 4.51/2$0.471.1x
DeepSeek V4 Pro1/3$0.521.2x
Qwen 3.6 Plus2/3$0.713.2x

See all 10 qualifying models →

Infrastructure & Utility

Mechanical competence at format conversion, metadata manipulation, prompt rewriting, translation; minimal domain expertise required.

ModelTasksAvg $/1M inOverpay
Gemini 3.1 Flash Lite 1/7$0.41best value
DeepSeek V4 Flash 1/5$0.44best value
MiniMax M3 2/6$0.63best value
DeepSeek V4 Pro 1/6$1.64best value
Gemini 3.5 Flash 7/9$3.243.2x

See all 13 qualifying models →

Methodology, briefly

Quality scores come from LLM-judge verdicts on production workloads, on a 0–10 scale. Only MEDIUM-or-better cells qualify. The quality bar is relative to the best-performing model on each task — adjust the slider above to see your own pipeline.

Full methodology →

What changed this week

What changed — 2026-07-07 → 2026-07-10

  • Tasks: 52 · Models: 22 (+1) · Evaluations: 662 (+41)

New models: Grok 4.5

Per-model movement (every model, 90% bar) — best value = tasks where the model is the best-value good-enough option; qualifying = tasks it’s good-enough on:

ModelBest ValueQualifying
DeepSeek V4 Flash1818 (-1)
Gemini 3.5 Flash736
MiniMax M3624
Qwen 3.5 Flash614 (-2)
Qwen 3.6 Flash5 (+3)16 (+2)
DeepSeek V4 Pro4 (-2)23
Qwen 3.7 Plus2 (+1)20 (+2)
Claude Haiku 4.519
GPT-5.4 Nano16
Gemini 3.1 Flash Lite1 (-2)5 (-1)
Gemini 3.1 Flash Image Preview1 (+1)1 (-1)
GPT-5.5029
Grok 4.5 (new)0 (new)26 (new)
Kimi K2.6024 (+1)
Qwen 3.6 Plus023
Claude Sonnet 5022
Claude Sonnet 4.6017 (-1)
Gemini 3.1 Pro Preview015
Claude Opus 4.808 (+1)
GPT-5.4 Mini05
Gemini 3 Pro Image Preview01 (-1)
GPT-image-20 (-1)1
  • Default quality bar 90% → 95%

Full changelog →