Cheapest LLM that's good enough.

We run over 150,000 large-language-model API calls per week and let the outputs be evaluated by a panel of leading models. Cost per token doesn't track quality — and when you say something like "give me the model that's reliably 95% as good as the best one but most cost-effective", the best-value model typically saves north of 90% of the cost.

This benchmark publishes our results so other builders can answer the same question for their own pipelines: which model(s) are good enough and the most cost-efficient for each category of tasks. The data also shows that the most expensive model — even when your budget can pay for it — is not always the best performer.

Have a look. Tell us what you think. We can also run tests for you — reach us at llmbench@kapualabs.com.

Snapshot: 2026-09-15

Cost mode:

Right-sized per step

Best-value model per capability, at your quality bar. Greyed cells mean no model qualifies on that category yet.

ModelFinancial Analysis & Trading DecisionsStructured Data & Fact ExtractionContent Summarization & SynthesisLong-form Content GenerationSocial & Promotional ContentRelevance, Classification & MatchingTopic Organization & ClusteringInfrastructure & UtilityClaim Verification & ReviewLong-form Segmentation & ChunkingSource-Grounded Dialogue GenerationQualifies on
Qwen 3.7 Plus Alibaba Cloud (DashScope)6.4x13x1.6x6.8x14x15x3x11x21/52
Qwen 3.8 Max Alibaba Cloud (DashScope)40x25x78x78x48x90x9/52
Qwen 3.8 Flash Alibaba Cloud (DashScope)5.3x8.8x3/52
Claude Sonnet 5 Anthropic15x5.3x 5.2x35x24x15x8.3x16x 25/52
Claude Opus 5 Anthropic43x19x18x28x38x 13x63x17/52
Claude Haiku 4.5 Anthropic8.6x3.3x18x9.3x2.9x17x8/52
DeepSeek V4 Pro DeepSeek14x29x5.5x22x19x18x10x23x22/52
DeepSeek V4 Flash DeepSeek5.2x 3.7x8.1x7.4x5.7x4x8.1x18/52
Gemini 3.5 Flash Gemini7.9x16x2.6x15x12x 16x16x18x35/52
Gemini 3.8 Flash Gemini1.6x4.8x4.9x9.2x5.6x7.9x15/52
Gemini 3.1 Flash Lite Gemini1.0x 1.3x 1.0x 5/52
Gemini 3.5 Flash Lite Gemini1.9x1.5x1.9x5/52
Moonshot Kimi K3 Moonshot AI54x34x19x46x66x110x68x73x23/52
GPT-5.6 Luna OpenAI1.0x 1.3x 1.1x 1.0x 1.0x 1.0x 1.0x 1.0x 31/52
GPT-5.6 Terra OpenAI14x10x4.6x9.6x6.5x10x9.1x10x30/52
GPT-5.6 Sol OpenAI16x13x13x16x12x17x38x17x29/52
GPT-5.4 Nano OpenAI2.9x4.6x1.0x 4.6x1.3x3.1x9/52
Thinking Machines Inkling OpenRouter11x2.2x4.8x8.6x 46x37x8.1x22x27/52
Thinking Machines Inkling Small OpenRouter7.6x 5.5x2.4x7.1x16x8.9x6.4x9.2x26/52
Meta Muse Spark 1.3 OpenRouter21x4.8x23x24x29x11x23x21/52
MiniMax M3 OpenRouter2.1x 3.1x1.1x 1.6x 1.0x 2.9x 1.0x 2.5x 20/52
Tencent Hy4 Preview OpenRouter25x13x34x52x47x11x68x18/52
Tencent Hy3 OpenRouter4.7x1.8x3.7x3.9x 1.4x 6.3x13/52
NVIDIA Nemotron-3 Ultra 550B OpenRouter8.4x1.3x6x17x11x6.4x11x12/52
NVIDIA Nemotron-3 Super 120B OpenRouter1.0x 2.7x6.9x7.7x8/52
NVIDIA Nemotron 3.5 Lightning OpenRouter4.1x6.5x3/52
NVIDIA Nemotron-3 Nano 30B-A3B OpenRouter1.3x2/52
Grok 4.6 xAI38x10x26x65x64x32x72x22/52
GLM-5.3 Z.AI25x12x31x51x41x33x80x22/52
GLM-5.3 Flash Z.AI4.1x1.1x 4.1x3.9x3.8x 3.9x6.5x19/52
Each cell shows how much you overpay vs the best-value good-enough model in that category: best value most expensive Cells show how much you overpay vs the best-value model clearing the bar (1.0x = this model is the best-value good-enough option, no overpayment; 3.2x = you pay ~3× more on average for the same quality). marks any model that is the best-value pick on at least one of the category's tasks (multiple stars per column are possible — they're the per-task cost winners). Greyed cells: model doesn't qualify on any task in that category at the current bar. Rows grouped by provider; click a model name for the full per-task breakdown.

By capability

Financial Analysis & Trading Decisions

Read business/financial docs deeply enough to make forward-looking judgments about opportunity and risk.

ModelTasksAvg $/1k runsOverpay
GPT-5.6 Luna 2/3$2.41best value
MiniMax M3 3/4$5.152.1x
GPT-5.4 Nano2/4$7.472.9x
GLM-5.3 Flash2/3$8.294.1x
Qwen 3.7 Plus3/3$11.346.4x

See all 21 qualifying models →

Structured Data & Fact Extraction

Precise pattern recognition and field-level information retrieval, strict schema adherence.

ModelTasksAvg $/1k runsOverpay
GPT-5.6 Luna 2/3$1.991.3x
NVIDIA Nemotron-3 Super 120B 1/1$2.03best value
Tencent Hy31/1$5.754.7x
MiniMax M32/3$5.833.1x
Qwen 3.8 Flash1/1$6.595.3x

See all 17 qualifying models →

Content Summarization & Synthesis

Compress long input into essential nuggets without losing material detail.

ModelTasksAvg $/1k runsOverpay
Gemini 3.1 Flash Lite 1/3$1.88best value
GPT-5.6 Luna 2/3$2.691.1x
GLM-5.3 Flash 2/2$3.601.1x
Tencent Hy31/2$3.741.8x
MiniMax M3 2/3$3.971.1x

See all 27 qualifying models →

Long-form Content Generation

Sustained compositional skill, voice consistency, coherent extended prose.

ModelTasksAvg $/1k runsOverpay
MiniMax M3 3/3$1.451.6x
GPT-5.6 Luna 4/4$1.83best value
GLM-5.3 Flash1/1$3.674.1x
GPT-5.4 Nano1/2$4.144.6x
NVIDIA Nemotron-3 Ultra 550B2/2$5.376x

See all 22 qualifying models →

Social & Promotional Content

Conciseness, platform-native conventions, engagement under tight character limits.

ModelTasksAvg $/1k runsOverpay
GPT-5.6 Luna 3/5$0.50best value
NVIDIA Nemotron 3.5 Lightning1/1$1.044.1x
Tencent Hy32/5$1.593.7x
NVIDIA Nemotron-3 Super 120B2/4$2.616.9x
GPT-5.6 Terra3/4$3.096.5x

See all 23 qualifying models →

Relevance, Classification & Matching

Semantic similarity judgment: does this thing belong in that bucket / match that target?

ModelTasksAvg $/1k runsOverpay
GPT-5.4 Nano2/8$0.151.3x
NVIDIA Nemotron-3 Nano 30B-A3B2/4$0.201.3x
Gemini 3.1 Flash Lite 3/9$0.211.3x
Gemini 3.5 Flash Lite3/7$0.231.5x
Tencent Hy3 3/10$0.453.9x

See all 30 qualifying models →

Topic Organization & Clustering

Discover natural groupings without external schema, name and order them coherently.

ModelTasksAvg $/1k runsOverpay
GLM-5.3 Flash1/1$0.983.9x
Tencent Hy3 3/3$2.681.4x
MiniMax M3 1/2$3.19best value
Gemini 3.8 Flash2/2$4.045.6x
GPT-5.6 Luna 2/3$4.15best value

See all 23 qualifying models →

Infrastructure & Utility

Mechanical competence at format conversion, metadata manipulation, prompt rewriting, translation; minimal domain expertise required.

ModelTasksAvg $/1k runsOverpay
Gemini 3.1 Flash Lite 1/7$0.53best value
Gemini 3.5 Flash Lite1/4$0.661.9x
GPT-5.6 Luna 6/8$0.89best value
MiniMax M3 3/5$2.542.5x
GLM-5.3 Flash4/4$2.806.5x

See all 26 qualifying models →

Claim Verification & Review

Cross-check assertions against supplied evidence and apply a fixed rubric or checklist to issue a cited approve/reject verdict on accuracy, support, and wording.

No model qualifies in this category at the current bar.

Long-form Segmentation & Chunking

Divide long inputs into ordered, semantically coherent segments that respect structure and size constraints without compressing content.

No model qualifies in this category at the current bar.

Source-Grounded Dialogue Generation

Transform source material into natural multi-voice spoken dialogue that preserves the source's facts, caveats, and sequence while hitting a target duration and style.

No model qualifies in this category at the current bar.

Methodology, briefly

Quality scores come from LLM-judge verdicts on production workloads, on a 0–10 scale. Only MEDIUM-or-better cells qualify. The quality bar is relative to the best-performing model on each task — adjust the slider above to see your own pipeline.

Full methodology →

What changed this week

What changed — 2026-09-05 → 2026-09-15

  • Tasks: 52 · Models: 30 (-6) · Model–task pairs: 833 (-195) · Graded samples: 57,538 (-7,724)

Dropped models: Gemini 3.6 Flash, Gemini 3.7 Flash, Grok 4.5, Meta Muse Spark 1.1, Meta Muse Spark 1.2, Alibaba Qwen3.7-Flash

Per-model movement (every model, 90% bar) — best value = tasks where the model is the cheapest good-enough option, counted separately under each cost mode; qualifying = tasks it’s good-enough on (quality only, so it is the same under both modes). A model with no batch API prices batch at its sync rate, so the two best-value columns diverge wherever a provider batch discount decides the winner:

ModelBest Value (Sync)Best Value (Async/Batch)Qualifying
GPT-5.6 Luna22 (+16)27 (+10)31 (+1)
MiniMax M311 (+4)8 (+4)20
GLM-5.3 Flash3 (-9)3 (-4)19 (+1)
Gemini 3.1 Flash Lite3 (+1)3 (+1)5
Tencent Hy33 (-2)2 (-2)13 (-2)
Thinking Machines Inkling Small2 (-1)1 (-1)26
Claude Sonnet 51 (+1)2 (+1)25 (-1)
Gemini 3.5 Flash1135
Thinking Machines Inkling1127
DeepSeek V4 Flash1 (-3)1 (-1)18
Claude Opus 51 (+1)1 (+1)17 (+1)
NVIDIA Nemotron-3 Super 120B118
Qwen 3.7 Plus1021
GPT-5.4 Nano019
NVIDIA Nemotron-3 Nano 30B-A3B10 (-1)2
GPT-5.6 Terra0030
GPT-5.6 Sol0029
Moonshot Kimi K30023
DeepSeek V4 Pro0022
GLM-5.30022 (+1)
Grok 4.60022
Meta Muse Spark 1.30021 (+1)
Tencent Hy4 Preview0018
Gemini 3.8 Flash0015
NVIDIA Nemotron-3 Ultra 550B0012
Qwen 3.8 Max009
Claude Haiku 4.5008
Gemini 3.5 Flash Lite005 (-1)
NVIDIA Nemotron 3.5 Lightning003
Qwen 3.8 Flash003
Alibaba Qwen3.7-Flash (dropped)— (was 7)— (was 7)— (was 9)
Gemini 3.6 Flash (dropped)— (was 0)— (was 0)— (was 18)
Gemini 3.7 Flash (dropped)— (was 0)— (was 0)— (was 13)
Grok 4.5 (dropped)— (was 0)— (was 0)— (was 37)
Meta Muse Spark 1.1 (dropped)— (was 1)— (was 1)— (was 29)
Meta Muse Spark 1.2 (dropped)— (was 0)— (was 0)— (was 21)

Full changelog →