Claude Sonnet 5 is strong, fast, and not the cheapest good-enough pick
Claude Sonnet 5 is a good model. The benchmark does not make that ambiguous.
It also does not make it the default model you would route to.
That is the useful tension in the 2026-07-06 llm-bench snapshot. Claude Sonnet 5 lands in the upper tier on quality, performs especially well on analyst-style and writing-heavy tasks, and is one of the fastest models in the current run. But once the benchmark asks the production question - which model is good enough at the lowest practical price? - Sonnet 5 stops looking like a value leader.
At the default 90% quality bar, Claude Sonnet 5 is good enough on 22 of 52 benchmark tasks. It has data on 34 of those tasks, so within the cells actually tested it clears the bar about 65% of the time. That is strong. It is also not dominant: Gemini 3.5 Flash clears 36 tasks, GPT-5.5 clears 29, MiniMax M3 clears 24, and DeepSeek V4 Pro, Kimi K2.6, and Qwen 3.6 Plus each clear 23.
The short version:
| Metric | Claude Sonnet 5 | Readout |
|---|---|---|
| Benchmark snapshot | 2026-07-06 | 52 tasks, 21 models |
| Tested task cells | 34 | Partial coverage, not full-grid |
| Good-enough tasks at 90% bar | 22/52 | 7th of 21 models |
| Good-enough rate on tested tasks | 64.7% | Strong but not universal |
| Average quality score | 8.04 | 8th of 21 models |
| Median quality score | 8.44 | Strong typical output |
| Best-quality task anchors | 2 | Clear task-level wins |
| Cheapest-good-enough wins | 0 | The central value weakness |
| Average blended cost | $10.03 | Expensive side of the board |
| Median blended cost | $5.05 | Still not cheap |
| Average latency | 12.2s | 2nd fastest by average latency |
| Median success rate | 99.5% | Usually reliable |
That shape matters more than a single leaderboard rank. Sonnet 5 is not a “bad” value because it lacks capability. It is a bad default value pick because enough cheaper models clear the same task bars often enough to take the star.
What “good enough” means here
The bench does not ask whether a model is globally best. For each task, it finds the highest trusted quality score - the task anchor - then asks which other models are within the selected quality bar. At the default 90% setting, a model qualifies if it is at least 90% as good as the task’s best anchored score, with enough confidence to trust the result.
That relative framing is important. A model can score well in absolute terms and still fail a task if the leader is much stronger. It can also be excellent and still not be the recommended routing choice if another model clears the same bar at a fraction of the cost.
For Sonnet 5, both things happen.
It posts two best-quality anchors:
- Subreddit Quality Vetting: 9.3685, the strongest score in that task.
- Onboarding Subject Analysis: 9.0542, also the task anchor.
Those are not marginal cells. They show the model doing exactly what Sonnet-class models are expected to do well: judgment, context reading, and structured analytical output.
But it has zero cheapest-good-enough wins. In other words, whenever Sonnet 5 clears the quality bar, another qualifying model is cheaper. That is the whole right-sizing thesis in one model page.
Where Sonnet 5 is genuinely strong
Sonnet 5’s best fit categories in the snapshot are:
- Financial Analysis & Trading Decisions
- Structured Data & Fact Extraction
- Long-form Content Generation
That is a coherent profile. It looks like a strong analyst-writer with good schema discipline, not a commodity utility model.
Financial analysis
In financial analysis, Sonnet 5 qualifies on 3 of 4 tested tasks. The strongest result is Onboarding Subject Analysis, where it is the task anchor at 9.0542. It also clears the bar on SEC S-1 Chunk Analysis and Investment Panel Voting.
The miss is SEC Filing Analysis, where Sonnet 5 scores 7.5732, below the 90% bar set by MiniMax M3’s stronger task score. That does not mean Sonnet 5 cannot read filings. It means that, in this particular harness, there are cheaper or better fits for that exact work package.
The practical readout: use Sonnet 5 when the financial task needs synthesis, judgment, and a polished analytical frame. Do not assume it is the best default for every filing-processing step.
Structured extraction
The extraction sample is small, but clean: Sonnet 5 qualifies on 2 of 2 tested extraction tasks, with an average score around 8.65. It clears both Geographic Region Identification and S-1 TOC Extraction.
That points to good instruction following and good structure preservation. It is not enough coverage to declare it the extraction champion, because DeepSeek V4 Pro and GPT-5.5 cover more extraction tasks in the snapshot. But where Sonnet 5 is tested, it behaves well.
Long-form writing
Long-form generation is the category where Sonnet 5 most looks like itself. It qualifies on 4 of 5 tested tasks and averages about 8.91 across those tasks.
Notable cells:
- Author Voice Generation: 9.3208
- Substack Newsletter: 9.1996
- Onboarding Chapter Outline Generation: 9.2138
- Theme Generation: 8.7451
The miss is Claim-Referenced Analyst Writing, where the score is respectable at 8.0895 but below the task bar. More importantly, that cell has a low success rate in the current data, which pushes its blended cost sharply higher.
The practical readout: Sonnet 5 is a credible premium model for writing that needs voice, continuity, structure, and judgment. It is not automatically the cheapest way to get acceptable prose.
Where the story is mixed
The middle of the Sonnet 5 profile is less clean. It can perform well on summarization, social copy, and classification, but it is not consistently the model you would choose first.
Summarization
Sonnet 5 qualifies on 2 of 3 tested summarization tasks. It performs well on Executive Summary Generation and Publication Title Generation, but misses Content Summarization. The miss matters because summarization is often high-volume work, and cheaper models can be very competitive there.
Social and promotional content
Sonnet 5 qualifies on 3 of 5 tested social/promotional tasks. It does well on Reddit Post Generation, Activity Feed Blurb Generation, and Social Post Promotion. It misses X.com Promotional Post Generation and Engagement Reply Review.
That is an interesting split. The model can write strong promotional prose, but short-form platform-native constraints appear less consistent than longer-form writing. In production terms, it may be a good drafter but not the cheapest or safest automatic router for every social task.
Relevance, classification, and matching
This category is the most revealing mix. Sonnet 5 anchors Subreddit Quality Vetting with a very strong 9.3685, and it clears Language Detection, X Post Selection, Engagement Triage, and the subreddit vetting task.
But it misses:
- Author Matching
- Author Living-Person Safety Check
- Content Domain Suggestion
So the model can make nuanced classification judgments, but the benchmark does not show a universal classification edge. Gemini 3.5 Flash, GPT-5.5, and DeepSeek V4 Pro cover more of this category at the 90% bar.
Where Sonnet 5 is weak
Two areas stand out: topic organization and cost-sensitive utility work.
Topic organization and clustering
Sonnet 5 qualifies on only 1 of 3 tested topic-organization tasks. The bad cell is not subtle: Topic Discovery Clustering scores -0.6644, ranking last among trusted cells for that task. It also misses Topic Sequence Ordering.
That makes topic discovery and clustering a poor fit in this snapshot. The bench has other models that are both better and cheaper for that category.
Utility and mechanical tasks
Sonnet 5 can be expensive and brittle on mechanical utility work.
Examples:
- Markdown Newline Repair: score 7.1749, success rate 8.3%, blended cost $68.46.
- Translation: score 8.4937, success rate 34.7%, blended cost $42.12.
- Claim Refinement: score 8.0570, success rate 10.0%, blended cost $18.03.
These are exactly the kinds of tasks where a right-sized harness should usually avoid a premium model unless the data proves it clears the bar cheaply. In this snapshot, Sonnet 5 does not.
The head-to-head picture
Against GPT-5.5, Sonnet 5 is more competitive than the overall ranking implies. On 28 common tasks, Sonnet 5 wins 15 direct quality matchups, and the median quality gap is essentially flat. It is also much cheaper than GPT-5.5 on those common tasks: about 0.26x the median cost.
That makes Sonnet 5 a plausible near-frontier substitute when GPT-5.5 is overkill.
Against Claude Opus 4.8, Sonnet 5 is clearly the cheaper model, but not the stronger one. On 10 common tasks, Sonnet 5 wins only 2 quality matchups. If you are choosing within Anthropic, Sonnet 5 looks like the economical high-quality model, while Opus remains the premium quality option where its edge is proven.
Against Claude Sonnet 4.6, Sonnet 5 looks like a real upgrade. It wins 15 of 23 common quality matchups, qualifies on more benchmark tasks, and is cheaper on the median common task.
Against Gemini 3.5 Flash, the picture is harsher. Gemini 3.5 Flash beats Sonnet 5 on broad coverage, average quality, cost, and cheapest-good-enough wins in this snapshot. It qualifies on 36 tasks and is best value on 7. Sonnet 5 qualifies on 22 and is best value on 0.
Against DeepSeek V4 Flash and DeepSeek V4 Pro, Sonnet 5 often has the more premium feel, but the cost case is difficult. DeepSeek V4 Flash is the board’s value monster with 18 cheapest-good-enough wins. DeepSeek V4 Pro clears 23 tasks and wins 6 cheapest-good-enough slots. For high-volume routing, those models are hard for Sonnet 5 to beat.
The reliability wrinkle
The typical Sonnet 5 cell is reliable: the median success rate is 99.5%.
The average success rate, however, falls to 77.8% because several cells are severe outliers. Markdown Newline Repair, Claim Refinement, Claim-Referenced Analyst Writing, Publication Title Generation, Translation, and Reddit Post Generation all show low success rates in the current data.
That distinction matters. Sonnet 5 is not generally flaky in the snapshot, but its failures are concentrated enough to affect blended cost on specific tasks. And because the bench prices useful work rather than raw calls, those reliability failures become real dollars.
This is why blended cost is such a useful unit. A model that looks high-quality but needs repeated attempts on a task is not cheap for that task.
The routing conclusion
Claude Sonnet 5 should not be read as a default model. It should be read as a targeted premium model.
It belongs in the candidate set for:
- analyst-style reasoning and synthesis;
- financial subject analysis;
- structured extraction where its coverage is proven;
- long-form writing that needs voice and continuity;
- GPT-5.5 replacement cases where the quality gap is small and cost matters.
It should be used cautiously for:
- topic discovery and clustering;
- mechanical formatting and utility tasks;
- high-volume summarization where cheaper models clear the bar;
- platform-specific social tasks with strict constraints;
- any task where current success-rate data shows retry-driven blended cost.
The headline is not “Sonnet 5 disappoints.” It does not. The headline is that Sonnet 5 is exactly the kind of model a task-level benchmark is built to place correctly: strong enough to deserve serious use, expensive enough that you should not use it reflexively, and uneven enough that a single global ranking would hide the jobs where it shines and the jobs where it quietly overcharges you.
The right role for Claude Sonnet 5 is not “run everything.” It is: keep it in the portfolio, route it to the work where its analyst and writing strengths matter, and let cheaper qualifiers take the jobs where the harness has already proven they are good enough.