Provider
Anthropic
Model name
claude-sonnet-5
Cost mode:
Qualifies on
24 / 52 tasks (at 90% bar)
Best value on
1 tasks

Cost vs quality across all tasks

0%25%50%75%100%78910Quality score (7–10)Cost-efficiency vs best value (1.0 = best value)Onboarding Subject Analysis — quality 8.85, 7.4x the cost of the best-value good-enough optionTheme Generation — quality 9.70, 13.6x the cost of the best-value good-enough optionEngagement Reply Review — quality 8.97, 30.5x the cost of the best-value good-enough optionActivity Feed Blurb Generation — quality 8.74, 25.1x the cost of the best-value good-enough optionMetadata Paragraph Rewriting — quality 9.51, this model IS the best-value good-enough optionSubstack Newsletter — quality 9.25, 18.9x the cost of the best-value good-enough optionInvestment Panel Voting — quality 8.67, 4.4x the cost of the best-value good-enough optionSocial Post Promotion — quality 8.22, 7.8x the cost of the best-value good-enough optionSubreddit Quality Vetting — quality 9.28, 7.7x the cost of the best-value good-enough optionS-1 TOC Extraction — quality 8.54, 9.2x the cost of the best-value good-enough optionLanguage Detection — quality 10.06, 34.9x the cost of the best-value good-enough optionExecutive Summary Generation — quality 7.60, 7.4x the cost of the best-value good-enough optionPublication Title Generation — quality 8.81, 70.0x the cost of the best-value good-enough optionTranslation — quality 8.89, 29.8x the cost of the best-value good-enough optionTopic Cluster Naming — quality 8.28, 8.7x the cost of the best-value good-enough optionEngagement Triage — quality 8.65, 28.1x the cost of the best-value good-enough optionOnboarding Chapter Outline Generation — quality 9.19, 9.7x the cost of the best-value good-enough optionX Post Selection — quality 8.33, 19.9x the cost of the best-value good-enough optionClaim Extraction — quality 7.38, 22.8x the cost of the best-value good-enough optionAuthor Voice Generation — quality 9.12, 11.8x the cost of the best-value good-enough optionSEC S-1 Chunk Analysis — quality 8.89, 11.4x the cost of the best-value good-enough optionClaim Refinement — quality 8.35, 14.8x the cost of the best-value good-enough optionImage Prompt Generation — quality 8.62, 21.3x the cost of the best-value good-enough optionReddit Post Generation — quality 8.07, 19.9x the cost of the best-value good-enough option

within ~1.3× of the best-value model · 1.3–2× · >2× · ★ this model is the best-value pick on that task. Top-right = best quadrant. Only tasks where this model qualifies at the 90% bar are plotted.

Per-task breakdown

TaskCategoryQuality (% of best)ConfidenceOverpay
Metadata Paragraph Rewriting bestInfrastructure & Utility100%RANKEDbest value
Investment Panel VotingFinancial Analysis & Trading Decisions95%RANKED4.4x
Onboarding Subject AnalysisFinancial Analysis & Trading Decisions96%HIGH7.4x
Executive Summary GenerationContent Summarization & Synthesis96%RANKED7.4x
Subreddit Quality VettingRelevance, Classification & Matching98%HIGH7.7x
Social Post PromotionSocial & Promotional Content93%RANKED7.8x
Topic Cluster NamingTopic Organization & Clustering95%RANKED8.7x
S-1 TOC ExtractionStructured Data & Fact Extraction91%HIGH9.2x
Onboarding Chapter Outline GenerationLong-form Content Generation94%RANKED9.7x
SEC S-1 Chunk AnalysisFinancial Analysis & Trading Decisions96%RANKED11x
Author Voice GenerationLong-form Content Generation94%RANKED12x
Theme Generation bestLong-form Content Generation100%RANKED14x
Claim RefinementInfrastructure & Utility99%HIGH15x
Substack Newsletter bestLong-form Content Generation100%RANKED19x
X Post SelectionRelevance, Classification & Matching95%RANKED20x
Reddit Post GenerationSocial & Promotional Content92%RANKED20x
Image Prompt GenerationInfrastructure & Utility94%RANKED21x
Claim ExtractionStructured Data & Fact Extraction94%MEDIUM23x
Activity Feed Blurb GenerationSocial & Promotional Content99%RANKED25x
Engagement TriageRelevance, Classification & Matching97%RANKED28x
TranslationInfrastructure & Utility90%RANKED30x
Engagement Reply ReviewRelevance, Classification & Matching91%MEDIUM30x
Language Detection bestRelevance, Classification & Matching100%RANKED35x
Publication Title GenerationContent Summarization & Synthesis98%RANKED70x

Overpay — how much more you pay by running this model instead of the best-value model that clears the quality bar on that task (marked ★). "16x" means you overpay 16× — the same output for 16× the best-value good-enough option; ★ means this model is that option (no overpayment). Confidence — how sure we are about the quality score (more judgments + more agreement = higher confidence): RANKED many independent judges scored this model's outputs and their agreement is very high (most confident) — HIGH many judges have scored it and they mostly agree (well-pinned) — MEDIUM enough judges have weighed in to publish, but they disagree more than we'd like (treat with a small grain of salt). LOW-confidence cells are hidden everywhere on the site. See the methodology for the exact thresholds.