Provider
Gemini
Model name
gemini-3.5-flash
Cost mode:
Qualifies on
38 / 52 tasks (at 90% bar)
Best value on
2 tasks

Cost vs quality across all tasks

0%25%50%75%100%78910Quality score (7–10)Cost-efficiency vs best value (1.0 = best value)Onboarding Subject Analysis — quality 8.63, 7.7x the cost of the best-value good-enough optionTopic Sequence Ordering — quality 8.96, 7.8x the cost of the best-value good-enough optionStructured Output Extraction — quality 9.70, 107.2x the cost of the best-value good-enough optionX Post Relevance Scoring — quality 7.39, 2.6x the cost of the best-value good-enough optionX.com Promotional Post Generation — quality 8.71, 7.2x the cost of the best-value good-enough optionActivity Feed Blurb Generation — quality 8.03, 21.1x the cost of the best-value good-enough optionClaim-Referenced Analyst Writing — quality 8.86, 6.7x the cost of the best-value good-enough optionResearch Query Validation — quality 8.37, 5.5x the cost of the best-value good-enough optionMetadata Paragraph Rewriting — quality 9.10, 3.5x the cost of the best-value good-enough optionSubstack Newsletter — quality 9.12, 24.1x the cost of the best-value good-enough optionSocial Post Promotion — quality 8.88, 12.9x the cost of the best-value good-enough optionSubreddit Quality Vetting — quality 8.70, 7.3x the cost of the best-value good-enough optionSocial Post Relevance Scoring — quality 7.89, this model IS the best-value good-enough optionTrading Recommendation — quality 8.67, 2.7x the cost of the best-value good-enough optionSEC Filing Analysis — quality 8.87, 6.0x the cost of the best-value good-enough optionS-1 TOC Extraction — quality 9.38, 13.3x the cost of the best-value good-enough optionAuthor Living-Person Safety Check — quality 8.94, 18.3x the cost of the best-value good-enough optionLanguage Detection — quality 9.98, 34.6x the cost of the best-value good-enough optionExecutive Summary Generation — quality 7.57, 5.6x the cost of the best-value good-enough optionTopic Report Relevance Scoring — quality 9.30, 1.3x the cost of the best-value good-enough optionPublication Title Generation — quality 8.80, 16.1x the cost of the best-value good-enough optionMarkdown Newline Repair — quality 9.33, 32.5x the cost of the best-value good-enough optionTranslation — quality 9.31, 21.1x the cost of the best-value good-enough optionTopic Cluster Naming — quality 8.15, 4.8x the cost of the best-value good-enough optionTopic Grouping and Client Matching — quality 8.51, 12.4x the cost of the best-value good-enough optionEngagement Triage — quality 8.56, 23.0x the cost of the best-value good-enough optionPrompt Adaptation — quality 8.71, 15.8x the cost of the best-value good-enough optionX Post Selection — quality 8.48, 37.2x the cost of the best-value good-enough optionClaim Extraction — quality 7.87, 14.1x the cost of the best-value good-enough optionAuthor Voice Generation — quality 9.47, 12.6x the cost of the best-value good-enough optionAuthor Matching — quality 8.33, 2.6x the cost of the best-value good-enough optionDirect Browse Content Synthesis — quality 9.08, this model IS the best-value good-enough optionSEC S-1 Chunk Analysis — quality 8.97, 6.0x the cost of the best-value good-enough optionClaim Refinement — quality 8.31, 13.3x the cost of the best-value good-enough optionImage Prompt Generation — quality 8.85, 29.8x the cost of the best-value good-enough optionReddit Post Generation — quality 8.73, 6.2x the cost of the best-value good-enough optionEngagement Reply Draft — quality 7.75, 6.5x the cost of the best-value good-enough optionTopic-to-Section Assignment — quality 8.99, 18.2x the cost of the best-value good-enough option

within ~1.3× of the best-value model · 1.3–2× · >2× · ★ this model is the best-value pick on that task. Top-right = best quadrant. Only tasks where this model qualifies at the 90% bar are plotted.

Per-task breakdown

TaskCategoryQuality (% of best)ConfidenceOverpay
Direct Browse Content Synthesis bestContent Summarization & Synthesis100%RANKEDbest value
Social Post Relevance Scoring bestRelevance, Classification & Matching100%MEDIUMbest value
Topic Report Relevance Scoring bestRelevance, Classification & Matching100%MEDIUM1.3x
Author MatchingRelevance, Classification & Matching92%MEDIUM2.6x
X Post Relevance Scoring bestRelevance, Classification & Matching100%MEDIUM2.6x
Trading Recommendation bestFinancial Analysis & Trading Decisions100%RANKED2.7x
Metadata Paragraph RewritingInfrastructure & Utility96%HIGH3.5x
Topic Cluster NamingTopic Organization & Clustering93%RANKED4.8x
Research Query ValidationInfrastructure & Utility97%HIGH5.5x
Executive Summary GenerationContent Summarization & Synthesis95%MEDIUM5.6x
SEC Filing AnalysisFinancial Analysis & Trading Decisions91%RANKED6x
SEC S-1 Chunk AnalysisFinancial Analysis & Trading Decisions97%RANKED6x
Reddit Post Generation bestSocial & Promotional Content100%RANKED6.2x
Engagement Reply DraftSocial & Promotional Content92%HIGH6.5x
Claim-Referenced Analyst WritingLong-form Content Generation91%RANKED6.7x
X.com Promotional Post Generation bestSocial & Promotional Content100%RANKED7.2x
Subreddit Quality VettingRelevance, Classification & Matching92%MEDIUM7.3x
Onboarding Subject AnalysisFinancial Analysis & Trading Decisions94%RANKED7.7x
Topic Sequence Ordering bestTopic Organization & Clustering100%HIGH7.8x
Topic Grouping and Client MatchingRelevance, Classification & Matching99%RANKED12x
Author Voice GenerationLong-form Content Generation98%RANKED13x
Social Post Promotion bestSocial & Promotional Content100%RANKED13x
Claim RefinementInfrastructure & Utility98%HIGH13x
S-1 TOC Extraction bestStructured Data & Fact Extraction100%HIGH13x
Claim Extraction bestStructured Data & Fact Extraction100%RANKED14x
Prompt AdaptationInfrastructure & Utility98%RANKED16x
Publication Title GenerationContent Summarization & Synthesis98%RANKED16x
Topic-to-Section Assignment bestTopic Organization & Clustering100%MEDIUM18x
Author Living-Person Safety CheckRelevance, Classification & Matching96%RANKED18x
TranslationInfrastructure & Utility94%MEDIUM21x
Activity Feed Blurb GenerationSocial & Promotional Content91%RANKED21x
Engagement TriageRelevance, Classification & Matching96%RANKED23x
Substack NewsletterLong-form Content Generation99%RANKED24x
Image Prompt GenerationInfrastructure & Utility96%RANKED30x
Markdown Newline Repair bestInfrastructure & Utility100%RANKED32x
Language DetectionRelevance, Classification & Matching99%RANKED35x
X Post SelectionRelevance, Classification & Matching97%RANKED37x
Structured Output ExtractionStructured Data & Fact Extraction99%RANKED107x

Overpay — how much more you pay by running this model instead of the best-value model that clears the quality bar on that task (marked ★). "16x" means you overpay 16× — the same output for 16× the best-value good-enough option; ★ means this model is that option (no overpayment). Confidence — how sure we are about the quality score (more judgments + more agreement = higher confidence): RANKED many independent judges scored this model's outputs and their agreement is very high (most confident) — HIGH many judges have scored it and they mostly agree (well-pinned) — MEDIUM enough judges have weighed in to publish, but they disagree more than we'd like (treat with a small grain of salt). LOW-confidence cells are hidden everywhere on the site. See the methodology for the exact thresholds.