Snapshot 2026-09-15
What changed — 2026-09-05 → 2026-09-15
- Tasks: 52 · Models: 30 (-6) · Model–task pairs: 833 (-195) · Graded samples: 57,538 (-7,724)
Dropped models: Gemini 3.6 Flash, Gemini 3.7 Flash, Grok 4.5, Meta Muse Spark 1.1, Meta Muse Spark 1.2, Alibaba Qwen3.7-Flash
Per-model movement (every model, 90% bar) — best value = tasks where the model is the cheapest good-enough option, counted separately under each cost mode; qualifying = tasks it’s good-enough on (quality only, so it is the same under both modes). A model with no batch API prices batch at its sync rate, so the two best-value columns diverge wherever a provider batch discount decides the winner:
| Model | Best Value (Sync) | Best Value (Async/Batch) | Qualifying |
|---|
| GPT-5.6 Luna | 22 (+16) | 27 (+10) | 31 (+1) |
| MiniMax M3 | 11 (+4) | 8 (+4) | 20 |
| GLM-5.3 Flash | 3 (-9) | 3 (-4) | 19 (+1) |
| Gemini 3.1 Flash Lite | 3 (+1) | 3 (+1) | 5 |
| Tencent Hy3 | 3 (-2) | 2 (-2) | 13 (-2) |
| Thinking Machines Inkling Small | 2 (-1) | 1 (-1) | 26 |
| Claude Sonnet 5 | 1 (+1) | 2 (+1) | 25 (-1) |
| Gemini 3.5 Flash | 1 | 1 | 35 |
| Thinking Machines Inkling | 1 | 1 | 27 |
| DeepSeek V4 Flash | 1 (-3) | 1 (-1) | 18 |
| Claude Opus 5 | 1 (+1) | 1 (+1) | 17 (+1) |
| NVIDIA Nemotron-3 Super 120B | 1 | 1 | 8 |
| Qwen 3.7 Plus | 1 | 0 | 21 |
| GPT-5.4 Nano | 0 | 1 | 9 |
| NVIDIA Nemotron-3 Nano 30B-A3B | 1 | 0 (-1) | 2 |
| GPT-5.6 Terra | 0 | 0 | 30 |
| GPT-5.6 Sol | 0 | 0 | 29 |
| Moonshot Kimi K3 | 0 | 0 | 23 |
| DeepSeek V4 Pro | 0 | 0 | 22 |
| GLM-5.3 | 0 | 0 | 22 (+1) |
| Grok 4.6 | 0 | 0 | 22 |
| Meta Muse Spark 1.3 | 0 | 0 | 21 (+1) |
| Tencent Hy4 Preview | 0 | 0 | 18 |
| Gemini 3.8 Flash | 0 | 0 | 15 |
| NVIDIA Nemotron-3 Ultra 550B | 0 | 0 | 12 |
| Qwen 3.8 Max | 0 | 0 | 9 |
| Claude Haiku 4.5 | 0 | 0 | 8 |
| Gemini 3.5 Flash Lite | 0 | 0 | 5 (-1) |
| NVIDIA Nemotron 3.5 Lightning | 0 | 0 | 3 |
| Qwen 3.8 Flash | 0 | 0 | 3 |
| Alibaba Qwen3.7-Flash (dropped) | — (was 7) | — (was 7) | — (was 9) |
| Gemini 3.6 Flash (dropped) | — (was 0) | — (was 0) | — (was 18) |
| Gemini 3.7 Flash (dropped) | — (was 0) | — (was 0) | — (was 13) |
| Grok 4.5 (dropped) | — (was 0) | — (was 0) | — (was 37) |
| Meta Muse Spark 1.1 (dropped) | — (was 1) | — (was 1) | — (was 29) |
| Meta Muse Spark 1.2 (dropped) | — (was 0) | — (was 0) | — (was 21) |
Snapshot 2026-09-05
What changed — 2026-08-22 → 2026-09-05
- Tasks: 52 (+1) · Models: 36 (+10) · Model–task pairs: 1,028 (+222) · Graded samples: 65,262 (+10,483)
New capabilities (50):
- Atomic Fact Claim Extraction —
atomic-fact-claim-extraction (extraction_and_parsing) - Authoritative Source Selection —
authoritative-source-selection (relevance_and_classification) - Batch Text Translation —
batch-text-translation (infrastructure_and_utility) - Catalyst And Scenario Analysis —
catalyst-and-scenario-analysis (financial_analysis_and_trading) - Community Content Promotion Generation —
community-content-promotion-generation (social_and_promotional) - Community Policy Evaluation —
community-policy-evaluation (relevance_and_classification) - Content Domain Suggestion —
content-domain-suggestion (relevance_and_classification) - Content Set Relevance Scoring —
content-set-relevance-scoring (relevance_and_classification) - Content To Image Prompt Generation —
content-to-image-prompt-generation (infrastructure_and_utility) - Engagement Opportunity Triage —
engagement-opportunity-triage (relevance_and_classification) - Evidence Grounded Claim Generation —
evidence-grounded-claim-generation (summarization_and_synthesis) - Factual Claim Refinement —
factual-claim-refinement (infrastructure_and_utility) - Generated Content Relevance Scoring —
generated-content-relevance-scoring (relevance_and_classification) - Known Url Content Extraction —
known-url-content-extraction (summarization_and_synthesis) - Language Identification —
language-identification (relevance_and_classification) - Metadata Paragraph Rewriting —
metadata-paragraph-rewriting (infrastructure_and_utility) - Model Specific Prompt Adaptation —
model-specific-prompt-adaptation (infrastructure_and_utility) - Multi Perspective Decision Synthesis —
multi-perspective-decision-synthesis (financial_analysis_and_trading) - Newsletter Copy Generation —
newsletter-copy-generation (long_form_writing) - Persona Policy Eligibility Check —
persona-policy-eligibility-check (relevance_and_classification) - Profile Pool Matching —
profile-pool-matching (relevance_and_classification) - Profiled Document Analysis —
profiled-document-analysis (financial_analysis_and_trading) - Profiled Document Section Analysis —
profiled-document-section-analysis (financial_analysis_and_trading) - Profiled Document Structure Extraction —
profiled-document-structure-extraction (extraction_and_parsing) - Promotional Campaign Generation —
promotional-campaign-generation (social_and_promotional) - Promotional Message Generation —
promotional-message-generation (social_and_promotional) - Public Engagement Reply Review —
public-engagement-reply-review (relevance_and_classification) - Public Response Generation —
public-response-generation (social_and_promotional) - Publication Section Taxonomy Generation —
publication-section-taxonomy-generation (long_form_writing) - Publication Title Package Generation —
publication-title-package-generation (summarization_and_synthesis) - Reference Preserving Analytical Writing —
reference-preserving-analytical-writing (long_form_writing) - Report Executive Summary Generation —
report-executive-summary-generation (summarization_and_synthesis) - Report Outline Generation —
report-outline-generation (long_form_writing) - Research Community Selection —
research-community-selection (relevance_and_classification) - Research Query Generation —
research-query-generation (infrastructure_and_utility) - Research Query Validation —
research-query-validation (infrastructure_and_utility) - Research Region Identification —
research-region-identification (extraction_and_parsing) - Section Prompt Generation —
section-prompt-generation (infrastructure_and_utility) - Short Post Batch Relevance Scoring —
short-post-batch-relevance-scoring (relevance_and_classification) - Short Promotional Teaser Generation —
short-promotional-teaser-generation (social_and_promotional) - Social Post Portfolio Selection —
social-post-portfolio-selection (relevance_and_classification) - Structured Content Summarization —
structured-content-summarization (summarization_and_synthesis) - Subject Configuration Generation —
subject-configuration-generation (financial_analysis_and_trading) - Thematic Topic Discovery —
thematic-topic-discovery (topic_organization) - Topic Cluster Labeling —
topic-cluster-labeling (topic_organization) - Topic Section Assignment —
topic-section-assignment (topic_organization) - Topic Sequence Optimization —
topic-sequence-optimization (topic_organization) - Topic Taxonomy Matching —
topic-taxonomy-matching (relevance_and_classification) - Visual Theme Configuration Generation —
visual-theme-configuration-generation (long_form_writing) - Voice Profile Generation —
voice-profile-generation (long_form_writing)
Removed capabilities (49):
- Activity Feed Blurb Generation —
activity_promo_generation - Content Domain Suggestion —
at_content_domain_suggest - Author Living-Person Safety Check —
author_living_check - Author Matching —
author_matching - Author Voice Generation —
author_soul_generation - Reddit Post Generation —
auto_reddit_post_generation - Bench Claim Generation —
bench_claim_generation - Claim Extraction —
claim_extraction - Claim-Referenced Analyst Writing —
claim_referenced_analyst_writing - Claim Refinement —
claim_refinement - Content Summarization —
content-summarization - Engagement Reply Draft —
engagement_reply - Engagement Reply Review —
engagement_review - Engagement Triage —
engagement_triage - Executive Summary Generation —
executive_summary_generation - Image Prompt Generation —
image_prompt_generation - Language Detection —
language-detection - Metadata Paragraph Rewriting —
metadata_paragraph_improvement - Onboarding Chapter Outline Generation —
onboarding_chapter_generation - Onboarding Chapter Prompt Adaptation —
onboarding_chapter_prompt_generation - Onboarding Subject Analysis —
onboarding_prospect_analysis - Prompt Adaptation —
prompt-adaptation - Research Query Generation —
query_generation - Research Query Validation —
query_validation - Geographic Region Identification —
region_identification - Social Post Relevance Scoring —
relevance_scoring_post - Topic Report Relevance Scoring —
relevance_scoring_topic_report - X Post Relevance Scoring —
relevance_scoring_x_post - S-1 TOC Extraction —
s1-toc-extraction - SEC Filing Analysis —
sec-filling-analysis - SEC S-1 Chunk Analysis —
sec-s1-chunk-analysis - Landing Page Section Generation —
section_generation - Social Post Promotion —
social_post_promo - Subreddit Selection for Research —
subreddit_selection - Subreddit Quality Vetting —
subreddit_vetting - Substack Newsletter —
substack_newsletter - Investment Panel Voting —
synthesis_analysis - Publication Title Generation —
synthesis_of_titles_for_publication - Theme Generation —
theme_generation - Topic Grouping and Client Matching —
topic_client_matching - Topic Cluster Naming —
topic_cluster_naming - Topic-to-Section Assignment —
topic_clustering_assign_sections - Topic Discovery Clustering —
topic_discovery_clustering - Topic Sequence Ordering —
topic_sequence_determination - Trading Recommendation —
trading_recommendation - Translation —
translation - Vetted News Site Selection —
vetted_site_selection - X.com Promotional Post Generation —
x_com_messages_for_promotion - X Post Selection —
x_post_selection
New models: Claude Opus 5, Gemini 3.7 Flash, Gemini 3.8 Flash, GLM-5.3, GLM-5.3 Flash, Grok 4.6, Meta Muse Spark 1.2, Meta Muse Spark 1.3, NVIDIA Nemotron 3.5 Lightning, Qwen 3.8 Flash, Qwen 3.8 Max, Tencent Hy4 Preview
Dropped models: Claude Opus 4.8, GPT-5.5
Per-model movement (every model, 90% bar) — best value = tasks where the model is the cheapest good-enough option, counted separately under each cost mode; qualifying = tasks it’s good-enough on (quality only, so it is the same under both modes). A model with no batch API prices batch at its sync rate, so the two best-value columns diverge wherever a provider batch discount decides the winner:
| Model | Best Value (Sync) | Best Value (Async/Batch) | Qualifying |
|---|
| GPT-5.6 Luna | 6 (-6) | 17 (-2) | 30 (-2) |
| GLM-5.3 Flash (new) | 12 (new) | 7 (new) | 18 (new) |
| Alibaba Qwen3.7-Flash | 7 | 7 (+1) | 9 |
| MiniMax M3 | 7 | 4 (-1) | 20 |
| Tencent Hy3 | 5 (-2) | 4 | 15 |
| DeepSeek V4 Flash | 4 (+2) | 2 (+1) | 18 (-2) |
| Thinking Machines Inkling Small | 3 (+1) | 2 | 26 (+1) |
| Gemini 3.1 Flash Lite | 2 (+1) | 2 (+1) | 5 |
| Gemini 3.5 Flash | 1 (+1) | 1 | 35 |
| Meta Muse Spark 1.1 | 1 (-2) | 1 (-1) | 29 (-3) |
| Thinking Machines Inkling | 1 | 1 | 27 (+3) |
| NVIDIA Nemotron-3 Super 120B | 1 (-5) | 1 (-4) | 8 (+1) |
| NVIDIA Nemotron-3 Nano 30B-A3B | 1 | 1 (+1) | 2 (-1) |
| Claude Sonnet 5 | 0 | 1 (+1) | 26 |
| Qwen 3.7 Plus | 1 | 0 (-1) | 21 |
| GPT-5.4 Nano | 0 | 1 | 9 (+1) |
| Grok 4.5 | 0 | 0 | 37 (+2) |
| GPT-5.6 Terra | 0 (-1) | 0 (-1) | 30 (-2) |
| GPT-5.6 Sol | 0 | 0 | 29 (-1) |
| Moonshot Kimi K3 | 0 | 0 | 23 (+2) |
| DeepSeek V4 Pro | 0 | 0 | 22 (+2) |
| Grok 4.6 (new) | 0 (new) | 0 (new) | 22 (new) |
| GLM-5.3 (new) | 0 (new) | 0 (new) | 21 (new) |
| Meta Muse Spark 1.2 (new) | 0 (new) | 0 (new) | 21 (new) |
| Meta Muse Spark 1.3 (new) | 0 (new) | 0 (new) | 20 (new) |
| Gemini 3.6 Flash | 0 | 0 | 18 (-3) |
| Tencent Hy4 Preview (new) | 0 (new) | 0 (new) | 18 (new) |
| Claude Opus 5 (new) | 0 (new) | 0 (new) | 16 (new) |
| Gemini 3.8 Flash (new) | 0 (new) | 0 (new) | 15 (new) |
| Gemini 3.7 Flash (new) | 0 (new) | 0 (new) | 13 (new) |
| NVIDIA Nemotron-3 Ultra 550B | 0 | 0 | 12 (-1) |
| Qwen 3.8 Max (new) | 0 (new) | 0 (new) | 9 (new) |
| Claude Haiku 4.5 | 0 | 0 | 8 (-2) |
| Gemini 3.5 Flash Lite | 0 | 0 (-1) | 6 |
| NVIDIA Nemotron 3.5 Lightning (new) | 0 (new) | 0 (new) | 3 (new) |
| Qwen 3.8 Flash (new) | 0 (new) | 0 (new) | 3 (new) |
| Claude Opus 4.8 (dropped) | — (was 0) | — (was 0) | — (was 7) |
| GPT-5.5 (dropped) | — (was 0) | — (was 0) | — (was 27) |
Snapshot 2026-08-22
What changed — 2026-08-21 → 2026-08-22
- Tasks: 51 · Models: 26 · Model–task pairs: 806 · Graded samples: 54,779 (+186)
Per-model movement (every model, 90% bar) — best value = tasks where the model is the cheapest good-enough option, counted separately under each cost mode; qualifying = tasks it’s good-enough on (quality only, so it is the same under both modes). A model with no batch API prices batch at its sync rate, so the two best-value columns diverge wherever a provider batch discount decides the winner:
| Model | Best Value (Sync) | Best Value (Async/Batch) | Qualifying |
|---|
| GPT-5.6 Luna | 12 (-1) | 19 (+1) | 32 |
| Alibaba Qwen3.7-Flash | 7 | 6 (-1) | 9 |
| MiniMax M3 | 7 | 5 | 20 |
| Tencent Hy3 | 7 | 4 | 15 |
| NVIDIA Nemotron-3 Super 120B | 6 (+1) | 5 | 7 (+1) |
| Meta Muse Spark 1.1 | 3 | 2 | 32 |
| Thinking Machines Inkling Small | 2 | 2 | 25 |
| DeepSeek V4 Flash | 2 | 1 | 20 |
| GPT-5.6 Terra | 1 | 1 | 32 |
| Thinking Machines Inkling | 1 | 1 | 24 |
| Qwen 3.7 Plus | 1 | 1 | 21 |
| Gemini 3.1 Flash Lite | 1 | 1 | 5 |
| Gemini 3.5 Flash | 0 | 1 | 35 |
| GPT-5.4 Nano | 0 | 1 | 8 |
| Gemini 3.5 Flash Lite | 0 | 1 | 6 |
| NVIDIA Nemotron-3 Nano 30B-A3B | 1 | 0 | 3 |
| Grok 4.5 | 0 | 0 | 35 |
| GPT-5.6 Sol | 0 | 0 | 30 |
| GPT-5.5 | 0 | 0 | 27 |
| Claude Sonnet 5 | 0 | 0 | 26 |
| Gemini 3.6 Flash | 0 | 0 | 21 |
| Moonshot Kimi K3 | 0 | 0 | 21 |
| DeepSeek V4 Pro | 0 | 0 | 20 |
| NVIDIA Nemotron-3 Ultra 550B | 0 | 0 | 13 |
| Claude Haiku 4.5 | 0 | 0 | 10 |
| Claude Opus 4.8 | 0 | 0 | 7 |
Snapshot 2026-08-21
What changed — 2026-08-16 → 2026-08-21
- Tasks: 51 (-1) · Models: 26 · Model–task pairs: 806 (-104) · Graded samples: 54,593 (-34)
Removed capabilities (1):
- Direct Browse Content Synthesis —
direct-browse-content-synthesis
Per-model movement (every model, 90% bar) — best value = tasks where the model is the cheapest good-enough option, counted separately under each cost mode; qualifying = tasks it’s good-enough on (quality only, so it is the same under both modes). A model with no batch API prices batch at its sync rate, so the two best-value columns diverge wherever a provider batch discount decides the winner:
| Model | Best Value (Sync) | Best Value (Async/Batch) | Qualifying |
|---|
| GPT-5.6 Luna | 13 (+4) | 18 | 32 (-2) |
| Alibaba Qwen3.7-Flash | 7 (+2) | 7 (+3) | 9 (-2) |
| MiniMax M3 | 7 (-3) | 5 (-3) | 20 (-5) |
| Tencent Hy3 | 7 (+2) | 4 (+1) | 15 (-3) |
| NVIDIA Nemotron-3 Super 120B | 5 (+3) | 5 (+3) | 6 (-3) |
| Meta Muse Spark 1.1 | 3 | 2 | 32 (-3) |
| Thinking Machines Inkling Small | 2 (+1) | 2 (+1) | 25 (-3) |
| DeepSeek V4 Flash | 2 (-10) | 1 (-6) | 20 |
| GPT-5.6 Terra | 1 | 1 | 32 (-1) |
| Thinking Machines Inkling | 1 (+1) | 1 (+1) | 24 (-3) |
| Qwen 3.7 Plus | 1 | 1 | 21 (-3) |
| Gemini 3.1 Flash Lite | 1 (-1) | 1 (-1) | 5 (-2) |
| Gemini 3.5 Flash | 0 | 1 | 35 (-2) |
| GPT-5.4 Nano | 0 | 1 | 8 |
| Gemini 3.5 Flash Lite | 0 | 1 | 6 |
| NVIDIA Nemotron-3 Nano 30B-A3B | 1 | 0 | 3 (-1) |
| Grok 4.5 | 0 | 0 | 35 (+1) |
| GPT-5.6 Sol | 0 | 0 | 30 (-3) |
| GPT-5.5 | 0 | 0 | 27 (-4) |
| Claude Sonnet 5 | 0 | 0 | 26 (-5) |
| Gemini 3.6 Flash | 0 | 0 | 21 (+1) |
| Moonshot Kimi K3 | 0 | 0 | 21 (-5) |
| DeepSeek V4 Pro | 0 | 0 | 20 (-4) |
| NVIDIA Nemotron-3 Ultra 550B | 0 | 0 | 13 (-6) |
| Claude Haiku 4.5 | 0 | 0 | 10 |
| Claude Opus 4.8 | 0 | 0 | 7 (-7) |
Snapshot 2026-08-16
No changes recorded for this snapshot.