Snapshot 2026-09-15

What changed — 2026-09-05 → 2026-09-15

  • Tasks: 52 · Models: 30 (-6) · Model–task pairs: 833 (-195) · Graded samples: 57,538 (-7,724)

Dropped models: Gemini 3.6 Flash, Gemini 3.7 Flash, Grok 4.5, Meta Muse Spark 1.1, Meta Muse Spark 1.2, Alibaba Qwen3.7-Flash

Per-model movement (every model, 90% bar) — best value = tasks where the model is the cheapest good-enough option, counted separately under each cost mode; qualifying = tasks it’s good-enough on (quality only, so it is the same under both modes). A model with no batch API prices batch at its sync rate, so the two best-value columns diverge wherever a provider batch discount decides the winner:

ModelBest Value (Sync)Best Value (Async/Batch)Qualifying
GPT-5.6 Luna22 (+16)27 (+10)31 (+1)
MiniMax M311 (+4)8 (+4)20
GLM-5.3 Flash3 (-9)3 (-4)19 (+1)
Gemini 3.1 Flash Lite3 (+1)3 (+1)5
Tencent Hy33 (-2)2 (-2)13 (-2)
Thinking Machines Inkling Small2 (-1)1 (-1)26
Claude Sonnet 51 (+1)2 (+1)25 (-1)
Gemini 3.5 Flash1135
Thinking Machines Inkling1127
DeepSeek V4 Flash1 (-3)1 (-1)18
Claude Opus 51 (+1)1 (+1)17 (+1)
NVIDIA Nemotron-3 Super 120B118
Qwen 3.7 Plus1021
GPT-5.4 Nano019
NVIDIA Nemotron-3 Nano 30B-A3B10 (-1)2
GPT-5.6 Terra0030
GPT-5.6 Sol0029
Moonshot Kimi K30023
DeepSeek V4 Pro0022
GLM-5.30022 (+1)
Grok 4.60022
Meta Muse Spark 1.30021 (+1)
Tencent Hy4 Preview0018
Gemini 3.8 Flash0015
NVIDIA Nemotron-3 Ultra 550B0012
Qwen 3.8 Max009
Claude Haiku 4.5008
Gemini 3.5 Flash Lite005 (-1)
NVIDIA Nemotron 3.5 Lightning003
Qwen 3.8 Flash003
Alibaba Qwen3.7-Flash (dropped)— (was 7)— (was 7)— (was 9)
Gemini 3.6 Flash (dropped)— (was 0)— (was 0)— (was 18)
Gemini 3.7 Flash (dropped)— (was 0)— (was 0)— (was 13)
Grok 4.5 (dropped)— (was 0)— (was 0)— (was 37)
Meta Muse Spark 1.1 (dropped)— (was 1)— (was 1)— (was 29)
Meta Muse Spark 1.2 (dropped)— (was 0)— (was 0)— (was 21)

Snapshot 2026-09-05

What changed — 2026-08-22 → 2026-09-05

  • Tasks: 52 (+1) · Models: 36 (+10) · Model–task pairs: 1,028 (+222) · Graded samples: 65,262 (+10,483)

New capabilities (50):

  • Atomic Fact Claim Extraction — atomic-fact-claim-extraction (extraction_and_parsing)
  • Authoritative Source Selection — authoritative-source-selection (relevance_and_classification)
  • Batch Text Translation — batch-text-translation (infrastructure_and_utility)
  • Catalyst And Scenario Analysis — catalyst-and-scenario-analysis (financial_analysis_and_trading)
  • Community Content Promotion Generation — community-content-promotion-generation (social_and_promotional)
  • Community Policy Evaluation — community-policy-evaluation (relevance_and_classification)
  • Content Domain Suggestion — content-domain-suggestion (relevance_and_classification)
  • Content Set Relevance Scoring — content-set-relevance-scoring (relevance_and_classification)
  • Content To Image Prompt Generation — content-to-image-prompt-generation (infrastructure_and_utility)
  • Engagement Opportunity Triage — engagement-opportunity-triage (relevance_and_classification)
  • Evidence Grounded Claim Generation — evidence-grounded-claim-generation (summarization_and_synthesis)
  • Factual Claim Refinement — factual-claim-refinement (infrastructure_and_utility)
  • Generated Content Relevance Scoring — generated-content-relevance-scoring (relevance_and_classification)
  • Known Url Content Extraction — known-url-content-extraction (summarization_and_synthesis)
  • Language Identification — language-identification (relevance_and_classification)
  • Metadata Paragraph Rewriting — metadata-paragraph-rewriting (infrastructure_and_utility)
  • Model Specific Prompt Adaptation — model-specific-prompt-adaptation (infrastructure_and_utility)
  • Multi Perspective Decision Synthesis — multi-perspective-decision-synthesis (financial_analysis_and_trading)
  • Newsletter Copy Generation — newsletter-copy-generation (long_form_writing)
  • Persona Policy Eligibility Check — persona-policy-eligibility-check (relevance_and_classification)
  • Profile Pool Matching — profile-pool-matching (relevance_and_classification)
  • Profiled Document Analysis — profiled-document-analysis (financial_analysis_and_trading)
  • Profiled Document Section Analysis — profiled-document-section-analysis (financial_analysis_and_trading)
  • Profiled Document Structure Extraction — profiled-document-structure-extraction (extraction_and_parsing)
  • Promotional Campaign Generation — promotional-campaign-generation (social_and_promotional)
  • Promotional Message Generation — promotional-message-generation (social_and_promotional)
  • Public Engagement Reply Review — public-engagement-reply-review (relevance_and_classification)
  • Public Response Generation — public-response-generation (social_and_promotional)
  • Publication Section Taxonomy Generation — publication-section-taxonomy-generation (long_form_writing)
  • Publication Title Package Generation — publication-title-package-generation (summarization_and_synthesis)
  • Reference Preserving Analytical Writing — reference-preserving-analytical-writing (long_form_writing)
  • Report Executive Summary Generation — report-executive-summary-generation (summarization_and_synthesis)
  • Report Outline Generation — report-outline-generation (long_form_writing)
  • Research Community Selection — research-community-selection (relevance_and_classification)
  • Research Query Generation — research-query-generation (infrastructure_and_utility)
  • Research Query Validation — research-query-validation (infrastructure_and_utility)
  • Research Region Identification — research-region-identification (extraction_and_parsing)
  • Section Prompt Generation — section-prompt-generation (infrastructure_and_utility)
  • Short Post Batch Relevance Scoring — short-post-batch-relevance-scoring (relevance_and_classification)
  • Short Promotional Teaser Generation — short-promotional-teaser-generation (social_and_promotional)
  • Social Post Portfolio Selection — social-post-portfolio-selection (relevance_and_classification)
  • Structured Content Summarization — structured-content-summarization (summarization_and_synthesis)
  • Subject Configuration Generation — subject-configuration-generation (financial_analysis_and_trading)
  • Thematic Topic Discovery — thematic-topic-discovery (topic_organization)
  • Topic Cluster Labeling — topic-cluster-labeling (topic_organization)
  • Topic Section Assignment — topic-section-assignment (topic_organization)
  • Topic Sequence Optimization — topic-sequence-optimization (topic_organization)
  • Topic Taxonomy Matching — topic-taxonomy-matching (relevance_and_classification)
  • Visual Theme Configuration Generation — visual-theme-configuration-generation (long_form_writing)
  • Voice Profile Generation — voice-profile-generation (long_form_writing)

Removed capabilities (49):

  • Activity Feed Blurb Generation — activity_promo_generation
  • Content Domain Suggestion — at_content_domain_suggest
  • Author Living-Person Safety Check — author_living_check
  • Author Matching — author_matching
  • Author Voice Generation — author_soul_generation
  • Reddit Post Generation — auto_reddit_post_generation
  • Bench Claim Generation — bench_claim_generation
  • Claim Extraction — claim_extraction
  • Claim-Referenced Analyst Writing — claim_referenced_analyst_writing
  • Claim Refinement — claim_refinement
  • Content Summarization — content-summarization
  • Engagement Reply Draft — engagement_reply
  • Engagement Reply Review — engagement_review
  • Engagement Triage — engagement_triage
  • Executive Summary Generation — executive_summary_generation
  • Image Prompt Generation — image_prompt_generation
  • Language Detection — language-detection
  • Metadata Paragraph Rewriting — metadata_paragraph_improvement
  • Onboarding Chapter Outline Generation — onboarding_chapter_generation
  • Onboarding Chapter Prompt Adaptation — onboarding_chapter_prompt_generation
  • Onboarding Subject Analysis — onboarding_prospect_analysis
  • Prompt Adaptation — prompt-adaptation
  • Research Query Generation — query_generation
  • Research Query Validation — query_validation
  • Geographic Region Identification — region_identification
  • Social Post Relevance Scoring — relevance_scoring_post
  • Topic Report Relevance Scoring — relevance_scoring_topic_report
  • X Post Relevance Scoring — relevance_scoring_x_post
  • S-1 TOC Extraction — s1-toc-extraction
  • SEC Filing Analysis — sec-filling-analysis
  • SEC S-1 Chunk Analysis — sec-s1-chunk-analysis
  • Landing Page Section Generation — section_generation
  • Social Post Promotion — social_post_promo
  • Subreddit Selection for Research — subreddit_selection
  • Subreddit Quality Vetting — subreddit_vetting
  • Substack Newsletter — substack_newsletter
  • Investment Panel Voting — synthesis_analysis
  • Publication Title Generation — synthesis_of_titles_for_publication
  • Theme Generation — theme_generation
  • Topic Grouping and Client Matching — topic_client_matching
  • Topic Cluster Naming — topic_cluster_naming
  • Topic-to-Section Assignment — topic_clustering_assign_sections
  • Topic Discovery Clustering — topic_discovery_clustering
  • Topic Sequence Ordering — topic_sequence_determination
  • Trading Recommendation — trading_recommendation
  • Translation — translation
  • Vetted News Site Selection — vetted_site_selection
  • X.com Promotional Post Generation — x_com_messages_for_promotion
  • X Post Selection — x_post_selection

New models: Claude Opus 5, Gemini 3.7 Flash, Gemini 3.8 Flash, GLM-5.3, GLM-5.3 Flash, Grok 4.6, Meta Muse Spark 1.2, Meta Muse Spark 1.3, NVIDIA Nemotron 3.5 Lightning, Qwen 3.8 Flash, Qwen 3.8 Max, Tencent Hy4 Preview

Dropped models: Claude Opus 4.8, GPT-5.5

Per-model movement (every model, 90% bar) — best value = tasks where the model is the cheapest good-enough option, counted separately under each cost mode; qualifying = tasks it’s good-enough on (quality only, so it is the same under both modes). A model with no batch API prices batch at its sync rate, so the two best-value columns diverge wherever a provider batch discount decides the winner:

ModelBest Value (Sync)Best Value (Async/Batch)Qualifying
GPT-5.6 Luna6 (-6)17 (-2)30 (-2)
GLM-5.3 Flash (new)12 (new)7 (new)18 (new)
Alibaba Qwen3.7-Flash77 (+1)9
MiniMax M374 (-1)20
Tencent Hy35 (-2)415
DeepSeek V4 Flash4 (+2)2 (+1)18 (-2)
Thinking Machines Inkling Small3 (+1)226 (+1)
Gemini 3.1 Flash Lite2 (+1)2 (+1)5
Gemini 3.5 Flash1 (+1)135
Meta Muse Spark 1.11 (-2)1 (-1)29 (-3)
Thinking Machines Inkling1127 (+3)
NVIDIA Nemotron-3 Super 120B1 (-5)1 (-4)8 (+1)
NVIDIA Nemotron-3 Nano 30B-A3B11 (+1)2 (-1)
Claude Sonnet 501 (+1)26
Qwen 3.7 Plus10 (-1)21
GPT-5.4 Nano019 (+1)
Grok 4.50037 (+2)
GPT-5.6 Terra0 (-1)0 (-1)30 (-2)
GPT-5.6 Sol0029 (-1)
Moonshot Kimi K30023 (+2)
DeepSeek V4 Pro0022 (+2)
Grok 4.6 (new)0 (new)0 (new)22 (new)
GLM-5.3 (new)0 (new)0 (new)21 (new)
Meta Muse Spark 1.2 (new)0 (new)0 (new)21 (new)
Meta Muse Spark 1.3 (new)0 (new)0 (new)20 (new)
Gemini 3.6 Flash0018 (-3)
Tencent Hy4 Preview (new)0 (new)0 (new)18 (new)
Claude Opus 5 (new)0 (new)0 (new)16 (new)
Gemini 3.8 Flash (new)0 (new)0 (new)15 (new)
Gemini 3.7 Flash (new)0 (new)0 (new)13 (new)
NVIDIA Nemotron-3 Ultra 550B0012 (-1)
Qwen 3.8 Max (new)0 (new)0 (new)9 (new)
Claude Haiku 4.5008 (-2)
Gemini 3.5 Flash Lite00 (-1)6
NVIDIA Nemotron 3.5 Lightning (new)0 (new)0 (new)3 (new)
Qwen 3.8 Flash (new)0 (new)0 (new)3 (new)
Claude Opus 4.8 (dropped)— (was 0)— (was 0)— (was 7)
GPT-5.5 (dropped)— (was 0)— (was 0)— (was 27)

Snapshot 2026-08-22

What changed — 2026-08-21 → 2026-08-22

  • Tasks: 51 · Models: 26 · Model–task pairs: 806 · Graded samples: 54,779 (+186)

Per-model movement (every model, 90% bar) — best value = tasks where the model is the cheapest good-enough option, counted separately under each cost mode; qualifying = tasks it’s good-enough on (quality only, so it is the same under both modes). A model with no batch API prices batch at its sync rate, so the two best-value columns diverge wherever a provider batch discount decides the winner:

ModelBest Value (Sync)Best Value (Async/Batch)Qualifying
GPT-5.6 Luna12 (-1)19 (+1)32
Alibaba Qwen3.7-Flash76 (-1)9
MiniMax M37520
Tencent Hy37415
NVIDIA Nemotron-3 Super 120B6 (+1)57 (+1)
Meta Muse Spark 1.13232
Thinking Machines Inkling Small2225
DeepSeek V4 Flash2120
GPT-5.6 Terra1132
Thinking Machines Inkling1124
Qwen 3.7 Plus1121
Gemini 3.1 Flash Lite115
Gemini 3.5 Flash0135
GPT-5.4 Nano018
Gemini 3.5 Flash Lite016
NVIDIA Nemotron-3 Nano 30B-A3B103
Grok 4.50035
GPT-5.6 Sol0030
GPT-5.50027
Claude Sonnet 50026
Gemini 3.6 Flash0021
Moonshot Kimi K30021
DeepSeek V4 Pro0020
NVIDIA Nemotron-3 Ultra 550B0013
Claude Haiku 4.50010
Claude Opus 4.8007

Snapshot 2026-08-21

What changed — 2026-08-16 → 2026-08-21

  • Tasks: 51 (-1) · Models: 26 · Model–task pairs: 806 (-104) · Graded samples: 54,593 (-34)

Removed capabilities (1):

  • Direct Browse Content Synthesis — direct-browse-content-synthesis

Per-model movement (every model, 90% bar) — best value = tasks where the model is the cheapest good-enough option, counted separately under each cost mode; qualifying = tasks it’s good-enough on (quality only, so it is the same under both modes). A model with no batch API prices batch at its sync rate, so the two best-value columns diverge wherever a provider batch discount decides the winner:

ModelBest Value (Sync)Best Value (Async/Batch)Qualifying
GPT-5.6 Luna13 (+4)1832 (-2)
Alibaba Qwen3.7-Flash7 (+2)7 (+3)9 (-2)
MiniMax M37 (-3)5 (-3)20 (-5)
Tencent Hy37 (+2)4 (+1)15 (-3)
NVIDIA Nemotron-3 Super 120B5 (+3)5 (+3)6 (-3)
Meta Muse Spark 1.13232 (-3)
Thinking Machines Inkling Small2 (+1)2 (+1)25 (-3)
DeepSeek V4 Flash2 (-10)1 (-6)20
GPT-5.6 Terra1132 (-1)
Thinking Machines Inkling1 (+1)1 (+1)24 (-3)
Qwen 3.7 Plus1121 (-3)
Gemini 3.1 Flash Lite1 (-1)1 (-1)5 (-2)
Gemini 3.5 Flash0135 (-2)
GPT-5.4 Nano018
Gemini 3.5 Flash Lite016
NVIDIA Nemotron-3 Nano 30B-A3B103 (-1)
Grok 4.50035 (+1)
GPT-5.6 Sol0030 (-3)
GPT-5.50027 (-4)
Claude Sonnet 50026 (-5)
Gemini 3.6 Flash0021 (+1)
Moonshot Kimi K30021 (-5)
DeepSeek V4 Pro0020 (-4)
NVIDIA Nemotron-3 Ultra 550B0013 (-6)
Claude Haiku 4.50010
Claude Opus 4.8007 (-7)

Snapshot 2026-08-16

No changes recorded for this snapshot.