The right-sized model portfolio: current routing evidence

Production routing is a portfolio problem. The current evidence covers 54,627 scored judgments across 910 measured model-and-task combinations, covering 26 models on 52 tasks; the figures were published on a single date, and the measurement behind them is continuous. No single model wins enough of these tasks to serve as a universal default. The analysis here assumes a production-usable quality bar: conditional quality around 8 of 10 or higher where the task supports that comparison, and completed calls are treated as a co-requisite. A model that earns a high score but returns nothing on many attempts, or is weak on the specific step a workflow needs, is not a production default.

Bounded extraction and document analysis are the strongest ground

Document extraction and chunk-level analysis show the most dependable quality. Qwen3.7-plus is a strong default for structured facts, document sections, and synthesis analysis: 9.67 ±0.15 at $0.02 per run, 8.65 ±0.16 at $0.0096, and 8.64 ±0.15 at $0.02 [^16][^4][^3]. DeepSeek-v4-flash offers nearly the same structured-output quality at far lower cost—9.54 at $0.0004 per run [^4]. For quality-sensitive work, gpt-5.6-sol scores 9.74 [^4] and gemini-3.5-flash scores 9.78 while answering on 77% of attempts [^4]. Deepseek-v4-pro scores 9.54 with an 80% answer rate [^4], and grok-4.5 scores 9.33 ±0.19 at $0.03 [^4]. Deepseek-v4-pro also scores 8.45 for naming topic clusters with an 84% answer rate [^25]. Minimax/minimax-m3 scores 9.36 ±0.16 for structured facts and 9.0 ±0.08 for SEC filing sections [^4][^16], but its 0% answer rates on most tested tasks make it unsuitable for unattended use [^4].

Chunk-level preprocessing is a separate strength. Gpt-5.4-nano is well suited to batch preprocessing: 9.46 at $0.0069 batch for structured output extraction and 8.41 at $0.0042 for analyzing filing chunks [^4][^16], but its broader SEC-filing analysis score is only 6.02 at $0.002 batch [^20]. Gpt-5.6-terra scores 8.84 for structured facts at $0.03 batch and 8.76 for SEC filing chunks [^4][^16]; Claude Sonnet 5 scores 8.92 at $0.04 batch [^16][^20]; Full Inkling scores 9.0 at $0.03 [^16][^20]. Gpt-5.6-luna also posts 9.36 ±0.17 at $0.0005 batch for generating sections [^18][^26], and gpt-5.6-sol scores 9.0 ±0.39 at $0.0079 batch for table-of-contents extraction [^9].

Topic organization depends on the operation

Topic work must be divided by operation. Kimi-k3 is the strongest broad option for organizing and interpreting topics: 8.47 ±0.35 at $0.08 for determining topic order, 8.98 ±0.17 at $0.05 for naming clusters, and 8.35 ±0.38 at $0.22 for matching topics to clients [^31][^25][^11]. For bounded topic operations, minimax/minimax-m3 is a lower-cost choice: 8.59 ±0.18 at $0.0005 for validating queries, 8.23 ±0.14 at $0.0006 for naming topic clusters, and 8.63 ±0.03 at $0.0011 for theme generation [^14][^25][^27]; its theme generation beats Meta/Muse-Spark-1.1’s 6.06 at $0.05 [^27], but its client-topic matching falls to 6.63 ±0.19 at $0.0046 [^11]. Tencent/hy3 is another inexpensive option: 8.28 ±0.31 at $0.0011 for sequencing, 8.08 ±0.18 at $0.0057 for cluster naming, and 8.4 ±0.49 at $0.008 for relevance scoring [^31][^25][^10]; its client-topic matching is 6.68 ±0.47 at $0.0046 [^11]. For assigning document sections to topics, tencent/hy3 scores 8.18 at $0.0004 and gpt-5.5 scores 8.55 at $0.01; their intervals overlap, so neither is decisive [^5][^25]. Gemini-3.1-flash-lite scores 7.71 at $0.0004 for assigning sections but only 4.53 at $0.02 for clustering topics [^5][^19].

Open-ended discovery is not the same as organization. Claude Sonnet 5 scores 1.46 for discovering topic clusters despite scoring 8.35 for naming clusters and 8.72 for synthesis analysis [^19][^25][^3][^15]. Content-domain suggestion is equally weak: qwen3.7-plus 3.8 ±0.14 [^22], Kimi-k3 3.82 ±0.32 at $0.0076 [^22], gpt-5.6-sol 3.84 ±0.18 at $0.0078 [^22], qwen3.7-flash 3.63 ±0.12 at $0.0003 [^22], thinkingmachines/inkling 4.33 ±0.45 at $0.0083 [^22], and tencent/hy3 5.21 ±0.41 at $0.0004 [^22]. Tencent/hy3 is the best measured, but still below a production bar. Grok 4.5 posts strong conditional scores—9.33 for structured output [^4], 8.74 for SEC-filing analysis [^16][^20], 8.91 for report relevance [^10], and 8.97 ±0.15 at $0.06 for referenced analyst writing [^29]—but only 5.23 for content-domain suggestion [^22], 6.93 for theme generation [^27], and 8.9 for topic sequences while answering on only 29% of attempts [^31]. No cost or latency is supplied for Grok 4.5, so it cannot anchor a budget-sensitive decision.

Creative generation and promotion split by audience

Creative generation also routes by task. Qwen3.7-plus scores 9.78 ±0.05 at $0.0055 for generating onboarding chapters, ahead of Claude Sonnet 5’s 9.19 ±0.17 at $0.04 [^30]. Tencent/hy3 is the low-cost synchronous option for newsletters and activity promotion: 8.52 at $0.0003 and 8.72 at $0.0004 [^34][^24]. Inkling-small scores 8.99 at $0.0061 for generating themes and 8.96 at $0.0026 for writing a newsletter [^27][^34]. Gpt-5.6-terra scores 9.57 at $0.0072 synchronous for generating sections [^18][^26][^32]. Gpt-5.6-luna scores 8.56 at $0.0012 synchronous for promotional posts [^28]. For promotional messages on X, grok-4.5 scores 7.97 at $0.34 versus tencent/hy3’s 6.6 at $0.0036, so Grok is justified only when that measured quality is worth the listed cost [^21].

Reliability can overturn a quality score

Completion behavior changes routing. Claude Sonnet 5 scores 9.51 for generating sections but returns answers on only 5% of attempts [^18]; it scores 9.19 for onboarding chapters at a 57% answer rate [^30], and 8.74 for activity promotion at a 49% answer rate [^24]. For claim-referenced analyst writing, Claude Opus 4-8 scores 9.02 ±0.21 and answers 72% of attempts, while Claude Sonnet 5 scores 8.85 but answers 14% [^29]. Kimi-k3 is strong for analyst writing at 9.21 ±0.15 at $0.19, promotional posts at 8.58 ±0.24 at $0.03, and translation at 9.4 ±0.17 at $0.03 [^29][^28][^17]; its synthesis analysis is 9.09 ±0.23 at $0.18 [^3][^15]. But no answer rate is supplied for those Kimi-k3 results, so its default status depends on confirming returned-call behavior in the application.

The 0% answer-rate models must be handled explicitly. Minimax/minimax-m3 reports 9.36 at $0.0045 for structured facts but answers on 0% of attempts for nearly all tested tasks, including subreddit vetting, query validation, and X-post selection [^4][^6][^13][^14][^32][^8]. Meta/Muse-Spark-1.1 likewise reports 0% answer rates for every listed task despite a conditional score of 9.15 for onboarding-prospect analysis [^12][^30]. These models require a fallback or an application that can tolerate unanswered requests. Similarly, Claude Sonnet 5’s low answer rates and Grok 4.5’s 29% answer rate on topic sequencing show why returned-call behavior belongs in the routing decision, not just the conditional score.

Routine routing and relevance have a cost floor

For repetitive batch classification and routine translation, gpt-5.6-luna scores 8.72 at $0.0002 batch for engagement triage and 8.78 at $0.0005 batch for translation [^35][^17]. Its post-relevance score is 6.33 [^2], so relevance-sensitive work should move to gpt-5.6-terra, which scores 9.14 at $0.02 batch for topic-report relevance [^10]. For X-post selection, gpt-5.4-nano scores 7.69 at $0.0002 batch versus gpt-5.6-terra’s 8.59 at $0.0018 batch [^8]. Deepseek-v4-flash scores 9.54 for structured-output extraction at $0.0004 and 8.28 for matching topics to clients at $0.0047 [^4][^11], but only 5.77 for topic sequencing at $0.0014 per run; that sequencing cost is estimated because the cell has no qualifying usage history of its own [^31]. Deepseek-v4-pro scores 8.34 at $0.04 for report relevance and 7.89 at $0.03 for executive summaries [^2][^10][^23][^26].

Claims and relevance decisions need particular caution. Gpt-5.4-nano scores 4.42 at $0.0017 for claim extraction [^7]; qwen3.7-flash scores 4.3 ±0.5 at $0.0004 for query validation [^14]; Claude Haiku 4.5 scores 4.61 for X-post relevance and 4.37 for subreddit selection [^2][^10][^23][^6]; minimax/minimax-m3 scores 6.63 for client-topic matching [^11]. These steps should move to stronger models or human review when misclassification is costly.

Batch versus synchronous serving changes the decision

Batch mode materially changes cost. Gpt-5.6-sol’s newsletter-writing cost is $0.0084 in batch versus $0.02 synchronously [^34]; gpt-5.6-terra’s topic-report relevance cost is $0.02 versus $0.05 [^2][^10][^23]; Claude Sonnet 5’s theme-generation cost is $0.03 versus $0.05 [^1][^27]; Claude Opus 4-8 costs $0.18 versus $0.38 for claim-referenced analyst writing and $0.14 versus $0.30 for synthesis analysis [^7][^29][^33][^3]; Gemini batch synthesis analysis is $0.03 versus $0.08 [^3]; and Claude Haiku 4.5’s claim-refinement cost is $0.0025 versus $0.0059 [^7][^33]. Use batch mode when latency can be deferred, but compare quality within the serving configuration the application actually requires.

A concrete routing portfolio

The following routing table separates activity categories, default routes, fallbacks, and the quality bar the portfolio assumes.

Activity categoryDefault routeFallback / escalationQuality bar assumed
High-volume batch classification, triage, routine translationgpt-5.6-lunagpt-5.6-terra for relevance-sensitive work; DeepSeek-v4-flash for extraction≈8+ with low batch cost and completed calls
Structured fact extraction and document section analysisDeepSeek-v4-flash (economical) or qwen3.7-plus (higher-confidence)gpt-5.6-sol; gemini-3.5-flash≈8.5+ conditional, answer rate checked
Batch chunk-level preprocessinggpt-5.4-nanoqwen3.7-plus; gpt-5.6-terra≈8+ chunk quality, avoid end-to-end filing if weaker
Topic ordering, cluster naming, client matchingKimi-k3 for quality; tencent/hy3 for costgpt-5.5 for section assignment when quality matters≈8+ when cost/latency allow; completion required
Open-ended discovery and content-domain suggestiontencent/hy3 best measured, but route to human reviewhuman review because no model clears the baraccept lower conditional score with mandatory review
Newsletters and activity promotiontencent/hy3Inkling-small≈8+ with low cost
Onboarding chapters and structured creative generationqwen3.7-plusgpt-5.6-terra≈9 for public-facing content
Claim-referenced analyst writing and high-consequence synthesisClaude Opus 4-8 (batch)Kimi-k3high quality plus dependable completion
Report relevance and X-post selectiongpt-5.6-terraDeepSeek-v4-pro≈8.5+ due misclassification cost
Premium translation and downstream synthesisKimi-k3 for premium quality; gpt-5.6-luna for routinegpt-5.6-lunahigh quality when semantic errors are costly

This table uses thirteen distinct models as explicit defaults or fallbacks: gpt-5.6-luna, gpt-5.6-terra, DeepSeek-v4-flash, qwen3.7-plus, gpt-5.4-nano, gpt-5.6-sol, gemini-3.5-flash, tencent/hy3, Kimi-k3, gpt-5.5, Inkling-small, Claude Opus 4-8, and DeepSeek-v4-pro. It excludes several models that are conditionally strong but unreliable or unevaluated on cost—Full Inkling, Claude Sonnet 5, minimax/minimax-m3, grok-4.5, Meta/Muse-Spark-1.1, qwen3.7-flash, and thinkingmachines/inkling—because their answer rates, missing cost data, or weak discovery results would make them unsafe defaults.

Why not fewer? The evidence does not permit a one-model portfolio. Gpt-5.6-luna is inexpensive for triage and batch translation, but its post-relevance score is 6.33 [^2] and it cannot cover premium analyst writing or topic organization. Qwen3.7-plus is excellent for extraction and onboarding chapters, but it scores 3.8 on content-domain suggestion [^22]. Kimi-k3 or Claude Opus 4-8 would improve high-value outputs, but they carry far higher per-task costs and would impose unnecessary spend on triage, extraction, and routine creative work. A single-model portfolio would therefore either accept measurable quality drops on the steps that particular model scores poorly, or pay premium prices on high-volume work that cheaper models complete adequately.

Several routing decisions are close enough to flip next snapshot. DeepSeek-v4-flash and qwen3.7-plus are separated by only 0.13 quality points for structured extraction, with a large cost difference; if flash’s returned-call behavior is weaker than qwen3.7-plus’s, the default should flip. Tencent/hy3 and gpt-5.5 overlap for section assignment [^5][^25]. Minimax/minimax-m3 would challenge on bounded topic work if its 0% answer-rate behavior improves. Claude Opus 4-8 and Kimi-k3 are close for analyst writing, but Kimi-k3 currently lacks a supplied answer rate. Qwen3.7-plus versus gpt-5.6-terra for onboarding and section generation may also shift if completion data become available. These routes should remain policy-controlled rather than hard-coded as permanent rankings.

What the evidence cannot yet settle

Some results carry no confidence intervals, and several intervals overlap where they are supplied. The evidence does not state answer rates for every model, so completion risk is incompletely measured for some routes. No cost or latency figures are supplied for grok-4.5. DeepSeek-v4-flash’s topic-sequencing cost is estimated because the cell has no qualifying usage history of its own [^31]. Batch and synchronous prices are not standardized to workload or token assumptions; engineers should compare effective cost per accepted clean output, including retries and review, rather than the headline per-task price.

The weakest measured steps are query validation, claim extraction, X-post relevance, and subreddit selection: gpt-5.4-nano scores 4.42 [^7], qwen3.7-flash 4.3 [^14], Claude Haiku 4.5 4.61 and 4.37 [^2][^10][^23][^6]. Open-ended discovery and content-domain suggestion also fail to clear a production bar for every model measured [^22]. These should stay with human review until a model clears the bar.

The live benchmark is available at https://llm-bench.kapualabs.com/, where engineers can set the quality bar their workflow actually requires and compare batch against synchronous pricing themselves. The numbers here are a snapshot of a continuous measurement programme; the routing should be rechecked when those numbers or completion rates change.