The Right-Sized Model Portfolio
The right-sized model portfolio: current routing evidence
Production routing is a portfolio problem. The current evidence covers 54,627 scored judgments across 910 measured model-and-task combinations, covering 26 models on 52 tasks; the figures were published on a single date, and the measurement behind them is continuous. No single model wins enough of these tasks to serve as a universal default. The analysis here assumes a production-usable quality bar: conditional quality around 8 of 10 or higher where the task supports that comparison, and completed calls are treated as a co-requisite. A model that earns a high score but returns nothing on many attempts, or is weak on the specific step a workflow needs, is not a production default.
Bounded extraction and document analysis are the strongest ground
Document extraction and chunk-level analysis show the most dependable quality. Qwen3.7-plus is a strong default for structured facts, document sections, and synthesis analysis: 9.67 ±0.15 at $0.02 per run, 8.65 ±0.16 at $0.0096, and 8.64 ±0.15 at $0.02 [^16][^4][^3]. DeepSeek-v4-flash offers nearly the same structured-output quality at far lower cost—9.54 at $0.0004 per run [^4]. For quality-sensitive work, gpt-5.6-sol scores 9.74 [^4] and gemini-3.5-flash scores 9.78 while answering on 77% of attempts [^4]. Deepseek-v4-pro scores 9.54 with an 80% answer rate [^4], and grok-4.5 scores 9.33 ±0.19 at $0.03 [^4]. Deepseek-v4-pro also scores 8.45 for naming topic clusters with an 84% answer rate [^25]. Minimax/minimax-m3 scores 9.36 ±0.16 for structured facts and 9.0 ±0.08 for SEC filing sections [^4][^16], but its 0% answer rates on most tested tasks make it unsuitable for unattended use [^4].
Chunk-level preprocessing is a separate strength. Gpt-5.4-nano is well suited to batch preprocessing: 9.46 at $0.0069 batch for structured output extraction and 8.41 at $0.0042 for analyzing filing chunks [^4][^16], but its broader SEC-filing analysis score is only 6.02 at $0.002 batch [^20]. Gpt-5.6-terra scores 8.84 for structured facts at $0.03 batch and 8.76 for SEC filing chunks [^4][^16]; Claude Sonnet 5 scores 8.92 at $0.04 batch [^16][^20]; Full Inkling scores 9.0 at $0.03 [^16][^20]. Gpt-5.6-luna also posts 9.36 ±0.17 at $0.0005 batch for generating sections [^18][^26], and gpt-5.6-sol scores 9.0 ±0.39 at $0.0079 batch for table-of-contents extraction [^9].
Topic organization depends on the operation
Topic work must be divided by operation. Kimi-k3 is the strongest broad option for organizing and interpreting topics: 8.47 ±0.35 at $0.08 for determining topic order, 8.98 ±0.17 at $0.05 for naming clusters, and 8.35 ±0.38 at $0.22 for matching topics to clients [^31][^25][^11]. For bounded topic operations, minimax/minimax-m3 is a lower-cost choice: 8.59 ±0.18 at $0.0005 for validating queries, 8.23 ±0.14 at $0.0006 for naming topic clusters, and 8.63 ±0.03 at $0.0011 for theme generation [^14][^25][^27]; its theme generation beats Meta/Muse-Spark-1.1’s 6.06 at $0.05 [^27], but its client-topic matching falls to 6.63 ±0.19 at $0.0046 [^11]. Tencent/hy3 is another inexpensive option: 8.28 ±0.31 at $0.0011 for sequencing, 8.08 ±0.18 at $0.0057 for cluster naming, and 8.4 ±0.49 at $0.008 for relevance scoring [^31][^25][^10]; its client-topic matching is 6.68 ±0.47 at $0.0046 [^11]. For assigning document sections to topics, tencent/hy3 scores 8.18 at $0.0004 and gpt-5.5 scores 8.55 at $0.01; their intervals overlap, so neither is decisive [^5][^25]. Gemini-3.1-flash-lite scores 7.71 at $0.0004 for assigning sections but only 4.53 at $0.02 for clustering topics [^5][^19].
Open-ended discovery is not the same as organization. Claude Sonnet 5 scores 1.46 for discovering topic clusters despite scoring 8.35 for naming clusters and 8.72 for synthesis analysis [^19][^25][^3][^15]. Content-domain suggestion is equally weak: qwen3.7-plus 3.8 ±0.14 [^22], Kimi-k3 3.82 ±0.32 at $0.0076 [^22], gpt-5.6-sol 3.84 ±0.18 at $0.0078 [^22], qwen3.7-flash 3.63 ±0.12 at $0.0003 [^22], thinkingmachines/inkling 4.33 ±0.45 at $0.0083 [^22], and tencent/hy3 5.21 ±0.41 at $0.0004 [^22]. Tencent/hy3 is the best measured, but still below a production bar. Grok 4.5 posts strong conditional scores—9.33 for structured output [^4], 8.74 for SEC-filing analysis [^16][^20], 8.91 for report relevance [^10], and 8.97 ±0.15 at $0.06 for referenced analyst writing [^29]—but only 5.23 for content-domain suggestion [^22], 6.93 for theme generation [^27], and 8.9 for topic sequences while answering on only 29% of attempts [^31]. No cost or latency is supplied for Grok 4.5, so it cannot anchor a budget-sensitive decision.
Creative generation and promotion split by audience
Creative generation also routes by task. Qwen3.7-plus scores 9.78 ±0.05 at $0.0055 for generating onboarding chapters, ahead of Claude Sonnet 5’s 9.19 ±0.17 at $0.04 [^30]. Tencent/hy3 is the low-cost synchronous option for newsletters and activity promotion: 8.52 at $0.0003 and 8.72 at $0.0004 [^34][^24]. Inkling-small scores 8.99 at $0.0061 for generating themes and 8.96 at $0.0026 for writing a newsletter [^27][^34]. Gpt-5.6-terra scores 9.57 at $0.0072 synchronous for generating sections [^18][^26][^32]. Gpt-5.6-luna scores 8.56 at $0.0012 synchronous for promotional posts [^28]. For promotional messages on X, grok-4.5 scores 7.97 at $0.34 versus tencent/hy3’s 6.6 at $0.0036, so Grok is justified only when that measured quality is worth the listed cost [^21].
Reliability can overturn a quality score
Completion behavior changes routing. Claude Sonnet 5 scores 9.51 for generating sections but returns answers on only 5% of attempts [^18]; it scores 9.19 for onboarding chapters at a 57% answer rate [^30], and 8.74 for activity promotion at a 49% answer rate [^24]. For claim-referenced analyst writing, Claude Opus 4-8 scores 9.02 ±0.21 and answers 72% of attempts, while Claude Sonnet 5 scores 8.85 but answers 14% [^29]. Kimi-k3 is strong for analyst writing at 9.21 ±0.15 at $0.19, promotional posts at 8.58 ±0.24 at $0.03, and translation at 9.4 ±0.17 at $0.03 [^29][^28][^17]; its synthesis analysis is 9.09 ±0.23 at $0.18 [^3][^15]. But no answer rate is supplied for those Kimi-k3 results, so its default status depends on confirming returned-call behavior in the application.
The 0% answer-rate models must be handled explicitly. Minimax/minimax-m3 reports 9.36 at $0.0045 for structured facts but answers on 0% of attempts for nearly all tested tasks, including subreddit vetting, query validation, and X-post selection [^4][^6][^13][^14][^32][^8]. Meta/Muse-Spark-1.1 likewise reports 0% answer rates for every listed task despite a conditional score of 9.15 for onboarding-prospect analysis [^12][^30]. These models require a fallback or an application that can tolerate unanswered requests. Similarly, Claude Sonnet 5’s low answer rates and Grok 4.5’s 29% answer rate on topic sequencing show why returned-call behavior belongs in the routing decision, not just the conditional score.
Routine routing and relevance have a cost floor
For repetitive batch classification and routine translation, gpt-5.6-luna scores 8.72 at $0.0002 batch for engagement triage and 8.78 at $0.0005 batch for translation [^35][^17]. Its post-relevance score is 6.33 [^2], so relevance-sensitive work should move to gpt-5.6-terra, which scores 9.14 at $0.02 batch for topic-report relevance [^10]. For X-post selection, gpt-5.4-nano scores 7.69 at $0.0002 batch versus gpt-5.6-terra’s 8.59 at $0.0018 batch [^8]. Deepseek-v4-flash scores 9.54 for structured-output extraction at $0.0004 and 8.28 for matching topics to clients at $0.0047 [^4][^11], but only 5.77 for topic sequencing at $0.0014 per run; that sequencing cost is estimated because the cell has no qualifying usage history of its own [^31]. Deepseek-v4-pro scores 8.34 at $0.04 for report relevance and 7.89 at $0.03 for executive summaries [^2][^10][^23][^26].
Claims and relevance decisions need particular caution. Gpt-5.4-nano scores 4.42 at $0.0017 for claim extraction [^7]; qwen3.7-flash scores 4.3 ±0.5 at $0.0004 for query validation [^14]; Claude Haiku 4.5 scores 4.61 for X-post relevance and 4.37 for subreddit selection [^2][^10][^23][^6]; minimax/minimax-m3 scores 6.63 for client-topic matching [^11]. These steps should move to stronger models or human review when misclassification is costly.
Batch versus synchronous serving changes the decision
Batch mode materially changes cost. Gpt-5.6-sol’s newsletter-writing cost is $0.0084 in batch versus $0.02 synchronously [^34]; gpt-5.6-terra’s topic-report relevance cost is $0.02 versus $0.05 [^2][^10][^23]; Claude Sonnet 5’s theme-generation cost is $0.03 versus $0.05 [^1][^27]; Claude Opus 4-8 costs $0.18 versus $0.38 for claim-referenced analyst writing and $0.14 versus $0.30 for synthesis analysis [^7][^29][^33][^3]; Gemini batch synthesis analysis is $0.03 versus $0.08 [^3]; and Claude Haiku 4.5’s claim-refinement cost is $0.0025 versus $0.0059 [^7][^33]. Use batch mode when latency can be deferred, but compare quality within the serving configuration the application actually requires.
A concrete routing portfolio
The following routing table separates activity categories, default routes, fallbacks, and the quality bar the portfolio assumes.
| Activity category | Default route | Fallback / escalation | Quality bar assumed |
|---|---|---|---|
| High-volume batch classification, triage, routine translation | gpt-5.6-luna | gpt-5.6-terra for relevance-sensitive work; DeepSeek-v4-flash for extraction | ≈8+ with low batch cost and completed calls |
| Structured fact extraction and document section analysis | DeepSeek-v4-flash (economical) or qwen3.7-plus (higher-confidence) | gpt-5.6-sol; gemini-3.5-flash | ≈8.5+ conditional, answer rate checked |
| Batch chunk-level preprocessing | gpt-5.4-nano | qwen3.7-plus; gpt-5.6-terra | ≈8+ chunk quality, avoid end-to-end filing if weaker |
| Topic ordering, cluster naming, client matching | Kimi-k3 for quality; tencent/hy3 for cost | gpt-5.5 for section assignment when quality matters | ≈8+ when cost/latency allow; completion required |
| Open-ended discovery and content-domain suggestion | tencent/hy3 best measured, but route to human review | human review because no model clears the bar | accept lower conditional score with mandatory review |
| Newsletters and activity promotion | tencent/hy3 | Inkling-small | ≈8+ with low cost |
| Onboarding chapters and structured creative generation | qwen3.7-plus | gpt-5.6-terra | ≈9 for public-facing content |
| Claim-referenced analyst writing and high-consequence synthesis | Claude Opus 4-8 (batch) | Kimi-k3 | high quality plus dependable completion |
| Report relevance and X-post selection | gpt-5.6-terra | DeepSeek-v4-pro | ≈8.5+ due misclassification cost |
| Premium translation and downstream synthesis | Kimi-k3 for premium quality; gpt-5.6-luna for routine | gpt-5.6-luna | high quality when semantic errors are costly |
This table uses thirteen distinct models as explicit defaults or fallbacks: gpt-5.6-luna, gpt-5.6-terra, DeepSeek-v4-flash, qwen3.7-plus, gpt-5.4-nano, gpt-5.6-sol, gemini-3.5-flash, tencent/hy3, Kimi-k3, gpt-5.5, Inkling-small, Claude Opus 4-8, and DeepSeek-v4-pro. It excludes several models that are conditionally strong but unreliable or unevaluated on cost—Full Inkling, Claude Sonnet 5, minimax/minimax-m3, grok-4.5, Meta/Muse-Spark-1.1, qwen3.7-flash, and thinkingmachines/inkling—because their answer rates, missing cost data, or weak discovery results would make them unsafe defaults.
Why not fewer? The evidence does not permit a one-model portfolio. Gpt-5.6-luna is inexpensive for triage and batch translation, but its post-relevance score is 6.33 [^2] and it cannot cover premium analyst writing or topic organization. Qwen3.7-plus is excellent for extraction and onboarding chapters, but it scores 3.8 on content-domain suggestion [^22]. Kimi-k3 or Claude Opus 4-8 would improve high-value outputs, but they carry far higher per-task costs and would impose unnecessary spend on triage, extraction, and routine creative work. A single-model portfolio would therefore either accept measurable quality drops on the steps that particular model scores poorly, or pay premium prices on high-volume work that cheaper models complete adequately.
Several routing decisions are close enough to flip next snapshot. DeepSeek-v4-flash and qwen3.7-plus are separated by only 0.13 quality points for structured extraction, with a large cost difference; if flash’s returned-call behavior is weaker than qwen3.7-plus’s, the default should flip. Tencent/hy3 and gpt-5.5 overlap for section assignment [^5][^25]. Minimax/minimax-m3 would challenge on bounded topic work if its 0% answer-rate behavior improves. Claude Opus 4-8 and Kimi-k3 are close for analyst writing, but Kimi-k3 currently lacks a supplied answer rate. Qwen3.7-plus versus gpt-5.6-terra for onboarding and section generation may also shift if completion data become available. These routes should remain policy-controlled rather than hard-coded as permanent rankings.
What the evidence cannot yet settle
Some results carry no confidence intervals, and several intervals overlap where they are supplied. The evidence does not state answer rates for every model, so completion risk is incompletely measured for some routes. No cost or latency figures are supplied for grok-4.5. DeepSeek-v4-flash’s topic-sequencing cost is estimated because the cell has no qualifying usage history of its own [^31]. Batch and synchronous prices are not standardized to workload or token assumptions; engineers should compare effective cost per accepted clean output, including retries and review, rather than the headline per-task price.
The weakest measured steps are query validation, claim extraction, X-post relevance, and subreddit selection: gpt-5.4-nano scores 4.42 [^7], qwen3.7-flash 4.3 [^14], Claude Haiku 4.5 4.61 and 4.37 [^2][^10][^23][^6]. Open-ended discovery and content-domain suggestion also fail to clear a production bar for every model measured [^22]. These should stay with human review until a model clears the bar.
The live benchmark is available at https://llm-bench.kapualabs.com/, where engineers can set the quality bar their workflow actually requires and compare batch against synchronous pricing themselves. The numbers here are a snapshot of a continuous measurement programme; the routing should be rechecked when those numbers or completion rates change.