What This Measurement Can—and Cannot—Support

The current evidence does not support a universal model ranking. It supports routing decisions by workflow stage, quality bar, serving mode, cost, and the probability that a request produces a usable answer. A conditional quality score is not enough when the model frequently fails to answer, returns output that cannot be used downstream, or carries a cost and latency burden the production system cannot absorb.

For this comparison, “good enough” means clearing the quality threshold for the task while also meeting the workflow’s operational requirements: dependable completion, acceptable output format, a tolerable response time, and sustainable cost. The evidence does not define one universal threshold. It does show that 12 of 68 tasks use task-specific rubrics rather than the default, with qualification bars at 95%, 90%, and 80% [^11]. The appropriate bar is therefore a deployment choice, not a benchmark-wide ranking rule.

The measurement’s shape and its limits

The overall measurement covers 26 models on 52 tasks, with 910 measured model-and-task combinations out of 1,352 possible combinations in a complete 26-by-52 grid. It rests on 54,627 scored judgments. The supplied material does not break those 910 combinations down by model or task, however, so it cannot establish which models are measured broadly, which are measured narrowly, or which tasks have the thinnest coverage. A model’s absence from a task is not evidence of weakness, and a result on one task should not be treated as a model-wide capability claim.

The cost-related findings use a separate reported denominator of 68 tasks: 50 are described as batch-eligible, while the other 18 have no valid batch-priced comparison [^11]. The material does not explain how that 68-task cost scope relates to the 52-task overall measurement scope, so those figures should not be combined into a single coverage rate. What the cost evidence does establish is that serving mode matters: batch is not merely a pricing detail, but an operating choice that can materially change the economics of a route.

The current evidence also cannot quantify how much published cost data was freshly observed, how much was repriced from an older observation, or how much was modelled from an estimator. It reports sync and batch figures, and identifies some estimated values—including gpt-5.5’s $0.05 synchronous newsletter cost [^19]—but it does not provide the requested freshness breakdown. Cost comparisons should therefore be read as the figures attached to the reported task and mode, not as a complete statement about current or future operating cost.

Coverage also limits some direct comparisons. For example, the supplied evidence gives batch costs for topic sequencing, cluster naming, and client-topic matching, but does not provide an equally complete quality-and-reliability picture for every model in those comparisons. It likewise provides cost differences between batch and sync for several tasks without showing that the cheaper mode clears the same quality bar: gpt-5.5 is listed at $0.27 in batch versus $0.56 synchronously for trading recommendations [^32]; gemini-3.5-flash at $0.01 versus $0.14 for scoring document relevance [^1][^12][^25]; gpt-5.6-terra at $0.02 versus $0.05 [^1][^12][^25]; and gpt-5.6-sol at $0.0084 versus $0.02 for newsletters [^19]. The evidence explicitly does not supply the quality comparison needed to conclude that the cheaper mode is adequate in those cases [^2][^31][^4][^20][^19].

Where the models are strong

The clearest strength is task specialization. qwen3.7-plus is strong on structured extraction and document analysis, scoring 9.67 out of 10 (±0.15) for structured extraction [^8], 8.65 out of 10 (±0.16) for analyzing document sections [^13], and 8.64 out of 10 (±0.15) for synthesizing analysis [^3]. Those results do not generalize to discovery or content planning: the same model scores 5.89 out of 10 (±0.13) for generating themes [^9] and 3.8 out of 10 (±0.14) for suggesting content domains [^27]. More importantly, only 40% of extraction attempts are answered [^13]. It is a candidate for workflows that can manage retries or accept that completion rate, not an automatic default for every document task.

Other models show a similar pattern. deepseek-v4-flash scores 9.54 out of 10 for structured information extraction and 9.9 out of 10 for language detection [^13][^23], but falls to 7.43 out of 10 for source-grounded claim extraction and 7.1 out of 10 for summarization [^4][^24][^15]. gpt-5.4-nano reaches 9.46 out of 10 for structured output extraction and 9.33 out of 10 for summarization [^13][^15], while its claim-extraction score is only 4.42 out of 10 [^4]. Its summary answers arrive on 87% of attempts [^15], which may be acceptable for a recoverable pipeline but does not amount to guaranteed completion.

minimax/minimax-m3 is another strong structured-analysis candidate on nominal quality, scoring 9.36 out of 10 for structured fact extraction [^13] and 9.0 out of 10 (±0.08) for analyzing SEC filing sections [^8]. Yet it answers on 0% of attempts for theme generation despite a conditional score of 8.63 [^9], and likewise records a 9.74 out of 10 conditional score for author-voice generation with 0% answers [^26]. Those results make it suitable only where the measured task and a fallback strategy align; they do not support using it as a general content-generation route.

Claude-sonnet-5 is strong for several writing and organization stages. It scores 9.7 out of 10 (±0.09) for generating themes at $0.03 per task run in batch mode [^9], 8.35 out of 10 (±0.2) for naming topic clusters [^16], and 8.72 out of 10 (±0.13) for synthesis [^3][^7]. It also appears competitive for copy-related work [^9]. That strength has a sharp boundary: it scores only 1.46 out of 10 (±0.09) for discovering and clustering topics at $0.20 per task run synchronously [^10]. A production pipeline should therefore use a separate discovery method and route organization or writing to a model measured for those later stages.

Kimi-k3 is the strongest quality-oriented option among the supplied topic-organization results, scoring 8.47 out of 10 (±0.35) for determining topic order, 8.98 out of 10 (±0.17) for naming topic clusters, and 8.35 out of 10 (±0.38) for matching topics to clients [^21][^16][^5]. It also scores 9.21 for claim-referenced analyst writing, although it answers only 42% of attempts [^18], and answers 79% of cluster-naming attempts [^16]. The model is therefore a quality-sensitive option for multi-step organization when cost and incomplete responses can be tolerated, not a universally dependable route.

Mixed results and operational trade-offs

The central trade-off is visible in social and editorial generation. qwen3.7-plus costs $0.0098 per synchronous task run and scores 8.04 out of 10 (±0.26) for writing an automatic Reddit post [^17]. Claude-sonnet-5 scores 8.37 out of 10 (±0.16), but costs $0.08 per task run in batch [^17]. The figures describe different serving modes, so neither establishes a like-for-like winner. They instead show why cost, mode, and quality must be fixed before comparing routes.

Batch can substantially reduce price where it is available. For claim-referenced analyst writing, claude-opus-4-8 costs $0.18 per task in batch versus $0.38 synchronously [^4][^18][^24]; for synthesis analysis, it costs $0.14 versus $0.30 [^3]. Similar savings appear in the Gemini, GPT-5.5, terra, and sol comparisons above. But cheaper batch execution is not automatically the right choice: the supplied figures do not show whether each lower-cost mode preserves the quality, answer rate, or clean-output requirements of the synchronous route.

Hy3, served by Tencent, is the cost-sensitive alternative for some topic-organization stages. It costs $0.0011 for topic sequencing, $0.0057 for cluster naming, and $0.0046 for client-topic matching [^21][^16][^5], compared with kimi-k3 at $0.08, $0.05, and $0.22 for those respective tasks [^21][^16][^5]. The price advantage is not sufficient for every stage: hy3’s client-topic matching score is 6.68 out of 10 [^5], making that step a candidate for a stronger model or human review.

The confidence intervals also limit claims of separation. For author-referenced writing, Claude Opus 4-8 scores 9.02 out of 10 (±0.21), while Claude Sonnet 5 scores 8.85 out of 10 (±0.14) [^18]. Those intervals overlap, so the evidence does not establish a decisive winner. The same discipline applies wherever reported intervals overlap: a small numerical lead should not be converted into a ranking the measurement cannot support.

Where quality is not enough

Reliability frequently changes the practical conclusion. gpt-5.6-sol scores 8.88 out of 10 for grouping related topics and 9.13 out of 10 for generating themes when it answers, but answers on only 73% and 62% of attempts respectively [^10][^9]. Grok-4.5 scores 9.33 out of 10 for structured output [^13] and 8.74 out of 10 for SEC-filing analysis [^8][^30], but answers on only 50% of filing-analysis attempts [^30]. In both cases, the conditional quality result is useful evidence about answered requests, but it is not evidence of dependable end-to-end production performance.

The same warning is even clearer in discovery and onboarding. gemini-3.1-flash-lite scores 4.53 out of 10 (±0.29) for discovery [^10][^16], while gemini-3.5-flash scores 8.96 out of 10 (±0.21) for determining sequence [^21] and gemini-3.6-flash scores 8.27 out of 10 (±0.19) for naming clusters [^16]. NVIDIA-served nemotron-3-super-120b-a12b records a conditional onboarding-chapter-generation score of 9.09 out of 10 but answers on 0% of attempts [^29]. Muse-spark-1.1 reports 9.15 out of 10 for analyzing onboarding prospects [^6][^29], yet answers on 0% of attempts for every listed task [^33]. Neither result supports production use without a material change in answer behavior or a separate recovery path.

The supplied results also identify tasks where weak relevance, matching, claim extraction, query validation, or subreddit selection require validation or human review rather than blind routing [^1][^5][^28][^14][^4][^22]. This is not a minor qualification: a model that scores well on one stage can still fail the specific output contract needed by the next stage.

Head-to-head implications and routing

The practical choice is therefore stage-specific. For structured extraction and document analysis, qwen3.7-plus, gemini-3.5-flash, gpt-5.4-nano, and minimax/minimax-m3 are all candidates under different reliability requirements. Gemini-3.5-flash scores 9.78 out of 10 for structured extraction with 77% answers and 9.02 out of 10 for filing analysis with 86% answers [^13][^8]. Minimax/minimax-m3 offers high conditional quality on those tasks but needs a fallback for its 0% answer result on the relevant generation workload [^13][^8][^13]. Gpt-5.4-nano is attractive for extraction and summarization when its 87% summary answer rate meets the system’s requirement [^13][^15]. Qwen3.7-plus is compelling on quality and measured document work, but its 40% extraction answer rate must be treated as a primary routing constraint [^8][^13].

For topic discovery, do not transfer Claude Sonnet 5’s strong theme-generation result to discovery and clustering: its 1.46 score there argues for a separate discovery route [^9][^10]. For sequencing and cluster naming, hy3 is the economical option [^21][^16], while kimi-k3 is the quality-oriented option [^21][^16]. Client-topic matching is less settled: kimi-k3’s 8.35 score comes with cost and answer-behavior concerns [^5], while hy3 is much cheaper but scores 6.68 [^5]. That comparison supports validation or human review rather than a universal recommendation.

For writing and synthesis, Claude Sonnet 5 is a strong choice for theme generation, copy, cluster naming, and synthesis [^9][^16][^3][^7]. Claude Opus 4-8 may be preferable when its measured completion behavior or task-specific requirements matter, but the overlapping author-referenced-writing intervals do not establish a decisive quality winner [^18]. Kimi-k3 is reserved for quality-sensitive analyst writing or synthesis where its cost and incomplete-answer behavior are acceptable [^18].

An engineer should consequently build routing around the workflow’s actual gates: use a strong extraction model only when its answer rate and output contract are adequate; use separate models for discovery, organization, and generation; choose hy3 for cost-sensitive sequencing or cluster naming; choose kimi-k3 when those stages justify higher cost and more variable completion; and add validation or human review to weak relevance, matching, claim-extraction, query-validation, and subreddit-selection steps. A synchronous-only requirement, a higher quality bar, no tolerance for unanswered calls, or a requirement for clean downstream output can all change the preferred route, and the current material does not provide one universal latency measure that resolves those choices.

The current coverage therefore supports routing hypotheses, not universal procurement conclusions. It cannot establish model breadth, thinly covered tasks, cost-data freshness, or a definitive choice for unmeasured work; those comparisons require evidence not supplied here.

Check the live figures at llm-bench.kapualabs.com and set the quality bar your production workflow actually requires. You can also compare batch against synchronous pricing there before selecting a route.