For Infrastructure & Utility Work—work that names topic clusters, determines topic sequence, scores report relevance, analyzes filing chunks and full filings, writes analyst content tied to cited claims, refines claims, summarizes content, generates executive summaries, performs synthesis analysis, extracts table-of-contents, translates, creates image prompts, validates queries, identifies regions, analyzes onboarding prospects, extracts direct claims, and reviews engagement—the evidence does not define a numeric category-wide quality bar, and it does not state how many tasks sit in this category or how many models are measured specifically on them. The defensible operating standard is therefore task-specific: a model must deliver sufficient quality for the particular workflow while completing the pipeline reliably enough for production and meeting its cost or latency constraint. These results rest on 54,627 scored judgments across 910 measured model-and-task combinations, covering 26 models on 52 tasks; the figures were published on a single date, but the measurement behind them is continuous and ongoing.

Claude Sonnet 5 is the strongest broadly evidenced choice across this set. For filing interpretation, it scores 8.92 out of 10 (±0.11) at $0.04 per task run in batch mode for analyzing an SEC filing chunk [^10][^14], and 8.19 out of 10 (±0.3) at $0.03 per task run in sync mode for analyzing an SEC filing [^10]. It leads topic-level discovery at 8.35 out of 10 (±0.2) at $0.03 per task run in batch mode for naming topic clusters [^11], 7.01 out of 10 (±0.23) at $0.0095 per task run for determining topic sequence [^16], and 8.48 out of 10 (±0.48) at $0.09 per task run in sync mode for scoring a topic report’s relevance [^7]. For downstream content production it records 8.85 out of 10 (±0.14) at $0.37 per task run in batch mode for writing analyst content tied to cited claims [^17], 8.35 out of 10 (±0.26) at $0.02 per task run in sync mode for refining claims [^1], 8.05 out of 10 (±0.12) at $0.02 per task run in batch mode for summarizing content [^6], 7.6 out of 10 (±0.1) at $0.04 per task run in batch mode for generating executive summaries [^4], and 8.72 out of 10 (±0.13) at $0.13 per task run in sync mode for synthesis analysis [^2]. It also scores 8.41 out of 10 (±0.24) at $0.01 per task run in sync mode for extracting a filing’s table of contents [^18], 8.94 out of 10 (±0.13) at $0.01 per task run in batch mode for translation [^3], 8.68 out of 10 (±0.11) at $0.02 per task run in sync mode for generating image prompts [^15], 8.44 out of 10 (±0.18) at $0.0045 per task run in batch mode for identifying regions [^13], and 8.85 out of 10 (±0.25) at $0.0082 per task run in batch mode for analyzing onboarding prospects [^12]. At the lower-cost end, query validation scores 6.75 out of 10 (±0.48) at $0.0038 per task run in batch mode [^9], which is weak for this category.

Where results are mixed or weaker, they are not globally decisive but task-specific. Query validation is clearly the weakest cited result for Sonnet 5 at 6.75 [^9]. Topic sequencing is modest at 7.01 [^16]. Relevance scoring is strong in point estimate (8.48) but carries a wide interval (±0.48) at $0.09 [^7], so the evidence does not tightly pin that result. Direct claim extraction is essentially a tie rather than a clear win: Sonnet 5 scores 7.55 out of 10 (±0.42) at $0.01 per task run in batch mode [^5], while Claude Opus 4.8 scores 7.51 out of 10 (±0.4) at $0.03 per task run in sync mode [^5]; the quality intervals overlap and the cost modes differ, so neither is demonstrably superior, and neither is particularly high. The prior period analysis for this category also emphasized that task-level routing matters more than a single champion; the current evidence reaffirms Sonnet 5’s breadth but confirms that claim extraction and query validation remain weaker or inconclusive points rather than decisive wins.

Against the most relevant alternative measured in the current sources, Claude Opus 4.8, the comparison is narrow. The only direct overlap is claim extraction, where the results are indistinguishable given overlapping confidence intervals and differing batch-vs-sync reporting [^5]. Opus 4.8 does hold a result Sonnet 5 lacks: 8.53 out of 10 (±0.34) at $0.01 per task run in batch mode for reviewing engagement [^8][^19]. Because Sonnet 5 has no cited result on that specific task, Opus 4.8 cannot be dismissed as uncompetitive for it, but the reverse is also true: Sonnet 5’s cited breadth across filing analysis, synthesis, content generation, and extraction is substantially wider. No other models are measured in the supplied evidence, so any claim that Sonnet 5 is categorically dominant must be limited to the tasks on which it is actually reported.

Plain-language taskModel / providerQuality (±CI)Cost / runModeRef
Claim extractionClaude Sonnet 57.55 (±0.42)$0.01Batch[^5]
Claim extractionClaude Opus 4.87.51 (±0.4)$0.03Sync[^5]
Engagement reviewClaude Opus 4.88.53 (±0.34)$0.01Batch[^8][^19]
Filing chunk analysisClaude Sonnet 58.92 (±0.11)$0.04Batch[^10][^14]
Filing analysisClaude Sonnet 58.19 (±0.3)$0.03Sync[^10]
TranslationClaude Sonnet 58.94 (±0.13)$0.01Batch[^3]

Reliability caveats are substantive. The confidence intervals for claim extraction overlap directly, and the cost modes differ, so comparing $0.01 batch to $0.03 sync as if they were the same pipeline is misleading [^5]. Several of Sonnet 5’s strongest results are reported in sync mode at higher per-task cost (for example, relevance at $0.09 [^7] and synthesis analysis at $0.13 [^2]), while its cheapest results—identifying regions at $0.0045 [^13], query validation at $0.0038 [^9], and topic sequencing at $0.0095 [^16]—carry lower quality estimates. The evidence does not state sample sizes, measurement volumes, or a category-level aggregation; each result rests on its own body of scored measurements, and that body is not visible here. Because the benchmark is continuous, a publication date marks when figures were released, not when the work was measured.

What “good enough” means depends on the specific pipeline stage. At a quality bar of roughly 8.0 out of 10, Claude Sonnet 5 is the quality leader, with its highest cited results at 8.94 for translation [^3] and 8.92 for filing chunk analysis [^10][^14]. The cheapest cited option that clears that bar is identifying regions at 8.44 (±0.18) for $0.0045 per task run in batch mode [^13]. The gap between those two is about 0.50 quality points and roughly $0.0055 per task in cost ($0.01 versus $0.0045) if comparing the 8.94 and 8.44 results, or 0.48 quality points with a $0.0355 per task gap ($0.04 versus $0.0045) if comparing filing analysis [^10][^14] against region identification [^13]. By either comparison, moving from the cheapest qualifying route to the quality-first route buys a modest score increase at a multi-fold cost increase, and the decision should not be reduced to a single number.

The routing rule for this category is concrete: start with Claude Sonnet 5 in batch mode for filing interpretation, synthesis, content generation, and structured extraction when latency permits, because its broad evidence base is the most complete in the current sources; route routine region identification, prospect analysis, or low-cost extraction to Sonnet 5’s cheaper batch configurations only if their lower but acceptable scores meet the task’s bar [^3][^12][^13][^18]; treat claim extraction as an unresolved comparison that requires internal validation rather than automatic selection [^5]; avoid relying on Sonnet 5 for query validation unless the target bar is low [^9]; and assign reviewing engagement to Claude Opus 4.8 or to a separate evaluation, since Sonnet 5 has no cited result there and Opus 4.8 scores 8.53 [^8][^19]. Do not treat any single score as a universal ranking; the evidence supports routing by task, not by model.

Check the live benchmark at https://llm-bench.kapualabs.com/ to set the quality bar your workflow actually needs. There you can compare the relevant results and check batch against synchronous pricing yourself.