Long-form Content Generation: Route by Document Type, Quality Bar, and Output Reliability

There is no defensible single winner across the long-form content measurements supplied here. The category-level task and model counts are not stated in the available evidence, so this comparison cannot provide them without inventing a figure. The results instead support a practical quality bar of 9.0 out of 10 for quality-sensitive long-form documents, with lower-scoring models considered only when the document is less demanding, cost is more important, or the workflow can tolerate repair and retries.

The quality leader is task-dependent

The highest reported quality result is qwen3.7-plus at 9.78 out of 10 (±0.05) for onboarding-chapter writing, at $0.0055 per synchronous task run, ahead of claude-sonnet-5 at 9.19 out of 10 (±0.17) and $0.04 [^6]. On those reported figures, qwen3.7-plus leads quality by 0.59 points while costing $0.0345 less per synchronous task run than claude-sonnet-5. That is the clearest quality-and-cost advantage in a directly comparable onboarding result, although it is not a universal category verdict.

A separate onboarding finding places nvidia/nemotron-3-ultra-550b-a55 at 9.68 out of 10, narrowly ahead of kimi-k3 at 9.66 out of 10 [^6]. Because the reported onboarding results identify different leaders, the evidence does not establish one model as the uncontested best for every onboarding chapter. Engineers should treat qwen3.7-plus as the best reported result on quality alone, while validating it against the conflicting nemotron result before making onboarding the sole basis for a production-wide route.

Kimi-k3 has the broadest consistently high-quality profile across the supplied document types. It scores 9.34 out of 10 (±0.08) for section generation, 9.4 out of 10 (±0.16) for newsletter writing, and 9.21 out of 10 (±0.15) for claim-referenced analyst writing [^6] [^1][^2] [^5] [^4]. That breadth makes it the strongest candidate when one model must cover several forms of long-form work rather than maximize a single task result.

Claude-sonnet-5 is also strong across several writing tasks: 9.51 out of 10 (±0.36) for section generation, 9.25 out of 10 (±0.02) for Substack newsletters, 9.19 out of 10 (±0.17) for onboarding chapters, and 9.12 out of 10 (±0.04) for author-voice generation [^2] [^5] [^6] [^7]. Its executive-summary score is materially lower at 7.6 out of 10 (±0.1) [^1], so its strong long-form profile should not be generalized to every document type.

The cheapest model that clears the bar depends on the task

The most economical cited model that clears the 9.0 bar is gpt-5.6-luna for section generation: 9.36 out of 10 at $0.0005 per task in batch mode [^1][^2]. It is not a category-wide cheapest winner, because that result is for sections rather than onboarding chapters, newsletters, or analyst writing. For author-voice generation, minimax/minimax-m3 scores 9.74 out of 10 (±0.03) at $0.0014 per synchronous task run, beating both grok-4.5 at 9.4 and $0.01, and claude-sonnet-5 at 9.12 and $0.02 [^7]. It also scores 9.08 out of 10 (±0.03) for newsletter writing at $0.0007 [^5].

Those minimax results carry an important qualification: separate findings report 0% answer rates for its author-voice and newsletter calls [^7] [^5]. Its returned-call quality scores therefore do not establish dependable automation without retries or a fallback. In practice, gpt-5.6-luna is the clearest low-cost section route, while minimax/minimax-m3 is a high-scoring but operationally conditional choice for author voice and newsletters.

For sections, qwen3.7-plus scores 9.12 out of 10 at $0.0026 per task run, but that price is estimated because the cell has no qualifying usage history of its own [^2]. Kimi-k3 scores higher at 9.34, while gpt-5.6-luna is both cheaper and slightly higher-scoring in the cited section comparison [^6] [^1][^2]. For newsletters, qwen3.7-plus scores 8.88 out of 10 at $0.0045, below the quality bar, while kimi-k3 scores 9.4 at $0.02 and grok-4.5 scores 9.2 at $0.01 [^5]. The intervals for kimi-k3 and grok-4.5 overlap, so the evidence does not decisively separate those two newsletter options [^5].

Because the cheapest qualifying model and the best quality model are often measured on different tasks or in different serving modes, there is no honest category-wide cost-quality gap to report. The strongest directly comparable gap supplied is the onboarding comparison between qwen3.7-plus and claude-sonnet-5: 0.59 quality points and $0.0345 per synchronous task run in qwen3.7-plus’s favor [^6]. The section comparison shows a different trade-off: gpt-5.6-luna scores 0.22 points higher than qwen3.7-plus while costing $0.0021 less per task, although the cited prices and serving modes are not presented as a universal cross-task price list [^6] [^1][^2].

Where the results disagree

The individual tasks do not produce one coherent category leaderboard. Onboarding chapters favor qwen3.7-plus in one result and nvidia/nemotron-3-ultra-550b-a55 in another [^6]. Section generation favors kimi-k3 at 9.34, while gpt-5.6-luna reaches 9.36 in a separate comparison and claude-sonnet-5 reaches 9.51 in another [^6] [^1][^2]. These figures should not be collapsed into a single ranking without assuming that the task cells are directly interchangeable.

Newsletter writing presents the same pattern. Kimi-k3 scores 9.4, grok-4.5 9.2, and qwen3.7-plus 8.88 in one set of results [^1][^2] [^5], while minimax/minimax-m3 reaches 9.08 at a much lower cited cost but with a 0% answer rate in the relevant calls [^5]. Other reported newsletter results place gpt-5.5 at 8.95 out of 10 (±0.22) with an estimated $0.05 synchronous cost, gpt-5.6-luna at 8.53 out of 10 (±0.15) and $0.0003 in batch mode, and gpt-5.6-terra at 8.4 out of 10 (±0.17) and $0.0048 synchronously [^5]. Against a 9.0 quality bar, those models are not competitive for the cited newsletter task, even though the lower cost of luna may justify it for less demanding output.

Thinkingmachines/inkling has a document-centered profile rather than a clear category-wide lead. It scores 9.0 out of 10 for regulatory-filing analysis, 8.97 for section generation, 9.06 for newsletters, and 8.87 for referenced analyst claims [^3] [^2][^7] [^5] [^4]. It is therefore competitive for filing analysis and newsletters, but it falls just below the stated bar for sections and referenced analyst claims. Its results illustrate why routing by document type is more useful than assigning a universal label to a model.

Reliability changes the recommendation

Quality scores apply only to calls that return an answer when answer rates are low. Claude Opus 4-8 scores 9.02 out of 10 for referenced analyst writing and answers 72% of attempts, whereas Claude Sonnet 5 scores 8.85 but answers only 14% [^4]. Sonnet 5 reaches 9.51 for section generation but answers only 5% of attempts, making it suitable only when unanswered calls are acceptable or a fallback is available [^2]. These answer-rate differences outweigh small quality-score gaps in any workflow that requires predictable completion.

The confidence intervals also limit some conclusions. The onboarding result for qwen3.7-plus is 9.78 (±0.05), while claude-sonnet-5 is 9.19 (±0.17) [^6]. The separate newsletter comparison between kimi-k3 and grok-4.5 has overlapping intervals, so it is a tie for practical decision-making rather than a decisive kimi-k3 victory [^5]. The available evidence does not settle whether one model is categorically superior across all long-form document types, and estimated prices should not be treated as measured usage costs where the source explicitly identifies them as estimates [^2].

Routing rule

For production routing, use qwen3.7-plus for onboarding chapters when 9.78 quality and $0.0055 synchronous cost fit the requirement, but validate that route against the conflicting nvidia/nemotron-3-ultra-550b-a55 onboarding result; use gpt-5.6-luna for section generation when 9.36 clears the bar and batch execution is available; use kimi-k3 as the broad-coverage default for chapters, sections, newsletters, and analyst writing; use minimax/minimax-m3 for author voice or newsletters only with retries or a fallback; and prefer thinkingmachines/inkling for filing analysis when its document-specific profile fits. Do not choose claude-sonnet-5 for unattended section generation despite its high returned-call score unless the workflow can absorb its 5% answer rate, and do not choose the sub-9.0 newsletter options when the 9.0 bar is mandatory.

For current figures, visit llm-bench and set the quality bar your application actually needs. The live benchmark also lets engineers compare batch and synchronous pricing themselves before finalizing a route.