Cost mode:

Category: Long-form Content Generation · Rail: absolute · Typical I/O: 1210→2371 tokens

Models

Frontier on this task: Qwen 3.7 Plus at 9.78 / 10. Quality bar at 90%: 8.80.

point-estimate floor (CI low) · upper CI (less certain) · Bars sorted by blended cost; best-value model first. Greyed rows are MEDIUM+ models whose point estimate clears the bar but whose CI low does not.

ModelQuality scoreCI lowCost / 1k runsvs best value
NVIDIA Nemotron-3 Super 120B9.09 / 108.92$1.33best value
Tencent Hy38.91 / 108.46$2.141.6x more expensive
Qwen 3.7 Plus9.78 / 109.73$3.802.9x more expensive
Qwen 3.6 Flash9.61 / 109.55$4.433.3x more expensive
GPT-5.6 Luna9.32 / 109.23$6.084.6x more expensive
NVIDIA Nemotron-3 Ultra 550B9.68 / 109.61$9.266.9x more expensive
Claude Sonnet 59.19 / 109.03$12.879.7x more expensive
GPT-5.6 Terra9.46 / 109.37$14.4911x more expensive
Meta Muse Spark 1.18.88 / 108.67$15.5312x more expensive
Grok 4.59.59 / 109.55$17.2913x more expensive
GPT-5.6 Sol9.67 / 109.60$38.0729x more expensive
NVIDIA Nemotron-3 Nano 30B-A3B8.27 / 108.01$0.5360% cheaper

Cost breakdown

ModelQualityConfidenceCost / 1k runsOverpayMode
NVIDIA Nemotron-3 Super 120B OpenRouter9.09 / 10 CI [8.92, 9.26]RANKED$1.33best valuebatch
Tencent Hy3 OpenRouter8.91 / 10 CI [8.46, 9.37]MEDIUM$2.141.6xbatch
Qwen 3.7 Plus best Alibaba Cloud (DashScope)9.78 / 10 CI [9.73, 9.83]RANKED$3.802.9xbatch
Qwen 3.6 Flash Alibaba Cloud (DashScope)9.61 / 10 CI [9.55, 9.66]RANKED$4.433.3xbatch
GPT-5.6 Luna OpenAI9.32 / 10 CI [9.23, 9.42]RANKED$6.084.6xbatch
NVIDIA Nemotron-3 Ultra 550B OpenRouter9.68 / 10 CI [9.61, 9.75]RANKED$9.266.9xbatch
Claude Sonnet 5 Anthropic9.19 / 10 CI [9.03, 9.36]RANKED$12.879.7xbatch
GPT-5.6 Terra OpenAI9.46 / 10 CI [9.37, 9.54]RANKED$14.4911xbatch
Meta Muse Spark 1.1 Meta8.88 / 10 CI [8.67, 9.09]HIGH$15.5312xbatch
Grok 4.5 xAI9.59 / 10 CI [9.55, 9.63]RANKED$17.2913xbatch
GPT-5.6 Sol OpenAI9.67 / 10 CI [9.60, 9.74]RANKED$38.0729xbatch

Overpay shows how much more you pay than the best-value model that clears the quality bar (marked ★) — the best-value good-enough option. "16x" means you overpay 16× — 16× that reference for no quality benefit above the bar. Typical call shape for this task: 1210 input tokens → 2371 output tokens, EMA-tracked from production traffic. Cost is the observed, all-in $ per 1,000 task runs: each model's own measured usage on this task — output verbosity, thinking/reasoning tokens, cache reads and writes, and the spend on its billed failures — priced at current list rates and adjusted by the billing overhead we actually reconcile against provider invoices. Models that answer tersely cost what they actually cost; models that think at length pay for it. Not comparable to providers' advertised $/1M list rates — this is what running the task costs, not a per-token price.

Prompt templates

The system + user template pair used for this task.

GENERATE_CHAPTERS_SYSTEM + GENERATE_CHAPTERS_USER (629 calls in window)

System prompt

You are an expert report structure designer for a market intelligence system.

Your task is to design report chapters that will provide comprehensive coverage of a subject. The number of chapters should match the complexity and scope of the subject - typically 5-10 chapters, but could be more for complex subjects requiring deeper analysis.

Each chapter MUST be comprehensive and detailed. The chapter definitions are used downstream to:
- Generate search queries for research
- Guide relevance scoring of content
- Customize synthesis prompts

Each chapter should include:
1. **chapter_title**: Clear, descriptive title
2. **chapter_code**: snake_case identifier (max 30 chars)
3. **user_requirement**: DETAILED description (3-5 sentences) explaining:
   - What specific topics/aspects to cover
   - What questions should be answered
   - What type of information to gather (data, opinions, trends, etc.)
   - Key entities, metrics, or concepts to focus on

Example of a GOOD user_requirement:
"Analyze the competitive landscape including market share data, key competitors' strategies, recent product launches, and competitive advantages. Focus on direct competitors in the same market segment, their pricing strategies, technology differentiators, and customer acquisition approaches. Include any recent M&A activity or partnerships that affect competitive dynamics."

Example of a BAD user_requirement (too vague):
"Cover the competition and market dynamics."

Standard chapter types to consider (adapt based on subject):
- Market/Industry Overview (size, growth, key players, trends)
- Key Developments & News (recent announcements, events, milestones)
- Competitive Analysis (competitors, market share, positioning)
- Technology & Innovation (R&D, patents, product development)
- Regulatory & Policy (compliance, government actions, legal issues)
- Financial Analysis (revenue, margins, investments - for corporations)
- Consumer/Customer Insights (sentiment, behavior, preferences)
- Supply Chain & Operations (logistics, manufacturing, partnerships)
- Expert Perspectives (analyst opinions, industry expert views)
- Outlook & Predictions (forecasts, scenarios, risks)

Choose and customize chapters based on what's most relevant and actionable for the specific subject.

{schema_json_string}

User prompt

Design report chapters for the following subject:

Subject Name: {subject_name}
Subject Type: {subject_type}
Industry: {industry}
Focus Areas: {focus_areas}
Regions: {regions}

Requirements Description:
{requirement_description}

Create chapters that will provide comprehensive and actionable intelligence for this subject. If a requirements description is provided, use it as the primary guide for chapter design — the chapters should cover what the description asks for. Include as many chapters as needed to thoroughly cover all relevant aspects - don't artificially limit the number.