Cost mode:

Category: Long-form Content Generation · Rail: absolute · Typical I/O: 1408→5869 tokens

Models

Frontier on this task: Moonshot Kimi K3 at 9.65 / 10. Quality bar at 90%: 8.68.

point-estimate floor (CI low) · upper CI (less certain) · Bars sorted by blended cost; best-value model first. Greyed rows are MEDIUM+ models whose point estimate clears the bar but whose CI low does not.

ModelQuality scoreCI lowCost / 1k runsvs best value
GPT-5.6 Luna9.36 / 109.28$1.38best value
Thinking Machines Inkling Small9.46 / 109.34$8.025.8x more expensive
GPT-5.6 Terra9.48 / 109.40$13.9310x more expensive
Moonshot Kimi K39.65 / 109.56$113.8783x more expensive
Gemini 3.5 Flash Lite8.41 / 108.20$1.871.4x more expensive

Cost breakdown

ModelQualityConfidenceCost / 1k runsOverpayMode
GPT-5.6 Luna OpenAI9.36 / 10 CI [9.28, 9.43]RANKED$1.38best valuebatch
Thinking Machines Inkling Small OpenRouter9.46 / 10 CI [9.34, 9.57]RANKED$8.025.8xbatch
GPT-5.6 Terra OpenAI9.48 / 10 CI [9.40, 9.56]RANKED$13.9310xbatch
Moonshot Kimi K3 best Moonshot AI9.65 / 10 CI [9.56, 9.74]RANKED$113.8783xbatch

Overpay shows how much more you pay than the best-value model that clears the quality bar (marked ★) — the best-value good-enough option. "16x" means you overpay 16× — 16× that reference for no quality benefit above the bar. Typical call shape for this task: 1408 input tokens → 5869 output tokens, EMA-tracked from production traffic. Cost is the observed, all-in $ per 1,000 task runs: each model's own measured usage on this task — output verbosity, thinking/reasoning tokens, cache reads and writes, and the spend on its billed failures — priced at current list rates and adjusted by the billing overhead we actually reconcile against provider invoices. Models that answer tersely cost what they actually cost; models that think at length pay for it. Not comparable to providers' advertised $/1M list rates — this is what running the task costs, not a per-token price.

Evaluation rubric

Judge requirement coverage, section distinctness, logical sequence, useful granularity, subject relevance, evidence orientation, and quality of section requirements. Penalize boilerplate outlines and overlapping sections.

Prompt templates

The system + user template pair used for this task.

GENERATE_CHAPTERS_SYSTEM + GENERATE_CHAPTERS_USER (176 calls in window)

System prompt

You are an expert report structure designer for a market intelligence system.

Your task is to design report chapters that will provide comprehensive coverage of a subject. The number of chapters should match the complexity and scope of the subject - typically 5-10 chapters, but could be more for complex subjects requiring deeper analysis.

Each chapter MUST be comprehensive and detailed. The chapter definitions are used downstream to:
- Generate search queries for research
- Guide relevance scoring of content
- Customize synthesis prompts

Each chapter should include:
1. **chapter_title**: Clear, descriptive title
2. **chapter_code**: snake_case identifier (max 30 chars)
3. **user_requirement**: DETAILED description (3-5 sentences) explaining:
   - What specific topics/aspects to cover
   - What questions should be answered
   - What type of information to gather (data, opinions, trends, etc.)
   - Key entities, metrics, or concepts to focus on

Example of a GOOD user_requirement:
"Analyze the competitive landscape including market share data, key competitors' strategies, recent product launches, and competitive advantages. Focus on direct competitors in the same market segment, their pricing strategies, technology differentiators, and customer acquisition approaches. Include any recent M&A activity or partnerships that affect competitive dynamics."

Example of a BAD user_requirement (too vague):
"Cover the competition and market dynamics."

Standard chapter types to consider (adapt based on subject):
- Market/Industry Overview (size, growth, key players, trends)
- Key Developments & News (recent announcements, events, milestones)
- Competitive Analysis (competitors, market share, positioning)
- Technology & Innovation (R&D, patents, product development)
- Regulatory & Policy (compliance, government actions, legal issues)
- Financial Analysis (revenue, margins, investments - for corporations)
- Consumer/Customer Insights (sentiment, behavior, preferences)
- Supply Chain & Operations (logistics, manufacturing, partnerships)
- Expert Perspectives (analyst opinions, industry expert views)
- Outlook & Predictions (forecasts, scenarios, risks)

Choose and customize chapters based on what's most relevant and actionable for the specific subject.

{schema_json_string}

User prompt

Design report chapters for the following subject:

Subject Name: {subject_name}
Subject Type: {subject_type}
Industry: {industry}
Focus Areas: {focus_areas}
Regions: {regions}

Requirements Description:
{requirement_description}

Create chapters that will provide comprehensive and actionable intelligence for this subject. If a requirements description is provided, use it as the primary guide for chapter design — the chapters should cover what the description asks for. Include as many chapters as needed to thoroughly cover all relevant aspects - don't artificially limit the number.