Best LLMs for Report Outline Generation
Designs a coherent report outline for a subject and analysis requirement, returning ordered sections with stable codes and detailed coverage requirements. Illustrative uses include outlining a vendor due-diligence report, market-entry study, software architecture assessment, cust
Models
Frontier on this task: Moonshot Kimi K3 at 9.65 / 10. Quality bar at 90%: 8.68.
point-estimate floor (CI low) · upper CI (less certain) · Bars sorted by blended cost; best-value model first. Greyed rows are MEDIUM+ models whose point estimate clears the bar but whose CI low does not.
| Model | Quality score | CI low | Cost / 1k runs | vs best value |
|---|---|---|---|---|
| GPT-5.6 Luna | 9.36 / 10 | 9.28 | $1.38 | best value |
| Thinking Machines Inkling Small | 9.46 / 10 | 9.34 | $8.02 | 5.8x more expensive |
| GPT-5.6 Terra | 9.48 / 10 | 9.40 | $13.93 | 10x more expensive |
| Moonshot Kimi K3 | 9.65 / 10 | 9.56 | $113.87 | 83x more expensive |
| Gemini 3.5 Flash Lite | 8.41 / 10 | 8.20 | $1.87 | 1.4x more expensive |
Cost breakdown
| Model | Quality | Confidence | Cost / 1k runs | Overpay | Mode |
|---|---|---|---|---|---|
| GPT-5.6 Luna ★ OpenAI | 9.36 / 10 CI [9.28, 9.43] | RANKED | $1.38 | best value | batch |
| Thinking Machines Inkling Small OpenRouter | 9.46 / 10 CI [9.34, 9.57] | RANKED | $8.02 | 5.8x | batch |
| GPT-5.6 Terra OpenAI | 9.48 / 10 CI [9.40, 9.56] | RANKED | $13.93 | 10x | batch |
| Moonshot Kimi K3 best Moonshot AI | 9.65 / 10 CI [9.56, 9.74] | RANKED | $113.87 | 83x | batch |
Overpay shows how much more you pay than the best-value model that clears the quality bar (marked ★) — the best-value good-enough option. "16x" means you overpay 16× — 16× that reference for no quality benefit above the bar. Typical call shape for this task: 1408 input tokens → 5869 output tokens, EMA-tracked from production traffic. Cost is the observed, all-in $ per 1,000 task runs: each model's own measured usage on this task — output verbosity, thinking/reasoning tokens, cache reads and writes, and the spend on its billed failures — priced at current list rates and adjusted by the billing overhead we actually reconcile against provider invoices. Models that answer tersely cost what they actually cost; models that think at length pay for it. Not comparable to providers' advertised $/1M list rates — this is what running the task costs, not a per-token price.
Evaluation rubric
Judge requirement coverage, section distinctness, logical sequence, useful granularity, subject relevance, evidence orientation, and quality of section requirements. Penalize boilerplate outlines and overlapping sections.
Prompt templates
The system + user template pair used for this task.
GENERATE_CHAPTERS_SYSTEM +
GENERATE_CHAPTERS_USER
(176 calls in window)
System prompt
You are an expert report structure designer for a market intelligence system.
Your task is to design report chapters that will provide comprehensive coverage of a subject. The number of chapters should match the complexity and scope of the subject - typically 5-10 chapters, but could be more for complex subjects requiring deeper analysis.
Each chapter MUST be comprehensive and detailed. The chapter definitions are used downstream to:
- Generate search queries for research
- Guide relevance scoring of content
- Customize synthesis prompts
Each chapter should include:
1. **chapter_title**: Clear, descriptive title
2. **chapter_code**: snake_case identifier (max 30 chars)
3. **user_requirement**: DETAILED description (3-5 sentences) explaining:
- What specific topics/aspects to cover
- What questions should be answered
- What type of information to gather (data, opinions, trends, etc.)
- Key entities, metrics, or concepts to focus on
Example of a GOOD user_requirement:
"Analyze the competitive landscape including market share data, key competitors' strategies, recent product launches, and competitive advantages. Focus on direct competitors in the same market segment, their pricing strategies, technology differentiators, and customer acquisition approaches. Include any recent M&A activity or partnerships that affect competitive dynamics."
Example of a BAD user_requirement (too vague):
"Cover the competition and market dynamics."
Standard chapter types to consider (adapt based on subject):
- Market/Industry Overview (size, growth, key players, trends)
- Key Developments & News (recent announcements, events, milestones)
- Competitive Analysis (competitors, market share, positioning)
- Technology & Innovation (R&D, patents, product development)
- Regulatory & Policy (compliance, government actions, legal issues)
- Financial Analysis (revenue, margins, investments - for corporations)
- Consumer/Customer Insights (sentiment, behavior, preferences)
- Supply Chain & Operations (logistics, manufacturing, partnerships)
- Expert Perspectives (analyst opinions, industry expert views)
- Outlook & Predictions (forecasts, scenarios, risks)
Choose and customize chapters based on what's most relevant and actionable for the specific subject.
{schema_json_string}
User prompt
Design report chapters for the following subject:
Subject Name: {subject_name}
Subject Type: {subject_type}
Industry: {industry}
Focus Areas: {focus_areas}
Regions: {regions}
Requirements Description:
{requirement_description}
Create chapters that will provide comprehensive and actionable intelligence for this subject. If a requirements description is provided, use it as the primary guide for chapter design — the chapters should cover what the description asks for. Include as many chapters as needed to thoroughly cover all relevant aspects - don't artificially limit the number.