Cost mode:

Category: Infrastructure & Utility · Rail: absolute · Typical I/O: 3313→4456 tokens

Models

Frontier on this task: GPT-5.5 at 9.05 / 10. Quality bar at 90%: 8.15.

point-estimate floor (CI low) · upper CI (less certain) · Bars sorted by blended cost; best-value model first. Greyed rows are MEDIUM+ models whose point estimate clears the bar but whose CI low does not.

ModelQuality scoreCI lowCost / 1k runsvs best value
DeepSeek V4 Flash8.27 / 108.00$3.62best value
GPT-5.4 Nano8.17 / 107.96$8.382.3x more expensive
DeepSeek V4 Pro8.34 / 108.02$19.825.5x more expensive
Qwen 3.6 Plus8.66 / 108.48$43.2312x more expensive
Kimi K2.68.92 / 108.78$73.1720x more expensive
Claude Sonnet 4.68.66 / 108.45$147.2741x more expensive
GPT-5.59.05 / 108.90$324.7990x more expensive
Gemini 3.1 Flash Lite6.23 / 105.94$3.2510% cheaper
Claude Haiku 4.57.97 / 107.73$48.0613x more expensive
Qwen 3.5 Flash8.09 / 107.75$5.321.5x more expensive
GPT-5.4 Mini7.88 / 107.65$20.245.6x more expensive
Gemini 3.1 Pro Preview7.89 / 107.67$30.908.5x more expensive
Qwen 3.7 Plus7.36 / 107.13$13.913.8x more expensive
Qwen 3.6 Flash7.21 / 106.97$14.133.9x more expensive
Grok 4.57.39 / 107.22$41.2811x more expensive
Gemini 3.5 Flash4.77 / 104.49$26.437.3x more expensive

Cost breakdown

ModelQualityConfidenceCost / 1k runsOverpayMode
DeepSeek V4 Flash DeepSeek8.27 / 10 CI [8.00, 8.54]HIGH$3.62best valuebatch
GPT-5.4 Nano OpenAI8.17 / 10 CI [7.96, 8.38]HIGH$8.382.3xbatch
DeepSeek V4 Pro DeepSeek8.34 / 10 CI [8.02, 8.66]MEDIUM$19.825.5xbatch
Qwen 3.6 Plus Alibaba Cloud (DashScope)8.66 / 10 CI [8.48, 8.83]RANKED$43.2312xbatch
Kimi K2.6 Moonshot AI8.92 / 10 CI [8.78, 9.05]RANKED$73.1720xbatch
Claude Sonnet 4.6 Anthropic8.66 / 10 CI [8.45, 8.86]HIGH$147.2741xbatch
GPT-5.5 best OpenAI9.05 / 10 CI [8.90, 9.20]RANKED$324.7990xbatch

Overpay shows how much more you pay than the best-value model that clears the quality bar (marked ★) — the best-value good-enough option. "16x" means you overpay 16× — 16× that reference for no quality benefit above the bar. Typical call shape for this task: 3313 input tokens → 4456 output tokens, EMA-tracked from production traffic. Cost is the observed, all-in $ per 1,000 task runs: each model's own measured usage on this task — output verbosity, thinking/reasoning tokens, cache reads and writes, and the spend on its billed failures — priced at current list rates and adjusted by the billing overhead we actually reconcile against provider invoices. Models that answer tersely cost what they actually cost; models that think at length pay for it. Not comparable to providers' advertised $/1M list rates — this is what running the task costs, not a per-token price.

Prompt templates

This is a pooled capability — 2 prompt families share it. The pair shown first is the most frequently used in production.

CHAPTER_PROMPT_GENERATOR_SYSTEM_TEMPLATE + CHAPTER_PROMPT_GENERATOR_USER_TEMPLATE (1004 calls in window)

System prompt

You are an expert at customizing financial analysis prompts for specific companies and subjects.

Your task is to take a TEMPLATE chapter prompt (system + user) and customize it for a specific subject by:

1. **Replacing generic terms** with subject-specific terminology
   - "the company" → "{subject_name}"
   - "key metrics" → specific KPIs relevant to this subject
   - "competitive landscape" → actual competitors by name

2. **Adding subject-specific focus areas**
   - Example Tesla: EV technology, battery supply chain, autonomous driving, energy products
   - Example Oklo: Nuclear regulations, SMR technology, data center partnerships
   - Example crypto companies: Blockchain infrastructure, regulatory environment, token economics

3. **Including relevant context**
   - Industry-specific metrics and benchmarks
   - Key competitors, partners, suppliers
   - Regulatory considerations
   - Technology trends

4. **Maintaining structure and intent**
   - Keep the same analysis depth and requirements
   - Preserve required output format
   - Don't change variable placeholders like datetime_from or relevance_threshold

**Output Format:**
Return customized system_prompt and user_prompt as separate fields.
Include reasoning explaining key customizations made.

User prompt

Customize the following chapter prompt for a specific subject:

**Subject Information:**
- Name: {subject_name}
- Code: {subject_code}
- Description: {subject_description}
- Industry: {industry}
- Key Focus Areas: {focus_areas}

**Chapter Information:**
- Chapter Code: {chapter_code}
- Chapter Name: {chapter_name}
- User Requirement: {user_requirement}

**Template System Prompt:**
{template_system_prompt}

**Template User Prompt:**
{template_user_prompt}

**Instructions:**
Generate customized system_prompt and user_prompt that are specifically tailored for {subject_name}.

Focus on:
1. What makes {subject_name} unique in its industry
2. What specific analysis would be most valuable for investors analyzing {subject_name}
3. What risks, opportunities, or trends are particularly relevant to {subject_name}
4. What competitive dynamics or market forces {subject_name} faces

Output the customized prompts as a valid JSON object matching the schema.
CHAPTER_PROMPT_GEN_SYSTEM + CHAPTER_PROMPT_GEN_USER (441 calls in window)

System prompt

You are an expert prompt engineer specializing in creating synthesis prompts for intelligence reports.

Your task is to generate system and user prompts that will guide an LLM to synthesize source material into a specific chapter of an intelligence report.

The prompts you create must:
1. Be specific to the chapter's topic and requirements
2. Guide the LLM to extract relevant insights from source material
3. Produce structured, professional output suitable for business intelligence
4. Include clear instructions on what to analyze and how to present findings

For the system prompt:
- Define the analyst's role and expertise relevant to the chapter topic
- Specify the analysis approach and methodology
- Set quality standards and output expectations
- Keep it concise but comprehensive (150-250 words)

For the user prompt:
- Include a {content} placeholder where source material will be inserted
- Reference the chapter's specific requirements
- Provide clear instructions on what aspects to analyze
- Request structured output with key findings and takeaways
- Keep it actionable and specific (100-200 words)

Output your prompts in the required JSON schema format.

## Required Output Format
Your response MUST be a single, valid JSON object conforming to this schema:
```json
{schema_json_string}
```

User prompt

Generate synthesis prompts for the following report chapter:

Chapter Title: {chapter_title}
Chapter Code: {chapter_code}

Chapter Requirements:
{user_requirement}

Create prompts that will guide an LLM to analyze source material and produce the "{chapter_title}" section of an intelligence report. The prompts should be tailored to the specific requirements above.

Remember:
- The system prompt defines the analyst's role and approach
- The user prompt must include {content} placeholder for source material
- Include relevant keywords that indicate content is relevant to this chapter

The required JSON output schema is provided in the system prompt.