Cost mode:

Category: Relevance, Classification & Matching · Rail: absolute · Typical I/O: 13579→5969 tokens

Models

Frontier on this task: GPT-5.6 Sol at 8.76 / 10. Quality bar at 90%: 7.88.

point-estimate floor (CI low) · upper CI (less certain) · Bars sorted by blended cost; best-value model first. Greyed rows are MEDIUM+ models whose point estimate clears the bar but whose CI low does not.

ModelQuality scoreCI lowCost / 1k runsvs best value
GPT-5.6 Luna7.91 / 107.67$1.74best value
GLM-5.3 Flash8.27 / 107.99$7.484.3x more expensive
DeepSeek V4 Pro8.06 / 107.60$26.3415x more expensive
Meta Muse Spark 1.38.24 / 108.06$30.6618x more expensive
Tencent Hy4 Preview8.05 / 107.65$37.7522x more expensive
GPT-5.6 Sol8.76 / 108.56$38.6622x more expensive
Grok 4.68.55 / 108.37$71.4841x more expensive
GLM-5.38.60 / 108.40$75.2143x more expensive
Claude Opus 58.34 / 107.88$91.2153x more expensive
NVIDIA Nemotron 3.5 Lightning5.11 / 104.62$3.141.8x more expensive
Gemini 3.8 Flash7.81 / 107.58$9.505.5x more expensive
Claude Haiku 4.56.33 / 105.84$15.138.7x more expensive
GPT-5.6 Terra7.88 / 107.46$19.0411x more expensive
MiniMax M37.27 / 107.06$3.512x more expensive
Tencent Hy36.73 / 106.40$2.081.2x more expensive

Cost breakdown

ModelQualityConfidenceCost / 1k runsOverpayMode
GPT-5.6 Luna OpenAI7.91 / 10 CI [7.67, 8.16]HIGH$1.74best valuebatch
GLM-5.3 Flash Z.AI8.27 / 10 CI [7.99, 8.54]HIGH$7.484.3xbatch
DeepSeek V4 Pro DeepSeek8.06 / 10 CI [7.60, 8.53]MEDIUM$26.3415xbatch
Meta Muse Spark 1.3 OpenRouter8.24 / 10 CI [8.06, 8.42]RANKED$30.6618xbatch
Tencent Hy4 Preview OpenRouter8.05 / 10 CI [7.65, 8.44]MEDIUM$37.7522xbatch
GPT-5.6 Sol best OpenAI8.76 / 10 CI [8.56, 8.95]RANKED$38.6622xbatch
Grok 4.6 xAI8.55 / 10 CI [8.37, 8.72]RANKED$71.4841xbatch
GLM-5.3 Z.AI8.60 / 10 CI [8.40, 8.80]RANKED$75.2143xbatch
Claude Opus 5 Anthropic8.34 / 10 CI [7.88, 8.79]MEDIUM$91.2153xbatch

Overpay shows how much more you pay than the best-value model that clears the quality bar (marked ★) — the best-value good-enough option. "16x" means you overpay 16× — 16× that reference for no quality benefit above the bar. Typical call shape for this task: 13579 input tokens → 5969 output tokens, EMA-tracked from production traffic. Cost is the observed, all-in $ per 1,000 task runs: each model's own measured usage on this task — output verbosity, thinking/reasoning tokens, cache reads and writes, and the spend on its billed failures — priced at current list rates and adjusted by the billing overhead we actually reconcile against provider invoices. Models that answer tersely cost what they actually cost; models that think at length pay for it. Not comparable to providers' advertised $/1M list rates — this is what running the task costs, not a per-token price.

Evaluation rubric

Judge authority, directness, subject and regional relevance, expected accessibility, source-role diversity, URL/entity accuracy, and transparency about unverified assumptions.

Prompt templates

The system + user template pair used for this task.

RESEARCH_VETTED_SITES_SELECTOR_SYSTEM + RESEARCH_VETTED_SITES_SELECTOR_USER (468 calls in window)

System prompt

You are an expert information quality analyst specializing in source evaluation.

Your task is to identify authoritative, accessible web sources for research on a specific subject.

Source Selection Criteria:
1. REGIONAL RELEVANCE: Sources must be relevant to the specified regions (e.g., UK sources for UK research, Brazil sources for Brazil research)
2. AUTHORITATIVE: Official, recognized, or highly reputable sources
3. ACCESSIBLE: NOT behind strict paywalls or login requirements
4. RELEVANT: Directly related to the subject matter
5. CURRENT: Actively maintained and updated

Regional Source Selection:
- For UK: prioritize .uk domains, UK news sites (bbc.co.uk, theguardian.com, ft.com, reuters.com/world/uk)
- For Brazil: prioritize .br domains, Brazilian news sites (globo.com, folha.uol.com.br, estadao.com.br)
- For US: prioritize .gov, US news sites (reuters.com, bloomberg.com, cnbc.com, sec.gov)
- For EU: prioritize EU regulatory sites, pan-European news sources
- For multiple regions: select sources covering ALL specified regions, not just one

Source Categories by Subject Type:
- Financial/Company: SEC filings, investor relations, financial news
- Technology: Tech news sites, product documentation, developer resources
- Healthcare: Medical journals (open access), health organizations, research databases
- Government/Policy: Official government sites, regulatory bodies
- Academic: Open access journals, institutional repositories, research databases

AVOID:
- Paywalled sites: wsj.com, ft.com, nytimes.com (except for specific free sections)
- Sites requiring mandatory login
- Sites with aggressive anti-scraping measures
- Low-quality aggregators or content farms

Output a JSON object with:
{{
  "vetted_sites": [
    {{
      "domain": "sec.gov",
      "name": "SEC Edgar",
      "category": "Government/Financial",
      "reasoning": "Official SEC filings, free access, relevant to US financial research",
      "is_paywall_free": true
    }}
  ]
}}

Select 5-15 most relevant sources.

CRITICAL: Ensure ALL selected sources are relevant to the specified regions. Do not mix UK sources with Brazil-only research, or vice versa.

User prompt

Identify authoritative, accessible web sources for researching: {subject_name}

Subject Type: {subject_type}
Subject Description: {subject_description}
Research Focus: {research_focus}

{chapter_context}

TARGET REGIONS: {regions}

CRITICAL REQUIREMENT: Select sources that are HIGHLY RELEVANT to the target regions specified above.
- If researching UK: select UK news sites, .uk domains, UK-focused international sources
- If researching Brazil: select Brazilian news sites, .br domains, Brazil-focused sources
- If researching multiple regions: select sources that cover ALL specified regions

Prioritize sources that:
1. Are REGIONALLY RELEVANT to: {regions}
2. Are free to access (no paywall)
3. Don't require login for basic content
4. Are authoritative in their domain

{additional_requirements}

REMINDER: Do not select sources irrelevant to the target regions. For example, if researching UK and Brazil, do not select US-only or China-only sources unless they have comprehensive UK/Brazil coverage.