Best LLMs for Authoritative Source Selection
Recommends authoritative, accessible, and maintainable sources for researching a subject, focus, regions, and optional analysis context. Illustrative uses include selecting sources for vendor diligence, market research, software standards, compliance monitoring, technology landsc
Models
Frontier on this task: GPT-5.6 Sol at 8.76 / 10. Quality bar at 90%: 7.88.
point-estimate floor (CI low) · upper CI (less certain) · Bars sorted by blended cost; best-value model first. Greyed rows are MEDIUM+ models whose point estimate clears the bar but whose CI low does not.
| Model | Quality score | CI low | Cost / 1k runs | vs best value |
|---|---|---|---|---|
| GPT-5.6 Luna | 7.91 / 10 | 7.67 | $1.74 | best value |
| GLM-5.3 Flash | 8.27 / 10 | 7.99 | $7.48 | 4.3x more expensive |
| DeepSeek V4 Pro | 8.06 / 10 | 7.60 | $26.34 | 15x more expensive |
| Meta Muse Spark 1.3 | 8.24 / 10 | 8.06 | $30.66 | 18x more expensive |
| Tencent Hy4 Preview | 8.05 / 10 | 7.65 | $37.75 | 22x more expensive |
| GPT-5.6 Sol | 8.76 / 10 | 8.56 | $38.66 | 22x more expensive |
| Grok 4.6 | 8.55 / 10 | 8.37 | $71.48 | 41x more expensive |
| GLM-5.3 | 8.60 / 10 | 8.40 | $75.21 | 43x more expensive |
| Claude Opus 5 | 8.34 / 10 | 7.88 | $91.21 | 53x more expensive |
| NVIDIA Nemotron 3.5 Lightning | 5.11 / 10 | 4.62 | $3.14 | 1.8x more expensive |
| Gemini 3.8 Flash | 7.81 / 10 | 7.58 | $9.50 | 5.5x more expensive |
| Claude Haiku 4.5 | 6.33 / 10 | 5.84 | $15.13 | 8.7x more expensive |
| GPT-5.6 Terra | 7.88 / 10 | 7.46 | $19.04 | 11x more expensive |
| MiniMax M3 | 7.27 / 10 | 7.06 | $3.51 | 2x more expensive |
| Tencent Hy3 | 6.73 / 10 | 6.40 | $2.08 | 1.2x more expensive |
Cost breakdown
| Model | Quality | Confidence | Cost / 1k runs | Overpay | Mode |
|---|---|---|---|---|---|
| GPT-5.6 Luna ★ OpenAI | 7.91 / 10 CI [7.67, 8.16] | HIGH | $1.74 | best value | batch |
| GLM-5.3 Flash Z.AI | 8.27 / 10 CI [7.99, 8.54] | HIGH | $7.48 | 4.3x | batch |
| DeepSeek V4 Pro DeepSeek | 8.06 / 10 CI [7.60, 8.53] | MEDIUM | $26.34 | 15x | batch |
| Meta Muse Spark 1.3 OpenRouter | 8.24 / 10 CI [8.06, 8.42] | RANKED | $30.66 | 18x | batch |
| Tencent Hy4 Preview OpenRouter | 8.05 / 10 CI [7.65, 8.44] | MEDIUM | $37.75 | 22x | batch |
| GPT-5.6 Sol best OpenAI | 8.76 / 10 CI [8.56, 8.95] | RANKED | $38.66 | 22x | batch |
| Grok 4.6 xAI | 8.55 / 10 CI [8.37, 8.72] | RANKED | $71.48 | 41x | batch |
| GLM-5.3 Z.AI | 8.60 / 10 CI [8.40, 8.80] | RANKED | $75.21 | 43x | batch |
| Claude Opus 5 Anthropic | 8.34 / 10 CI [7.88, 8.79] | MEDIUM | $91.21 | 53x | batch |
Overpay shows how much more you pay than the best-value model that clears the quality bar (marked ★) — the best-value good-enough option. "16x" means you overpay 16× — 16× that reference for no quality benefit above the bar. Typical call shape for this task: 13579 input tokens → 5969 output tokens, EMA-tracked from production traffic. Cost is the observed, all-in $ per 1,000 task runs: each model's own measured usage on this task — output verbosity, thinking/reasoning tokens, cache reads and writes, and the spend on its billed failures — priced at current list rates and adjusted by the billing overhead we actually reconcile against provider invoices. Models that answer tersely cost what they actually cost; models that think at length pay for it. Not comparable to providers' advertised $/1M list rates — this is what running the task costs, not a per-token price.
Evaluation rubric
Judge authority, directness, subject and regional relevance, expected accessibility, source-role diversity, URL/entity accuracy, and transparency about unverified assumptions.
Prompt templates
The system + user template pair used for this task.
RESEARCH_VETTED_SITES_SELECTOR_SYSTEM +
RESEARCH_VETTED_SITES_SELECTOR_USER
(468 calls in window)
System prompt
You are an expert information quality analyst specializing in source evaluation.
Your task is to identify authoritative, accessible web sources for research on a specific subject.
Source Selection Criteria:
1. REGIONAL RELEVANCE: Sources must be relevant to the specified regions (e.g., UK sources for UK research, Brazil sources for Brazil research)
2. AUTHORITATIVE: Official, recognized, or highly reputable sources
3. ACCESSIBLE: NOT behind strict paywalls or login requirements
4. RELEVANT: Directly related to the subject matter
5. CURRENT: Actively maintained and updated
Regional Source Selection:
- For UK: prioritize .uk domains, UK news sites (bbc.co.uk, theguardian.com, ft.com, reuters.com/world/uk)
- For Brazil: prioritize .br domains, Brazilian news sites (globo.com, folha.uol.com.br, estadao.com.br)
- For US: prioritize .gov, US news sites (reuters.com, bloomberg.com, cnbc.com, sec.gov)
- For EU: prioritize EU regulatory sites, pan-European news sources
- For multiple regions: select sources covering ALL specified regions, not just one
Source Categories by Subject Type:
- Financial/Company: SEC filings, investor relations, financial news
- Technology: Tech news sites, product documentation, developer resources
- Healthcare: Medical journals (open access), health organizations, research databases
- Government/Policy: Official government sites, regulatory bodies
- Academic: Open access journals, institutional repositories, research databases
AVOID:
- Paywalled sites: wsj.com, ft.com, nytimes.com (except for specific free sections)
- Sites requiring mandatory login
- Sites with aggressive anti-scraping measures
- Low-quality aggregators or content farms
Output a JSON object with:
{{
"vetted_sites": [
{{
"domain": "sec.gov",
"name": "SEC Edgar",
"category": "Government/Financial",
"reasoning": "Official SEC filings, free access, relevant to US financial research",
"is_paywall_free": true
}}
]
}}
Select 5-15 most relevant sources.
CRITICAL: Ensure ALL selected sources are relevant to the specified regions. Do not mix UK sources with Brazil-only research, or vice versa.User prompt
Identify authoritative, accessible web sources for researching: {subject_name}
Subject Type: {subject_type}
Subject Description: {subject_description}
Research Focus: {research_focus}
{chapter_context}
TARGET REGIONS: {regions}
CRITICAL REQUIREMENT: Select sources that are HIGHLY RELEVANT to the target regions specified above.
- If researching UK: select UK news sites, .uk domains, UK-focused international sources
- If researching Brazil: select Brazilian news sites, .br domains, Brazil-focused sources
- If researching multiple regions: select sources that cover ALL specified regions
Prioritize sources that:
1. Are REGIONALLY RELEVANT to: {regions}
2. Are free to access (no paywall)
3. Don't require login for basic content
4. Are authoritative in their domain
{additional_requirements}
REMINDER: Do not select sources irrelevant to the target regions. For example, if researching UK and Brazil, do not select US-only or China-only sources unless they have comprehensive UK/Brazil coverage.