Best LLMs for Factual Claim Refinement
Reviews extracted factual claims and either minimally refines them into self-contained, verifiable statements or drops them when they are non-factual, unsupported, duplicate, or not useful. Illustrative uses include refining audit observations, customer-interview findings, softwa
Models
Frontier on this task: GLM-5.3 at 8.63 / 10. Quality bar at 90%: 7.77.
point-estimate floor (CI low) · upper CI (less certain) · Bars sorted by blended cost; best-value model first. Greyed rows are MEDIUM+ models whose point estimate clears the bar but whose CI low does not.
| Model | Quality score | CI low | Cost / 1k runs | vs best value |
|---|---|---|---|---|
| Gemini 3.1 Flash Lite | 7.88 / 10 | 7.75 | $0.53 | best value |
| GLM-5.3 Flash | 8.35 / 10 | 7.95 | $4.87 | 9.2x more expensive |
| Qwen 3.7 Plus | 8.32 / 10 | 8.11 | $4.96 | 9.3x more expensive |
| Tencent Hy3 | 8.27 / 10 | 7.99 | $5.19 | 9.8x more expensive |
| Thinking Machines Inkling Small | 8.33 / 10 | 8.00 | $7.86 | 15x more expensive |
| Gemini 3.5 Flash | 8.31 / 10 | 8.05 | $7.87 | 15x more expensive |
| Claude Sonnet 5 | 8.35 / 10 | 8.11 | $8.66 | 16x more expensive |
| NVIDIA Nemotron-3 Ultra 550B | 8.08 / 10 | 7.64 | $9.69 | 18x more expensive |
| Thinking Machines Inkling | 7.99 / 10 | 7.53 | $18.38 | 35x more expensive |
| Meta Muse Spark 1.3 | 8.44 / 10 | 8.07 | $19.94 | 38x more expensive |
| Claude Opus 5 | 8.06 / 10 | 7.61 | $26.17 | 49x more expensive |
| GLM-5.3 | 8.63 / 10 | 8.21 | $67.59 | 127x more expensive |
| DeepSeek V4 Flash | 7.71 / 10 | 7.50 | $4.86 | 9.1x more expensive |
| GPT-5.6 Luna | 7.71 / 10 | 7.46 | $0.99 | 1.9x more expensive |
| DeepSeek V4 Pro | 7.22 / 10 | 6.88 | $9.82 | 18x more expensive |
| Gemini 3.5 Flash Lite | 7.51 / 10 | 7.16 | $1.21 | 2.3x more expensive |
| GPT-5.4 Nano | 7.22 / 10 | 6.92 | $0.42 | 20% cheaper |
| NVIDIA Nemotron-3 Nano 30B-A3B | 7.71 / 10 | 7.45 | $1.47 | 2.8x more expensive |
| NVIDIA Nemotron-3 Super 120B | 7.29 / 10 | 6.84 | $3.99 | 7.5x more expensive |
| MiniMax M3 | 7.48 / 10 | 7.16 | $5.91 | 11x more expensive |
| Claude Haiku 4.5 | 6.95 / 10 | 6.64 | $2.53 | 4.8x more expensive |
Cost breakdown
| Model | Quality | Confidence | Cost / 1k runs | Overpay | Mode |
|---|---|---|---|---|---|
| Gemini 3.1 Flash Lite ★ Gemini | 7.88 / 10 CI [7.75, 8.02] | RANKED | $0.53 | best value | batch |
| GLM-5.3 Flash Z.AI | 8.35 / 10 CI [7.95, 8.74] | MEDIUM | $4.87 | 9.2x | batch |
| Qwen 3.7 Plus Alibaba Cloud (DashScope) | 8.32 / 10 CI [8.11, 8.53] | HIGH | $4.96 | 9.3x | batch |
| Tencent Hy3 OpenRouter | 8.27 / 10 CI [7.99, 8.56] | HIGH | $5.19 | 9.8x | batch |
| Thinking Machines Inkling Small OpenRouter | 8.33 / 10 CI [8.00, 8.66] | MEDIUM | $7.86 | 15x | batch |
| Gemini 3.5 Flash Gemini | 8.31 / 10 CI [8.05, 8.57] | HIGH | $7.87 | 15x | batch |
| Claude Sonnet 5 Anthropic | 8.35 / 10 CI [8.11, 8.60] | HIGH | $8.66 | 16x | batch |
| NVIDIA Nemotron-3 Ultra 550B OpenRouter | 8.08 / 10 CI [7.64, 8.52] | MEDIUM | $9.69 | 18x | batch |
| Thinking Machines Inkling OpenRouter | 7.99 / 10 CI [7.53, 8.46] | MEDIUM | $18.38 | 35x | batch |
| Meta Muse Spark 1.3 OpenRouter | 8.44 / 10 CI [8.07, 8.82] | MEDIUM | $19.94 | 38x | batch |
| Claude Opus 5 Anthropic | 8.06 / 10 CI [7.61, 8.52] | MEDIUM | $26.17 | 49x | batch |
| GLM-5.3 best Z.AI | 8.63 / 10 CI [8.21, 9.05] | MEDIUM | $67.59 | 127x | batch |
Overpay shows how much more you pay than the best-value model that clears the quality bar (marked ★) — the best-value good-enough option. "16x" means you overpay 16× — 16× that reference for no quality benefit above the bar. Typical call shape for this task: 2593 input tokens → 236 output tokens, EMA-tracked from production traffic. Cost is the observed, all-in $ per 1,000 task runs: each model's own measured usage on this task — output verbosity, thinking/reasoning tokens, cache reads and writes, and the spend on its billed failures — priced at current list rates and adjusted by the billing overhead we actually reconcile against provider invoices. Models that answer tersely cost what they actually cost; models that think at length pay for it. Not comparable to providers' advertised $/1M list rates — this is what running the task costs, not a per-token price.
Evaluation rubric
Judge semantic preservation, improved self-containment, minimality of edits, correct drop decisions, source support, qualifier preservation, and reason accuracy. Penalize stylistic rewrites that alter meaning.
Prompt templates
This is a pooled capability — 3 prompt families share it. The pair shown first is the most frequently used in production.
LLMB_FACTUAL_CLAIM_REFINEMENT_SYSTEM +
LLMB_FACTUAL_CLAIM_REFINEMENT_USER
(57773 calls in window)
System prompt
Preserve the original meaning and evidence boundary. Refine only to add an unambiguous subject, expand a source-defined abbreviation, restore a necessary qualifier, or resolve a local reference. Never add a fact from general knowledge. Drop claims that cannot be made atomic and self-contained without inference, and provide the schema-defined reason. Treat empty optional values as absent and return only the requested result. Your response must conform exactly to this output schema: {schema_json_string}.
User prompt
Inputs — summary_text: {summary_text}; chapter_descriptions: {chapter_descriptions}; claims_list: {claims_list}. Use only these inputs to complete the task defined by the system prompt.
CLAIM_REFINEMENT_SYSTEM_PROMPT +
CLAIM_REFINEMENT_USER_PROMPT
(171 calls in window)
System prompt
You are a quality reviewer for extracted factual claims. You receive a list of claims that were extracted from a research summary, along with the original summary text.
Your job is to review each claim and either REFINE it or DROP it.
## REFINE a claim when:
- It is a valid factual claim but lacks context — add the missing subject, entity name, or qualifier from the original summary so the claim is self-contained
- It uses abbreviations or ticker symbols without full names — expand them (e.g. "MU" → "Micron Technology (MU)")
- It references "the post", "the author", "the article" — rewrite as a standalone fact
- It has ambiguous references ("the fund", "this strategy", "the product") — resolve from the summary
## DROP a claim when:
- It reports absence of information ("no ESG data is provided", "no macro drivers are stated")
- It describes methodology, tools, platforms, simulation setup, or backtest configuration rather than the financial subject
- It is promotional or marketing content (trials, subscriptions, community invitations)
- It is a mere mention, association, or classification without substantive content
- It is a meta-description of the source (its topic, audience, format, tone, or purpose)
- It is a user comment, opinion, or community reaction rather than a factual data point
- It cannot be made self-contained because the summary lacks sufficient context — a vague claim has no value
- It is a vague procedural statement with no concrete information
## Output rules:
- Return ONLY the refined claims that pass quality review
- Each claim must be understandable on its own without the original summary
- Each claim must convey actionable information about the financial subject being analyzed
- Preserve the original category and source_reference — only modify claim_text
- Preserve numerical precision exactly as stated
## Required Output Format
Your response MUST be a single, valid JSON object conforming to this schema:
```json
{schema_json_string}
```User prompt
## Original Summary
{summary_text}
## Report Chapter Context
{chapter_descriptions}
## Extracted Claims to Review
{claims_list}
## Instructions
Review each claim above against the original summary. For each claim:
1. If it is a valid factual claim about the financial subject — refine it to be self-contained and clear, then include it
2. If it is NOT a valid claim (methodology, meta-description, vague, promotional, mere mention, missing context that cannot be resolved) — drop it
Return only the claims that pass review, with refined claim_text where needed. Keep the original category and source_reference.
## Output Format
Return ONLY the fields defined in the schema below. Do not add extra fields.
The required JSON output schema is provided in the system prompt.
JSON_REPAIR_SYSTEM +
JSON_REPAIR_USER
(3 calls in window)
System prompt
You are a JSON repair tool. The user gives you malformed or partial model output and a JSON Schema. Return ONLY a single valid JSON object that satisfies the schema, salvaging as much real content from the input as possible. Do not invent data for fields the input doesn't support — use the schema's allowed empty/null values. Output the JSON object only: no prose, no markdown, no code fences.
User prompt
JSON Schema:
{schema_json}
Malformed output to repair:
{raw_text}
Return only the corrected JSON object.