Best LLMs for Evidence Grounded Claim Generation
Distils supplied source material or facts into a caller-configured number of atomic, quotable, evidence-verifiable claims. Illustrative uses include deriving verifiable claims from contracts, customer interviews, technical test results, market research, investigative reporting, s
Models
Frontier on this task: DeepSeek V4 Flash at 8.85 / 10. Quality bar at 90%: 7.96.
point-estimate floor (CI low) · upper CI (less certain) · Bars sorted by blended cost; best-value model first. Greyed rows are MEDIUM+ models whose point estimate clears the bar but whose CI low does not.
| Model | Quality score | CI low | Cost / 1k runs | vs best value |
|---|---|---|---|---|
| GPT-5.6 Luna | 8.03 / 10 | 7.62 | $2.65 | best value |
| Gemini 3.5 Flash Lite | 8.46 / 10 | 8.11 | $5.13 | 1.9x more expensive |
| NVIDIA Nemotron-3 Super 120B | 8.04 / 10 | 7.76 | $7.08 | 2.7x more expensive |
| DeepSeek V4 Flash | 8.85 / 10 | 8.67 | $9.77 | 3.7x more expensive |
| GPT-5.4 Nano | 7.03 / 10 | 6.58 | $2.38 | 10% cheaper |
Cost breakdown
| Model | Quality | Confidence | Cost / 1k runs | Overpay | Mode |
|---|---|---|---|---|---|
| GPT-5.6 Luna ★ OpenAI | 8.03 / 10 CI [7.62, 8.43] | MEDIUM | $2.65 | best value | batch |
| Gemini 3.5 Flash Lite Gemini | 8.46 / 10 CI [8.11, 8.80] | MEDIUM | $5.13 | 1.9x | batch |
| NVIDIA Nemotron-3 Super 120B OpenRouter | 8.04 / 10 CI [7.76, 8.33] | HIGH | $7.08 | 2.7x | batch |
| DeepSeek V4 Flash best DeepSeek | 8.85 / 10 CI [8.67, 9.02] | RANKED | $9.77 | 3.7x | batch |
Overpay shows how much more you pay than the best-value model that clears the quality bar (marked ★) — the best-value good-enough option. "16x" means you overpay 16× — 16× that reference for no quality benefit above the bar. Typical call shape for this task: 0 input tokens → 0 output tokens, EMA-tracked from production traffic. Cost is the observed, all-in $ per 1,000 task runs: each model's own measured usage on this task — output verbosity, thinking/reasoning tokens, cache reads and writes, and the spend on its billed failures — priced at current list rates and adjusted by the billing overhead we actually reconcile against provider invoices. Models that answer tersely cost what they actually cost; models that think at length pay for it. Not comparable to providers' advertised $/1M list rates — this is what running the task costs, not a per-token price.
Evaluation rubric
Judge entailment, atomicity, verifiability, specificity, preservation of qualifiers, source coverage, and non-duplication. Penalize inference beyond the source, compound claims, and claims with no precise evidence anchor.