Cost mode:

Category: Content Summarization & Synthesis · Rail: absolute · Typical I/O: 0→0 tokens

Models

Frontier on this task: DeepSeek V4 Flash at 8.85 / 10. Quality bar at 90%: 7.96.

point-estimate floor (CI low) · upper CI (less certain) · Bars sorted by blended cost; best-value model first. Greyed rows are MEDIUM+ models whose point estimate clears the bar but whose CI low does not.

ModelQuality scoreCI lowCost / 1k runsvs best value
GPT-5.6 Luna8.03 / 107.62$2.65best value
Gemini 3.5 Flash Lite8.46 / 108.11$5.131.9x more expensive
NVIDIA Nemotron-3 Super 120B8.04 / 107.76$7.082.7x more expensive
DeepSeek V4 Flash8.85 / 108.67$9.773.7x more expensive
GPT-5.4 Nano7.03 / 106.58$2.3810% cheaper

Cost breakdown

ModelQualityConfidenceCost / 1k runsOverpayMode
GPT-5.6 Luna OpenAI8.03 / 10 CI [7.62, 8.43]MEDIUM$2.65best valuebatch
Gemini 3.5 Flash Lite Gemini8.46 / 10 CI [8.11, 8.80]MEDIUM$5.131.9xbatch
NVIDIA Nemotron-3 Super 120B OpenRouter8.04 / 10 CI [7.76, 8.33]HIGH$7.082.7xbatch
DeepSeek V4 Flash best DeepSeek8.85 / 10 CI [8.67, 9.02]RANKED$9.773.7xbatch

Overpay shows how much more you pay than the best-value model that clears the quality bar (marked ★) — the best-value good-enough option. "16x" means you overpay 16× — 16× that reference for no quality benefit above the bar. Typical call shape for this task: 0 input tokens → 0 output tokens, EMA-tracked from production traffic. Cost is the observed, all-in $ per 1,000 task runs: each model's own measured usage on this task — output verbosity, thinking/reasoning tokens, cache reads and writes, and the spend on its billed failures — priced at current list rates and adjusted by the billing overhead we actually reconcile against provider invoices. Models that answer tersely cost what they actually cost; models that think at length pay for it. Not comparable to providers' advertised $/1M list rates — this is what running the task costs, not a per-token price.

Evaluation rubric

Judge entailment, atomicity, verifiability, specificity, preservation of qualifiers, source coverage, and non-duplication. Penalize inference beyond the source, compound claims, and claims with no precise evidence anchor.