Cost mode:

Category: Relevance, Classification & Matching · Rail: absolute · Typical I/O: 1177→1914 tokens

Models

Frontier on this task: GLM-5.3 Flash at 9.48 / 10. Quality bar at 90%: 8.53.

point-estimate floor (CI low) · upper CI (less certain) · Bars sorted by blended cost; best-value model first. Greyed rows are MEDIUM+ models whose point estimate clears the bar but whose CI low does not.

ModelQuality scoreCI lowCost / 1k runsvs best value
GLM-5.3 Flash9.48 / 109.00$2.37best value
Claude Sonnet 58.61 / 108.29$6.442.7x more expensive
Tencent Hy4 Preview9.11 / 108.61$25.4911x more expensive
Claude Haiku 4.57.14 / 106.69$1.7925% cheaper
Tencent Hy37.95 / 107.52$0.3386% cheaper
GPT-5.6 Terra8.31 / 107.97$3.171.3x more expensive
GPT-5.6 Luna7.97 / 107.56$0.3187% cheaper
MiniMax M37.34 / 106.90$0.9759% cheaper
GPT-5.6 Sol8.46 / 108.07$5.432.3x more expensive

Cost breakdown

ModelQualityConfidenceCost / 1k runsOverpayMode
GLM-5.3 Flash best Z.AI9.48 / 10 CI [9.00, 9.96]MEDIUM$2.37best valuebatch
Claude Sonnet 5 Anthropic8.61 / 10 CI [8.29, 8.93]MEDIUM$6.442.7xbatch
Tencent Hy4 Preview OpenRouter9.11 / 10 CI [8.61, 9.61]MEDIUM$25.4911xbatch

Overpay shows how much more you pay than the best-value model that clears the quality bar (marked ★) — the best-value good-enough option. "16x" means you overpay 16× — 16× that reference for no quality benefit above the bar. Typical call shape for this task: 1177 input tokens → 1914 output tokens, EMA-tracked from production traffic. Cost is the observed, all-in $ per 1,000 task runs: each model's own measured usage on this task — output verbosity, thinking/reasoning tokens, cache reads and writes, and the spend on its billed failures — priced at current list rates and adjusted by the billing overhead we actually reconcile against provider invoices. Models that answer tersely cost what they actually cost; models that think at length pay for it. Not comparable to providers' advertised $/1M list rates — this is what running the task costs, not a per-token price.

Evaluation rubric

Judge readiness-verdict correctness, evidence and instruction fidelity, response relevance, faithful application of channel and review policies, manipulation, privacy, conflict, and promotional-risk detection, severity and uncertainty calibration, and remediation usefulness. Penalize invented evidence or requirements, silent rewriting, and decisions that depend on external state. Schema validity is deterministic.