Cost mode:

Category: Relevance, Classification & Matching · Rail: absolute · Typical I/O: 384→83 tokens

Models

Frontier on this task: Claude Sonnet 5 at 10.06 / 10. Quality bar at 90%: 9.06.

point-estimate floor (CI low) · upper CI (less certain) · Bars sorted by blended cost; best-value model first.

ModelQuality scoreCI lowCost / 1k runsvs best value
NVIDIA Nemotron-3 Nano 30B-A3B9.59 / 109.16$0.04best value
Gemini 3.1 Flash Lite9.95 / 109.91$0.051.4x more expensive
DeepSeek V4 Flash9.89 / 109.79$0.051.4x more expensive
GPT-5.4 Nano9.92 / 109.88$0.061.5x more expensive
GPT-5.4 Mini9.91 / 109.85$0.102.7x more expensive
NVIDIA Nemotron-3 Super 120B9.86 / 109.62$0.123.1x more expensive
Tencent Hy39.90 / 109.76$0.143.8x more expensive
MiniMax M39.95 / 109.87$0.225.8x more expensive
DeepSeek V4 Pro9.90 / 109.84$0.236.1x more expensive
GPT-5.6 Luna9.95 / 109.82$0.277.2x more expensive
NVIDIA Nemotron-3 Ultra 550B9.98 / 109.88$0.5314x more expensive
Claude Haiku 4.59.94 / 109.90$0.6016x more expensive
Qwen 3.5 Flash9.99 / 109.86$0.6317x more expensive
GPT-5.6 Terra9.99 / 109.92$0.6717x more expensive
Gemini 3.5 Flash9.98 / 109.90$1.3235x more expensive
Claude Sonnet 510.06 / 1010.04$1.3335x more expensive
GPT-5.59.95 / 109.94$1.4337x more expensive
GPT-5.6 Sol10.00 / 109.97$1.4438x more expensive
Qwen 3.7 Plus9.94 / 109.79$1.4739x more expensive
Gemini 3.1 Pro Preview9.93 / 109.87$1.8248x more expensive
Qwen 3.6 Plus9.91 / 109.72$1.9350x more expensive
Claude Sonnet 4.69.96 / 109.95$1.9752x more expensive
Qwen 3.6 Flash9.99 / 109.88$2.5266x more expensive
Meta Muse Spark 1.19.84 / 109.63$2.8976x more expensive
Claude Opus 4.89.98 / 109.84$3.2986x more expensive
Grok 4.59.92 / 109.75$3.4189x more expensive
Kimi K2.69.91 / 109.85$5.47143x more expensive

Cost breakdown

ModelQualityConfidenceCost / 1k runsOverpayMode
NVIDIA Nemotron-3 Nano 30B-A3B OpenRouter9.59 / 10 CI [9.16, 10.00]MEDIUM$0.04best valuebatch
Gemini 3.1 Flash Lite Gemini9.95 / 10 CI [9.91, 9.98]RANKED$0.051.4xbatch
DeepSeek V4 Flash DeepSeek9.89 / 10 CI [9.79, 10.00]RANKED$0.051.4xbatch
GPT-5.4 Nano OpenAI9.92 / 10 CI [9.88, 9.97]RANKED$0.061.5xbatch
GPT-5.4 Mini OpenAI9.91 / 10 CI [9.85, 9.98]RANKED$0.102.7xbatch
NVIDIA Nemotron-3 Super 120B OpenRouter9.86 / 10 CI [9.62, 10.00]HIGH$0.123.1xbatch
Tencent Hy3 OpenRouter9.90 / 10 CI [9.76, 10.00]RANKED$0.143.8xbatch
MiniMax M3 MiniMax9.95 / 10 CI [9.87, 10.00]RANKED$0.225.8xbatch
DeepSeek V4 Pro DeepSeek9.90 / 10 CI [9.84, 9.97]RANKED$0.236.1xbatch
GPT-5.6 Luna OpenAI9.95 / 10 CI [9.82, 10.00]RANKED$0.277.2xbatch
NVIDIA Nemotron-3 Ultra 550B OpenRouter9.98 / 10 CI [9.88, 10.00]RANKED$0.5314xbatch
Claude Haiku 4.5 Anthropic9.94 / 10 CI [9.90, 9.98]RANKED$0.6016xbatch
Qwen 3.5 Flash Alibaba Cloud (DashScope)9.99 / 10 CI [9.86, 10.00]RANKED$0.6317xbatch
GPT-5.6 Terra OpenAI9.99 / 10 CI [9.92, 10.00]RANKED$0.6717xbatch
Gemini 3.5 Flash Gemini9.98 / 10 CI [9.90, 10.00]RANKED$1.3235xbatch
Claude Sonnet 5 best Anthropic10.06 / 10 CI [10.04, 10.00]RANKED$1.3335xbatch
GPT-5.5 OpenAI9.95 / 10 CI [9.94, 9.96]RANKED$1.4337xbatch
GPT-5.6 Sol OpenAI10.00 / 10 CI [9.97, 10.00]RANKED$1.4438xbatch
Qwen 3.7 Plus Alibaba Cloud (DashScope)9.94 / 10 CI [9.79, 10.00]RANKED$1.4739xbatch
Gemini 3.1 Pro Preview Gemini9.93 / 10 CI [9.87, 9.98]RANKED$1.8248xbatch
Qwen 3.6 Plus Alibaba Cloud (DashScope)9.91 / 10 CI [9.72, 10.00]RANKED$1.9350xbatch
Claude Sonnet 4.6 Anthropic9.96 / 10 CI [9.95, 9.97]RANKED$1.9752xbatch
Qwen 3.6 Flash Alibaba Cloud (DashScope)9.99 / 10 CI [9.88, 10.00]RANKED$2.5266xbatch
Meta Muse Spark 1.1 Meta9.84 / 10 CI [9.63, 10.00]HIGH$2.8976xbatch
Claude Opus 4.8 Anthropic9.98 / 10 CI [9.84, 10.00]RANKED$3.2986xbatch
Grok 4.5 xAI9.92 / 10 CI [9.75, 10.00]RANKED$3.4189xbatch
Kimi K2.6 Moonshot AI9.91 / 10 CI [9.85, 9.96]RANKED$5.47143xbatch

Overpay shows how much more you pay than the best-value model that clears the quality bar (marked ★) — the best-value good-enough option. "16x" means you overpay 16× — 16× that reference for no quality benefit above the bar. Typical call shape for this task: 384 input tokens → 83 output tokens, EMA-tracked from production traffic. Cost is the observed, all-in $ per 1,000 task runs: each model's own measured usage on this task — output verbosity, thinking/reasoning tokens, cache reads and writes, and the spend on its billed failures — priced at current list rates and adjusted by the billing overhead we actually reconcile against provider invoices. Models that answer tersely cost what they actually cost; models that think at length pay for it. Not comparable to providers' advertised $/1M list rates — this is what running the task costs, not a per-token price.

Prompt templates

This is a pooled capability — 2 prompt families share it. The pair shown first is the most frequently used in production.

URL_PARSER_LANGUAGE_DETECTION_SYSTEM + URL_PARSER_LANGUAGE_DETECTION_USER (194906 calls in window)

System prompt

You are a highly accurate language identification expert. Your sole task is to identify the primary language of the provided text snippet.

Respond ONLY with the two-letter ISO 639-1 code for the detected language (e.g., "en" for English, "es" for Spanish, "zh" for Chinese).

## Required Output Format
Your response MUST be a single, valid JSON object conforming to this schema:
```json
{schema_json_string}
```

User prompt

1. Text: {input_text}

2. Generate a comprehensive response as a single, well-formed JSON object that strictly adheres to the Pydantic schema provided below. Schema:
The required JSON output schema is provided in the system prompt.
JSON_REPAIR_SYSTEM + JSON_REPAIR_USER (7099 calls in window)

System prompt

You are a JSON repair tool. The user gives you malformed or partial model output and a JSON Schema. Return ONLY a single valid JSON object that satisfies the schema, salvaging as much real content from the input as possible. Do not invent data for fields the input doesn't support — use the schema's allowed empty/null values. Output the JSON object only: no prose, no markdown, no code fences.

User prompt

JSON Schema:
{schema_json}

Malformed output to repair:
{raw_text}

Return only the corrected JSON object.