Cost mode:

Category: Relevance, Classification & Matching · Rail: absolute · Typical I/O: 1448→1125 tokens

Models

Frontier on this task: Claude Sonnet 4.6 at 8.92 / 10. Quality bar at 90%: 8.03.

point-estimate floor (CI low) · upper CI (less certain) · Bars sorted by blended cost; best-value model first. Greyed rows are MEDIUM+ models whose point estimate clears the bar but whose CI low does not.

ModelQuality scoreCI lowCost / 1k runsvs best value
DeepSeek V4 Flash8.20 / 107.99$0.15best value
GPT-5.4 Nano8.42 / 108.27$0.231.6x more expensive
Gemini 3.1 Flash Lite8.38 / 108.18$0.231.6x more expensive
Tencent Hy38.39 / 108.06$0.261.8x more expensive
GPT-5.4 Mini8.26 / 108.05$0.412.8x more expensive
NVIDIA Nemotron-3 Super 120B8.33 / 108.02$0.563.8x more expensive
MiniMax M38.77 / 108.62$0.684.7x more expensive
GPT-5.6 Luna8.64 / 108.34$1.188.1x more expensive
Qwen 3.5 Flash8.35 / 108.16$1.339.1x more expensive
Claude Haiku 4.58.71 / 108.53$1.6311x more expensive
GPT-5.6 Terra8.71 / 108.51$2.1515x more expensive
Qwen 3.7 Plus8.78 / 108.64$2.4717x more expensive
Gemini 3.5 Flash8.56 / 108.41$3.3723x more expensive
Gemini 3.1 Pro Preview8.43 / 108.27$3.8226x more expensive
Claude Sonnet 58.65 / 108.48$4.1128x more expensive
NVIDIA Nemotron-3 Ultra 550B8.60 / 108.34$4.9234x more expensive
Claude Sonnet 4.68.92 / 108.77$5.0534x more expensive
GPT-5.6 Sol8.28 / 107.92$5.5738x more expensive
Qwen 3.6 Plus8.51 / 108.30$6.3643x more expensive
Grok 4.58.77 / 108.63$6.8046x more expensive
GPT-5.58.51 / 108.33$8.0655x more expensive
Meta Muse Spark 1.18.81 / 108.59$8.3357x more expensive
Claude Opus 4.88.70 / 108.55$9.3064x more expensive
Kimi K2.68.66 / 108.50$11.3177x more expensive
Qwen 3.6 Flash7.65 / 107.23$3.4023x more expensive
DeepSeek V4 Pro8.03 / 107.79$0.956.5x more expensive

Cost breakdown

ModelQualityConfidenceCost / 1k runsOverpayMode
DeepSeek V4 Flash DeepSeek8.20 / 10 CI [7.99, 8.40]HIGH$0.15best valuebatch
GPT-5.4 Nano OpenAI8.42 / 10 CI [8.27, 8.58]RANKED$0.231.6xbatch
Gemini 3.1 Flash Lite Gemini8.38 / 10 CI [8.18, 8.57]RANKED$0.231.6xbatch
Tencent Hy3 OpenRouter8.39 / 10 CI [8.06, 8.71]MEDIUM$0.261.8xbatch
GPT-5.4 Mini OpenAI8.26 / 10 CI [8.05, 8.47]HIGH$0.412.8xbatch
NVIDIA Nemotron-3 Super 120B OpenRouter8.33 / 10 CI [8.02, 8.63]MEDIUM$0.563.8xbatch
MiniMax M3 MiniMax8.77 / 10 CI [8.62, 8.92]RANKED$0.684.7xbatch
GPT-5.6 Luna OpenAI8.64 / 10 CI [8.34, 8.93]HIGH$1.188.1xbatch
Qwen 3.5 Flash Alibaba Cloud (DashScope)8.35 / 10 CI [8.16, 8.54]RANKED$1.339.1xbatch
Claude Haiku 4.5 Anthropic8.71 / 10 CI [8.53, 8.88]RANKED$1.6311xbatch
GPT-5.6 Terra OpenAI8.71 / 10 CI [8.51, 8.91]RANKED$2.1515xbatch
Qwen 3.7 Plus Alibaba Cloud (DashScope)8.78 / 10 CI [8.64, 8.93]RANKED$2.4717xbatch
Gemini 3.5 Flash Gemini8.56 / 10 CI [8.41, 8.72]RANKED$3.3723xbatch
Gemini 3.1 Pro Preview Gemini8.43 / 10 CI [8.27, 8.59]RANKED$3.8226xbatch
Claude Sonnet 5 Anthropic8.65 / 10 CI [8.48, 8.81]RANKED$4.1128xbatch
NVIDIA Nemotron-3 Ultra 550B OpenRouter8.60 / 10 CI [8.34, 8.86]HIGH$4.9234xbatch
Claude Sonnet 4.6 best Anthropic8.92 / 10 CI [8.77, 9.06]RANKED$5.0534xbatch
GPT-5.6 Sol OpenAI8.28 / 10 CI [7.92, 8.64]MEDIUM$5.5738xbatch
Qwen 3.6 Plus Alibaba Cloud (DashScope)8.51 / 10 CI [8.30, 8.72]HIGH$6.3643xbatch
Grok 4.5 xAI8.77 / 10 CI [8.63, 8.92]RANKED$6.8046xbatch
GPT-5.5 OpenAI8.51 / 10 CI [8.33, 8.69]RANKED$8.0655xbatch
Meta Muse Spark 1.1 Meta8.81 / 10 CI [8.59, 9.03]HIGH$8.3357xbatch
Claude Opus 4.8 Anthropic8.70 / 10 CI [8.55, 8.86]RANKED$9.3064xbatch
Kimi K2.6 Moonshot AI8.66 / 10 CI [8.50, 8.82]RANKED$11.3177xbatch

Overpay shows how much more you pay than the best-value model that clears the quality bar (marked ★) — the best-value good-enough option. "16x" means you overpay 16× — 16× that reference for no quality benefit above the bar. Typical call shape for this task: 1448 input tokens → 1125 output tokens, EMA-tracked from production traffic. Cost is the observed, all-in $ per 1,000 task runs: each model's own measured usage on this task — output verbosity, thinking/reasoning tokens, cache reads and writes, and the spend on its billed failures — priced at current list rates and adjusted by the billing overhead we actually reconcile against provider invoices. Models that answer tersely cost what they actually cost; models that think at length pay for it. Not comparable to providers' advertised $/1M list rates — this is what running the task costs, not a per-token price.

Prompt templates

This is a pooled capability — 2 prompt families share it. The pair shown first is the most frequently used in production.

ENGAGEMENT_TRIAGE_SYSTEM_PROMPT + ENGAGEMENT_TRIAGE_USER_PROMPT (35901 calls in window)

System prompt

You are the triage step for llm-bench's community-engagement system. llm-bench benchmarks LLMs on cost vs quality per task type and publishes model recommendations with an open methodology.

You are given one social-media post. Decide whether llm-bench should ENGAGE (reply).

ENGAGE only when ALL hold:
- The post is genuinely about LLM token cost/pricing, model selection/comparison, LLM evaluation/benchmarking, or inference cost/infra — topics where llm-bench can add real value.
- THE HOOK TEST: you can state in ONE sentence the specific, post-relevant thing a reply would add (answer this question, correct this misconception, give this data point). Topical adjacency is NOT enough — a post that merely mentions benchmarks, GPUs, cost, "LLM", or "AI" with no opening for a genuinely useful reply is IGNORE. If the only way to mention llm-bench is to change the subject, IGNORE.
- A good reply would not require any claim on the forbidden-claims list.

IGNORE examples (engage on NONE of these): a benchmark for non-LLM / agent / game tasks; a GPU- or hardware-comparison thread; an infra-scaling, funding, or company-announcement thread with no model-selection question; a product launch where a reply would be self-promotion; a hostile/flame thread; or anything where you cannot pass the hook test.

The `angle` MUST be specific to this post's content (reference what the author actually said). A generic angle like "mention cost vs quality" means there is no real hook — choose IGNORE instead.

Assign a risk tier:
- low: factual/supportive answer to a genuine question; neutral author.
- medium: opinionated or promotional angle; new/unknown author.
- high: contentious, critical, or sensitive thread where a misstep could harm the brand.

Return ONLY JSON matching the provided schema.

{operator_instructions}

## Required Output Format
Your response MUST be a single, valid JSON object conforming to this schema:
```json
{schema_json_string}
```

User prompt

Platform: {platform}
Author: @{author_handle}
Post:
"""
{post_text}
"""

llm-bench themes: {themes}
Approved claims: {approved_claims}
Forbidden claims: {forbidden_claims}

Decide engage/ignore, relevance (0..1 to the themes above), risk_tier, and a short angle for a helpful reply. Respond with JSON only, matching this schema:
The required JSON output schema is provided in the system prompt.
JSON_REPAIR_SYSTEM + JSON_REPAIR_USER (76 calls in window)

System prompt

You are a JSON repair tool. The user gives you malformed or partial model output and a JSON Schema. Return ONLY a single valid JSON object that satisfies the schema, salvaging as much real content from the input as possible. Do not invent data for fields the input doesn't support — use the schema's allowed empty/null values. Output the JSON object only: no prose, no markdown, no code fences.

User prompt

JSON Schema:
{schema_json}

Malformed output to repair:
{raw_text}

Return only the corrected JSON object.