Cost mode:

Category: Relevance, Classification & Matching · Rail: absolute · Typical I/O: 1408→873 tokens

Models

Frontier on this task: Thinking Machines Inkling at 8.87 / 10. Quality bar at 90%: 7.98.

point-estimate floor (CI low) · upper CI (less certain) · Bars sorted by blended cost; best-value model first. Greyed rows are MEDIUM+ models whose point estimate clears the bar but whose CI low does not.

ModelQuality scoreCI lowCost / 1k runsvs best value
GPT-5.6 Luna8.72 / 108.56$0.16best value
Gemini 3.1 Flash Lite8.40 / 108.25$0.231.4x more expensive
GPT-5.4 Nano8.45 / 108.32$0.241.5x more expensive
Gemini 3.5 Flash Lite8.26 / 108.04$0.301.9x more expensive
MiniMax M38.65 / 108.52$0.543.3x more expensive
DeepSeek V4 Flash8.36 / 108.19$0.744.6x more expensive
Tencent Hy38.54 / 108.34$0.774.7x more expensive
Gemini 3.8 Flash8.53 / 108.32$1.096.7x more expensive
NVIDIA Nemotron-3 Super 120B8.50 / 108.30$1.368.3x more expensive
Claude Haiku 4.58.71 / 108.58$1.6610x more expensive
GPT-5.6 Terra8.45 / 108.28$1.6810x more expensive
Thinking Machines Inkling Small8.78 / 108.60$1.9012x more expensive
Qwen 3.7 Plus8.69 / 108.53$2.4315x more expensive
NVIDIA Nemotron-3 Ultra 550B8.72 / 108.53$2.4615x more expensive
Claude Sonnet 58.55 / 108.40$3.1319x more expensive
Gemini 3.5 Flash8.59 / 108.45$3.2620x more expensive
GPT-5.6 Sol8.43 / 108.22$3.3220x more expensive
DeepSeek V4 Pro8.31 / 108.13$3.9224x more expensive
Thinking Machines Inkling8.87 / 108.68$6.0937x more expensive
GLM-5.38.73 / 108.52$6.8842x more expensive
Claude Opus 58.61 / 108.43$7.0343x more expensive
Grok 4.68.60 / 108.41$8.1650x more expensive
Qwen 3.8 Max8.47 / 108.26$9.2456x more expensive
Moonshot Kimi K38.80 / 108.57$15.0192x more expensive
NVIDIA Nemotron-3 Nano 30B-A3B7.49 / 107.00$0.271.6x more expensive

Cost breakdown

ModelQualityConfidenceCost / 1k runsOverpayMode
GPT-5.6 Luna OpenAI8.72 / 10 CI [8.56, 8.88]RANKED$0.16best valuebatch
Gemini 3.1 Flash Lite Gemini8.40 / 10 CI [8.25, 8.55]RANKED$0.231.4xbatch
GPT-5.4 Nano OpenAI8.45 / 10 CI [8.32, 8.58]RANKED$0.241.5xbatch
Gemini 3.5 Flash Lite Gemini8.26 / 10 CI [8.04, 8.49]HIGH$0.301.9xbatch
MiniMax M3 OpenRouter8.65 / 10 CI [8.52, 8.77]RANKED$0.543.3xbatch
DeepSeek V4 Flash DeepSeek8.36 / 10 CI [8.19, 8.53]RANKED$0.744.6xbatch
Tencent Hy3 OpenRouter8.54 / 10 CI [8.34, 8.73]RANKED$0.774.7xbatch
Gemini 3.8 Flash Gemini8.53 / 10 CI [8.32, 8.74]HIGH$1.096.7xbatch
NVIDIA Nemotron-3 Super 120B OpenRouter8.50 / 10 CI [8.30, 8.70]HIGH$1.368.3xbatch
Claude Haiku 4.5 Anthropic8.71 / 10 CI [8.58, 8.85]RANKED$1.6610xbatch
GPT-5.6 Terra OpenAI8.45 / 10 CI [8.28, 8.62]RANKED$1.6810xbatch
Thinking Machines Inkling Small OpenRouter8.78 / 10 CI [8.60, 8.96]RANKED$1.9012xbatch
Qwen 3.7 Plus Alibaba Cloud (DashScope)8.69 / 10 CI [8.53, 8.85]RANKED$2.4315xbatch
NVIDIA Nemotron-3 Ultra 550B OpenRouter8.72 / 10 CI [8.53, 8.92]RANKED$2.4615xbatch
Claude Sonnet 5 Anthropic8.55 / 10 CI [8.40, 8.70]RANKED$3.1319xbatch
Gemini 3.5 Flash Gemini8.59 / 10 CI [8.45, 8.72]RANKED$3.2620xbatch
GPT-5.6 Sol OpenAI8.43 / 10 CI [8.22, 8.64]HIGH$3.3220xbatch
DeepSeek V4 Pro DeepSeek8.31 / 10 CI [8.13, 8.49]RANKED$3.9224xbatch
Thinking Machines Inkling best OpenRouter8.87 / 10 CI [8.68, 9.05]RANKED$6.0937xbatch
GLM-5.3 Z.AI8.73 / 10 CI [8.52, 8.94]HIGH$6.8842xbatch
Claude Opus 5 Anthropic8.61 / 10 CI [8.43, 8.80]RANKED$7.0343xbatch
Grok 4.6 xAI8.60 / 10 CI [8.41, 8.79]RANKED$8.1650xbatch
Qwen 3.8 Max Alibaba Cloud (DashScope)8.47 / 10 CI [8.26, 8.68]HIGH$9.2456xbatch
Moonshot Kimi K3 Moonshot AI8.80 / 10 CI [8.57, 9.02]HIGH$15.0192xbatch

Overpay shows how much more you pay than the best-value model that clears the quality bar (marked ★) — the best-value good-enough option. "16x" means you overpay 16× — 16× that reference for no quality benefit above the bar. Typical call shape for this task: 1408 input tokens → 873 output tokens, EMA-tracked from production traffic. Cost is the observed, all-in $ per 1,000 task runs: each model's own measured usage on this task — output verbosity, thinking/reasoning tokens, cache reads and writes, and the spend on its billed failures — priced at current list rates and adjusted by the billing overhead we actually reconcile against provider invoices. Models that answer tersely cost what they actually cost; models that think at length pay for it. Not comparable to providers' advertised $/1M list rates — this is what running the task costs, not a per-token price.

Evaluation rubric

Judge decision correctness, calibration of risk, relevance, evidence feasibility of the angle, policy compliance, and clarity of rationale. Penalize engagement recommendations driven only by reach.

Prompt templates

This is a pooled capability — 3 prompt families share it. The pair shown first is the most frequently used in production.

ENGAGEMENT_TRIAGE_SYSTEM_PROMPT + ENGAGEMENT_TRIAGE_USER_PROMPT (360 calls in window)

System prompt

You are the triage step for llm-bench's community-engagement system. llm-bench benchmarks LLMs on cost vs quality per task type and publishes model recommendations with an open methodology.

You are given one social-media post. Decide whether llm-bench should ENGAGE (reply).

ENGAGE only when ALL hold:
- The post is genuinely about LLM token cost/pricing, model selection/comparison, LLM evaluation/benchmarking, or inference cost/infra — topics where llm-bench can add real value.
- THE HOOK TEST: you can state in ONE sentence the specific, post-relevant thing a reply would add (answer this question, correct this misconception, give this data point). Topical adjacency is NOT enough — a post that merely mentions benchmarks, GPUs, cost, "LLM", or "AI" with no opening for a genuinely useful reply is IGNORE. If the only way to mention llm-bench is to change the subject, IGNORE.
- A good reply would not require any claim on the forbidden-claims list.

IGNORE examples (engage on NONE of these): a benchmark for non-LLM / agent / game tasks; a GPU- or hardware-comparison thread; an infra-scaling, funding, or company-announcement thread with no model-selection question; a product launch where a reply would be self-promotion; a hostile/flame thread; or anything where you cannot pass the hook test.

The `angle` MUST be specific to this post's content (reference what the author actually said). A generic angle like "mention cost vs quality" means there is no real hook — choose IGNORE instead.

Assign a risk tier:
- low: factual/supportive answer to a genuine question; neutral author.
- medium: opinionated or promotional angle; new/unknown author.
- high: contentious, critical, or sensitive thread where a misstep could harm the brand.

Return ONLY JSON matching the provided schema.

{operator_instructions}

## Required Output Format
Your response MUST be a single, valid JSON object conforming to this schema:
```json
{schema_json_string}
```

User prompt

Platform: {platform}
Author: @{author_handle}
Post:
"""
{post_text}
"""

llm-bench themes: {themes}
Approved claims: {approved_claims}
Forbidden claims: {forbidden_claims}

Decide engage/ignore, relevance (0..1 to the themes above), risk_tier, and a short angle for a helpful reply. Respond with JSON only, matching this schema:
The required JSON output schema is provided in the system prompt.
LLMB_ENGAGEMENT_OPPORTUNITY_TRIAGE_SYSTEM + LLMB_ENGAGEMENT_OPPORTUNITY_TRIAGE_USER (342 calls in window)

System prompt

COVERAGE INDEX — the subjects we hold measured evidence about, and the kinds of fact held about each:
{bench_claim_index}

COVERED SUBJECTS — the only models that may be named or evaluated:
{covered_models}

Make an engage-or-ignore decision based on relevance, ability to add value, evidence availability, and risk. Treat controversy or popularity as context, not automatic reasons to engage. Evidence availability means the coverage index above: engage where a specific measured fact we hold bears on the post, and prefer ignoring over engaging on a subject the index does not list. Any proposed angle must be supportable by that evidence together with the approved claims, and must respect forbidden claims and operator instructions. In claim_subjects, name the index entries the post is actually about, copied verbatim and at most six — these decide which measured facts the reply may draw on, so copy them exactly or leave the list empty rather than approximating. Do not name or evaluate anything absent from the covered subjects list; address the underlying tradeoff generally instead. Explain the decisive factors compactly. Treat empty optional values as absent and return only the requested result. Your response must conform exactly to this output schema: {schema_json_string}.

User prompt

Inputs — operator_instructions: {operator_instructions}; platform: {platform}; author_handle: {author_handle}; post_text: {post_text}; themes: {themes}; approved_claims: {approved_claims}; forbidden_claims: {forbidden_claims}; channel_profile: {channel_profile}; policy_profile: {policy_profile}. Use only these inputs to complete the task defined by the system prompt.
JSON_REPAIR_SYSTEM + JSON_REPAIR_USER (3 calls in window)

System prompt

You are a JSON repair tool. The user gives you malformed or partial model output and a JSON Schema. Return ONLY a single valid JSON object that satisfies the schema, salvaging as much real content from the input as possible. Do not invent data for fields the input doesn't support — use the schema's allowed empty/null values. Output the JSON object only: no prose, no markdown, no code fences.

User prompt

JSON Schema:
{schema_json}

Malformed output to repair:
{raw_text}

Return only the corrected JSON object.