Best LLMs for Engagement Triage
Decide engage/ignore + risk + angle for one social post.
Models
Frontier on this task: Claude Sonnet 4.6 at 8.92 / 10. Quality bar at 90%: 8.03.
point-estimate floor (CI low) · upper CI (less certain) · Bars sorted by blended cost; best-value model first. Greyed rows are MEDIUM+ models whose point estimate clears the bar but whose CI low does not.
| Model | Quality score | CI low | Cost / 1k runs | vs best value |
|---|---|---|---|---|
| DeepSeek V4 Flash | 8.20 / 10 | 7.99 | $0.15 | best value |
| GPT-5.4 Nano | 8.42 / 10 | 8.27 | $0.23 | 1.6x more expensive |
| Gemini 3.1 Flash Lite | 8.38 / 10 | 8.18 | $0.23 | 1.6x more expensive |
| Tencent Hy3 | 8.39 / 10 | 8.06 | $0.26 | 1.8x more expensive |
| GPT-5.4 Mini | 8.26 / 10 | 8.05 | $0.41 | 2.8x more expensive |
| NVIDIA Nemotron-3 Super 120B | 8.33 / 10 | 8.02 | $0.56 | 3.8x more expensive |
| MiniMax M3 | 8.77 / 10 | 8.62 | $0.68 | 4.7x more expensive |
| GPT-5.6 Luna | 8.64 / 10 | 8.34 | $1.18 | 8.1x more expensive |
| Qwen 3.5 Flash | 8.35 / 10 | 8.16 | $1.33 | 9.1x more expensive |
| Claude Haiku 4.5 | 8.71 / 10 | 8.53 | $1.63 | 11x more expensive |
| GPT-5.6 Terra | 8.71 / 10 | 8.51 | $2.15 | 15x more expensive |
| Qwen 3.7 Plus | 8.78 / 10 | 8.64 | $2.47 | 17x more expensive |
| Gemini 3.5 Flash | 8.56 / 10 | 8.41 | $3.37 | 23x more expensive |
| Gemini 3.1 Pro Preview | 8.43 / 10 | 8.27 | $3.82 | 26x more expensive |
| Claude Sonnet 5 | 8.65 / 10 | 8.48 | $4.11 | 28x more expensive |
| NVIDIA Nemotron-3 Ultra 550B | 8.60 / 10 | 8.34 | $4.92 | 34x more expensive |
| Claude Sonnet 4.6 | 8.92 / 10 | 8.77 | $5.05 | 34x more expensive |
| GPT-5.6 Sol | 8.28 / 10 | 7.92 | $5.57 | 38x more expensive |
| Qwen 3.6 Plus | 8.51 / 10 | 8.30 | $6.36 | 43x more expensive |
| Grok 4.5 | 8.77 / 10 | 8.63 | $6.80 | 46x more expensive |
| GPT-5.5 | 8.51 / 10 | 8.33 | $8.06 | 55x more expensive |
| Meta Muse Spark 1.1 | 8.81 / 10 | 8.59 | $8.33 | 57x more expensive |
| Claude Opus 4.8 | 8.70 / 10 | 8.55 | $9.30 | 64x more expensive |
| Kimi K2.6 | 8.66 / 10 | 8.50 | $11.31 | 77x more expensive |
| Qwen 3.6 Flash | 7.65 / 10 | 7.23 | $3.40 | 23x more expensive |
| DeepSeek V4 Pro | 8.03 / 10 | 7.79 | $0.95 | 6.5x more expensive |
Cost breakdown
| Model | Quality | Confidence | Cost / 1k runs | Overpay | Mode |
|---|---|---|---|---|---|
| DeepSeek V4 Flash ★ DeepSeek | 8.20 / 10 CI [7.99, 8.40] | HIGH | $0.15 | best value | batch |
| GPT-5.4 Nano OpenAI | 8.42 / 10 CI [8.27, 8.58] | RANKED | $0.23 | 1.6x | batch |
| Gemini 3.1 Flash Lite Gemini | 8.38 / 10 CI [8.18, 8.57] | RANKED | $0.23 | 1.6x | batch |
| Tencent Hy3 OpenRouter | 8.39 / 10 CI [8.06, 8.71] | MEDIUM | $0.26 | 1.8x | batch |
| GPT-5.4 Mini OpenAI | 8.26 / 10 CI [8.05, 8.47] | HIGH | $0.41 | 2.8x | batch |
| NVIDIA Nemotron-3 Super 120B OpenRouter | 8.33 / 10 CI [8.02, 8.63] | MEDIUM | $0.56 | 3.8x | batch |
| MiniMax M3 MiniMax | 8.77 / 10 CI [8.62, 8.92] | RANKED | $0.68 | 4.7x | batch |
| GPT-5.6 Luna OpenAI | 8.64 / 10 CI [8.34, 8.93] | HIGH | $1.18 | 8.1x | batch |
| Qwen 3.5 Flash Alibaba Cloud (DashScope) | 8.35 / 10 CI [8.16, 8.54] | RANKED | $1.33 | 9.1x | batch |
| Claude Haiku 4.5 Anthropic | 8.71 / 10 CI [8.53, 8.88] | RANKED | $1.63 | 11x | batch |
| GPT-5.6 Terra OpenAI | 8.71 / 10 CI [8.51, 8.91] | RANKED | $2.15 | 15x | batch |
| Qwen 3.7 Plus Alibaba Cloud (DashScope) | 8.78 / 10 CI [8.64, 8.93] | RANKED | $2.47 | 17x | batch |
| Gemini 3.5 Flash Gemini | 8.56 / 10 CI [8.41, 8.72] | RANKED | $3.37 | 23x | batch |
| Gemini 3.1 Pro Preview Gemini | 8.43 / 10 CI [8.27, 8.59] | RANKED | $3.82 | 26x | batch |
| Claude Sonnet 5 Anthropic | 8.65 / 10 CI [8.48, 8.81] | RANKED | $4.11 | 28x | batch |
| NVIDIA Nemotron-3 Ultra 550B OpenRouter | 8.60 / 10 CI [8.34, 8.86] | HIGH | $4.92 | 34x | batch |
| Claude Sonnet 4.6 best Anthropic | 8.92 / 10 CI [8.77, 9.06] | RANKED | $5.05 | 34x | batch |
| GPT-5.6 Sol OpenAI | 8.28 / 10 CI [7.92, 8.64] | MEDIUM | $5.57 | 38x | batch |
| Qwen 3.6 Plus Alibaba Cloud (DashScope) | 8.51 / 10 CI [8.30, 8.72] | HIGH | $6.36 | 43x | batch |
| Grok 4.5 xAI | 8.77 / 10 CI [8.63, 8.92] | RANKED | $6.80 | 46x | batch |
| GPT-5.5 OpenAI | 8.51 / 10 CI [8.33, 8.69] | RANKED | $8.06 | 55x | batch |
| Meta Muse Spark 1.1 Meta | 8.81 / 10 CI [8.59, 9.03] | HIGH | $8.33 | 57x | batch |
| Claude Opus 4.8 Anthropic | 8.70 / 10 CI [8.55, 8.86] | RANKED | $9.30 | 64x | batch |
| Kimi K2.6 Moonshot AI | 8.66 / 10 CI [8.50, 8.82] | RANKED | $11.31 | 77x | batch |
Overpay shows how much more you pay than the best-value model that clears the quality bar (marked ★) — the best-value good-enough option. "16x" means you overpay 16× — 16× that reference for no quality benefit above the bar. Typical call shape for this task: 1448 input tokens → 1125 output tokens, EMA-tracked from production traffic. Cost is the observed, all-in $ per 1,000 task runs: each model's own measured usage on this task — output verbosity, thinking/reasoning tokens, cache reads and writes, and the spend on its billed failures — priced at current list rates and adjusted by the billing overhead we actually reconcile against provider invoices. Models that answer tersely cost what they actually cost; models that think at length pay for it. Not comparable to providers' advertised $/1M list rates — this is what running the task costs, not a per-token price.
Prompt templates
This is a pooled capability — 2 prompt families share it. The pair shown first is the most frequently used in production.
ENGAGEMENT_TRIAGE_SYSTEM_PROMPT +
ENGAGEMENT_TRIAGE_USER_PROMPT
(35901 calls in window)
System prompt
You are the triage step for llm-bench's community-engagement system. llm-bench benchmarks LLMs on cost vs quality per task type and publishes model recommendations with an open methodology.
You are given one social-media post. Decide whether llm-bench should ENGAGE (reply).
ENGAGE only when ALL hold:
- The post is genuinely about LLM token cost/pricing, model selection/comparison, LLM evaluation/benchmarking, or inference cost/infra — topics where llm-bench can add real value.
- THE HOOK TEST: you can state in ONE sentence the specific, post-relevant thing a reply would add (answer this question, correct this misconception, give this data point). Topical adjacency is NOT enough — a post that merely mentions benchmarks, GPUs, cost, "LLM", or "AI" with no opening for a genuinely useful reply is IGNORE. If the only way to mention llm-bench is to change the subject, IGNORE.
- A good reply would not require any claim on the forbidden-claims list.
IGNORE examples (engage on NONE of these): a benchmark for non-LLM / agent / game tasks; a GPU- or hardware-comparison thread; an infra-scaling, funding, or company-announcement thread with no model-selection question; a product launch where a reply would be self-promotion; a hostile/flame thread; or anything where you cannot pass the hook test.
The `angle` MUST be specific to this post's content (reference what the author actually said). A generic angle like "mention cost vs quality" means there is no real hook — choose IGNORE instead.
Assign a risk tier:
- low: factual/supportive answer to a genuine question; neutral author.
- medium: opinionated or promotional angle; new/unknown author.
- high: contentious, critical, or sensitive thread where a misstep could harm the brand.
Return ONLY JSON matching the provided schema.
{operator_instructions}
## Required Output Format
Your response MUST be a single, valid JSON object conforming to this schema:
```json
{schema_json_string}
```User prompt
Platform: {platform}
Author: @{author_handle}
Post:
"""
{post_text}
"""
llm-bench themes: {themes}
Approved claims: {approved_claims}
Forbidden claims: {forbidden_claims}
Decide engage/ignore, relevance (0..1 to the themes above), risk_tier, and a short angle for a helpful reply. Respond with JSON only, matching this schema:
The required JSON output schema is provided in the system prompt.JSON_REPAIR_SYSTEM +
JSON_REPAIR_USER
(76 calls in window)
System prompt
You are a JSON repair tool. The user gives you malformed or partial model output and a JSON Schema. Return ONLY a single valid JSON object that satisfies the schema, salvaging as much real content from the input as possible. Do not invent data for fields the input doesn't support — use the schema's allowed empty/null values. Output the JSON object only: no prose, no markdown, no code fences.
User prompt
JSON Schema:
{schema_json}
Malformed output to repair:
{raw_text}
Return only the corrected JSON object.