Best LLMs for Engagement Opportunity Triage
Classifies whether an incoming public message merits a response, identifies risk, and proposes a defensible angle using supplied themes, evidence boundaries, channel context, and policy. Illustrative uses include triaging a customer escalation, software security report, partner q
Models
Frontier on this task: Thinking Machines Inkling at 8.87 / 10. Quality bar at 90%: 7.98.
point-estimate floor (CI low) · upper CI (less certain) · Bars sorted by blended cost; best-value model first. Greyed rows are MEDIUM+ models whose point estimate clears the bar but whose CI low does not.
| Model | Quality score | CI low | Cost / 1k runs | vs best value |
|---|---|---|---|---|
| GPT-5.6 Luna | 8.72 / 10 | 8.56 | $0.16 | best value |
| Gemini 3.1 Flash Lite | 8.40 / 10 | 8.25 | $0.23 | 1.4x more expensive |
| GPT-5.4 Nano | 8.45 / 10 | 8.32 | $0.24 | 1.5x more expensive |
| Gemini 3.5 Flash Lite | 8.26 / 10 | 8.04 | $0.30 | 1.9x more expensive |
| MiniMax M3 | 8.65 / 10 | 8.52 | $0.54 | 3.3x more expensive |
| DeepSeek V4 Flash | 8.36 / 10 | 8.19 | $0.74 | 4.6x more expensive |
| Tencent Hy3 | 8.54 / 10 | 8.34 | $0.77 | 4.7x more expensive |
| Gemini 3.8 Flash | 8.53 / 10 | 8.32 | $1.09 | 6.7x more expensive |
| NVIDIA Nemotron-3 Super 120B | 8.50 / 10 | 8.30 | $1.36 | 8.3x more expensive |
| Claude Haiku 4.5 | 8.71 / 10 | 8.58 | $1.66 | 10x more expensive |
| GPT-5.6 Terra | 8.45 / 10 | 8.28 | $1.68 | 10x more expensive |
| Thinking Machines Inkling Small | 8.78 / 10 | 8.60 | $1.90 | 12x more expensive |
| Qwen 3.7 Plus | 8.69 / 10 | 8.53 | $2.43 | 15x more expensive |
| NVIDIA Nemotron-3 Ultra 550B | 8.72 / 10 | 8.53 | $2.46 | 15x more expensive |
| Claude Sonnet 5 | 8.55 / 10 | 8.40 | $3.13 | 19x more expensive |
| Gemini 3.5 Flash | 8.59 / 10 | 8.45 | $3.26 | 20x more expensive |
| GPT-5.6 Sol | 8.43 / 10 | 8.22 | $3.32 | 20x more expensive |
| DeepSeek V4 Pro | 8.31 / 10 | 8.13 | $3.92 | 24x more expensive |
| Thinking Machines Inkling | 8.87 / 10 | 8.68 | $6.09 | 37x more expensive |
| GLM-5.3 | 8.73 / 10 | 8.52 | $6.88 | 42x more expensive |
| Claude Opus 5 | 8.61 / 10 | 8.43 | $7.03 | 43x more expensive |
| Grok 4.6 | 8.60 / 10 | 8.41 | $8.16 | 50x more expensive |
| Qwen 3.8 Max | 8.47 / 10 | 8.26 | $9.24 | 56x more expensive |
| Moonshot Kimi K3 | 8.80 / 10 | 8.57 | $15.01 | 92x more expensive |
| NVIDIA Nemotron-3 Nano 30B-A3B | 7.49 / 10 | 7.00 | $0.27 | 1.6x more expensive |
Cost breakdown
| Model | Quality | Confidence | Cost / 1k runs | Overpay | Mode |
|---|---|---|---|---|---|
| GPT-5.6 Luna ★ OpenAI | 8.72 / 10 CI [8.56, 8.88] | RANKED | $0.16 | best value | batch |
| Gemini 3.1 Flash Lite Gemini | 8.40 / 10 CI [8.25, 8.55] | RANKED | $0.23 | 1.4x | batch |
| GPT-5.4 Nano OpenAI | 8.45 / 10 CI [8.32, 8.58] | RANKED | $0.24 | 1.5x | batch |
| Gemini 3.5 Flash Lite Gemini | 8.26 / 10 CI [8.04, 8.49] | HIGH | $0.30 | 1.9x | batch |
| MiniMax M3 OpenRouter | 8.65 / 10 CI [8.52, 8.77] | RANKED | $0.54 | 3.3x | batch |
| DeepSeek V4 Flash DeepSeek | 8.36 / 10 CI [8.19, 8.53] | RANKED | $0.74 | 4.6x | batch |
| Tencent Hy3 OpenRouter | 8.54 / 10 CI [8.34, 8.73] | RANKED | $0.77 | 4.7x | batch |
| Gemini 3.8 Flash Gemini | 8.53 / 10 CI [8.32, 8.74] | HIGH | $1.09 | 6.7x | batch |
| NVIDIA Nemotron-3 Super 120B OpenRouter | 8.50 / 10 CI [8.30, 8.70] | HIGH | $1.36 | 8.3x | batch |
| Claude Haiku 4.5 Anthropic | 8.71 / 10 CI [8.58, 8.85] | RANKED | $1.66 | 10x | batch |
| GPT-5.6 Terra OpenAI | 8.45 / 10 CI [8.28, 8.62] | RANKED | $1.68 | 10x | batch |
| Thinking Machines Inkling Small OpenRouter | 8.78 / 10 CI [8.60, 8.96] | RANKED | $1.90 | 12x | batch |
| Qwen 3.7 Plus Alibaba Cloud (DashScope) | 8.69 / 10 CI [8.53, 8.85] | RANKED | $2.43 | 15x | batch |
| NVIDIA Nemotron-3 Ultra 550B OpenRouter | 8.72 / 10 CI [8.53, 8.92] | RANKED | $2.46 | 15x | batch |
| Claude Sonnet 5 Anthropic | 8.55 / 10 CI [8.40, 8.70] | RANKED | $3.13 | 19x | batch |
| Gemini 3.5 Flash Gemini | 8.59 / 10 CI [8.45, 8.72] | RANKED | $3.26 | 20x | batch |
| GPT-5.6 Sol OpenAI | 8.43 / 10 CI [8.22, 8.64] | HIGH | $3.32 | 20x | batch |
| DeepSeek V4 Pro DeepSeek | 8.31 / 10 CI [8.13, 8.49] | RANKED | $3.92 | 24x | batch |
| Thinking Machines Inkling best OpenRouter | 8.87 / 10 CI [8.68, 9.05] | RANKED | $6.09 | 37x | batch |
| GLM-5.3 Z.AI | 8.73 / 10 CI [8.52, 8.94] | HIGH | $6.88 | 42x | batch |
| Claude Opus 5 Anthropic | 8.61 / 10 CI [8.43, 8.80] | RANKED | $7.03 | 43x | batch |
| Grok 4.6 xAI | 8.60 / 10 CI [8.41, 8.79] | RANKED | $8.16 | 50x | batch |
| Qwen 3.8 Max Alibaba Cloud (DashScope) | 8.47 / 10 CI [8.26, 8.68] | HIGH | $9.24 | 56x | batch |
| Moonshot Kimi K3 Moonshot AI | 8.80 / 10 CI [8.57, 9.02] | HIGH | $15.01 | 92x | batch |
Overpay shows how much more you pay than the best-value model that clears the quality bar (marked ★) — the best-value good-enough option. "16x" means you overpay 16× — 16× that reference for no quality benefit above the bar. Typical call shape for this task: 1408 input tokens → 873 output tokens, EMA-tracked from production traffic. Cost is the observed, all-in $ per 1,000 task runs: each model's own measured usage on this task — output verbosity, thinking/reasoning tokens, cache reads and writes, and the spend on its billed failures — priced at current list rates and adjusted by the billing overhead we actually reconcile against provider invoices. Models that answer tersely cost what they actually cost; models that think at length pay for it. Not comparable to providers' advertised $/1M list rates — this is what running the task costs, not a per-token price.
Evaluation rubric
Judge decision correctness, calibration of risk, relevance, evidence feasibility of the angle, policy compliance, and clarity of rationale. Penalize engagement recommendations driven only by reach.
Prompt templates
This is a pooled capability — 3 prompt families share it. The pair shown first is the most frequently used in production.
ENGAGEMENT_TRIAGE_SYSTEM_PROMPT +
ENGAGEMENT_TRIAGE_USER_PROMPT
(360 calls in window)
System prompt
You are the triage step for llm-bench's community-engagement system. llm-bench benchmarks LLMs on cost vs quality per task type and publishes model recommendations with an open methodology.
You are given one social-media post. Decide whether llm-bench should ENGAGE (reply).
ENGAGE only when ALL hold:
- The post is genuinely about LLM token cost/pricing, model selection/comparison, LLM evaluation/benchmarking, or inference cost/infra — topics where llm-bench can add real value.
- THE HOOK TEST: you can state in ONE sentence the specific, post-relevant thing a reply would add (answer this question, correct this misconception, give this data point). Topical adjacency is NOT enough — a post that merely mentions benchmarks, GPUs, cost, "LLM", or "AI" with no opening for a genuinely useful reply is IGNORE. If the only way to mention llm-bench is to change the subject, IGNORE.
- A good reply would not require any claim on the forbidden-claims list.
IGNORE examples (engage on NONE of these): a benchmark for non-LLM / agent / game tasks; a GPU- or hardware-comparison thread; an infra-scaling, funding, or company-announcement thread with no model-selection question; a product launch where a reply would be self-promotion; a hostile/flame thread; or anything where you cannot pass the hook test.
The `angle` MUST be specific to this post's content (reference what the author actually said). A generic angle like "mention cost vs quality" means there is no real hook — choose IGNORE instead.
Assign a risk tier:
- low: factual/supportive answer to a genuine question; neutral author.
- medium: opinionated or promotional angle; new/unknown author.
- high: contentious, critical, or sensitive thread where a misstep could harm the brand.
Return ONLY JSON matching the provided schema.
{operator_instructions}
## Required Output Format
Your response MUST be a single, valid JSON object conforming to this schema:
```json
{schema_json_string}
```User prompt
Platform: {platform}
Author: @{author_handle}
Post:
"""
{post_text}
"""
llm-bench themes: {themes}
Approved claims: {approved_claims}
Forbidden claims: {forbidden_claims}
Decide engage/ignore, relevance (0..1 to the themes above), risk_tier, and a short angle for a helpful reply. Respond with JSON only, matching this schema:
The required JSON output schema is provided in the system prompt.LLMB_ENGAGEMENT_OPPORTUNITY_TRIAGE_SYSTEM +
LLMB_ENGAGEMENT_OPPORTUNITY_TRIAGE_USER
(342 calls in window)
System prompt
COVERAGE INDEX — the subjects we hold measured evidence about, and the kinds of fact held about each:
{bench_claim_index}
COVERED SUBJECTS — the only models that may be named or evaluated:
{covered_models}
Make an engage-or-ignore decision based on relevance, ability to add value, evidence availability, and risk. Treat controversy or popularity as context, not automatic reasons to engage. Evidence availability means the coverage index above: engage where a specific measured fact we hold bears on the post, and prefer ignoring over engaging on a subject the index does not list. Any proposed angle must be supportable by that evidence together with the approved claims, and must respect forbidden claims and operator instructions. In claim_subjects, name the index entries the post is actually about, copied verbatim and at most six — these decide which measured facts the reply may draw on, so copy them exactly or leave the list empty rather than approximating. Do not name or evaluate anything absent from the covered subjects list; address the underlying tradeoff generally instead. Explain the decisive factors compactly. Treat empty optional values as absent and return only the requested result. Your response must conform exactly to this output schema: {schema_json_string}.
User prompt
Inputs — operator_instructions: {operator_instructions}; platform: {platform}; author_handle: {author_handle}; post_text: {post_text}; themes: {themes}; approved_claims: {approved_claims}; forbidden_claims: {forbidden_claims}; channel_profile: {channel_profile}; policy_profile: {policy_profile}. Use only these inputs to complete the task defined by the system prompt.
JSON_REPAIR_SYSTEM +
JSON_REPAIR_USER
(3 calls in window)
System prompt
You are a JSON repair tool. The user gives you malformed or partial model output and a JSON Schema. Return ONLY a single valid JSON object that satisfies the schema, salvaging as much real content from the input as possible. Do not invent data for fields the input doesn't support — use the schema's allowed empty/null values. Output the JSON object only: no prose, no markdown, no code fences.
User prompt
JSON Schema:
{schema_json}
Malformed output to repair:
{raw_text}
Return only the corrected JSON object.