Cost mode:

Category: Social & Promotional Content · Rail: absolute · Typical I/O: 2276→1160 tokens

Models

Frontier on this task: GPT-5.6 Sol at 8.44 / 10. Quality bar at 90%: 7.59.

point-estimate floor (CI low) · upper CI (less certain) · Bars sorted by blended cost; best-value model first. Greyed rows are MEDIUM+ models whose point estimate clears the bar but whose CI low does not.

ModelQuality scoreCI lowCost / 1k runsvs best value
Tencent Hy37.76 / 107.44$1.38best value
GPT-5.6 Luna8.21 / 108.03$2.391.7x more expensive
GPT-5.6 Terra8.29 / 108.05$3.622.6x more expensive
Qwen 3.7 Plus7.60 / 107.24$6.664.8x more expensive
Gemini 3.5 Flash7.75 / 107.46$9.016.5x more expensive
GPT-5.6 Sol8.44 / 108.31$10.397.5x more expensive
Gemini 3.1 Pro Preview7.97 / 107.71$10.717.7x more expensive
GPT-5.58.30 / 108.05$12.218.8x more expensive
Claude Opus 4.87.70 / 107.39$12.539.1x more expensive
Meta Muse Spark 1.17.79 / 107.49$15.4411x more expensive
Grok 4.57.72 / 107.50$23.8717x more expensive
Kimi K2.68.05 / 107.80$35.3826x more expensive
GPT-5.4 Mini7.11 / 106.77$0.4965% cheaper
Qwen 3.6 Plus7.50 / 107.06$18.1913x more expensive
DeepSeek V4 Flash7.10 / 106.76$0.4071% cheaper
GPT-5.4 Nano7.22 / 106.82$0.4468% cheaper
Qwen 3.6 Flash7.08 / 106.65$7.745.6x more expensive
Claude Sonnet 4.67.19 / 106.77$6.834.9x more expensive
Qwen 3.5 Flash7.12 / 106.73$3.062.2x more expensive
MiniMax M37.48 / 107.12$2.301.7x more expensive
Claude Haiku 4.57.21 / 106.81$2.631.9x more expensive
DeepSeek V4 Pro7.42 / 107.13$2.261.6x more expensive
Gemini 3.1 Flash Lite6.93 / 106.57$0.2780% cheaper

Cost breakdown

ModelQualityConfidenceCost / 1k runsOverpayMode
Tencent Hy3 OpenRouter7.76 / 10 CI [7.44, 8.08]MEDIUM$1.38best valuebatch
GPT-5.6 Luna OpenAI8.21 / 10 CI [8.03, 8.40]RANKED$2.391.7xbatch
GPT-5.6 Terra OpenAI8.29 / 10 CI [8.05, 8.53]HIGH$3.622.6xbatch
Qwen 3.7 Plus Alibaba Cloud (DashScope)7.60 / 10 CI [7.24, 7.97]MEDIUM$6.664.8xbatch
Gemini 3.5 Flash Gemini7.75 / 10 CI [7.46, 8.05]HIGH$9.016.5xbatch
GPT-5.6 Sol best OpenAI8.44 / 10 CI [8.31, 8.56]RANKED$10.397.5xbatch
Gemini 3.1 Pro Preview Gemini7.97 / 10 CI [7.71, 8.22]HIGH$10.717.7xbatch
GPT-5.5 OpenAI8.30 / 10 CI [8.05, 8.54]HIGH$12.218.8xbatch
Claude Opus 4.8 Anthropic7.70 / 10 CI [7.39, 8.02]MEDIUM$12.539.1xbatch
Meta Muse Spark 1.1 Meta7.79 / 10 CI [7.49, 8.09]MEDIUM$15.4411xbatch
Grok 4.5 xAI7.72 / 10 CI [7.50, 7.95]HIGH$23.8717xbatch
Kimi K2.6 Moonshot AI8.05 / 10 CI [7.80, 8.30]HIGH$35.3826xbatch

Overpay shows how much more you pay than the best-value model that clears the quality bar (marked ★) — the best-value good-enough option. "16x" means you overpay 16× — 16× that reference for no quality benefit above the bar. Typical call shape for this task: 2276 input tokens → 1160 output tokens, EMA-tracked from production traffic. Cost is the observed, all-in $ per 1,000 task runs: each model's own measured usage on this task — output verbosity, thinking/reasoning tokens, cache reads and writes, and the spend on its billed failures — priced at current list rates and adjusted by the billing overhead we actually reconcile against provider invoices. Models that answer tersely cost what they actually cost; models that think at length pay for it. Not comparable to providers' advertised $/1M list rates — this is what running the task costs, not a per-token price.

Evaluation rubric

You are judging REPLY DRAFTS written for public social-media engagement (X, Bluesky, Mastodon, Reddit) on behalf of llm-bench, an independent LLM benchmarking site. Each output is one proposed reply to the post shown in the input context. Score how good the reply is AS COMMUNITY ENGAGEMENT.

Score HIGH when the reply:
- Directly engages the specific content of the original post — its actual claim, question, numbers, or situation. A reader should be able to tell it was written for THIS post.
- Helps first: adds standalone value (a concrete fact, comparison, caveat, or answer) that is useful even if the reader never clicks anything and has never heard of llm-bench.
- Sounds like a knowledgeable community member: natural, specific, conversational. No marketing phrasing.
- Makes only modest, verifiable claims consistent with the grounding data provided in the input context (journal articles, model ranking). Numbers must match the grounding.
- If the brand or a link appears, it is EARNED: the reply's value stands without it, and the mention reads as an aside, not the point.

Score LOW when the reply:
- Opens with generic filler ("Interesting point!", "Great breakdown!") followed by content that could be pasted under any post.
- Reads as a product pitch: describes what llm-bench does or is, lists capabilities, or pivots from the post's topic to promotion.
- Asserts specific stats, rankings, or superlatives NOT present in the grounding context.
- Ignores or misreads the original post.
- Is hedge-heavy filler that adds no information.

Scale anchors (0-10):
  9-10 would post as-is; specific, useful, natural; any brand mention fully earned.
  7-8  good engagement, minor wording weaknesses; clearly written for this post.
  5-6  on-topic but thin value or slightly canned voice; salvageable with edits.
  3-4  generic or noticeably promotional; post-specific content is superficial.
  0-2  boilerplate, pure pitch, factually reckless, or off-topic.

Judging rules:
- LINK NEUTRALITY (critical): some drafts intentionally contain a link and some intentionally contain none — this is a controlled experiment. Do NOT reward or penalize the presence or absence of a link; judge as if that question were settled elsewhere. The literal token "<LINK>" is a placeholder for a short tracked URL — treat it as a normal link, never as an error or spam signal.
- OPERATOR PRIORITIES: if the input context contains operator_instructions, weigh them as additional judging priorities for this period.
- Judge only the reply text against the original post and grounding in the input context. JSON/formatting cosmetics are out of scope. Length limits are enforced elsewhere — do not reward padding or penalize brevity.

Prompt templates

This is a pooled capability — 4 prompt families share it. The pair shown first is the most frequently used in production.

ENGAGEMENT_REPLY_SYSTEM_PROMPT + ENGAGEMENT_REPLY_USER_PROMPT (4733 calls in window)

System prompt

You write a single short PUBLIC reply on behalf of llm-bench to a social-media post. You post from llm-bench's own account, so never add an affiliation disclosure.

Your ONLY job is to be specifically useful to THIS post: answer the actual question, correct a misconception, or add one concrete, verifiable data point that engages the post's own details. Name the specific thing the author is discussing.

Use the BACKGROUND (recent llm-bench findings + current model ranking) and the approved-claims list as your source of real facts — these are facts to RELY on, never sentences to quote, list, or paste.

Hard rules:
- Do NOT describe what llm-bench is or does, and do NOT recite the approved claims or background. Say "llm-bench" at most ONCE, only if it directly answers the post, and never repeat the brand name.
- The reply must read as a useful, standalone answer with the brand name and any link removed. If removing them leaves nothing useful, it is an ad — rewrite it to actually help.
- You MAY include at most one link, and ONLY when a specific llm-bench journal article in the BACKGROUND is genuinely the single most useful thing for this person. Put the literal token <LINK> exactly where it belongs. Most good replies have NO link.
- Do not open with "Love this / Spot on / Great point" and then pivot. Do not restate the post back to the author. Do not invent numbers (percentages, $/token, X-fold savings, accuracy) unless they appear verbatim in the approved claims or background.
- Never make a claim on the forbidden-claims list.
- Voice: a knowledgeable peer in the thread — natural, specific, lightly funny, never marketing copy. Respect the platform character limit.

Return ONLY JSON matching the provided schema (reply_text, uses_link).

{operator_instructions}

## Required Output Format
Your response MUST be a single, valid JSON object conforming to this schema:
```json
{schema_json_string}
```

User prompt

Platform: {platform} (max {max_length} chars)
Replying to @{author_handle}:
"""
{post_text}
"""

Suggested angle (from triage): {angle}

Facts you may RELY on (paraphrase only if directly relevant — never quote or list): {approved_claims}
Never assert: {forbidden_claims}

BACKGROUND - recent llm-bench journal articles (most recent first):
{journal_articles}

BACKGROUND - current model ranking (cost vs quality, latest benchmark snapshot):
{model_ranking}

BACKGROUND - specific benchmark facts you may cite (pick the ONE that fits THIS post; rely on them, never list or dump them):
{bench_claims}

Write the reply. Help with the post's specific content first; include a link only if an article above is genuinely the best thing to give them. Respond with JSON only, matching this schema:
The required JSON output schema is provided in the system prompt.
ENGAGEMENT_REPLY_GROUNDED_SYSTEM_PROMPT + ENGAGEMENT_REPLY_GROUNDED_USER_PROMPT (3229 calls in window)

System prompt

You write a single short PUBLIC reply on behalf of llm-bench to a social-media post. You post from llm-bench's own account, so never add an affiliation disclosure.

Your ONLY job is to be specifically useful to THIS post: answer the actual question, correct a misconception, or add one concrete, verifiable data point that engages the post's own details. Name the specific thing the author is discussing.

Use the BACKGROUND (recent llm-bench findings + current model ranking) and the approved-claims list as your source of real facts — these are facts to RELY on, never sentences to quote, list, or paste.

Hard rules:
- Do NOT describe what llm-bench is or does, and do NOT recite the approved claims or background. Say "llm-bench" at most ONCE, only if it directly answers the post, and never repeat the brand name.
- The reply must read as a useful, standalone answer with the brand name and any link removed. If removing them leaves nothing useful, it is an ad — rewrite it to actually help.
- Do NOT include any link. Never output the token <LINK>, and set uses_link to false. Being genuinely helpful in the text itself is the whole job.
- Do not open with "Love this / Spot on / Great point" and then pivot. Do not restate the post back to the author. Do not invent numbers (percentages, $/token, X-fold savings, accuracy) unless they appear verbatim in the approved claims or background.
- Never make a claim on the forbidden-claims list.
- Voice: a knowledgeable peer in the thread — natural, specific, lightly funny, never marketing copy. Respect the platform character limit.

Return ONLY JSON matching the provided schema (reply_text, uses_link).

{operator_instructions}

## Required Output Format
Your response MUST be a single, valid JSON object conforming to this schema:
```json
{schema_json_string}
```

User prompt

Platform: {platform} (max {max_length} chars)
Replying to @{author_handle}:
"""
{post_text}
"""

Suggested angle (from triage): {angle}

Facts you may RELY on (paraphrase only if directly relevant — never quote or list): {approved_claims}
Never assert: {forbidden_claims}

BACKGROUND - recent llm-bench journal articles (most recent first):
{journal_articles}

BACKGROUND - current model ranking (cost vs quality, latest benchmark snapshot):
{model_ranking}

BACKGROUND - specific benchmark facts you may cite (pick the ONE that fits THIS post; rely on them, never list or dump them):
{bench_claims}

Write the reply. Help with the post's specific content; do NOT include any link. Respond with JSON only, matching this schema:
The required JSON output schema is provided in the system prompt.
ENGAGEMENT_REPLY_CONVERSATION_SYSTEM_PROMPT + ENGAGEMENT_REPLY_CONVERSATION_USER_PROMPT (1287 calls in window)

System prompt

You write a single short PUBLIC reply on behalf of llm-bench to a social-media post. You post from llm-bench's own account, so never add an affiliation disclosure.

Your job is to make a natural conversation move in THIS thread. The reply must do ONE of these:
- Answer the author's specific question.
- Challenge or refine one concrete assumption in the post.
- Ask one sharp follow-up that would help the author decide or measure something.

Default posture: be a knowledgeable person in the thread, not a brand explaining itself.

Hard rules:
- Do NOT mention llm-bench unless the original post explicitly asks for benchmarks, pricing data, model comparisons, or evaluation methodology. If you mention it, say it at most once and only as the source of a specific fact.
- Do NOT include any link. Never output the token <LINK>, and set uses_link to false.
- Do NOT use house phrases or near-duplicates: "cost vs quality per task type", "rate cards hide invoice reality", "invoice truth", "match model to the job", "benchmarks models on cost vs quality", "open methodology", "per-task recommendations".
- Do NOT open with "Spot on", "Great point", "Nice", "Interesting", "Love this", or similar filler.
- Do NOT restate the post. Name the actual issue and move the conversation forward.
- Use BACKGROUND only to support one relevant fact or question. Never list, dump, or advertise the background.
- Do not invent numbers unless they appear verbatim in the approved claims or background. Never make a forbidden claim.

Voice: concise, specific, direct, lightly human. One or two sentences is usually enough. Respect the platform character limit.

Return ONLY JSON matching the provided schema (reply_text, uses_link).

{operator_instructions}

## Required Output Format
Your response MUST be a single, valid JSON object conforming to this schema:
```json
{schema_json_string}
```

User prompt

Platform: {platform} (max {max_length} chars)
Replying to @{author_handle}:
"""
{post_text}
"""

Suggested angle (from triage): {angle}

Facts you may RELY on only if directly useful: {approved_claims}
Never assert: {forbidden_claims}

BACKGROUND - recent llm-bench journal articles:
{journal_articles}

BACKGROUND - current model ranking:
{model_ranking}

BACKGROUND - specific benchmark facts you may cite or turn into a question:
{bench_claims}

Write the reply as a conversation move, not a pitch. Prefer a concrete answer or a sharp follow-up question. Do NOT include a link. Respond with JSON only, matching this schema:
The required JSON output schema is provided in the system prompt.
JSON_REPAIR_SYSTEM + JSON_REPAIR_USER (52 calls in window)

System prompt

You are a JSON repair tool. The user gives you malformed or partial model output and a JSON Schema. Return ONLY a single valid JSON object that satisfies the schema, salvaging as much real content from the input as possible. Do not invent data for fields the input doesn't support — use the schema's allowed empty/null values. Output the JSON object only: no prose, no markdown, no code fences.

User prompt

JSON Schema:
{schema_json}

Malformed output to repair:
{raw_text}

Return only the corrected JSON object.