Cost mode:

Category: Social & Promotional Content · Rail: absolute · Typical I/O: 1670→1099 tokens

Models

Frontier on this task: Tencent Hy3 at 8.81 / 10. Quality bar at 90%: 7.93.

point-estimate floor (CI low) · upper CI (less certain) · Bars sorted by blended cost; best-value model first. Greyed rows are MEDIUM+ models whose point estimate clears the bar but whose CI low does not.

ModelQuality scoreCI lowCost / 1k runsvs best value
DeepSeek V4 Flash8.38 / 108.07$0.28best value
NVIDIA Nemotron-3 Nano 30B-A3B8.00 / 107.71$0.301.1x more expensive
GPT-5.4 Mini8.64 / 108.47$0.521.9x more expensive
GPT-5.6 Luna8.36 / 108.21$0.782.8x more expensive
NVIDIA Nemotron-3 Super 120B8.30 / 108.06$0.853.1x more expensive
MiniMax M38.36 / 107.94$0.903.2x more expensive
Tencent Hy38.81 / 108.72$1.214.4x more expensive
DeepSeek V4 Pro8.58 / 108.33$1.244.5x more expensive
GPT-5.6 Terra8.38 / 108.24$1.866.7x more expensive
Qwen 3.5 Flash8.34 / 108.11$2.539.1x more expensive
GPT-5.6 Sol8.20 / 108.02$4.3316x more expensive
Gemini 3.1 Pro Preview8.18 / 108.02$4.9518x more expensive
NVIDIA Nemotron-3 Ultra 550B8.69 / 108.46$4.9618x more expensive
Claude Sonnet 4.68.19 / 107.80$5.0818x more expensive
Gemini 3.5 Flash8.03 / 107.85$5.8821x more expensive
Qwen 3.7 Plus8.33 / 108.18$5.9321x more expensive
GPT-5.58.65 / 108.47$6.0822x more expensive
Qwen 3.6 Flash7.94 / 107.74$6.7424x more expensive
Claude Sonnet 58.74 / 108.64$6.9825x more expensive
Qwen 3.6 Plus8.52 / 108.33$7.8528x more expensive
Claude Opus 4.88.76 / 108.57$9.6635x more expensive
Meta Muse Spark 1.18.47 / 108.29$10.1837x more expensive
Kimi K2.68.66 / 108.50$12.3344x more expensive
Grok 4.57.95 / 107.77$14.3752x more expensive
Gemini 3.1 Flash Lite7.64 / 107.38$0.351.2x more expensive

Cost breakdown

ModelQualityConfidenceCost / 1k runsOverpayMode
DeepSeek V4 Flash DeepSeek8.38 / 10 CI [8.07, 8.69]MEDIUM$0.28best valuebatch
NVIDIA Nemotron-3 Nano 30B-A3B OpenRouter8.00 / 10 CI [7.71, 8.29]HIGH$0.301.1xbatch
GPT-5.4 Mini OpenAI8.64 / 10 CI [8.47, 8.81]RANKED$0.521.9xbatch
GPT-5.6 Luna OpenAI8.36 / 10 CI [8.21, 8.50]RANKED$0.782.8xbatch
NVIDIA Nemotron-3 Super 120B OpenRouter8.30 / 10 CI [8.06, 8.54]HIGH$0.853.1xbatch
MiniMax M3 MiniMax8.36 / 10 CI [7.94, 8.78]MEDIUM$0.903.2xbatch
Tencent Hy3 best OpenRouter8.81 / 10 CI [8.72, 8.89]RANKED$1.214.4xbatch
DeepSeek V4 Pro DeepSeek8.58 / 10 CI [8.33, 8.83]HIGH$1.244.5xbatch
GPT-5.6 Terra OpenAI8.38 / 10 CI [8.24, 8.51]RANKED$1.866.7xbatch
Qwen 3.5 Flash Alibaba Cloud (DashScope)8.34 / 10 CI [8.11, 8.58]HIGH$2.539.1xbatch
GPT-5.6 Sol OpenAI8.20 / 10 CI [8.02, 8.38]RANKED$4.3316xbatch
Gemini 3.1 Pro Preview Gemini8.18 / 10 CI [8.02, 8.35]RANKED$4.9518xbatch
NVIDIA Nemotron-3 Ultra 550B OpenRouter8.69 / 10 CI [8.46, 8.92]HIGH$4.9618xbatch
Claude Sonnet 4.6 Anthropic8.19 / 10 CI [7.80, 8.59]MEDIUM$5.0818xbatch
Gemini 3.5 Flash Gemini8.03 / 10 CI [7.85, 8.22]RANKED$5.8821xbatch
Qwen 3.7 Plus Alibaba Cloud (DashScope)8.33 / 10 CI [8.18, 8.47]RANKED$5.9321xbatch
GPT-5.5 OpenAI8.65 / 10 CI [8.47, 8.84]RANKED$6.0822xbatch
Qwen 3.6 Flash Alibaba Cloud (DashScope)7.94 / 10 CI [7.74, 8.13]RANKED$6.7424xbatch
Claude Sonnet 5 Anthropic8.74 / 10 CI [8.64, 8.84]RANKED$6.9825xbatch
Qwen 3.6 Plus Alibaba Cloud (DashScope)8.52 / 10 CI [8.33, 8.71]RANKED$7.8528xbatch
Claude Opus 4.8 Anthropic8.76 / 10 CI [8.57, 8.95]RANKED$9.6635xbatch
Meta Muse Spark 1.1 Meta8.47 / 10 CI [8.29, 8.65]RANKED$10.1837xbatch
Kimi K2.6 Moonshot AI8.66 / 10 CI [8.50, 8.82]RANKED$12.3344xbatch
Grok 4.5 xAI7.95 / 10 CI [7.77, 8.13]RANKED$14.3752xbatch

Overpay shows how much more you pay than the best-value model that clears the quality bar (marked ★) — the best-value good-enough option. "16x" means you overpay 16× — 16× that reference for no quality benefit above the bar. Typical call shape for this task: 1670 input tokens → 1099 output tokens, EMA-tracked from production traffic. Cost is the observed, all-in $ per 1,000 task runs: each model's own measured usage on this task — output verbosity, thinking/reasoning tokens, cache reads and writes, and the spend on its billed failures — priced at current list rates and adjusted by the billing overhead we actually reconcile against provider invoices. Models that answer tersely cost what they actually cost; models that think at length pay for it. Not comparable to providers' advertised $/1M list rates — this is what running the task costs, not a per-token price.

Evaluation rubric

The output under evaluation is a short promotional blurb (1-2 sentences) for
a client's activity feed, generated from a published report's title and
content.

Judge ONLY the editorial quality of the blurb:
- Accuracy: it reflects the report's actual thesis; no invented facts,
  figures, or claims that the report does not support.
- Promotional fit: specific and engaging — it names or clearly evokes the
  client and gives a concrete reason to read the report; it reads as a hook,
  not a generic announcement.
- Style: clear, concise, professional finance/news voice; no clickbait
  phrasing, no filler.

OUT OF SCOPE — do not reward or penalize:
- Character/length-limit compliance. The length cap is enforced by a
  deterministic validator before any output is accepted, so every output you
  see already complies. Do NOT count characters or grade length.
- Output structure, schema, field naming, or formatting details.

Prompt templates

This is a pooled capability — 2 prompt families share it. The pair shown first is the most frequently used in production.

ACTIVITY_PROMO_SYSTEM_PROMPT + ACTIVITY_PROMO_USER_PROMPT (4741 calls in window)

System prompt

You are an expert copywriter for a financial research platform. Your task is to write a short, engaging promotional blurb for an activity feed.

The blurb should:
- Be 1-2 sentences, under 200 characters total
- Spark curiosity and encourage clicking through
- Reference the client/brand name naturally
- Feel like an editorial teaser, not an advertisement
- Avoid clickbait, hyperbole, or exclamation marks
- Use present tense

Output your result in the specified JSON format.

## Required Output Format
Your response MUST be a single, valid JSON object conforming to this schema:
```json
{schema_json_string}
```

User prompt

Client: {client_name}
Post title: {post_title}

Full report text:
{report_text}

Write a short promotional blurb for this post to appear in the activity feed of other clients' pages.

The required JSON output schema is provided in the system prompt.
JSON_REPAIR_SYSTEM + JSON_REPAIR_USER (9 calls in window)

System prompt

You are a JSON repair tool. The user gives you malformed or partial model output and a JSON Schema. Return ONLY a single valid JSON object that satisfies the schema, salvaging as much real content from the input as possible. Do not invent data for fields the input doesn't support — use the schema's allowed empty/null values. Output the JSON object only: no prose, no markdown, no code fences.

User prompt

JSON Schema:
{schema_json}

Malformed output to repair:
{raw_text}

Return only the corrected JSON object.