Cost mode:

Conciseness, platform-native conventions, engagement under tight character limits.

5 capabilities in this category.

Task-by-task breakdown

Activity Feed Blurb Generation

Writes a short, editorial-tone activity-feed blurb (1-2 sentences, under 200 chars) promoting a newly published article. Spark curiosity, no clickbait, no exclamation marks.

ModelQuality (% of best)ConfidenceOverpay
DeepSeek V4 Flash 95%MEDIUMbest value
NVIDIA Nemotron-3 Nano 30B-A3B91%HIGH1.1x
GPT-5.4 Mini98%RANKED1.9x
GPT-5.6 Luna95%RANKED2.8x
NVIDIA Nemotron-3 Super 120B94%HIGH3.1x
MiniMax M395%MEDIUM3.2x
Tencent Hy3 best100%RANKED4.4x
DeepSeek V4 Pro97%HIGH4.5x
GPT-5.6 Terra95%RANKED6.7x
Qwen 3.5 Flash95%HIGH9.1x
GPT-5.6 Sol93%RANKED16x
Gemini 3.1 Pro Preview93%RANKED18x
NVIDIA Nemotron-3 Ultra 550B99%HIGH18x
Claude Sonnet 4.693%MEDIUM18x
Gemini 3.5 Flash91%RANKED21x
Qwen 3.7 Plus95%RANKED21x
GPT-5.598%RANKED22x
Qwen 3.6 Flash90%RANKED24x
Claude Sonnet 599%RANKED25x
Qwen 3.6 Plus97%RANKED28x
Claude Opus 4.8100%RANKED35x
Meta Muse Spark 1.196%RANKED37x
Kimi K2.698%RANKED44x
Grok 4.590%RANKED52x

Task detail →

Social Post Promotion

Pooled TT for single-platform article promo posts (X, Bluesky, Mastodon, Threads). Same prompt skeleton, per-platform style addendum and char limit.

ModelQuality (% of best)ConfidenceOverpay
DeepSeek V4 Flash 90%RANKEDbest value
GPT-5.6 Luna96%RANKED3x
GPT-5.6 Terra97%RANKED4.9x
Qwen 3.5 Flash95%RANKED5.6x
DeepSeek V4 Pro92%HIGH5.7x
Claude Sonnet 593%RANKED7.8x
Claude Sonnet 4.690%HIGH9.9x
Qwen 3.7 Plus91%HIGH10x
GPT-5.592%HIGH13x
Gemini 3.5 Flash best100%RANKED13x
GPT-5.6 Sol97%HIGH14x
Qwen 3.6 Plus92%HIGH16x
Kimi K2.693%HIGH20x
Claude Opus 4.896%HIGH23x
Meta Muse Spark 1.199%RANKED24x
Grok 4.591%RANKED38x

Task detail →

Confidence — how sure we are about the quality score (more judgments + more agreement = higher confidence): RANKED many independent judges scored this model's outputs and their agreement is very high (most confident) — HIGH many judges have scored it and they mostly agree (well-pinned) — MEDIUM enough judges have weighed in to publish, but they disagree more than we'd like (treat with a small grain of salt). LOW-confidence cells are hidden everywhere on the site. See the methodology for the exact thresholds.