Best LLMs for Social & Promotional Content
Conciseness, platform-native conventions, engagement under tight character limits.
Conciseness, platform-native conventions, engagement under tight character limits.
Task-by-task breakdown
X.com Promotional Post Generation
Generates a 20-tweet X.com promotional campaign for a research report: 10 pre-release tweets building anticipation (no links, questions/teasers/urgency) plus 10 post-release tweets driving …
| Model | Quality (% of best) | Confidence | Overpay |
|---|---|---|---|
| MiniMax M3 ★ | 93% | RANKED | best value |
| Gemini 3.5 Flash best | 100% | RANKED | 7.2x |
Activity Feed Blurb Generation
Writes a short, editorial-tone activity-feed blurb (1-2 sentences, under 200 chars) promoting a newly published article. Spark curiosity, no clickbait, no exclamation marks.
| Model | Quality (% of best) | Confidence | Overpay |
|---|---|---|---|
| DeepSeek V4 Flash ★ | 95% | MEDIUM | best value |
| NVIDIA Nemotron-3 Nano 30B-A3B | 91% | HIGH | 1.1x |
| GPT-5.4 Mini | 98% | RANKED | 1.9x |
| GPT-5.6 Luna | 95% | RANKED | 2.8x |
| NVIDIA Nemotron-3 Super 120B | 94% | HIGH | 3.1x |
| MiniMax M3 | 95% | MEDIUM | 3.2x |
| Tencent Hy3 best | 100% | RANKED | 4.4x |
| DeepSeek V4 Pro | 97% | HIGH | 4.5x |
| GPT-5.6 Terra | 95% | RANKED | 6.7x |
| Qwen 3.5 Flash | 95% | HIGH | 9.1x |
| GPT-5.6 Sol | 93% | RANKED | 16x |
| Gemini 3.1 Pro Preview | 93% | RANKED | 18x |
| NVIDIA Nemotron-3 Ultra 550B | 99% | HIGH | 18x |
| Claude Sonnet 4.6 | 93% | MEDIUM | 18x |
| Gemini 3.5 Flash | 91% | RANKED | 21x |
| Qwen 3.7 Plus | 95% | RANKED | 21x |
| GPT-5.5 | 98% | RANKED | 22x |
| Qwen 3.6 Flash | 90% | RANKED | 24x |
| Claude Sonnet 5 | 99% | RANKED | 25x |
| Qwen 3.6 Plus | 97% | RANKED | 28x |
| Claude Opus 4.8 | 100% | RANKED | 35x |
| Meta Muse Spark 1.1 | 96% | RANKED | 37x |
| Kimi K2.6 | 98% | RANKED | 44x |
| Grok 4.5 | 90% | RANKED | 52x |
Social Post Promotion
Pooled TT for single-platform article promo posts (X, Bluesky, Mastodon, Threads). Same prompt skeleton, per-platform style addendum and char limit.
| Model | Quality (% of best) | Confidence | Overpay |
|---|---|---|---|
| DeepSeek V4 Flash ★ | 90% | RANKED | best value |
| GPT-5.6 Luna | 96% | RANKED | 3x |
| GPT-5.6 Terra | 97% | RANKED | 4.9x |
| Qwen 3.5 Flash | 95% | RANKED | 5.6x |
| DeepSeek V4 Pro | 92% | HIGH | 5.7x |
| Claude Sonnet 5 | 93% | RANKED | 7.8x |
| Claude Sonnet 4.6 | 90% | HIGH | 9.9x |
| Qwen 3.7 Plus | 91% | HIGH | 10x |
| GPT-5.5 | 92% | HIGH | 13x |
| Gemini 3.5 Flash best | 100% | RANKED | 13x |
| GPT-5.6 Sol | 97% | HIGH | 14x |
| Qwen 3.6 Plus | 92% | HIGH | 16x |
| Kimi K2.6 | 93% | HIGH | 20x |
| Claude Opus 4.8 | 96% | HIGH | 23x |
| Meta Muse Spark 1.1 | 99% | RANKED | 24x |
| Grok 4.5 | 91% | RANKED | 38x |
Reddit Post Generation
Decides per target subreddit whether a newly published report is worth posting and, if so, writes a full self-post tailored to that community's rules, tone, and audience. Skips when the topic is …
| Model | Quality (% of best) | Confidence | Overpay |
|---|---|---|---|
| MiniMax M3 ★ | 91% | RANKED | best value |
| DeepSeek V4 Pro | 92% | MEDIUM | 2.7x |
| Gemini 3.5 Flash best | 100% | RANKED | 6.2x |
| Grok 4.5 | 94% | RANKED | 8.2x |
| Kimi K2.6 | 95% | HIGH | 12x |
| Claude Sonnet 5 | 92% | RANKED | 20x |
Engagement Reply Draft
Draft a public reply for an engage-decision candidate.
| Model | Quality (% of best) | Confidence | Overpay |
|---|---|---|---|
| Tencent Hy3 ★ | 92% | MEDIUM | best value |
| GPT-5.6 Luna | 97% | RANKED | 1.7x |
| GPT-5.6 Terra | 98% | HIGH | 2.6x |
| Qwen 3.7 Plus | 90% | MEDIUM | 4.8x |
| Gemini 3.5 Flash | 92% | HIGH | 6.5x |
| GPT-5.6 Sol best | 100% | RANKED | 7.5x |
| Gemini 3.1 Pro Preview | 94% | HIGH | 7.7x |
| GPT-5.5 | 98% | HIGH | 8.8x |
| Claude Opus 4.8 | 91% | MEDIUM | 9.1x |
| Meta Muse Spark 1.1 | 92% | MEDIUM | 11x |
| Grok 4.5 | 92% | HIGH | 17x |
| Kimi K2.6 | 95% | HIGH | 26x |
Confidence — how sure we are about the quality score (more judgments + more agreement = higher confidence): RANKED many independent judges scored this model's outputs and their agreement is very high (most confident) — HIGH many judges have scored it and they mostly agree (well-pinned) — MEDIUM enough judges have weighed in to publish, but they disagree more than we'd like (treat with a small grain of salt). LOW-confidence cells are hidden everywhere on the site. See the methodology for the exact thresholds.