Best LLMs for Social Post Portfolio Selection
Selects the best fixed-size portfolio of candidate social posts using quality, expected audience value, topical diversity, timeliness, and caller-supplied constraints. Illustrative uses include selecting a balanced portfolio of customer updates, product messages, developer-commun
Models
Frontier on this task: DeepSeek V4 Flash at 8.78 / 10. Quality bar at 90%: 7.90.
point-estimate floor (CI low) · upper CI (less certain) · Bars sorted by blended cost; best-value model first. Greyed rows are MEDIUM+ models whose point estimate clears the bar but whose CI low does not.
| Model | Quality score | CI low | Cost / 1k runs | vs best value |
|---|---|---|---|---|
| Tencent Hy3 | 8.09 / 10 | 7.91 | $0.22 | best value |
| GPT-5.6 Luna | 8.46 / 10 | 8.32 | $0.23 | 1.1x more expensive |
| Gemini 3.5 Flash Lite | 8.60 / 10 | 8.35 | $0.30 | 1.4x more expensive |
| NVIDIA Nemotron-3 Nano 30B-A3B | 8.35 / 10 | 7.94 | $0.32 | 1.5x more expensive |
| Gemini 3.1 Flash Lite | 8.50 / 10 | 8.25 | $0.33 | 1.5x more expensive |
| MiniMax M3 | 8.40 / 10 | 8.33 | $0.87 | 3.9x more expensive |
| DeepSeek V4 Flash | 8.78 / 10 | 8.65 | $1.39 | 6.2x more expensive |
| Claude Haiku 4.5 | 8.06 / 10 | 7.67 | $1.57 | 7x more expensive |
| GLM-5.3 Flash | 8.70 / 10 | 8.53 | $1.78 | 8x more expensive |
| NVIDIA Nemotron 3.5 Lightning | 8.08 / 10 | 7.75 | $1.94 | 8.7x more expensive |
| GPT-5.6 Terra | 8.57 / 10 | 8.43 | $2.13 | 9.6x more expensive |
| NVIDIA Nemotron-3 Ultra 550B | 8.33 / 10 | 7.98 | $2.15 | 9.6x more expensive |
| NVIDIA Nemotron-3 Super 120B | 8.10 / 10 | 7.64 | $3.23 | 14x more expensive |
| DeepSeek V4 Pro | 8.47 / 10 | 8.22 | $3.23 | 14x more expensive |
| Qwen 3.8 Flash | 8.31 / 10 | 8.05 | $3.37 | 15x more expensive |
| Gemini 3.8 Flash | 8.72 / 10 | 8.49 | $3.63 | 16x more expensive |
| GPT-5.6 Sol | 8.48 / 10 | 8.34 | $3.93 | 18x more expensive |
| Claude Sonnet 5 | 8.33 / 10 | 8.21 | $4.09 | 18x more expensive |
| Thinking Machines Inkling Small | 8.74 / 10 | 8.52 | $4.58 | 21x more expensive |
| Qwen 3.7 Plus | 8.33 / 10 | 8.16 | $5.17 | 23x more expensive |
| Gemini 3.5 Flash | 8.49 / 10 | 8.39 | $7.79 | 35x more expensive |
| Claude Opus 5 | 8.64 / 10 | 8.37 | $10.71 | 48x more expensive |
| Meta Muse Spark 1.3 | 8.73 / 10 | 8.56 | $12.85 | 58x more expensive |
| GLM-5.3 | 8.64 / 10 | 8.42 | $14.21 | 64x more expensive |
| Thinking Machines Inkling | 8.58 / 10 | 8.31 | $19.72 | 88x more expensive |
| Qwen 3.8 Max | 8.17 / 10 | 7.94 | $30.07 | 135x more expensive |
| Tencent Hy4 Preview | 8.66 / 10 | 8.46 | $30.76 | 138x more expensive |
| Grok 4.6 | 8.66 / 10 | 8.42 | $34.90 | 157x more expensive |
| Moonshot Kimi K3 | 8.66 / 10 | 8.42 | $39.14 | 176x more expensive |
| GPT-5.4 Nano | 7.71 / 10 | 7.36 | $0.23 | 1x more expensive |
Cost breakdown
| Model | Quality | Confidence | Cost / 1k runs | Overpay | Mode |
|---|---|---|---|---|---|
| Tencent Hy3 ★ OpenRouter | 8.09 / 10 CI [7.91, 8.27] | RANKED | $0.22 | best value | batch |
| GPT-5.6 Luna OpenAI | 8.46 / 10 CI [8.32, 8.61] | RANKED | $0.23 | 1.1x | batch |
| Gemini 3.5 Flash Lite Gemini | 8.60 / 10 CI [8.35, 8.86] | HIGH | $0.30 | 1.4x | batch |
| NVIDIA Nemotron-3 Nano 30B-A3B OpenRouter | 8.35 / 10 CI [7.94, 8.77] | MEDIUM | $0.32 | 1.5x | batch |
| Gemini 3.1 Flash Lite Gemini | 8.50 / 10 CI [8.25, 8.75] | HIGH | $0.33 | 1.5x | batch |
| MiniMax M3 OpenRouter | 8.40 / 10 CI [8.33, 8.48] | RANKED | $0.87 | 3.9x | batch |
| DeepSeek V4 Flash best DeepSeek | 8.78 / 10 CI [8.65, 8.90] | RANKED | $1.39 | 6.2x | batch |
| Claude Haiku 4.5 Anthropic | 8.06 / 10 CI [7.67, 8.45] | MEDIUM | $1.57 | 7x | batch |
| GLM-5.3 Flash Z.AI | 8.70 / 10 CI [8.53, 8.87] | RANKED | $1.78 | 8x | batch |
| NVIDIA Nemotron 3.5 Lightning OpenRouter | 8.08 / 10 CI [7.75, 8.41] | MEDIUM | $1.94 | 8.7x | batch |
| GPT-5.6 Terra OpenAI | 8.57 / 10 CI [8.43, 8.71] | RANKED | $2.13 | 9.6x | batch |
| NVIDIA Nemotron-3 Ultra 550B OpenRouter | 8.33 / 10 CI [7.98, 8.68] | MEDIUM | $2.15 | 9.6x | batch |
| NVIDIA Nemotron-3 Super 120B OpenRouter | 8.10 / 10 CI [7.64, 8.57] | MEDIUM | $3.23 | 14x | batch |
| DeepSeek V4 Pro DeepSeek | 8.47 / 10 CI [8.22, 8.72] | HIGH | $3.23 | 14x | batch |
| Qwen 3.8 Flash Alibaba Cloud (DashScope) | 8.31 / 10 CI [8.05, 8.57] | HIGH | $3.37 | 15x | batch |
| Gemini 3.8 Flash Gemini | 8.72 / 10 CI [8.49, 8.95] | HIGH | $3.63 | 16x | batch |
| GPT-5.6 Sol OpenAI | 8.48 / 10 CI [8.34, 8.63] | RANKED | $3.93 | 18x | batch |
| Claude Sonnet 5 Anthropic | 8.33 / 10 CI [8.21, 8.46] | RANKED | $4.09 | 18x | batch |
| Thinking Machines Inkling Small OpenRouter | 8.74 / 10 CI [8.52, 8.97] | HIGH | $4.58 | 21x | batch |
| Qwen 3.7 Plus Alibaba Cloud (DashScope) | 8.33 / 10 CI [8.16, 8.49] | RANKED | $5.17 | 23x | batch |
| Gemini 3.5 Flash Gemini | 8.49 / 10 CI [8.39, 8.59] | RANKED | $7.79 | 35x | batch |
| Claude Opus 5 Anthropic | 8.64 / 10 CI [8.37, 8.91] | HIGH | $10.71 | 48x | batch |
| Meta Muse Spark 1.3 OpenRouter | 8.73 / 10 CI [8.56, 8.91] | RANKED | $12.85 | 58x | batch |
| GLM-5.3 Z.AI | 8.64 / 10 CI [8.42, 8.86] | HIGH | $14.21 | 64x | batch |
| Thinking Machines Inkling OpenRouter | 8.58 / 10 CI [8.31, 8.85] | HIGH | $19.72 | 88x | batch |
| Qwen 3.8 Max Alibaba Cloud (DashScope) | 8.17 / 10 CI [7.94, 8.41] | HIGH | $30.07 | 135x | batch |
| Tencent Hy4 Preview OpenRouter | 8.66 / 10 CI [8.46, 8.85] | RANKED | $30.76 | 138x | batch |
| Grok 4.6 xAI | 8.66 / 10 CI [8.42, 8.90] | HIGH | $34.90 | 157x | batch |
| Moonshot Kimi K3 Moonshot AI | 8.66 / 10 CI [8.42, 8.90] | HIGH | $39.14 | 176x | batch |
Overpay shows how much more you pay than the best-value model that clears the quality bar (marked ★) — the best-value good-enough option. "16x" means you overpay 16× — 16× that reference for no quality benefit above the bar. Typical call shape for this task: 1980 input tokens → 3091 output tokens, EMA-tracked from production traffic. Cost is the observed, all-in $ per 1,000 task runs: each model's own measured usage on this task — output verbosity, thinking/reasoning tokens, cache reads and writes, and the spend on its billed failures — priced at current list rates and adjusted by the billing overhead we actually reconcile against provider invoices. Models that answer tersely cost what they actually cost; models that think at length pay for it. Not comparable to providers' advertised $/1M list rates — this is what running the task costs, not a per-token price.
Evaluation rubric
Judge selected-item quality, portfolio diversity, timeliness, relevance, avoidance of redundancy, identifier fidelity, and consistency with the requested count and constraints.
Prompt templates
The system + user template pair used for this task.
X_POST_SELECTION_SYSTEM_PROMPT +
X_POST_SELECTION_USER_PROMPT
(426 calls in window)
System prompt
You are a social media strategist selecting the best X (Twitter) posts to publish within a daily budget.
Evaluate each candidate post and select exactly the requested number of posts (specified in the user message) that will maximize overall engagement and audience value.
Selection criteria (in order of importance):
1. Engagement potential — posts likely to generate clicks, replies, retweets
2. Topic diversity — avoid selecting multiple posts about the same topic
3. Content quality — clear, compelling, well-written posts
4. Timeliness — prefer posts about recent or trending topics
Return only the IDs of the selected posts in the specified JSON format.
## Required Output Format
Your response MUST be a single, valid JSON object conforming to this schema:
```json
{schema_json_string}
```User prompt
Select the best {select_count} posts from the {total_count} candidates below.
Candidates:
{candidates_text}
The required JSON output schema is provided in the system prompt.