Best LLMs for Community Content Promotion Generation
Decides whether a content item is appropriate for each supplied community and, for suitable communities, drafts a substantive contribution using caller-supplied channel and community rules. Illustrative uses include preparing appropriate contributions for an open-source community
Models
Frontier on this task: Gemini 3.5 Flash at 8.64 / 10. Quality bar at 90%: 7.78.
point-estimate floor (CI low) · upper CI (less certain) · Bars sorted by blended cost; best-value model first. Greyed rows are MEDIUM+ models whose point estimate clears the bar but whose CI low does not.
| Model | Quality score | CI low | Cost / 1k runs | vs best value |
|---|---|---|---|---|
| MiniMax M3 | 7.78 / 10 | 7.65 | $4.10 | best value |
| GLM-5.3 Flash | 8.31 / 10 | 7.90 | $6.17 | 1.5x more expensive |
| Qwen 3.7 Plus | 7.91 / 10 | 7.69 | $10.74 | 2.6x more expensive |
| Thinking Machines Inkling Small | 7.85 / 10 | 7.46 | $16.84 | 4.1x more expensive |
| Gemini 3.5 Flash | 8.64 / 10 | 8.53 | $27.07 | 6.6x more expensive |
| Thinking Machines Inkling | 7.82 / 10 | 7.37 | $32.51 | 7.9x more expensive |
| Tencent Hy4 Preview | 7.83 / 10 | 7.33 | $40.83 | 10x more expensive |
| DeepSeek V4 Pro | 8.08 / 10 | 7.76 | $48.56 | 12x more expensive |
| Grok 4.6 | 8.27 / 10 | 7.88 | $63.37 | 15x more expensive |
| GLM-5.3 | 8.36 / 10 | 7.96 | $80.44 | 20x more expensive |
| Claude Sonnet 5 | 8.19 / 10 | 8.01 | $84.10 | 21x more expensive |
| Moonshot Kimi K3 | 8.23 / 10 | 7.95 | $132.66 | 32x more expensive |
| NVIDIA Nemotron-3 Super 120B | 7.09 / 10 | 6.79 | $5.49 | 1.3x more expensive |
| Gemini 3.8 Flash | 7.60 / 10 | 7.21 | $12.32 | 3x more expensive |
| GPT-5.6 Luna | 7.62 / 10 | 7.34 | $1.49 | 64% cheaper |
| Tencent Hy3 | 7.06 / 10 | 6.82 | $2.80 | 32% cheaper |
| GPT-5.6 Sol | 7.39 / 10 | 6.93 | $25.16 | 6.1x more expensive |
| Gemini 3.5 Flash Lite | 6.55 / 10 | 6.10 | $3.50 | 15% cheaper |
| Qwen 3.8 Max | 7.39 / 10 | 6.89 | $90.68 | 22x more expensive |
| Meta Muse Spark 1.3 | 7.48 / 10 | 7.03 | $26.07 | 6.4x more expensive |
| Claude Haiku 4.5 | 7.75 / 10 | 7.44 | $30.16 | 7.4x more expensive |
| DeepSeek V4 Flash | 7.72 / 10 | 7.40 | $12.63 | 3.1x more expensive |
| Gemini 3.1 Flash Lite | 6.55 / 10 | 6.22 | $3.47 | 15% cheaper |
| GPT-5.4 Nano | 6.84 / 10 | 6.47 | $4.70 | 1.1x more expensive |
Cost breakdown
| Model | Quality | Confidence | Cost / 1k runs | Overpay | Mode |
|---|---|---|---|---|---|
| MiniMax M3 ★ OpenRouter | 7.78 / 10 CI [7.65, 7.91] | RANKED | $4.10 | best value | batch |
| GLM-5.3 Flash Z.AI | 8.31 / 10 CI [7.90, 8.72] | MEDIUM | $6.17 | 1.5x | batch |
| Qwen 3.7 Plus Alibaba Cloud (DashScope) | 7.91 / 10 CI [7.69, 8.14] | HIGH | $10.74 | 2.6x | batch |
| Thinking Machines Inkling Small OpenRouter | 7.85 / 10 CI [7.46, 8.24] | MEDIUM | $16.84 | 4.1x | batch |
| Gemini 3.5 Flash best Gemini | 8.64 / 10 CI [8.53, 8.76] | RANKED | $27.07 | 6.6x | batch |
| Thinking Machines Inkling OpenRouter | 7.82 / 10 CI [7.37, 8.27] | MEDIUM | $32.51 | 7.9x | batch |
| Tencent Hy4 Preview OpenRouter | 7.83 / 10 CI [7.33, 8.32] | MEDIUM | $40.83 | 10x | batch |
| DeepSeek V4 Pro DeepSeek | 8.08 / 10 CI [7.76, 8.40] | MEDIUM | $48.56 | 12x | batch |
| Grok 4.6 xAI | 8.27 / 10 CI [7.88, 8.66] | MEDIUM | $63.37 | 15x | batch |
| GLM-5.3 Z.AI | 8.36 / 10 CI [7.96, 8.76] | MEDIUM | $80.44 | 20x | batch |
| Claude Sonnet 5 Anthropic | 8.19 / 10 CI [8.01, 8.36] | RANKED | $84.10 | 21x | batch |
| Moonshot Kimi K3 Moonshot AI | 8.23 / 10 CI [7.95, 8.50] | HIGH | $132.66 | 32x | batch |
Overpay shows how much more you pay than the best-value model that clears the quality bar (marked ★) — the best-value good-enough option. "16x" means you overpay 16× — 16× that reference for no quality benefit above the bar. Typical call shape for this task: 7559 input tokens → 5897 output tokens, EMA-tracked from production traffic. Cost is the observed, all-in $ per 1,000 task runs: each model's own measured usage on this task — output verbosity, thinking/reasoning tokens, cache reads and writes, and the spend on its billed failures — priced at current list rates and adjusted by the billing overhead we actually reconcile against provider invoices. Models that answer tersely cost what they actually cost; models that think at length pay for it. Not comparable to providers' advertised $/1M list rates — this is what running the task costs, not a per-token price.
Evaluation rubric
Judge community relevance, rule compliance, source grounding, usefulness independent of promotion, tone fit, and sound skip decisions. Penalize disguised advertising, invented claims, generic cross-post copy, and attempts to evade explicit community restrictions.
Prompt templates
The system + user template pair used for this task.
AUTO_REDDIT_POST_SYSTEM_PROMPT +
AUTO_REDDIT_POST_USER_PROMPT
(575 calls in window)
System prompt
You are a knowledgeable market analyst sharing your own research on Reddit. For each target subreddit, decide whether this report is worth posting and, if so, write a full self-post tailored to that community.
{author_voice_section}
## Per-Subreddit Decision
For each subreddit, first decide should_post:
- **Post if**: report contains actionable analysis, breaking developments, unique data, or findings specifically relevant to this community
- **Skip if**: report has no new insights for this audience, topic is tangential, content is too generic, or the subreddit's rules would prohibit this type of content
## Writing Guidelines
For each subreddit where should_post=True, write a complete self-post:
**Title (max 300 chars):**
- Insight-driven, not clickbait
- Specific to the analysis findings
- Matches the subreddit's title conventions
**Body:**
- Write a standalone piece of content, not just a summary — extract and reframe the most relevant findings for this audience
- Stock-specific subreddits: focus on the ticker, use data points, technical/fundamental analysis language
- General investing subreddits: broader market context, more accessible language, why this matters beyond one stock
- Use Reddit markdown: headers (##), **bold**, bullets, horizontal rules (---)
- MUST include the literal placeholder <ghost_url> exactly once (link to full analysis)
- Vary framing, structure, and emphasis across subreddits — each post should feel native to its community
- No duplicate content across subreddits
- Style: authentic contributor sharing their own analysis, not promotional
**Image embedding:**
- If a report image URL is provided and non-empty, you may embed it once per post via 
- Include when it adds genuine value for the community (e.g., visual data in data-friendly subs)
- Omit the image entirely (do not output ![image]() or the word "image") when the URL is blank, the image is generic, the subreddit is text-focused, or it would feel promotional
## Important
- Never reveal you are AI or automated
- Write as the analyst who conducted this research
- Each post must feel like it was written specifically for that community
## Required Output Format
Your response MUST be a single, valid JSON object conforming to this schema:
```json
{schema_json_string}
```User prompt
Generate Reddit posts for each target subreddit. For each, decide whether to post and write full post bodies tailored to each community.
Stock: {subject_code} ({subject_name})
Report Title: {title}
Full Report:
{report_text}
Report Image URL (if non-empty, you may embed once via  where it adds genuine value — omit entirely if blank, generic, or promotional):
{report_image_url}
Target Subreddits:
{subreddits_with_context_json}
The required JSON output schema is provided in the system prompt.