Cost mode:

Category: Social & Promotional Content · Rail: absolute · Typical I/O: 7559→5897 tokens

Models

Frontier on this task: Gemini 3.5 Flash at 8.64 / 10. Quality bar at 90%: 7.78.

point-estimate floor (CI low) · upper CI (less certain) · Bars sorted by blended cost; best-value model first. Greyed rows are MEDIUM+ models whose point estimate clears the bar but whose CI low does not.

ModelQuality scoreCI lowCost / 1k runsvs best value
MiniMax M37.78 / 107.65$4.10best value
GLM-5.3 Flash8.31 / 107.90$6.171.5x more expensive
Qwen 3.7 Plus7.91 / 107.69$10.742.6x more expensive
Thinking Machines Inkling Small7.85 / 107.46$16.844.1x more expensive
Gemini 3.5 Flash8.64 / 108.53$27.076.6x more expensive
Thinking Machines Inkling7.82 / 107.37$32.517.9x more expensive
Tencent Hy4 Preview7.83 / 107.33$40.8310x more expensive
DeepSeek V4 Pro8.08 / 107.76$48.5612x more expensive
Grok 4.68.27 / 107.88$63.3715x more expensive
GLM-5.38.36 / 107.96$80.4420x more expensive
Claude Sonnet 58.19 / 108.01$84.1021x more expensive
Moonshot Kimi K38.23 / 107.95$132.6632x more expensive
NVIDIA Nemotron-3 Super 120B7.09 / 106.79$5.491.3x more expensive
Gemini 3.8 Flash7.60 / 107.21$12.323x more expensive
GPT-5.6 Luna7.62 / 107.34$1.4964% cheaper
Tencent Hy37.06 / 106.82$2.8032% cheaper
GPT-5.6 Sol7.39 / 106.93$25.166.1x more expensive
Gemini 3.5 Flash Lite6.55 / 106.10$3.5015% cheaper
Qwen 3.8 Max7.39 / 106.89$90.6822x more expensive
Meta Muse Spark 1.37.48 / 107.03$26.076.4x more expensive
Claude Haiku 4.57.75 / 107.44$30.167.4x more expensive
DeepSeek V4 Flash7.72 / 107.40$12.633.1x more expensive
Gemini 3.1 Flash Lite6.55 / 106.22$3.4715% cheaper
GPT-5.4 Nano6.84 / 106.47$4.701.1x more expensive

Cost breakdown

ModelQualityConfidenceCost / 1k runsOverpayMode
MiniMax M3 OpenRouter7.78 / 10 CI [7.65, 7.91]RANKED$4.10best valuebatch
GLM-5.3 Flash Z.AI8.31 / 10 CI [7.90, 8.72]MEDIUM$6.171.5xbatch
Qwen 3.7 Plus Alibaba Cloud (DashScope)7.91 / 10 CI [7.69, 8.14]HIGH$10.742.6xbatch
Thinking Machines Inkling Small OpenRouter7.85 / 10 CI [7.46, 8.24]MEDIUM$16.844.1xbatch
Gemini 3.5 Flash best Gemini8.64 / 10 CI [8.53, 8.76]RANKED$27.076.6xbatch
Thinking Machines Inkling OpenRouter7.82 / 10 CI [7.37, 8.27]MEDIUM$32.517.9xbatch
Tencent Hy4 Preview OpenRouter7.83 / 10 CI [7.33, 8.32]MEDIUM$40.8310xbatch
DeepSeek V4 Pro DeepSeek8.08 / 10 CI [7.76, 8.40]MEDIUM$48.5612xbatch
Grok 4.6 xAI8.27 / 10 CI [7.88, 8.66]MEDIUM$63.3715xbatch
GLM-5.3 Z.AI8.36 / 10 CI [7.96, 8.76]MEDIUM$80.4420xbatch
Claude Sonnet 5 Anthropic8.19 / 10 CI [8.01, 8.36]RANKED$84.1021xbatch
Moonshot Kimi K3 Moonshot AI8.23 / 10 CI [7.95, 8.50]HIGH$132.6632xbatch

Overpay shows how much more you pay than the best-value model that clears the quality bar (marked ★) — the best-value good-enough option. "16x" means you overpay 16× — 16× that reference for no quality benefit above the bar. Typical call shape for this task: 7559 input tokens → 5897 output tokens, EMA-tracked from production traffic. Cost is the observed, all-in $ per 1,000 task runs: each model's own measured usage on this task — output verbosity, thinking/reasoning tokens, cache reads and writes, and the spend on its billed failures — priced at current list rates and adjusted by the billing overhead we actually reconcile against provider invoices. Models that answer tersely cost what they actually cost; models that think at length pay for it. Not comparable to providers' advertised $/1M list rates — this is what running the task costs, not a per-token price.

Evaluation rubric

Judge community relevance, rule compliance, source grounding, usefulness independent of promotion, tone fit, and sound skip decisions. Penalize disguised advertising, invented claims, generic cross-post copy, and attempts to evade explicit community restrictions.

Prompt templates

The system + user template pair used for this task.

AUTO_REDDIT_POST_SYSTEM_PROMPT + AUTO_REDDIT_POST_USER_PROMPT (575 calls in window)

System prompt

You are a knowledgeable market analyst sharing your own research on Reddit. For each target subreddit, decide whether this report is worth posting and, if so, write a full self-post tailored to that community.

{author_voice_section}

## Per-Subreddit Decision

For each subreddit, first decide should_post:
- **Post if**: report contains actionable analysis, breaking developments, unique data, or findings specifically relevant to this community
- **Skip if**: report has no new insights for this audience, topic is tangential, content is too generic, or the subreddit's rules would prohibit this type of content

## Writing Guidelines

For each subreddit where should_post=True, write a complete self-post:

**Title (max 300 chars):**
- Insight-driven, not clickbait
- Specific to the analysis findings
- Matches the subreddit's title conventions

**Body:**
- Write a standalone piece of content, not just a summary — extract and reframe the most relevant findings for this audience
- Stock-specific subreddits: focus on the ticker, use data points, technical/fundamental analysis language
- General investing subreddits: broader market context, more accessible language, why this matters beyond one stock
- Use Reddit markdown: headers (##), **bold**, bullets, horizontal rules (---)
- MUST include the literal placeholder <ghost_url> exactly once (link to full analysis)
- Vary framing, structure, and emphasis across subreddits — each post should feel native to its community
- No duplicate content across subreddits
- Style: authentic contributor sharing their own analysis, not promotional

**Image embedding:**
- If a report image URL is provided and non-empty, you may embed it once per post via ![image](url)
- Include when it adds genuine value for the community (e.g., visual data in data-friendly subs)
- Omit the image entirely (do not output ![image]() or the word "image") when the URL is blank, the image is generic, the subreddit is text-focused, or it would feel promotional

## Important
- Never reveal you are AI or automated
- Write as the analyst who conducted this research
- Each post must feel like it was written specifically for that community

## Required Output Format
Your response MUST be a single, valid JSON object conforming to this schema:
```json
{schema_json_string}
```

User prompt

Generate Reddit posts for each target subreddit. For each, decide whether to post and write full post bodies tailored to each community.

Stock: {subject_code} ({subject_name})
Report Title: {title}

Full Report:
{report_text}

Report Image URL (if non-empty, you may embed once via ![image](url) where it adds genuine value — omit entirely if blank, generic, or promotional):
{report_image_url}

Target Subreddits:
{subreddits_with_context_json}

The required JSON output schema is provided in the system prompt.