Best LLMs for Report Image Generation
Generates a cover/social image from a report-derived image prompt. Wraps DALL-E / Imagen / Stable Diffusion / etc. behind a single task-type for batched dispatch.
Models
Frontier on this task: Gemini 3 Pro Image Preview at 8.15 / 10. Quality bar at 90%: 7.33.
point-estimate floor (CI low) · upper CI (less certain) · Bars sorted by blended cost; best-value model first.
| Model | Quality score | CI low | Cost / 1k runs | vs best value |
|---|---|---|---|---|
| Gemini 3.1 Flash Image Preview | 7.53 / 10 | 7.30 | $57.27 | best value |
| GPT-image-2 | 7.84 / 10 | 7.73 | $88.00 | 1.5x more expensive |
| Gemini 3 Pro Image Preview | 8.15 / 10 | 7.77 | $91.05 | 1.6x more expensive |
Cost breakdown
| Model | Quality | Confidence | Cost / 1k runs | Overpay | Mode |
|---|---|---|---|---|---|
| Gemini 3.1 Flash Image Preview ★ Gemini | 7.53 / 10 CI [7.30, 7.76] | HIGH | $57.27 | best value | batch |
| GPT-image-2 OpenAI | 7.84 / 10 CI [7.73, 7.96] | RANKED | $88.00 | 1.5x | batch |
| Gemini 3 Pro Image Preview best Gemini | 8.15 / 10 CI [7.77, 8.53] | MEDIUM | $91.05 | 1.6x | batch |
Overpay shows how much more you pay than the best-value model that clears the quality bar (marked ★) — the best-value good-enough option. "16x" means you overpay 16× — 16× that reference for no quality benefit above the bar. Typical call shape for this task: 196 input tokens → 2624 output tokens, EMA-tracked from production traffic. Cost is the observed, all-in $ per 1,000 task runs: each model's own measured usage on this task — output verbosity, thinking/reasoning tokens, cache reads and writes, and the spend on its billed failures — priced at current list rates and adjusted by the billing overhead we actually reconcile against provider invoices. Models that answer tersely cost what they actually cost; models that think at length pay for it. Not comparable to providers' advertised $/1M list rates — this is what running the task costs, not a per-token price.
Evaluation rubric
Grade the generated publication image (report cover / author portrait / topic card). Weigh, in order: 1. Prompt adherence — the image depicts the requested subject, scene, and mood. 2. Visual quality — coherent composition; no distortions, garbled shapes, duplicated limbs/fingers, or compression artifacts; adequate resolution. 3. Text legibility — any embedded text is spelled correctly and readable; penalize garbled or nonsensical lettering heavily. 4. Editorial fitness — appropriate, on-brand, and inoffensive for a finance/news publication; no watermarks, stock-photo overlays, or unsafe content. Score 0.0 (unusable) to 1.0 (excellent). Be neutral on subjective artistic style unless it harms the intended use.