Cost mode:

Category: Infrastructure & Utility · Rail: absolute · Typical I/O: 196→2624 tokens

Models

Frontier on this task: Gemini 3 Pro Image Preview at 8.15 / 10. Quality bar at 90%: 7.33.

point-estimate floor (CI low) · upper CI (less certain) · Bars sorted by blended cost; best-value model first.

ModelQuality scoreCI lowCost / 1k runsvs best value
Gemini 3.1 Flash Image Preview7.53 / 107.30$57.27best value
GPT-image-27.84 / 107.73$88.001.5x more expensive
Gemini 3 Pro Image Preview8.15 / 107.77$91.051.6x more expensive

Cost breakdown

ModelQualityConfidenceCost / 1k runsOverpayMode
Gemini 3.1 Flash Image Preview Gemini7.53 / 10 CI [7.30, 7.76]HIGH$57.27best valuebatch
GPT-image-2 OpenAI7.84 / 10 CI [7.73, 7.96]RANKED$88.001.5xbatch
Gemini 3 Pro Image Preview best Gemini8.15 / 10 CI [7.77, 8.53]MEDIUM$91.051.6xbatch

Overpay shows how much more you pay than the best-value model that clears the quality bar (marked ★) — the best-value good-enough option. "16x" means you overpay 16× — 16× that reference for no quality benefit above the bar. Typical call shape for this task: 196 input tokens → 2624 output tokens, EMA-tracked from production traffic. Cost is the observed, all-in $ per 1,000 task runs: each model's own measured usage on this task — output verbosity, thinking/reasoning tokens, cache reads and writes, and the spend on its billed failures — priced at current list rates and adjusted by the billing overhead we actually reconcile against provider invoices. Models that answer tersely cost what they actually cost; models that think at length pay for it. Not comparable to providers' advertised $/1M list rates — this is what running the task costs, not a per-token price.

Evaluation rubric

Grade the generated publication image (report cover / author portrait / topic card). Weigh, in order:
1. Prompt adherence — the image depicts the requested subject, scene, and mood.
2. Visual quality — coherent composition; no distortions, garbled shapes, duplicated limbs/fingers, or compression artifacts; adequate resolution.
3. Text legibility — any embedded text is spelled correctly and readable; penalize garbled or nonsensical lettering heavily.
4. Editorial fitness — appropriate, on-brand, and inoffensive for a finance/news publication; no watermarks, stock-photo overlays, or unsafe content.
Score 0.0 (unusable) to 1.0 (excellent). Be neutral on subjective artistic style unless it harms the intended use.