Cost mode:

Sustained compositional skill, voice consistency, coherent extended prose.

6 capabilities in this category.

Task-by-task breakdown

Reference Preserving Analytical Writing

Produces analytical prose from supplied claims and source material while preserving references according to a caller-supplied reference protocol. Illustrative uses include producing a cited vendor …

ModelQuality (% of best)ConfidenceOverpay
GPT-5.6 Luna 93%HIGHbest value
Gemini 3.8 Flash92%MEDIUM5.1x
Qwen 3.7 Plus91%HIGH5.3x
GPT-5.6 Terra95%HIGH5.8x
DeepSeek V4 Pro94%HIGH11x
Thinking Machines Inkling94%MEDIUM13x
GPT-5.6 Sol94%MEDIUM14x
Gemini 3.5 Flash93%HIGH17x
Grok 4.698%MEDIUM29x
Meta Muse Spark 1.396%MEDIUM33x
Moonshot Kimi K3 best100%RANKED53x
Claude Sonnet 597%RANKED90x

Task detail →

Visual Theme Configuration Generation

Produces an accessible color and typography configuration from visual references, audience context, and caller-supplied design-system roles and constraints. Illustrative uses include configuring the …

ModelQuality (% of best)ConfidenceOverpay
GPT-5.6 Luna 96%HIGHbest value
MiniMax M390%RANKED1.7x
NVIDIA Nemotron-3 Ultra 550B95%HIGH8.3x
Thinking Machines Inkling Small94%HIGH8.3x
GPT-5.6 Terra93%HIGH12x
GPT-5.6 Sol97%HIGH19x
Claude Sonnet 5 best100%RANKED30x

Task detail →

Voice Profile Generation

Produces a detailed, reusable voice and personality specification from identity context, background, expertise, interests, themes, audience, and writing-style guidance. Illustrative uses include …

ModelQuality (% of best)ConfidenceOverpay
GPT-5.6 Luna 92%RANKEDbest value
MiniMax M397%RANKED2.2x
GLM-5.3 Flash99%HIGH4.1x
Gemini 3.8 Flash95%HIGH4.5x
GPT-5.4 Nano90%RANKED4.6x
DeepSeek V4 Flash95%RANKED8.1x
Qwen 3.7 Plus92%RANKED8.3x
GPT-5.6 Terra91%RANKED11x
Claude Sonnet 591%RANKED11x
Gemini 3.5 Flash93%RANKED12x
Thinking Machines Inkling92%HIGH12x
Meta Muse Spark 1.398%HIGH13x
GPT-5.6 Sol93%RANKED15x
Claude Haiku 4.592%RANKED18x
Grok 4.6 best100%HIGH24x
GLM-5.398%HIGH31x
DeepSeek V4 Pro95%RANKED32x
Tencent Hy4 Preview99%HIGH34x
Qwen 3.8 Max97%HIGH78x

Task detail →

Confidence — how sure we are about the quality score (more judgments + more agreement = higher confidence): RANKED many independent judges scored this model's outputs and their agreement is very high (most confident) — HIGH many judges have scored it and they mostly agree (well-pinned) — MEDIUM enough judges have weighed in to publish, but they disagree more than we'd like (treat with a small grain of salt). LOW-confidence cells are hidden everywhere on the site. See the methodology for the exact thresholds.