Cost mode:

Mechanical competence at format conversion, metadata manipulation, prompt rewriting, translation; minimal domain expertise required.

10 capabilities in this category.

Task-by-task breakdown

Markdown Newline Repair

Re-inserts blank lines into a markdown document whose line breaks were stripped (headings, separators, and paragraphs collapsed onto one line). Must not modify any non-whitespace character — pure …

ModelQuality (% of best)ConfidenceOverpay
MiniMax M3 91%MEDIUMbest value
Gemini 3.5 Flash best100%RANKED32x

Task detail →

Prompt Adaptation

Renormalizes a pool/template prompt for a specific (capability x model). Never itself adapted (adaptation_exempt); recursion is also backstopped by the _ADAPTING re-entrancy guard. Outputs ARE …

ModelQuality (% of best)ConfidenceOverpay
Tencent Hy3 98%RANKEDbest value
DeepSeek V4 Flash91%MEDIUM2.7x
Qwen 3.5 Flash91%RANKED3.6x
GPT-5.6 Luna95%RANKED4.3x
DeepSeek V4 Pro96%RANKED4.5x
NVIDIA Nemotron-3 Ultra 550B93%MEDIUM5.2x
GPT-5.6 Terra93%HIGH8.2x
Claude Sonnet 4.6 best100%RANKED13x
Gemini 3.5 Flash98%RANKED16x
GPT-5.6 Sol95%HIGH21x
Gemini 3.1 Pro Preview98%RANKED24x
Qwen 3.6 Plus99%RANKED26x
Meta Muse Spark 1.198%HIGH27x
Kimi K2.698%HIGH47x
GPT-5.593%RANKED104x

Task detail →

Claim Refinement

Reviews extracted claims and either REFINES (adds missing subject, expands abbreviations, resolves ambiguous references) or DROPS (methodology mentions, promotional content, claims about the source …

ModelQuality (% of best)ConfidenceOverpay
Gemini 3.1 Flash Lite 100%RANKEDbest value
MiniMax M390%MEDIUM2.5x
Tencent Hy3 best100%HIGH3.2x
Qwen 3.5 Flash92%HIGH5.1x
GPT-5.6 Luna93%MEDIUM5.2x
Qwen 3.7 Plus98%RANKED7.9x
Gemini 3.5 Flash98%HIGH13x
Qwen 3.6 Flash96%MEDIUM14x
Claude Sonnet 599%HIGH15x
Qwen 3.6 Plus96%HIGH16x
Claude Opus 4.894%MEDIUM35x

Task detail →

Image Prompt Generation

Transforms report content into an image-generation prompt for DALL-E / Imagen / similar. Specifies visual style, subject focus, colour palette, composition, mood, and technical details suitable for a …

ModelQuality (% of best)ConfidenceOverpay
DeepSeek V4 Flash 93%RANKEDbest value
Tencent Hy391%HIGH2.9x
MiniMax M395%MEDIUM5x
Qwen 3.5 Flash92%RANKED5.3x
GPT-5.6 Luna98%RANKED6.5x
GPT-5.4 Mini91%RANKED7.2x
DeepSeek V4 Pro94%RANKED7.6x
NVIDIA Nemotron-3 Ultra 550B93%HIGH11x
Qwen 3.7 Plus91%RANKED13x
GPT-5.6 Terra best100%HIGH14x
Qwen 3.6 Flash92%RANKED16x
Qwen 3.6 Plus94%RANKED18x
Grok 4.592%RANKED19x
Claude Sonnet 594%RANKED21x
Meta Muse Spark 1.197%HIGH30x
Gemini 3.5 Flash96%RANKED30x
Kimi K2.692%RANKED33x
GPT-5.6 Sol99%HIGH34x
Claude Sonnet 4.690%RANKED51x
GPT-5.592%RANKED90x

Task detail →

Confidence — how sure we are about the quality score (more judgments + more agreement = higher confidence): RANKED many independent judges scored this model's outputs and their agreement is very high (most confident) — HIGH many judges have scored it and they mostly agree (well-pinned) — MEDIUM enough judges have weighed in to publish, but they disagree more than we'd like (treat with a small grain of salt). LOW-confidence cells are hidden everywhere on the site. See the methodology for the exact thresholds.