Cost mode:

Mechanical competence at format conversion, metadata manipulation, prompt rewriting, translation; minimal domain expertise required.

9 capabilities in this category.

Task-by-task breakdown

Markdown Newline Repair

Re-inserts blank lines into a markdown document whose line breaks were stripped (headings, separators, and paragraphs collapsed onto one line). Must not modify any non-whitespace character — pure …

ModelQuality (% of best)ConfidenceOverpay
MiniMax M3 94%MEDIUMbest value
Gemini 3.5 Flash best100%RANKED32x

Task detail →

Research Query Validation

Validates and minimally repairs a research query for a specified search platform and region without changing its intended information need. Illustrative uses include repairing queries for enterprise …

ModelQuality (% of best)ConfidenceOverpay
GPT-5.6 Luna 90%MEDIUMbest value
MiniMax M398%RANKED1.9x
GLM-5.3 Flash97%MEDIUM4.2x
Gemini 3.8 Flash best100%HIGH5.3x
Gemini 3.5 Flash97%HIGH14x
Meta Muse Spark 1.3100%MEDIUM19x
Grok 4.692%MEDIUM49x
GLM-5.395%MEDIUM54x
Qwen 3.8 Max98%MEDIUM60x

Task detail →

Factual Claim Refinement

Reviews extracted factual claims and either minimally refines them into self-contained, verifiable statements or drops them when they are non-factual, unsupported, duplicate, or not useful. …

ModelQuality (% of best)ConfidenceOverpay
Gemini 3.1 Flash Lite 91%RANKEDbest value
GLM-5.3 Flash97%MEDIUM9.2x
Qwen 3.7 Plus96%HIGH9.3x
Tencent Hy396%HIGH9.8x
Thinking Machines Inkling Small97%MEDIUM15x
Gemini 3.5 Flash96%HIGH15x
Claude Sonnet 597%HIGH16x
NVIDIA Nemotron-3 Ultra 550B94%MEDIUM18x
Thinking Machines Inkling93%MEDIUM35x
Meta Muse Spark 1.398%MEDIUM38x
Claude Opus 593%MEDIUM49x
GLM-5.3 best100%MEDIUM127x

Task detail →

Batch Text Translation

Translates one or more text items from a supplied or detected source language into a target language while preserving meaning, structure, identifiers, URLs, handles, hashtags, and protected inline …

ModelQuality (% of best)ConfidenceOverpay
GPT-5.6 Luna 93%RANKEDbest value
Gemini 3.5 Flash Lite91%MEDIUM1.9x
MiniMax M394%HIGH4.5x
GLM-5.3 Flash95%HIGH7.9x
Gemini 3.8 Flash96%HIGH11x
Qwen 3.7 Plus95%RANKED14x
GPT-5.6 Terra96%RANKED14x
DeepSeek V4 Flash91%HIGH15x
Thinking Machines Inkling Small91%MEDIUM15x
Meta Muse Spark 1.396%RANKED20x
GPT-5.6 Sol95%RANKED22x
Gemini 3.5 Flash best100%HIGH24x
DeepSeek V4 Pro91%MEDIUM29x
Claude Sonnet 595%RANKED31x
Thinking Machines Inkling92%MEDIUM37x
Claude Opus 596%RANKED56x
GLM-5.394%MEDIUM91x
Tencent Hy4 Preview92%MEDIUM92x
Grok 4.696%RANKED95x
Moonshot Kimi K399%RANKED113x
Qwen 3.8 Max93%HIGH120x

Task detail →

Content To Image Prompt Generation

Converts supplied content and optional visual context into a detailed, model-agnostic image-generation brief describing subject, composition, style, palette, mood, and constraints. Illustrative uses …

ModelQuality (% of best)ConfidenceOverpay
GPT-5.6 Luna 98%RANKEDbest value
DeepSeek V4 Flash94%RANKED2.8x
Tencent Hy392%RANKED5.1x
Thinking Machines Inkling Small98%MEDIUM6.1x
GPT-5.6 Terra100%RANKED9x
Qwen 3.7 Plus94%RANKED9.7x
Thinking Machines Inkling99%HIGH14x
Claude Sonnet 596%RANKED15x
GPT-5.6 Sol best100%RANKED16x
DeepSeek V4 Pro95%RANKED21x
Gemini 3.5 Flash97%RANKED21x
Moonshot Kimi K397%RANKED53x

Task detail →

Section Prompt Generation

Generates or adapts system and user prompts for a report section using the section requirement, subject context, dependencies, source variables, and template constraints. Illustrative uses include …

ModelQuality (% of best)ConfidenceOverpay
GPT-5.6 Luna 99%MEDIUMbest value
GPT-5.4 Nano94%HIGH3.1x
Thinking Machines Inkling Small95%MEDIUM4.3x
DeepSeek V4 Flash96%HIGH5.9x
GPT-5.6 Terra99%MEDIUM12x
Thinking Machines Inkling98%MEDIUM12x
Claude Haiku 4.594%HIGH17x
GPT-5.6 Sol best100%MEDIUM20x
DeepSeek V4 Pro96%MEDIUM31x

Task detail →

Model Specific Prompt Adaptation

Adapts caller-supplied prompt content for a target model while preserving the task contract, required placeholders, output schema, evaluation semantics, and safety boundaries. Illustrative uses …

ModelQuality (% of best)ConfidenceOverpay
GPT-5.6 Luna 96%HIGHbest value
NVIDIA Nemotron-3 Ultra 550B94%MEDIUM3.8x
Tencent Hy398%RANKED4.1x
Thinking Machines Inkling Small96%MEDIUM4.8x
GPT-5.6 Terra94%MEDIUM5.6x
DeepSeek V4 Flash94%MEDIUM8.7x
GPT-5.6 Sol97%RANKED10x
DeepSeek V4 Pro96%RANKED13x
Meta Muse Spark 1.3 best100%MEDIUM14x
Thinking Machines Inkling97%MEDIUM15x
Moonshot Kimi K398%MEDIUM53x

Task detail →

Research Query Generation

Generates diverse, high-signal research queries for a subject, purpose, region, and target search platform while respecting platform syntax and caller requirements. Illustrative uses include …

ModelQuality (% of best)ConfidenceOverpay
GPT-5.6 Luna 94%HIGHbest value
GLM-5.3 Flash100%RANKED4.5x
Gemini 3.8 Flash97%HIGH7.8x
Thinking Machines Inkling Small90%MEDIUM9.8x
GPT-5.6 Terra best100%HIGH11x
GPT-5.6 Sol100%HIGH18x
Meta Muse Spark 1.394%HIGH25x
Tencent Hy4 Preview93%HIGH43x
GLM-5.397%HIGH50x
Claude Opus 590%MEDIUM82x

Task detail →

Confidence — how sure we are about the quality score (more judgments + more agreement = higher confidence): RANKED many independent judges scored this model's outputs and their agreement is very high (most confident) — HIGH many judges have scored it and they mostly agree (well-pinned) — MEDIUM enough judges have weighed in to publish, but they disagree more than we'd like (treat with a small grain of salt). LOW-confidence cells are hidden everywhere on the site. See the methodology for the exact thresholds.