Best LLMs for Infrastructure & Utility
Mechanical competence at format conversion, metadata manipulation, prompt rewriting, translation; minimal domain expertise required.
Mechanical competence at format conversion, metadata manipulation, prompt rewriting, translation; minimal domain expertise required.
Task-by-task breakdown
Markdown Newline Repair
Re-inserts blank lines into a markdown document whose line breaks were stripped (headings, separators, and paragraphs collapsed onto one line). Must not modify any non-whitespace character — pure …
| Model | Quality (% of best) | Confidence | Overpay |
|---|---|---|---|
| MiniMax M3 ★ | 94% | MEDIUM | best value |
| Gemini 3.5 Flash best | 100% | RANKED | 32x |
Research Query Validation
Validates and minimally repairs a research query for a specified search platform and region without changing its intended information need. Illustrative uses include repairing queries for enterprise …
| Model | Quality (% of best) | Confidence | Overpay |
|---|---|---|---|
| GPT-5.6 Luna ★ | 90% | MEDIUM | best value |
| MiniMax M3 | 98% | RANKED | 1.9x |
| GLM-5.3 Flash | 97% | MEDIUM | 4.2x |
| Gemini 3.8 Flash best | 100% | HIGH | 5.3x |
| Gemini 3.5 Flash | 97% | HIGH | 14x |
| Meta Muse Spark 1.3 | 100% | MEDIUM | 19x |
| Grok 4.6 | 92% | MEDIUM | 49x |
| GLM-5.3 | 95% | MEDIUM | 54x |
| Qwen 3.8 Max | 98% | MEDIUM | 60x |
Metadata Paragraph Rewriting
Rewrites a factual metadata paragraph for clarity and fluency while preserving all supplied dates, counts, source types, qualifications, and approximate length. Illustrative uses include polishing a …
| Model | Quality (% of best) | Confidence | Overpay |
|---|---|---|---|
| Claude Sonnet 5 ★ best | 100% | RANKED | best value |
| Gemini 3.5 Flash | 97% | HIGH | 3.5x |
Factual Claim Refinement
Reviews extracted factual claims and either minimally refines them into self-contained, verifiable statements or drops them when they are non-factual, unsupported, duplicate, or not useful. …
| Model | Quality (% of best) | Confidence | Overpay |
|---|---|---|---|
| Gemini 3.1 Flash Lite ★ | 91% | RANKED | best value |
| GLM-5.3 Flash | 97% | MEDIUM | 9.2x |
| Qwen 3.7 Plus | 96% | HIGH | 9.3x |
| Tencent Hy3 | 96% | HIGH | 9.8x |
| Thinking Machines Inkling Small | 97% | MEDIUM | 15x |
| Gemini 3.5 Flash | 96% | HIGH | 15x |
| Claude Sonnet 5 | 97% | HIGH | 16x |
| NVIDIA Nemotron-3 Ultra 550B | 94% | MEDIUM | 18x |
| Thinking Machines Inkling | 93% | MEDIUM | 35x |
| Meta Muse Spark 1.3 | 98% | MEDIUM | 38x |
| Claude Opus 5 | 93% | MEDIUM | 49x |
| GLM-5.3 best | 100% | MEDIUM | 127x |
Batch Text Translation
Translates one or more text items from a supplied or detected source language into a target language while preserving meaning, structure, identifiers, URLs, handles, hashtags, and protected inline …
| Model | Quality (% of best) | Confidence | Overpay |
|---|---|---|---|
| GPT-5.6 Luna ★ | 93% | RANKED | best value |
| Gemini 3.5 Flash Lite | 91% | MEDIUM | 1.9x |
| MiniMax M3 | 94% | HIGH | 4.5x |
| GLM-5.3 Flash | 95% | HIGH | 7.9x |
| Gemini 3.8 Flash | 96% | HIGH | 11x |
| Qwen 3.7 Plus | 95% | RANKED | 14x |
| GPT-5.6 Terra | 96% | RANKED | 14x |
| DeepSeek V4 Flash | 91% | HIGH | 15x |
| Thinking Machines Inkling Small | 91% | MEDIUM | 15x |
| Meta Muse Spark 1.3 | 96% | RANKED | 20x |
| GPT-5.6 Sol | 95% | RANKED | 22x |
| Gemini 3.5 Flash best | 100% | HIGH | 24x |
| DeepSeek V4 Pro | 91% | MEDIUM | 29x |
| Claude Sonnet 5 | 95% | RANKED | 31x |
| Thinking Machines Inkling | 92% | MEDIUM | 37x |
| Claude Opus 5 | 96% | RANKED | 56x |
| GLM-5.3 | 94% | MEDIUM | 91x |
| Tencent Hy4 Preview | 92% | MEDIUM | 92x |
| Grok 4.6 | 96% | RANKED | 95x |
| Moonshot Kimi K3 | 99% | RANKED | 113x |
| Qwen 3.8 Max | 93% | HIGH | 120x |
Content To Image Prompt Generation
Converts supplied content and optional visual context into a detailed, model-agnostic image-generation brief describing subject, composition, style, palette, mood, and constraints. Illustrative uses …
| Model | Quality (% of best) | Confidence | Overpay |
|---|---|---|---|
| GPT-5.6 Luna ★ | 98% | RANKED | best value |
| DeepSeek V4 Flash | 94% | RANKED | 2.8x |
| Tencent Hy3 | 92% | RANKED | 5.1x |
| Thinking Machines Inkling Small | 98% | MEDIUM | 6.1x |
| GPT-5.6 Terra | 100% | RANKED | 9x |
| Qwen 3.7 Plus | 94% | RANKED | 9.7x |
| Thinking Machines Inkling | 99% | HIGH | 14x |
| Claude Sonnet 5 | 96% | RANKED | 15x |
| GPT-5.6 Sol best | 100% | RANKED | 16x |
| DeepSeek V4 Pro | 95% | RANKED | 21x |
| Gemini 3.5 Flash | 97% | RANKED | 21x |
| Moonshot Kimi K3 | 97% | RANKED | 53x |
Section Prompt Generation
Generates or adapts system and user prompts for a report section using the section requirement, subject context, dependencies, source variables, and template constraints. Illustrative uses include …
| Model | Quality (% of best) | Confidence | Overpay |
|---|---|---|---|
| GPT-5.6 Luna ★ | 99% | MEDIUM | best value |
| GPT-5.4 Nano | 94% | HIGH | 3.1x |
| Thinking Machines Inkling Small | 95% | MEDIUM | 4.3x |
| DeepSeek V4 Flash | 96% | HIGH | 5.9x |
| GPT-5.6 Terra | 99% | MEDIUM | 12x |
| Thinking Machines Inkling | 98% | MEDIUM | 12x |
| Claude Haiku 4.5 | 94% | HIGH | 17x |
| GPT-5.6 Sol best | 100% | MEDIUM | 20x |
| DeepSeek V4 Pro | 96% | MEDIUM | 31x |
Model Specific Prompt Adaptation
Adapts caller-supplied prompt content for a target model while preserving the task contract, required placeholders, output schema, evaluation semantics, and safety boundaries. Illustrative uses …
| Model | Quality (% of best) | Confidence | Overpay |
|---|---|---|---|
| GPT-5.6 Luna ★ | 96% | HIGH | best value |
| NVIDIA Nemotron-3 Ultra 550B | 94% | MEDIUM | 3.8x |
| Tencent Hy3 | 98% | RANKED | 4.1x |
| Thinking Machines Inkling Small | 96% | MEDIUM | 4.8x |
| GPT-5.6 Terra | 94% | MEDIUM | 5.6x |
| DeepSeek V4 Flash | 94% | MEDIUM | 8.7x |
| GPT-5.6 Sol | 97% | RANKED | 10x |
| DeepSeek V4 Pro | 96% | RANKED | 13x |
| Meta Muse Spark 1.3 best | 100% | MEDIUM | 14x |
| Thinking Machines Inkling | 97% | MEDIUM | 15x |
| Moonshot Kimi K3 | 98% | MEDIUM | 53x |
Research Query Generation
Generates diverse, high-signal research queries for a subject, purpose, region, and target search platform while respecting platform syntax and caller requirements. Illustrative uses include …
| Model | Quality (% of best) | Confidence | Overpay |
|---|---|---|---|
| GPT-5.6 Luna ★ | 94% | HIGH | best value |
| GLM-5.3 Flash | 100% | RANKED | 4.5x |
| Gemini 3.8 Flash | 97% | HIGH | 7.8x |
| Thinking Machines Inkling Small | 90% | MEDIUM | 9.8x |
| GPT-5.6 Terra best | 100% | HIGH | 11x |
| GPT-5.6 Sol | 100% | HIGH | 18x |
| Meta Muse Spark 1.3 | 94% | HIGH | 25x |
| Tencent Hy4 Preview | 93% | HIGH | 43x |
| GLM-5.3 | 97% | HIGH | 50x |
| Claude Opus 5 | 90% | MEDIUM | 82x |
Confidence — how sure we are about the quality score (more judgments + more agreement = higher confidence): RANKED many independent judges scored this model's outputs and their agreement is very high (most confident) — HIGH many judges have scored it and they mostly agree (well-pinned) — MEDIUM enough judges have weighed in to publish, but they disagree more than we'd like (treat with a small grain of salt). LOW-confidence cells are hidden everywhere on the site. See the methodology for the exact thresholds.