Best LLMs for Topic Organization & Clustering
Discover natural groupings without external schema, name and order them coherently.
Discover natural groupings without external schema, name and order them coherently.
Task-by-task breakdown
Topic Sequence Ordering
Decides the presentation order of topics in a comprehensive report. Optimises for logical flow, building complexity, narrative arc, topic dependencies, and reader engagement; also rewrites each raw …
| Model | Quality (% of best) | Confidence | Overpay |
|---|---|---|---|
| Tencent Hy3 ★ | 91% | MEDIUM | best value |
| NVIDIA Nemotron-3 Ultra 550B | 90% | MEDIUM | 2.4x |
| Grok 4.5 | 99% | RANKED | 4.8x |
| Meta Muse Spark 1.1 | 98% | HIGH | 7.6x |
| Gemini 3.5 Flash best | 100% | HIGH | 7.8x |
Topic Discovery Clustering
Pooled TT for thematic topic discovery from a batch of items: topic_clustering_batch (content-driven) and topic_clustering_claims (claim-driven). Same output schema, same capability.
| Model | Quality (% of best) | Confidence | Overpay |
|---|---|---|---|
| Meta Muse Spark 1.1 ★ | 97% | MEDIUM | best value |
| GPT-5.6 Sol best | 100% | MEDIUM | 2.1x |
Topic Cluster Naming
Assigns a concise descriptive name (3-7 words) and a brief description (≤500 chars) to a cluster of semantically similar claims that has already been synthesised. Specific over generic — no 'Market …
| Model | Quality (% of best) | Confidence | Overpay |
|---|---|---|---|
| GPT-5.4 Mini ★ | 94% | MEDIUM | best value |
| Tencent Hy3 | 92% | RANKED | 1.1x |
| MiniMax M3 | 94% | RANKED | 1.2x |
| DeepSeek V4 Flash | 93% | HIGH | 1.6x |
| GPT-5.6 Luna | 96% | HIGH | 1.8x |
| DeepSeek V4 Pro best | 100% | HIGH | 2.1x |
| Qwen 3.5 Flash | 90% | HIGH | 2.2x |
| Qwen 3.7 Plus | 96% | RANKED | 2.6x |
| Qwen 3.6 Plus | 97% | HIGH | 2.6x |
| Gemini 3.1 Pro Preview | 91% | HIGH | 3x |
| Claude Haiku 4.5 | 96% | HIGH | 3.1x |
| NVIDIA Nemotron-3 Ultra 550B | 91% | MEDIUM | 4.2x |
| Qwen 3.6 Flash | 95% | RANKED | 4.3x |
| GPT-5.6 Terra | 95% | MEDIUM | 4.5x |
| Gemini 3.5 Flash | 93% | RANKED | 4.8x |
| Kimi K2.6 | 100% | RANKED | 5.7x |
| Claude Sonnet 4.6 | 98% | HIGH | 6.7x |
| Claude Sonnet 5 | 95% | RANKED | 8.7x |
| Meta Muse Spark 1.1 | 96% | MEDIUM | 9.1x |
| GPT-5.5 | 99% | RANKED | 9.8x |
| Grok 4.5 | 99% | RANKED | 10x |
| GPT-5.6 Sol | 96% | HIGH | 12x |
Topic-to-Section Assignment
Assigns each merged topic in a batch to 1-3 publication sections, ordered by descending relevance. Index 0 is the PRIMARY section (where the topic appears on the client landing page); secondaries …
| Model | Quality (% of best) | Confidence | Overpay |
|---|---|---|---|
| DeepSeek V4 Flash ★ | 91% | MEDIUM | best value |
| Qwen 3.5 Flash | 92% | HIGH | 2.7x |
| Tencent Hy3 | 91% | MEDIUM | 2.8x |
| Qwen 3.6 Plus | 92% | MEDIUM | 6.8x |
| Gemini 3.1 Pro Preview | 90% | MEDIUM | 8x |
| Claude Sonnet 4.6 | 90% | HIGH | 13x |
| Kimi K2.6 | 91% | MEDIUM | 14x |
| Gemini 3.5 Flash best | 100% | MEDIUM | 18x |
| GPT-5.5 | 95% | HIGH | 23x |
| Meta Muse Spark 1.1 | 93% | MEDIUM | 36x |
| GPT-5.6 Sol | 96% | MEDIUM | 46x |
Confidence — how sure we are about the quality score (more judgments + more agreement = higher confidence): RANKED many independent judges scored this model's outputs and their agreement is very high (most confident) — HIGH many judges have scored it and they mostly agree (well-pinned) — MEDIUM enough judges have weighed in to publish, but they disagree more than we'd like (treat with a small grain of salt). LOW-confidence cells are hidden everywhere on the site. See the methodology for the exact thresholds.