Best LLMs for Relevance, Classification & Matching
Semantic similarity judgment: does this thing belong in that bucket / match that target?
Semantic similarity judgment: does this thing belong in that bucket / match that target?
Task-by-task breakdown
Short Post Batch Relevance Scoring
Scores a batch of short-form posts against a supplied topic, query, or scoring context, returning an item-level relevance score for each post. It is intended for compact posts such as those from X, …
| Model | Quality (% of best) | Confidence | Overpay |
|---|---|---|---|
| MiniMax M3 ★ | 96% | RANKED | best value |
| Gemini 3.5 Flash best | 100% | MEDIUM | 9.2x |
| GPT-5.6 Terra | 95% | MEDIUM | 13x |
| Meta Muse Spark 1.3 | 91% | MEDIUM | 19x |
| Grok 4.6 | 95% | MEDIUM | 33x |
Content Set Relevance Scoring
Scores every item in one supplied set of retrieved content for relevance to a query or analysis objective, returning a calibrated item-level score with concise evidence for each. Illustrative uses …
| Model | Quality (% of best) | Confidence | Overpay |
|---|---|---|---|
| GPT-5.6 Luna ★ | 90% | MEDIUM | best value |
| Thinking Machines Inkling Small | 95% | MEDIUM | 2.3x |
| GPT-5.6 Terra | 99% | HIGH | 5.8x |
| Gemini 3.5 Flash best | 100% | HIGH | 6x |
| GPT-5.6 Sol | 98% | MEDIUM | 9.5x |
Community Policy Evaluation
Determines whether supplied community rules permit a proposed action under caller-defined channel and decision policies, returning a conservative decision and evidence. Illustrative uses include …
| Model | Quality (% of best) | Confidence | Overpay |
|---|---|---|---|
| GPT-5.6 Luna ★ | 100% | MEDIUM | best value |
| MiniMax M3 | 91% | HIGH | 1.8x |
| GPT-5.6 Terra | 100% | MEDIUM | 7.8x |
| GPT-5.6 Sol | 100% | MEDIUM | 12x |
| Gemini 3.5 Flash | 93% | MEDIUM | 14x |
| Claude Sonnet 5 best | 100% | MEDIUM | 14x |
Research Community Selection
Selects relevant, accessible communities as research sources for a subject, purpose, regions, and platform profile. Illustrative uses include finding developer forums, customer user groups, …
| Model | Quality (% of best) | Confidence | Overpay |
|---|---|---|---|
| GPT-5.6 Luna ★ | 93% | HIGH | best value |
| GPT-5.6 Terra best | 100% | HIGH | 12x |
| Qwen 3.7 Plus | 94% | MEDIUM | 13x |
| Gemini 3.5 Flash | 91% | HIGH | 18x |
Engagement Opportunity Triage
Classifies whether an incoming public message merits a response, identifies risk, and proposes a defensible angle using supplied themes, evidence boundaries, channel context, and policy. Illustrative …
| Model | Quality (% of best) | Confidence | Overpay |
|---|---|---|---|
| GPT-5.6 Luna ★ | 98% | RANKED | best value |
| Gemini 3.1 Flash Lite | 95% | RANKED | 1.4x |
| GPT-5.4 Nano | 95% | RANKED | 1.5x |
| Gemini 3.5 Flash Lite | 93% | HIGH | 1.9x |
| MiniMax M3 | 98% | RANKED | 3.3x |
| DeepSeek V4 Flash | 94% | RANKED | 4.6x |
| Tencent Hy3 | 96% | RANKED | 4.7x |
| Gemini 3.8 Flash | 96% | HIGH | 6.7x |
| NVIDIA Nemotron-3 Super 120B | 96% | HIGH | 8.3x |
| Claude Haiku 4.5 | 98% | RANKED | 10x |
| GPT-5.6 Terra | 95% | RANKED | 10x |
| Thinking Machines Inkling Small | 99% | RANKED | 12x |
| Qwen 3.7 Plus | 98% | RANKED | 15x |
| NVIDIA Nemotron-3 Ultra 550B | 98% | RANKED | 15x |
| Claude Sonnet 5 | 96% | RANKED | 19x |
| Gemini 3.5 Flash | 97% | RANKED | 20x |
| GPT-5.6 Sol | 95% | HIGH | 20x |
| DeepSeek V4 Pro | 94% | RANKED | 24x |
| Thinking Machines Inkling best | 100% | RANKED | 37x |
| GLM-5.3 | 98% | HIGH | 42x |
| Claude Opus 5 | 97% | RANKED | 43x |
| Grok 4.6 | 97% | RANKED | 50x |
| Qwen 3.8 Max | 96% | HIGH | 56x |
| Moonshot Kimi K3 | 99% | HIGH | 92x |
Profile Pool Matching
Matches a source object to the most suitable candidates in a supplied profile pool using caller-defined dimensions, eligibility rules, thresholds, and creation policy. When permitted, it may propose a …
| Model | Quality (% of best) | Confidence | Overpay |
|---|---|---|---|
| GLM-5.3 Flash ★ | 99% | MEDIUM | best value |
| Gemini 3.5 Flash | 94% | MEDIUM | 4.4x |
| DeepSeek V4 Pro | 92% | HIGH | 5.8x |
| Thinking Machines Inkling | 93% | MEDIUM | 6.2x |
| Meta Muse Spark 1.3 | 97% | MEDIUM | 6.6x |
| Tencent Hy4 Preview best | 100% | MEDIUM | 7.6x |
| GLM-5.3 | 96% | MEDIUM | 11x |
| Grok 4.6 | 95% | MEDIUM | 14x |
| Claude Opus 5 | 96% | MEDIUM | 15x |
Language Identification
Identifies the primary language of supplied text and returns its normalized language code in the configured response schema. Illustrative uses include routing multilingual support tickets, contracts, …
| Model | Quality (% of best) | Confidence | Overpay |
|---|---|---|---|
| Gemini 3.1 Flash Lite ★ | 99% | RANKED | best value |
| GPT-5.4 Nano | 99% | RANKED | 1.1x |
| NVIDIA Nemotron-3 Nano 30B-A3B | 98% | RANKED | 1.1x |
| GPT-5.6 Luna | 99% | RANKED | 1.2x |
| Gemini 3.5 Flash Lite | 98% | RANKED | 1.3x |
| DeepSeek V4 Flash | 98% | RANKED | 2.3x |
| Qwen 3.8 Flash | 99% | RANKED | 2.4x |
| NVIDIA Nemotron-3 Super 120B | 99% | RANKED | 3.2x |
| GLM-5.3 Flash | 97% | HIGH | 3.2x |
| NVIDIA Nemotron 3.5 Lightning | 95% | MEDIUM | 4.4x |
| MiniMax M3 | 99% | RANKED | 4.5x |
| Thinking Machines Inkling Small | 98% | RANKED | 4.6x |
| Tencent Hy3 | 99% | RANKED | 5.9x |
| Gemini 3.8 Flash | 98% | MEDIUM | 6.4x |
| Claude Haiku 4.5 | 99% | RANKED | 11x |
| GPT-5.6 Terra | 99% | RANKED | 11x |
| NVIDIA Nemotron-3 Ultra 550B | 99% | RANKED | 13x |
| DeepSeek V4 Pro | 98% | RANKED | 15x |
| GPT-5.6 Sol | 99% | RANKED | 19x |
| Claude Sonnet 5 | 100% | RANKED | 23x |
| Gemini 3.5 Flash | 99% | RANKED | 24x |
| Thinking Machines Inkling | 96% | MEDIUM | 24x |
| GLM-5.3 | 98% | MEDIUM | 25x |
| Qwen 3.7 Plus | 99% | RANKED | 27x |
| Tencent Hy4 Preview | 98% | MEDIUM | 36x |
| Claude Opus 5 | 97% | MEDIUM | 41x |
| Qwen 3.8 Max | 100% | RANKED | 43x |
| Meta Muse Spark 1.3 | 98% | HIGH | 54x |
| Moonshot Kimi K3 best | 100% | RANKED | 63x |
| Grok 4.6 | 96% | MEDIUM | 77x |
Content Domain Suggestion
Suggests one to three concise subject-domain labels for a content or analysis specification using its title, description, subject context, structure, and any existing domain vocabulary. Illustrative …
| Model | Quality (% of best) | Confidence | Overpay |
|---|---|---|---|
| Claude Opus 5 ★ best | 100% | MEDIUM | best value |
Social Post Portfolio Selection
Selects the best fixed-size portfolio of candidate social posts using quality, expected audience value, topical diversity, timeliness, and caller-supplied constraints. Illustrative uses include …
| Model | Quality (% of best) | Confidence | Overpay |
|---|---|---|---|
| Tencent Hy3 ★ | 92% | RANKED | best value |
| GPT-5.6 Luna | 96% | RANKED | 1.1x |
| Gemini 3.5 Flash Lite | 98% | HIGH | 1.4x |
| NVIDIA Nemotron-3 Nano 30B-A3B | 95% | MEDIUM | 1.5x |
| Gemini 3.1 Flash Lite | 97% | HIGH | 1.5x |
| MiniMax M3 | 96% | RANKED | 3.9x |
| DeepSeek V4 Flash best | 100% | RANKED | 6.2x |
| Claude Haiku 4.5 | 92% | MEDIUM | 7x |
| GLM-5.3 Flash | 99% | RANKED | 8x |
| NVIDIA Nemotron 3.5 Lightning | 92% | MEDIUM | 8.7x |
| GPT-5.6 Terra | 98% | RANKED | 9.6x |
| NVIDIA Nemotron-3 Ultra 550B | 95% | MEDIUM | 9.6x |
| NVIDIA Nemotron-3 Super 120B | 92% | MEDIUM | 14x |
| DeepSeek V4 Pro | 97% | HIGH | 14x |
| Qwen 3.8 Flash | 95% | HIGH | 15x |
| Gemini 3.8 Flash | 99% | HIGH | 16x |
| GPT-5.6 Sol | 97% | RANKED | 18x |
| Claude Sonnet 5 | 95% | RANKED | 18x |
| Thinking Machines Inkling Small | 100% | HIGH | 21x |
| Qwen 3.7 Plus | 95% | RANKED | 23x |
| Gemini 3.5 Flash | 97% | RANKED | 35x |
| Claude Opus 5 | 98% | HIGH | 48x |
| Meta Muse Spark 1.3 | 100% | RANKED | 58x |
| GLM-5.3 | 98% | HIGH | 64x |
| Thinking Machines Inkling | 98% | HIGH | 88x |
| Qwen 3.8 Max | 93% | HIGH | 135x |
| Tencent Hy4 Preview | 99% | RANKED | 138x |
| Grok 4.6 | 99% | HIGH | 157x |
| Moonshot Kimi K3 | 99% | HIGH | 176x |
Authoritative Source Selection
Recommends authoritative, accessible, and maintainable sources for researching a subject, focus, regions, and optional analysis context. Illustrative uses include selecting sources for vendor …
| Model | Quality (% of best) | Confidence | Overpay |
|---|---|---|---|
| GPT-5.6 Luna ★ | 90% | HIGH | best value |
| GLM-5.3 Flash | 94% | HIGH | 4.3x |
| DeepSeek V4 Pro | 92% | MEDIUM | 15x |
| Meta Muse Spark 1.3 | 94% | RANKED | 18x |
| Tencent Hy4 Preview | 92% | MEDIUM | 22x |
| GPT-5.6 Sol best | 100% | RANKED | 22x |
| Grok 4.6 | 98% | RANKED | 41x |
| GLM-5.3 | 98% | RANKED | 43x |
| Claude Opus 5 | 95% | MEDIUM | 53x |
Topic Taxonomy Matching
Maps transient input topics to persistent taxonomy topics, reusing existing topics where semantically appropriate and proposing new ones only for durable coverage gaps. Illustrative uses include …
| Model | Quality (% of best) | Confidence | Overpay |
|---|---|---|---|
| GPT-5.6 Luna ★ | 90% | MEDIUM | best value |
| DeepSeek V4 Flash | 98% | RANKED | 7.1x |
| Qwen 3.7 Plus | 94% | MEDIUM | 7.8x |
| Gemini 3.8 Flash best | 100% | MEDIUM | 9.6x |
| DeepSeek V4 Pro | 98% | RANKED | 15x |
| GPT-5.6 Terra | 98% | HIGH | 17x |
| GPT-5.6 Sol | 99% | HIGH | 19x |
| Gemini 3.5 Flash | 100% | RANKED | 22x |
| Meta Muse Spark 1.3 | 98% | MEDIUM | 24x |
Public Engagement Reply Review
Reviews a caller-supplied public response against its source message, allowed evidence, forbidden claims, operator instructions, channel rules, and review policy, returning an advisory readiness …
| Model | Quality (% of best) | Confidence | Overpay |
|---|---|---|---|
| GLM-5.3 Flash ★ best | 100% | MEDIUM | best value |
| Claude Sonnet 5 | 91% | MEDIUM | 2.7x |
| Tencent Hy4 Preview | 96% | MEDIUM | 11x |
Generated Content Relevance Scoring
Scores every item in a supplied set of generated or intermediate content against a target context, objective, or specification, returning a calibrated item-level relevance assessment. One current …
| Model | Quality (% of best) | Confidence | Overpay |
|---|---|---|---|
| GPT-5.6 Luna ★ | 96% | MEDIUM | best value |
| Gemini 3.5 Flash best | 100% | MEDIUM | 4.2x |
| Qwen 3.7 Plus | 99% | HIGH | 4.4x |
| NVIDIA Nemotron-3 Super 120B | 93% | MEDIUM | 4.7x |
| Thinking Machines Inkling Small | 91% | MEDIUM | 5.2x |
| NVIDIA Nemotron-3 Ultra 550B | 98% | MEDIUM | 5.6x |
| GPT-5.6 Terra | 100% | HIGH | 7.1x |
| DeepSeek V4 Flash | 92% | HIGH | 9.6x |
| GPT-5.6 Sol | 100% | HIGH | 12x |
| Claude Sonnet 5 | 92% | MEDIUM | 15x |
| DeepSeek V4 Pro | 91% | MEDIUM | 39x |
Persona Policy Eligibility Check
Determines whether a proposed persona or identity use satisfies a caller-supplied eligibility policy and evidence standard. Illustrative uses include checking a branded virtual spokesperson, …
| Model | Quality (% of best) | Confidence | Overpay |
|---|---|---|---|
| GPT-5.6 Luna ★ | 90% | MEDIUM | best value |
| DeepSeek V4 Flash best | 100% | HIGH | 4.2x |
| GLM-5.3 Flash | 98% | MEDIUM | 5.3x |
| Gemini 3.8 Flash | 95% | HIGH | 7.2x |
| GPT-5.6 Terra | 95% | HIGH | 12x |
| GPT-5.6 Sol | 98% | RANKED | 17x |
| Gemini 3.5 Flash | 99% | RANKED | 21x |
| Meta Muse Spark 1.3 | 97% | HIGH | 26x |
| Thinking Machines Inkling | 92% | MEDIUM | 26x |
| GLM-5.3 | 98% | MEDIUM | 62x |
| Tencent Hy4 Preview | 99% | HIGH | 65x |
| Claude Opus 5 | 98% | MEDIUM | 67x |
| Grok 4.6 | 95% | HIGH | 76x |
Confidence — how sure we are about the quality score (more judgments + more agreement = higher confidence): RANKED many independent judges scored this model's outputs and their agreement is very high (most confident) — HIGH many judges have scored it and they mostly agree (well-pinned) — MEDIUM enough judges have weighed in to publish, but they disagree more than we'd like (treat with a small grain of salt). LOW-confidence cells are hidden everywhere on the site. See the methodology for the exact thresholds.