Best Models for Relevance, Classification & Matching
Relevance, Classification & Matching: Route by Task, Not Category
There is no defensible universal winner across relevance scoring, classification, prioritization, and matching. The supplied results do not state how many tasks are in this category or how many models are measured on them, so those counts cannot be reported without inventing evidence. They do show a clear production rule, however: use an application-specific quality bar—here, 8.0 out of 10 as a practical guide—then check completion reliability, cost mode, and the exact workload rather than treating category-level scores as interchangeable.
Quality leaders and the limits of a category-wide price comparison
For report-to-topic relevance, qwen3.7-plus is the strongest reported model at 9.17 out of 10, with an interval of ±0.27 [^3]. Other high-scoring relevance results include gpt-5.6-terra at 9.14 out of 10 for $0.02 per batch run [^3], grok-4.5 at 8.91 out of 10 [^3], and nvidia/nemotron-3-super-120b-a12b at 8.7 out of 10 for $0.0055 per task run [^3]. Deepseek-v4-flash reaches 8.41 out of 10 [^1][^3].
The evidence does not support a single, literal cost-and-quality gap between the best model and the cheapest model that clears the bar. Qwen’s 9.17 relevance result has no corresponding relevance cost in the supplied material, while the lowest listed cost attached to a result above the 8.0 bar is nvidia/nemotron-3-super-120b-a12b’s estimated $0.0005 per task run for assigning sections to topic clusters, where it scores 8.28 [^7]. That is a different task, and the estimate is explicitly qualified because the cell has no qualifying usage history of its own [^7]. The figures therefore establish a possible low-cost route within the category, not a head-to-head cost comparison with qwen3.7-plus on relevance.
On selection work, gpt-5.6-terra scores 8.59 out of 10 for $0.0018 per batch run [^2], while gpt-5.6-sol scores 8.48 out of 10 for $0.01 per synchronous run [^2]. Kimi-k3 is higher at 8.65 out of 10, but its listed batch cost is $0.03 [^2]. Thus, when the bar is 8.0 and the task is X-post selection, gpt-5.6-terra is the cheapest listed option among these results; the figures are not transferable to relevance or matching, and batch and synchronous prices should not be treated as identical modes.
Where the results are strong
Classification and prioritization produce some of the clearest economical choices. Kimi-k3 scores 9.22 out of 10, ±0.19, on engagement triage [^4][^8] and 8.98 out of 10, ±0.17, on naming topic clusters [^9]. For assigning sections to topic clusters, nvidia/nemotron-3-super-120b-a12b scores 8.28 out of 10 at an estimated $0.0005 per task run [^7]. These results make kimi-k3 the quality leader on the two cited classification and naming tasks, while Nemotron is the low-cost option reported for section assignment.
Kimi-k3 also leads the cited X-post selection comparison at 8.65 out of 10 [^2]. Gpt-5.6-terra is close behind at 8.59 out of 10, but its $0.0018 batch price is much lower than kimi-k3’s listed $0.03 batch cost [^2]. Gpt-5.6-sol records 8.48 out of 10 for $0.01 per synchronous run [^2], and grok-4.5 records 8.4 out of 10 [^2]. The quality ordering and the value ordering therefore differ: kimi-k3 leads on the reported score, while gpt-5.6-terra is the more economical listed choice when batch execution fits the system.
Matching is similarly split. Grok-4.5 reaches 8.44 out of 10, ±0.15, at $0.07 per task [^5], while qwen3.7-plus scores 8.17 out of 10, ±0.24 [^5], reports the same 8.17 score with an 82% answer rate [^5], and costs $0.02 per task run [^5]. Kimi-k3 scores 8.35 out of 10 [^5], but answers only 51% of attempts [^5]. Deepseek-v4-flash scores 8.28 out of 10 for matching topics to clients [^5]. Qwen is therefore the lower-cost option in the direct qwen-versus-grok comparison, while grok-4.5 is the quality choice if its 8.44 score justifies the higher $0.07 cost. A requirement that every request complete could change that decision because the supplied answer-rate evidence is materially different.
Mixed and contradictory task-level evidence
The individual tasks disagree with one another often enough to rule out category-wide routing. Qwen leads the reported report-relevance result at 9.17 [^3], yet gpt-5.6-terra is nearly as high at 9.14 for $0.02 per batch run [^3], and Nemotron offers 8.7 at $0.0055 [^3]. On another relevance measurement, gpt-5.6-luna scores only 6.33 out of 10 for scoring a post’s relevance at $0.0067 per batch run [^1], while gpt-5.5 scores 6.03 [^1][^3]. The task wording matters: report-to-topic relevance and post relevance do not produce the same model ordering.
The selection results also reverse the simple quality-versus-cost assumption. Kimi-k3 has the highest cited selection score at 8.65 [^2], but gpt-5.6-terra is only 0.06 points lower at 8.59 and is listed at $0.0018 per batch run [^2]. Matching reverses the preference again: grok-4.5 leads qwen3.7-plus on quality, while qwen3.7-plus is cheaper [^5]. The relevance, selection, classification, and matching measurements should consequently be read as workload-specific evidence rather than as one common leaderboard.
Weak or noncompetitive choices
Several models are poor candidates for high-precision relevance judgments without human review. Claude Haiku 4.5 scores 4.61 out of 10 for scoring X-post relevance [^1][^3][^6]; minimax/minimax-m3 scores 6.17 [^1]; gpt-5.5 scores 6.03 [^1][^3]; and gpt-5.6-luna scores 6.33 [^1]. Those scores are below the 8.0 guide bar used here. Claude Sonnet 5 records 8.48 out of 10 on report relevance, but answered only 31% of attempts [^3], so its nominal quality does not by itself make it a dependable production route.
Kimi-k3 is strong on engagement triage, topic-cluster naming, and X-post selection, but its 51% answer rate on matching [^5] makes it unsuitable for a workflow that requires every matching request to complete unless a fallback path is in place. The same operational caveat applies to Claude Sonnet 5’s 31% answer rate in the cited relevance measurement [^3]. These completion results can outweigh a modest quality advantage when unusable outputs trigger retries, manual review, or missed downstream actions.
Reliability and routing conclusion
The current evidence supports a conditional routing policy. Use qwen3.7-plus for the highest reported report-relevance quality and for cost-sensitive topic-to-client matching; use grok-4.5 for matching when 8.44 out of 10 justifies $0.07 per task over qwen3.7-plus’s 8.17 and $0.02 [^5]. Use kimi-k3 for engagement triage or topic-cluster naming when its high measured scores are the priority [^4][^8][^9], and use it for X-post selection only when its $0.03 batch cost is acceptable [^2]. Use gpt-5.6-terra for economical batch selection [^2] or high-scoring relevance at $0.02 per batch run [^3]; use gpt-5.6-sol when synchronous selection at $0.01 per synchronous run or its 8.36 matching score [^2][^5] better fits the system. Use Nemotron for section assignment at the estimated $0.0005 cost with the stated qualification [^7], and for relevance at 8.7 and $0.0055 [^3]. Use deepseek-v4-flash for relevance or matching when 8.41 and 8.28 respectively meet the application’s bar [^1][^3][^5]. Keep low-scoring relevance models behind review or fallback paths, especially where completion is mandatory [^1][^3][^6][^1][^3][^1][^3][^5].
In short: route each workload to the cheapest measured option that clears that workload’s quality bar and completion requirement; escalate to qwen3.7-plus, gpt-5.6-terra, grok-4.5, or kimi-k3 when the task-specific evidence—not the category label—justifies the additional cost or operational complexity.
Check the live benchmark at https://llm-bench.kapualabs.com/ to set the quality bar your application actually needs and compare the task-level numbers. It also lets you compare batch against synchronous pricing yourself before fixing a production route.