Models — The Right-Sized LLM Benchmark
Per-model performance across all capabilities.
One page per LLM in the benchmark — its performance across every task, where it’s the cheapest qualifier at the default quality bar, and which categories it does and doesn’t qualify on.
- Claude Haiku 4.5
Good enough on 8/52 tasks at the 90% bar. Best value on 0 tasks. Doesn't qualify on any: Structured Data & Fact Extraction, Social & Promotional Content.
- Claude Opus 5
Good enough on 17/52 tasks at the 90% bar. Best value on 1 task. Best fit for: Infrastructure & Utility, Relevance, Classification & Matching, Social & Promotional Content, Content Summarization & Syn
- Claude Sonnet 5
Good enough on 25/52 tasks at the 90% bar. Best value on 2 tasks. Best fit for: Financial Analysis & Trading Decisions, Infrastructure & Utility, Long-form Content Generation, Relevance, Classificatio
- DeepSeek V4 Flash
Good enough on 18/52 tasks at the 90% bar. Best value on 1 task. Best fit for: Structured Data & Fact Extraction, Infrastructure & Utility, Topic Organization & Clustering. Doesn't qualify on any: Fin
- DeepSeek V4 Pro
Good enough on 22/52 tasks at the 90% bar. Best value on 0 tasks. Best fit for: Long-form Content Generation.
- Gemini 3.1 Flash Lite
Good enough on 5/52 tasks at the 90% bar. Best value on 3 tasks. Doesn't qualify on any: Financial Analysis & Trading Decisions, Long-form Content Generation, Social & Promotional Content, Topic Organ
- Gemini 3.5 Flash
Good enough on 35/52 tasks at the 90% bar. Best value on 1 task. Best fit for: Structured Data & Fact Extraction, Financial Analysis & Trading Decisions, Infrastructure & Utility, Long-form Content Ge
- Gemini 3.5 Flash Lite
Good enough on 5/52 tasks at the 90% bar. Best value on 0 tasks. Doesn't qualify on any: Financial Analysis & Trading Decisions, Long-form Content Generation, Social & Promotional Content, Topic Organ
- Gemini 3.8 Flash
Good enough on 15/52 tasks at the 90% bar. Best value on 0 tasks. Best fit for: Infrastructure & Utility, Long-form Content Generation, Topic Organization & Clustering. Doesn't qualify on any: Financi
- GLM-5.3
Good enough on 22/52 tasks at the 90% bar. Best value on 0 tasks. Best fit for: Financial Analysis & Trading Decisions, Infrastructure & Utility, Relevance, Classification & Matching, Social & Promoti
- GLM-5.3 Flash
Good enough on 19/52 tasks at the 90% bar. Best value on 3 tasks. Best fit for: Infrastructure & Utility, Relevance, Classification & Matching, Social & Promotional Content, Content Summarization & Sy
- GPT-5.4 Nano
Good enough on 9/52 tasks at the 90% bar. Best value on 1 task. Doesn't qualify on any: Social & Promotional Content, Topic Organization & Clustering.
- GPT-5.6 Luna
Good enough on 31/52 tasks at the 90% bar. Best value on 27 tasks. Best fit for: Infrastructure & Utility, Long-form Content Generation.
- GPT-5.6 Sol
Good enough on 29/52 tasks at the 90% bar. Best value on 0 tasks. Best fit for: Structured Data & Fact Extraction, Financial Analysis & Trading Decisions, Infrastructure & Utility, Long-form Content G
- GPT-5.6 Terra
Good enough on 30/52 tasks at the 90% bar. Best value on 0 tasks. Best fit for: Structured Data & Fact Extraction, Long-form Content Generation, Social & Promotional Content, Content Summarization & S
- Grok 4.6
Good enough on 22/52 tasks at the 90% bar. Best value on 0 tasks. Best fit for: Infrastructure & Utility, Long-form Content Generation, Relevance, Classification & Matching, Social & Promotional Conte
- Meta Muse Spark 1.3
Good enough on 21/52 tasks at the 90% bar. Best value on 0 tasks. Best fit for: Infrastructure & Utility, Long-form Content Generation, Relevance, Classification & Matching.
- MiniMax M3
Good enough on 20/52 tasks at the 90% bar. Best value on 8 tasks. Best fit for: Financial Analysis & Trading Decisions, Long-form Content Generation.
- Moonshot Kimi K3
Good enough on 23/52 tasks at the 90% bar. Best value on 0 tasks. Best fit for: Financial Analysis & Trading Decisions, Infrastructure & Utility, Long-form Content Generation, Relevance, Classificatio
- NVIDIA Nemotron 3.5 Lightning
Good enough on 3/52 tasks at the 90% bar. Best value on 0 tasks. Doesn't qualify on any: Financial Analysis & Trading Decisions.
- NVIDIA Nemotron-3 Nano 30B-A3B
Good enough on 2/52 tasks at the 90% bar. Best value on 0 tasks. Doesn't qualify on any: Financial Analysis & Trading Decisions, Infrastructure & Utility, Social & Promotional Content, Content Summari
- NVIDIA Nemotron-3 Super 120B
Good enough on 8/52 tasks at the 90% bar. Best value on 1 task. Doesn't qualify on any: Financial Analysis & Trading Decisions, Infrastructure & Utility.
- NVIDIA Nemotron-3 Ultra 550B
Good enough on 12/52 tasks at the 90% bar. Best value on 0 tasks. Best fit for: Long-form Content Generation, Relevance, Classification & Matching.
- Qwen 3.7 Plus
Good enough on 21/52 tasks at the 90% bar. Best value on 0 tasks. Best fit for: Financial Analysis & Trading Decisions, Long-form Content Generation, Social & Promotional Content.
- Qwen 3.8 Flash
Good enough on 3/52 tasks at the 90% bar. Best value on 0 tasks. Doesn't qualify on any: Financial Analysis & Trading Decisions, Infrastructure & Utility, Social & Promotional Content, Content Summari
- Qwen 3.8 Max
Good enough on 9/52 tasks at the 90% bar. Best value on 0 tasks. Doesn't qualify on any: Social & Promotional Content.
- Tencent Hy3
Good enough on 13/52 tasks at the 90% bar. Best value on 2 tasks. Best fit for: Topic Organization & Clustering. Doesn't qualify on any: Financial Analysis & Trading Decisions, Long-form Content Gener
- Tencent Hy4 Preview
Good enough on 18/52 tasks at the 90% bar. Best value on 0 tasks. Best fit for: Infrastructure & Utility, Relevance, Classification & Matching, Social & Promotional Content, Content Summarization & Sy
- Thinking Machines Inkling
Good enough on 27/52 tasks at the 90% bar. Best value on 1 task. Best fit for: Financial Analysis & Trading Decisions, Infrastructure & Utility, Long-form Content Generation, Social & Promotional Cont
- Thinking Machines Inkling Small
Good enough on 26/52 tasks at the 90% bar. Best value on 1 task. Best fit for: Structured Data & Fact Extraction, Financial Analysis & Trading Decisions, Infrastructure & Utility, Social & Promotional