Best LLMs for Financial Analysis & Trading Decisions
Read business/financial docs deeply enough to make forward-looking judgments about opportunity and risk.
Read business/financial docs deeply enough to make forward-looking judgments about opportunity and risk.
Task-by-task breakdown
Catalyst And Scenario Analysis
Identifies and analyzes catalysts, scenarios, risks, timing, and potential effects for a subject using supplied sources and an analytical profile. Illustrative uses include analyzing scenarios around …
| Model | Quality (% of best) | Confidence | Overpay |
|---|---|---|---|
| Thinking Machines Inkling Small ★ | 91% | MEDIUM | best value |
| GPT-5.4 Nano | 90% | HIGH | 1.1x |
| Thinking Machines Inkling | 91% | MEDIUM | 2.9x |
| Gemini 3.5 Flash best | 100% | RANKED | 3x |
| GLM-5.3 | 90% | MEDIUM | 7.2x |
| Claude Haiku 4.5 | 91% | MEDIUM | 8.6x |
| Moonshot Kimi K3 | 93% | MEDIUM | 14x |
Multi Perspective Decision Synthesis
Synthesizes multiple analyses and adversarial perspectives into a structured decision, confidence assessment, and optional constrained action plan. Illustrative uses include synthesizing evidence for …
| Model | Quality (% of best) | Confidence | Overpay |
|---|---|---|---|
| GPT-5.6 Luna ★ | 92% | HIGH | best value |
| GLM-5.3 Flash | 97% | HIGH | 2.9x |
| Qwen 3.7 Plus | 94% | RANKED | 5x |
| NVIDIA Nemotron-3 Ultra 550B | 92% | MEDIUM | 8.4x |
| GPT-5.6 Terra | 92% | HIGH | 9.7x |
| Thinking Machines Inkling Small | 96% | RANKED | 12x |
| Thinking Machines Inkling | 98% | HIGH | 14x |
| DeepSeek V4 Pro | 91% | RANKED | 14x |
| Tencent Hy4 Preview | 96% | RANKED | 15x |
| Meta Muse Spark 1.3 | 92% | RANKED | 17x |
| GPT-5.6 Sol | 92% | HIGH | 18x |
| Claude Sonnet 5 | 94% | MEDIUM | 18x |
| GLM-5.3 | 99% | RANKED | 26x |
| Grok 4.6 | 92% | HIGH | 27x |
| Qwen 3.8 Max | 92% | HIGH | 40x |
| Claude Opus 5 best | 100% | RANKED | 43x |
| Moonshot Kimi K3 | 100% | RANKED | 57x |
Profiled Document Analysis
Analyzes a supplied document using caller-provided document conventions, evidence rules, analytical dimensions, materiality, and decision context. Illustrative uses include analyzing contract risk, …
| Model | Quality (% of best) | Confidence | Overpay |
|---|---|---|---|
| MiniMax M3 ★ | 95% | MEDIUM | best value |
| Claude Sonnet 5 | 94% | HIGH | 1.4x |
| GPT-5.6 Sol | 96% | MEDIUM | 1.4x |
| Gemini 3.5 Flash best | 100% | RANKED | 1.4x |
| Thinking Machines Inkling | 97% | HIGH | 1.9x |
Subject Configuration Generation
Converts a free-form subject or organization brief into a normalized research configuration containing identity, type, industry, focus areas, regions, identifiers, and suggested cadence. Illustrative …
| Model | Quality (% of best) | Confidence | Overpay |
|---|---|---|---|
| MiniMax M3 ★ | 97% | RANKED | best value |
| Qwen 3.7 Plus | 98% | RANKED | 4.6x |
| Claude Sonnet 5 best | 100% | HIGH | 8.1x |
| Gemini 3.5 Flash | 96% | RANKED | 8.7x |
| Thinking Machines Inkling | 98% | HIGH | 11x |
Profiled Document Section Analysis
Analyzes one document section using supplied document conventions, analytical dimensions, evidence boundaries, and decision context. Illustrative uses include analyzing a contract clause, software …
| Model | Quality (% of best) | Confidence | Overpay |
|---|---|---|---|
| GPT-5.6 Luna ★ | 95% | RANKED | best value |
| MiniMax M3 | 97% | RANKED | 4.2x |
| GPT-5.4 Nano | 90% | HIGH | 4.6x |
| GLM-5.3 Flash best | 100% | RANKED | 5.4x |
| Thinking Machines Inkling Small | 96% | RANKED | 9.6x |
| Qwen 3.7 Plus | 92% | MEDIUM | 9.7x |
| GPT-5.6 Terra | 96% | RANKED | 18x |
| Gemini 3.5 Flash | 94% | HIGH | 18x |
| Meta Muse Spark 1.3 | 95% | HIGH | 26x |
| Thinking Machines Inkling | 98% | RANKED | 26x |
| GPT-5.6 Sol | 97% | RANKED | 29x |
| Claude Sonnet 5 | 96% | MEDIUM | 34x |
| Tencent Hy4 Preview | 98% | MEDIUM | 36x |
| GLM-5.3 | 97% | HIGH | 41x |
| Grok 4.6 | 98% | MEDIUM | 49x |
| Moonshot Kimi K3 | 99% | RANKED | 90x |
Confidence — how sure we are about the quality score (more judgments + more agreement = higher confidence): RANKED many independent judges scored this model's outputs and their agreement is very high (most confident) — HIGH many judges have scored it and they mostly agree (well-pinned) — MEDIUM enough judges have weighed in to publish, but they disagree more than we'd like (treat with a small grain of salt). LOW-confidence cells are hidden everywhere on the site. See the methodology for the exact thresholds.