Best LLMs for S-1 TOC Extraction
Extracts the table of contents from an SEC S-1 registration statement. Recognises both formal TOC blocks and item-numbered standard sections (Item 1. Business, Item 1A. Risk Factors, MD&A, Financial Statements, etc.). Returns structured JSON.
Models
Frontier on this task: Gemini 3.5 Flash at 9.38 / 10. Quality bar at 90%: 8.45.
point-estimate floor (CI low) · upper CI (less certain) · Bars sorted by blended cost; best-value model first. Greyed rows are MEDIUM+ models whose point estimate clears the bar but whose CI low does not.
| Model | Quality score | CI low | Cost / 1k runs | vs best value |
|---|---|---|---|---|
| DeepSeek V4 Flash | 8.58 / 10 | 8.13 | $0.70 | best value |
| Qwen 3.5 Flash | 8.77 / 10 | 8.39 | $0.75 | 1.1x more expensive |
| Gemini 3.1 Flash Lite | 8.61 / 10 | 8.14 | $0.83 | 1.2x more expensive |
| NVIDIA Nemotron-3 Super 120B | 8.81 / 10 | 8.43 | $1.39 | 2x more expensive |
| MiniMax M3 | 9.04 / 10 | 8.92 | $1.43 | 2x more expensive |
| DeepSeek V4 Pro | 8.93 / 10 | 8.61 | $2.01 | 2.9x more expensive |
| Qwen 3.6 Flash | 8.52 / 10 | 8.34 | $3.66 | 5.2x more expensive |
| GPT-5.6 Terra | 8.93 / 10 | 8.57 | $4.86 | 6.9x more expensive |
| Claude Sonnet 5 | 8.54 / 10 | 8.33 | $6.49 | 9.2x more expensive |
| Qwen 3.6 Plus | 8.69 / 10 | 8.24 | $7.10 | 10x more expensive |
| Gemini 3.1 Pro Preview | 9.06 / 10 | 8.77 | $7.45 | 11x more expensive |
| GPT-5.6 Sol | 9.13 / 10 | 8.79 | $8.14 | 12x more expensive |
| Gemini 3.5 Flash | 9.38 / 10 | 9.10 | $9.40 | 13x more expensive |
| Kimi K2.6 | 8.61 / 10 | 8.11 | $15.31 | 22x more expensive |
| GPT-5.5 | 8.46 / 10 | 7.97 | $17.49 | 25x more expensive |
| Claude Opus 4.8 | 8.57 / 10 | 8.10 | $18.78 | 27x more expensive |
| Qwen 3.7 Plus | 8.19 / 10 | 7.89 | $1.47 | 2.1x more expensive |
| Grok 4.5 | 7.52 / 10 | 7.40 | $7.57 | 11x more expensive |
Cost breakdown
| Model | Quality | Confidence | Cost / 1k runs | Overpay | Mode |
|---|---|---|---|---|---|
| DeepSeek V4 Flash ★ DeepSeek | 8.58 / 10 CI [8.13, 9.03] | MEDIUM | $0.70 | best value | batch |
| Qwen 3.5 Flash Alibaba Cloud (DashScope) | 8.77 / 10 CI [8.39, 9.16] | MEDIUM | $0.75 | 1.1x | batch |
| Gemini 3.1 Flash Lite Gemini | 8.61 / 10 CI [8.14, 9.08] | MEDIUM | $0.83 | 1.2x | batch |
| NVIDIA Nemotron-3 Super 120B OpenRouter | 8.81 / 10 CI [8.43, 9.20] | MEDIUM | $1.39 | 2x | batch |
| MiniMax M3 MiniMax | 9.04 / 10 CI [8.92, 9.16] | RANKED | $1.43 | 2x | batch |
| DeepSeek V4 Pro DeepSeek | 8.93 / 10 CI [8.61, 9.25] | MEDIUM | $2.01 | 2.9x | batch |
| Qwen 3.6 Flash Alibaba Cloud (DashScope) | 8.52 / 10 CI [8.34, 8.69] | RANKED | $3.66 | 5.2x | batch |
| GPT-5.6 Terra OpenAI | 8.93 / 10 CI [8.57, 9.29] | MEDIUM | $4.86 | 6.9x | batch |
| Claude Sonnet 5 Anthropic | 8.54 / 10 CI [8.33, 8.76] | HIGH | $6.49 | 9.2x | batch |
| Qwen 3.6 Plus Alibaba Cloud (DashScope) | 8.69 / 10 CI [8.24, 9.15] | MEDIUM | $7.10 | 10x | batch |
| Gemini 3.1 Pro Preview Gemini | 9.06 / 10 CI [8.77, 9.36] | HIGH | $7.45 | 11x | batch |
| GPT-5.6 Sol OpenAI | 9.13 / 10 CI [8.79, 9.47] | MEDIUM | $8.14 | 12x | batch |
| Gemini 3.5 Flash best Gemini | 9.38 / 10 CI [9.10, 9.67] | HIGH | $9.40 | 13x | batch |
| Kimi K2.6 Moonshot AI | 8.61 / 10 CI [8.11, 9.10] | MEDIUM | $15.31 | 22x | batch |
| GPT-5.5 OpenAI | 8.46 / 10 CI [7.97, 8.95] | MEDIUM | $17.49 | 25x | batch |
| Claude Opus 4.8 Anthropic | 8.57 / 10 CI [8.10, 9.05] | MEDIUM | $18.78 | 27x | batch |
Overpay shows how much more you pay than the best-value model that clears the quality bar (marked ★) — the best-value good-enough option. "16x" means you overpay 16× — 16× that reference for no quality benefit above the bar. Typical call shape for this task: 3904 input tokens → 744 output tokens, EMA-tracked from production traffic. Cost is the observed, all-in $ per 1,000 task runs: each model's own measured usage on this task — output verbosity, thinking/reasoning tokens, cache reads and writes, and the spend on its billed failures — priced at current list rates and adjusted by the billing overhead we actually reconcile against provider invoices. Models that answer tersely cost what they actually cost; models that think at length pay for it. Not comparable to providers' advertised $/1M list rates — this is what running the task costs, not a per-token price.
Prompt templates
The system + user template pair used for this task.
SEC_S1_TOC_SYSTEM_PROMPT +
SEC_S1_TOC_USER_PROMPT
(1061 calls in window)
System prompt
You are an expert at analyzing SEC filings, particularly S-1 registration statements.
Your task is to find and extract the table of contents from the provided S-1 filing text. Look for:
1. A formal "TABLE OF CONTENTS" or "INDEX" section
2. Chapter/section listings with titles and page numbers
3. Standard S-1 sections like:
- PART I: INFORMATION REQUIRED IN PROSPECTUS
- Item 1. Business
- Item 1A. Risk Factors
- Item 2. Use of Proceeds
- Management's Discussion and Analysis
- Financial Statements
- PART II: INFORMATION NOT REQUIRED IN PROSPECTUS
- Signatures
- Exhibit Index
Extract the chapter titles exactly as they appear, along with any page numbers provided. If no formal table of contents exists, look for section headers throughout the text and identify the main sections.
Provide your analysis as a structured JSON object.
## Required Output Format
Your response MUST be a single, valid JSON object conforming to this schema:
```json
{schema_json_string}
```User prompt
Analyze the following S-1 filing text and extract the table of contents information:
**S-1 Filing Text (First 50,000 characters):**
```
{sample_text}
```
**Instructions:**
1. Look for a table of contents section in the text
2. Extract all chapter/section titles and their page numbers
3. If no formal TOC exists, identify major section headers throughout the text
4. Focus on investment-relevant sections like Business, Risk Factors, Financial Data, etc.
5. Provide the exact titles as they appear in the document
**JSON Output Format:**
The required JSON output schema is provided in the system prompt.