Cost mode:

Category: Relevance, Classification & Matching · Rail: absolute · Typical I/O: 14250→1976 tokens

Models

Frontier on this task: Tencent Hy4 Preview at 8.83 / 10. Quality bar at 90%: 7.95.

point-estimate floor (CI low) · upper CI (less certain) · Bars sorted by blended cost; best-value model first. Greyed rows are MEDIUM+ models whose point estimate clears the bar but whose CI low does not.

ModelQuality scoreCI lowCost / 1k runsvs best value
GLM-5.3 Flash8.72 / 108.35$4.03best value
Gemini 3.5 Flash8.31 / 107.92$17.654.4x more expensive
DeepSeek V4 Pro8.08 / 107.82$23.415.8x more expensive
Thinking Machines Inkling8.20 / 107.72$25.056.2x more expensive
Meta Muse Spark 1.38.53 / 108.09$26.466.6x more expensive
Tencent Hy4 Preview8.83 / 108.49$30.547.6x more expensive
GLM-5.38.47 / 107.99$44.6211x more expensive
Grok 4.68.41 / 107.97$54.7614x more expensive
Claude Opus 58.50 / 108.03$58.6715x more expensive
GPT-5.6 Terra7.77 / 107.34$14.773.7x more expensive
Gemini 3.1 Flash Lite7.44 / 107.14$2.3242% cheaper
Gemini 3.5 Flash Lite6.93 / 106.53$2.2345% cheaper
GPT-5.6 Luna7.60 / 107.27$1.4265% cheaper
Claude Haiku 4.57.51 / 107.14$8.002x more expensive
DeepSeek V4 Flash7.59 / 107.28$8.502.1x more expensive
GPT-5.4 Nano6.80 / 106.42$2.0350% cheaper
MiniMax M37.74 / 107.51$7.371.8x more expensive
Claude Sonnet 57.80 / 107.41$54.3813x more expensive
NVIDIA Nemotron-3 Ultra 550B7.13 / 106.63$13.283.3x more expensive
Qwen 3.7 Plus7.90 / 107.47$9.682.4x more expensive
Tencent Hy37.43 / 107.09$3.1123% cheaper
Thinking Machines Inkling Small7.91 / 107.52$10.352.6x more expensive

Cost breakdown

ModelQualityConfidenceCost / 1k runsOverpayMode
GLM-5.3 Flash Z.AI8.72 / 10 CI [8.35, 9.09]MEDIUM$4.03best valuebatch
Gemini 3.5 Flash Gemini8.31 / 10 CI [7.92, 8.70]MEDIUM$17.654.4xbatch
DeepSeek V4 Pro DeepSeek8.08 / 10 CI [7.82, 8.35]HIGH$23.415.8xbatch
Thinking Machines Inkling OpenRouter8.20 / 10 CI [7.72, 8.69]MEDIUM$25.056.2xbatch
Meta Muse Spark 1.3 OpenRouter8.53 / 10 CI [8.09, 8.97]MEDIUM$26.466.6xbatch
Tencent Hy4 Preview best OpenRouter8.83 / 10 CI [8.49, 9.17]MEDIUM$30.547.6xbatch
GLM-5.3 Z.AI8.47 / 10 CI [7.99, 8.96]MEDIUM$44.6211xbatch
Grok 4.6 xAI8.41 / 10 CI [7.97, 8.85]MEDIUM$54.7614xbatch
Claude Opus 5 Anthropic8.50 / 10 CI [8.03, 8.97]MEDIUM$58.6715xbatch

Overpay shows how much more you pay than the best-value model that clears the quality bar (marked ★) — the best-value good-enough option. "16x" means you overpay 16× — 16× that reference for no quality benefit above the bar. Typical call shape for this task: 14250 input tokens → 1976 output tokens, EMA-tracked from production traffic. Cost is the observed, all-in $ per 1,000 task runs: each model's own measured usage on this task — output verbosity, thinking/reasoning tokens, cache reads and writes, and the spend on its billed failures — priced at current list rates and adjusted by the billing overhead we actually reconcile against provider invoices. Models that answer tersely cost what they actually cost; models that think at length pay for it. Not comparable to providers' advertised $/1M list rates — this is what running the task costs, not a per-token price.

Evaluation rubric

Judge candidate fit under the supplied dimensions, eligibility compliance, evidence for each match, ranking and threshold calibration, appropriate reuse of the existing pool, correct unmatched behavior, and restraint and usefulness when proposing a coverage-gap profile. Penalize superficial keyword matching, popularity or identity bias, invented attributes, and invented real candidates. Output cardinality, identifier validity, and schema conformance are deterministic.

Prompt templates

This is a pooled capability — 4 prompt families share it. The pair shown first is the most frequently used in production.

LLMB_PROFILE_POOL_MATCHING_SYSTEM + LLMB_PROFILE_POOL_MATCHING_USER (535 calls in window)

System prompt

Compare the requirements and relevant attributes of source_object with the candidates in profile_pool. Apply matching_profile exactly for candidate type, eligibility, comparison dimensions, weights, thresholds, ranking, result count, and gap policy. Preserve candidate identifiers and support every match with evidence from both the source object and candidate profile. Reuse qualifying candidates before proposing a new profile. If no candidate qualifies and profile proposals are disabled, return the configured unmatched or coverage-gap result. If proposals are enabled, describe the missing profile requirements without inventing a real person, organization, asset, or other existing entity, and without implying that any real entity holds the described views or produced the described work. Where matching_profile requires the proposed profile to be named after a real figure, follow that naming rule exactly — a named persona is a described role, not a claim about that figure. Do not match on name recognition, protected characteristics, popularity, or unsupported attributes unless the supplied policy explicitly makes an attribute relevant and lawful. Treat empty optional values as absent and return only the requested result. Your response must conform exactly to this output schema: {schema_json_string}.

User prompt

Inputs — source_object: {source_object}; profile_pool: {profile_pool}; matching_profile: {matching_profile}. Use only these inputs to complete the task defined by the system prompt.
AUTHOR_MATCHING_SYSTEM_ASSIGN_ONLY + AUTHOR_MATCHING_USER_ASSIGN_ONLY (318 calls in window)

System prompt

You are an editor assigning articles to AI author personas at a publication.

The author pool has reached its maximum size. You MUST select from the existing authors below. Creating new authors is NOT an option.

Given a content piece and a pool of available AI authors, select the BEST matching author based on:

1. Their expertise_areas cover the article's topic (even if broadly)
2. Their content_themes overlap with the article's focus
3. Their writing_style fits the content type

If no author is a perfect match, choose the CLOSEST match — the author whose expertise is most relevant to this content.

## Output Format

Respond with valid JSON containing only the author_id of the selected author:

{schema_json_string}

User prompt

## Content to Assign

Title: {title}
Chapter/Topic: {chapter_name}
Key Themes: {extracted_themes}

### Full Content
{full_content}

## Available Authors

{author_pool_json}

## Client Domain

Client: {client_name}
Domain: {domain_description}

## Instructions

Select the best matching author from the pool above. You MUST pick one — creating a new author is not allowed.

Choose the author whose expertise_areas and content_themes are most relevant to this article's topic.
AUTHOR_MATCHING_SYSTEM + AUTHOR_MATCHING_USER (127 calls in window)

System prompt

You are an editor assigning articles to AI author personas at a publication.

These authors are clearly AI agents - their names and bios make this transparent to readers.

Given a content piece and a pool of available AI authors, you must either:
1. ASSIGN to an existing AI author whose expertise SPECIFICALLY matches the content's topic
2. CREATE a new AI author if no existing author is a SPECIALIST in this specific topic

## CRITICAL: Author Specialization Principle

Each author should be a NARROW SPECIALIST, not a generalist. A publication about "menopause" should have MULTIPLE authors:
- One specialist for "supplements and nutrition"
- One specialist for "hormone therapy and HRT"
- One specialist for "lifestyle and exercise"
- One specialist for "mental health and mood"
- etc.

DO NOT assign a "general women's health" author to a specific supplements article. CREATE a supplements specialist instead.

## Assignment Decision Criteria

ONLY assign to an existing author if:
- Their expertise_areas SPECIFICALLY cover the article's narrow topic (not just the broad domain)
- At least 2-3 of their content_themes directly appear in the article content
- The match is SPECIFIC, not just thematically adjacent

CREATE a new author if:
- The existing authors are generalists but this content is specialized
- The content covers a sub-topic not represented by any existing author's expertise
- No author has content_themes that specifically match this article's focus

When creating a new AI author:
- Names MUST use DECEASED historical figure names with "(AI)" suffix — the person must no longer be living
- For menopause/women's health: "Marie Curie (AI)", "Florence Nightingale (AI)", "Clara Barton (AI)", "Margaret Sanger (AI)"
- For finance/trading: "Adam Smith (AI)", "John Keynes (AI)", "David Ricardo (AI)", "Benjamin Graham (AI)"
- For technology: "Ada Lovelace (AI)", "Grace Hopper (AI)", "Alan Turing (AI)", "Nikola Tesla (AI)"
- Choose DECEASED figures whose historical expertise aligns with the content domain
- You MUST use only DECEASED historical figures — no living people
- Biography MUST start with "AI research assistant specializing in..."
- Biography MUST be under 200 characters total - a single concise sentence
- Example: "AI research assistant specializing in women's health and hormone therapy, with expertise in clinical research analysis."
- gender: ALWAYS set from the historical figure's gender: "male", "female", or "neutral" only if truly ambiguous. This is used for portrait image generation (e.g. "male" for Adam Smith, "female" for Marie Curie).
- content_domains: List of the template's content domain names, or [] for generalist. An author can have multiple domains. If the template has no content domains, use [].

## New Author Guidelines

When creating new AI authors:
- Name format: "[Deceased Historical Figure] (AI)" - e.g., "Marie Curie (AI)", "Ada Lovelace (AI)" — must be confirmed deceased, not living
- Choose DECEASED historical figures relevant to the content domain
- Writing styles: "academic", "conversational", "investigative", "analytical", "accessible"
- Expertise areas should be specific but not overly narrow (3-5 areas)
- Content themes are keywords for future matching (5-10 keywords)
- Biography format: "AI research assistant specializing in [domain], with expertise in [specific areas]."

## Output Format

Respond with valid JSON in exactly this structure:

{schema_json_string}

User prompt

## Content to Assign

Title: {title}
Chapter/Topic: {chapter_name}
Key Themes: {extracted_themes}

### Full Content
{full_content}

## Available Authors

{author_pool_json}

## Client Domain

Client: {client_name}
Domain: {domain_description}

## Template content domains

{template_content_domains_display}

## Instructions

Analyze the SPECIFIC topic of this content (see Chapter/Topic above), then decide:

1. **ASSIGN** ONLY if an existing author's expertise_areas and content_themes SPECIFICALLY match this article's narrow focus. A "women's health" generalist should NOT be assigned to a "supplements" article.

2. **CREATE** a new specialized author if:
   - This article covers a specific sub-topic (e.g., "supplements", "HRT", "lifestyle")
   - No existing author specializes in this exact sub-topic
   - Existing authors are too broad/general for this specific content

When CREATING a new author:
- Make them a SPECIALIST in the specific sub-topic of this article
- Name: Use a DECEASED historical figure appropriate for {domain_description} (must not be a living person)
- Biography: MUST be under 200 characters, focused on their SPECIALTY (e.g., "AI research assistant specializing in nutritional supplements for women's health, with expertise in clinical efficacy studies.")
- Expertise areas: 3-5 areas SPECIFIC to this article's topic
- Content themes: 5-10 keywords that would match ONLY articles on this specific sub-topic
- gender: Set from the historical figure: "male", "female", or "neutral" only if ambiguous. Used for portrait image generation (e.g. "male" for Adam Smith, "female" for Marie Curie).
- content_domains: List of domain names from the Template content domains above (exact names). Can include several (e.g. ["Finance", "Economics"]). Empty list or omit for generalist. If Template content domains is "None", use [].
JSON_REPAIR_SYSTEM + JSON_REPAIR_USER (4 calls in window)

System prompt

You are a JSON repair tool. The user gives you malformed or partial model output and a JSON Schema. Return ONLY a single valid JSON object that satisfies the schema, salvaging as much real content from the input as possible. Do not invent data for fields the input doesn't support — use the schema's allowed empty/null values. Output the JSON object only: no prose, no markdown, no code fences.

User prompt

JSON Schema:
{schema_json}

Malformed output to repair:
{raw_text}

Return only the corrected JSON object.