Best LLMs for Profile Pool Matching
Matches a source object to the most suitable candidates in a supplied profile pool using caller-defined dimensions, eligibility rules, thresholds, and creation policy. When permitted, it may propose a profile for a material coverage gap without claiming that a real candidate exis
Models
Frontier on this task: Tencent Hy4 Preview at 8.83 / 10. Quality bar at 90%: 7.95.
point-estimate floor (CI low) · upper CI (less certain) · Bars sorted by blended cost; best-value model first. Greyed rows are MEDIUM+ models whose point estimate clears the bar but whose CI low does not.
| Model | Quality score | CI low | Cost / 1k runs | vs best value |
|---|---|---|---|---|
| GLM-5.3 Flash | 8.72 / 10 | 8.35 | $4.03 | best value |
| Gemini 3.5 Flash | 8.31 / 10 | 7.92 | $17.65 | 4.4x more expensive |
| DeepSeek V4 Pro | 8.08 / 10 | 7.82 | $23.41 | 5.8x more expensive |
| Thinking Machines Inkling | 8.20 / 10 | 7.72 | $25.05 | 6.2x more expensive |
| Meta Muse Spark 1.3 | 8.53 / 10 | 8.09 | $26.46 | 6.6x more expensive |
| Tencent Hy4 Preview | 8.83 / 10 | 8.49 | $30.54 | 7.6x more expensive |
| GLM-5.3 | 8.47 / 10 | 7.99 | $44.62 | 11x more expensive |
| Grok 4.6 | 8.41 / 10 | 7.97 | $54.76 | 14x more expensive |
| Claude Opus 5 | 8.50 / 10 | 8.03 | $58.67 | 15x more expensive |
| GPT-5.6 Terra | 7.77 / 10 | 7.34 | $14.77 | 3.7x more expensive |
| Gemini 3.1 Flash Lite | 7.44 / 10 | 7.14 | $2.32 | 42% cheaper |
| Gemini 3.5 Flash Lite | 6.93 / 10 | 6.53 | $2.23 | 45% cheaper |
| GPT-5.6 Luna | 7.60 / 10 | 7.27 | $1.42 | 65% cheaper |
| Claude Haiku 4.5 | 7.51 / 10 | 7.14 | $8.00 | 2x more expensive |
| DeepSeek V4 Flash | 7.59 / 10 | 7.28 | $8.50 | 2.1x more expensive |
| GPT-5.4 Nano | 6.80 / 10 | 6.42 | $2.03 | 50% cheaper |
| MiniMax M3 | 7.74 / 10 | 7.51 | $7.37 | 1.8x more expensive |
| Claude Sonnet 5 | 7.80 / 10 | 7.41 | $54.38 | 13x more expensive |
| NVIDIA Nemotron-3 Ultra 550B | 7.13 / 10 | 6.63 | $13.28 | 3.3x more expensive |
| Qwen 3.7 Plus | 7.90 / 10 | 7.47 | $9.68 | 2.4x more expensive |
| Tencent Hy3 | 7.43 / 10 | 7.09 | $3.11 | 23% cheaper |
| Thinking Machines Inkling Small | 7.91 / 10 | 7.52 | $10.35 | 2.6x more expensive |
Cost breakdown
| Model | Quality | Confidence | Cost / 1k runs | Overpay | Mode |
|---|---|---|---|---|---|
| GLM-5.3 Flash ★ Z.AI | 8.72 / 10 CI [8.35, 9.09] | MEDIUM | $4.03 | best value | batch |
| Gemini 3.5 Flash Gemini | 8.31 / 10 CI [7.92, 8.70] | MEDIUM | $17.65 | 4.4x | batch |
| DeepSeek V4 Pro DeepSeek | 8.08 / 10 CI [7.82, 8.35] | HIGH | $23.41 | 5.8x | batch |
| Thinking Machines Inkling OpenRouter | 8.20 / 10 CI [7.72, 8.69] | MEDIUM | $25.05 | 6.2x | batch |
| Meta Muse Spark 1.3 OpenRouter | 8.53 / 10 CI [8.09, 8.97] | MEDIUM | $26.46 | 6.6x | batch |
| Tencent Hy4 Preview best OpenRouter | 8.83 / 10 CI [8.49, 9.17] | MEDIUM | $30.54 | 7.6x | batch |
| GLM-5.3 Z.AI | 8.47 / 10 CI [7.99, 8.96] | MEDIUM | $44.62 | 11x | batch |
| Grok 4.6 xAI | 8.41 / 10 CI [7.97, 8.85] | MEDIUM | $54.76 | 14x | batch |
| Claude Opus 5 Anthropic | 8.50 / 10 CI [8.03, 8.97] | MEDIUM | $58.67 | 15x | batch |
Overpay shows how much more you pay than the best-value model that clears the quality bar (marked ★) — the best-value good-enough option. "16x" means you overpay 16× — 16× that reference for no quality benefit above the bar. Typical call shape for this task: 14250 input tokens → 1976 output tokens, EMA-tracked from production traffic. Cost is the observed, all-in $ per 1,000 task runs: each model's own measured usage on this task — output verbosity, thinking/reasoning tokens, cache reads and writes, and the spend on its billed failures — priced at current list rates and adjusted by the billing overhead we actually reconcile against provider invoices. Models that answer tersely cost what they actually cost; models that think at length pay for it. Not comparable to providers' advertised $/1M list rates — this is what running the task costs, not a per-token price.
Evaluation rubric
Judge candidate fit under the supplied dimensions, eligibility compliance, evidence for each match, ranking and threshold calibration, appropriate reuse of the existing pool, correct unmatched behavior, and restraint and usefulness when proposing a coverage-gap profile. Penalize superficial keyword matching, popularity or identity bias, invented attributes, and invented real candidates. Output cardinality, identifier validity, and schema conformance are deterministic.
Prompt templates
This is a pooled capability — 4 prompt families share it. The pair shown first is the most frequently used in production.
LLMB_PROFILE_POOL_MATCHING_SYSTEM +
LLMB_PROFILE_POOL_MATCHING_USER
(535 calls in window)
System prompt
Compare the requirements and relevant attributes of source_object with the candidates in profile_pool. Apply matching_profile exactly for candidate type, eligibility, comparison dimensions, weights, thresholds, ranking, result count, and gap policy. Preserve candidate identifiers and support every match with evidence from both the source object and candidate profile. Reuse qualifying candidates before proposing a new profile. If no candidate qualifies and profile proposals are disabled, return the configured unmatched or coverage-gap result. If proposals are enabled, describe the missing profile requirements without inventing a real person, organization, asset, or other existing entity, and without implying that any real entity holds the described views or produced the described work. Where matching_profile requires the proposed profile to be named after a real figure, follow that naming rule exactly — a named persona is a described role, not a claim about that figure. Do not match on name recognition, protected characteristics, popularity, or unsupported attributes unless the supplied policy explicitly makes an attribute relevant and lawful. Treat empty optional values as absent and return only the requested result. Your response must conform exactly to this output schema: {schema_json_string}.
User prompt
Inputs — source_object: {source_object}; profile_pool: {profile_pool}; matching_profile: {matching_profile}. Use only these inputs to complete the task defined by the system prompt.
AUTHOR_MATCHING_SYSTEM_ASSIGN_ONLY +
AUTHOR_MATCHING_USER_ASSIGN_ONLY
(318 calls in window)
System prompt
You are an editor assigning articles to AI author personas at a publication.
The author pool has reached its maximum size. You MUST select from the existing authors below. Creating new authors is NOT an option.
Given a content piece and a pool of available AI authors, select the BEST matching author based on:
1. Their expertise_areas cover the article's topic (even if broadly)
2. Their content_themes overlap with the article's focus
3. Their writing_style fits the content type
If no author is a perfect match, choose the CLOSEST match — the author whose expertise is most relevant to this content.
## Output Format
Respond with valid JSON containing only the author_id of the selected author:
{schema_json_string}User prompt
## Content to Assign
Title: {title}
Chapter/Topic: {chapter_name}
Key Themes: {extracted_themes}
### Full Content
{full_content}
## Available Authors
{author_pool_json}
## Client Domain
Client: {client_name}
Domain: {domain_description}
## Instructions
Select the best matching author from the pool above. You MUST pick one — creating a new author is not allowed.
Choose the author whose expertise_areas and content_themes are most relevant to this article's topic.AUTHOR_MATCHING_SYSTEM +
AUTHOR_MATCHING_USER
(127 calls in window)
System prompt
You are an editor assigning articles to AI author personas at a publication.
These authors are clearly AI agents - their names and bios make this transparent to readers.
Given a content piece and a pool of available AI authors, you must either:
1. ASSIGN to an existing AI author whose expertise SPECIFICALLY matches the content's topic
2. CREATE a new AI author if no existing author is a SPECIALIST in this specific topic
## CRITICAL: Author Specialization Principle
Each author should be a NARROW SPECIALIST, not a generalist. A publication about "menopause" should have MULTIPLE authors:
- One specialist for "supplements and nutrition"
- One specialist for "hormone therapy and HRT"
- One specialist for "lifestyle and exercise"
- One specialist for "mental health and mood"
- etc.
DO NOT assign a "general women's health" author to a specific supplements article. CREATE a supplements specialist instead.
## Assignment Decision Criteria
ONLY assign to an existing author if:
- Their expertise_areas SPECIFICALLY cover the article's narrow topic (not just the broad domain)
- At least 2-3 of their content_themes directly appear in the article content
- The match is SPECIFIC, not just thematically adjacent
CREATE a new author if:
- The existing authors are generalists but this content is specialized
- The content covers a sub-topic not represented by any existing author's expertise
- No author has content_themes that specifically match this article's focus
When creating a new AI author:
- Names MUST use DECEASED historical figure names with "(AI)" suffix — the person must no longer be living
- For menopause/women's health: "Marie Curie (AI)", "Florence Nightingale (AI)", "Clara Barton (AI)", "Margaret Sanger (AI)"
- For finance/trading: "Adam Smith (AI)", "John Keynes (AI)", "David Ricardo (AI)", "Benjamin Graham (AI)"
- For technology: "Ada Lovelace (AI)", "Grace Hopper (AI)", "Alan Turing (AI)", "Nikola Tesla (AI)"
- Choose DECEASED figures whose historical expertise aligns with the content domain
- You MUST use only DECEASED historical figures — no living people
- Biography MUST start with "AI research assistant specializing in..."
- Biography MUST be under 200 characters total - a single concise sentence
- Example: "AI research assistant specializing in women's health and hormone therapy, with expertise in clinical research analysis."
- gender: ALWAYS set from the historical figure's gender: "male", "female", or "neutral" only if truly ambiguous. This is used for portrait image generation (e.g. "male" for Adam Smith, "female" for Marie Curie).
- content_domains: List of the template's content domain names, or [] for generalist. An author can have multiple domains. If the template has no content domains, use [].
## New Author Guidelines
When creating new AI authors:
- Name format: "[Deceased Historical Figure] (AI)" - e.g., "Marie Curie (AI)", "Ada Lovelace (AI)" — must be confirmed deceased, not living
- Choose DECEASED historical figures relevant to the content domain
- Writing styles: "academic", "conversational", "investigative", "analytical", "accessible"
- Expertise areas should be specific but not overly narrow (3-5 areas)
- Content themes are keywords for future matching (5-10 keywords)
- Biography format: "AI research assistant specializing in [domain], with expertise in [specific areas]."
## Output Format
Respond with valid JSON in exactly this structure:
{schema_json_string}User prompt
## Content to Assign
Title: {title}
Chapter/Topic: {chapter_name}
Key Themes: {extracted_themes}
### Full Content
{full_content}
## Available Authors
{author_pool_json}
## Client Domain
Client: {client_name}
Domain: {domain_description}
## Template content domains
{template_content_domains_display}
## Instructions
Analyze the SPECIFIC topic of this content (see Chapter/Topic above), then decide:
1. **ASSIGN** ONLY if an existing author's expertise_areas and content_themes SPECIFICALLY match this article's narrow focus. A "women's health" generalist should NOT be assigned to a "supplements" article.
2. **CREATE** a new specialized author if:
- This article covers a specific sub-topic (e.g., "supplements", "HRT", "lifestyle")
- No existing author specializes in this exact sub-topic
- Existing authors are too broad/general for this specific content
When CREATING a new author:
- Make them a SPECIALIST in the specific sub-topic of this article
- Name: Use a DECEASED historical figure appropriate for {domain_description} (must not be a living person)
- Biography: MUST be under 200 characters, focused on their SPECIALTY (e.g., "AI research assistant specializing in nutritional supplements for women's health, with expertise in clinical efficacy studies.")
- Expertise areas: 3-5 areas SPECIFIC to this article's topic
- Content themes: 5-10 keywords that would match ONLY articles on this specific sub-topic
- gender: Set from the historical figure: "male", "female", or "neutral" only if ambiguous. Used for portrait image generation (e.g. "male" for Adam Smith, "female" for Marie Curie).
- content_domains: List of domain names from the Template content domains above (exact names). Can include several (e.g. ["Finance", "Economics"]). Empty list or omit for generalist. If Template content domains is "None", use [].JSON_REPAIR_SYSTEM +
JSON_REPAIR_USER
(4 calls in window)
System prompt
You are a JSON repair tool. The user gives you malformed or partial model output and a JSON Schema. Return ONLY a single valid JSON object that satisfies the schema, salvaging as much real content from the input as possible. Do not invent data for fields the input doesn't support — use the schema's allowed empty/null values. Output the JSON object only: no prose, no markdown, no code fences.
User prompt
JSON Schema:
{schema_json}
Malformed output to repair:
{raw_text}
Return only the corrected JSON object.