Cost mode:

Semantic similarity judgment: does this thing belong in that bucket / match that target?

14 capabilities in this category.

Task-by-task breakdown

Content Domain Suggestion

Picks 1-3 subject/vertical domain tags for an analysis template (e.g. 'Menopause', 'Femtech', 'Biotech') from its name, description, and chapter content. Names the WHAT, not the HOW — explicitly …

ModelQuality (% of best)ConfidenceOverpay
Meta Muse Spark 1.1 best100%MEDIUMbest value

Task detail →

Author Living-Person Safety Check

Postmortem-publicity-rights safety check for AI author personas. Determines whether the real figure behind a persona name has been deceased for ≥100 years (the threshold safely clears CA §3344.1, TN …

ModelQuality (% of best)ConfidenceOverpay
DeepSeek V4 Flash 96%MEDIUMbest value
DeepSeek V4 Pro best100%RANKED5.8x
Qwen 3.6 Plus93%MEDIUM11x
GPT-5.6 Terra95%MEDIUM12x
Gemini 3.1 Pro Preview94%HIGH16x
Gemini 3.5 Flash96%RANKED18x
Kimi K2.696%MEDIUM23x
GPT-5.6 Sol97%HIGH23x
Meta Muse Spark 1.198%HIGH25x
Grok 4.593%RANKED43x
GPT-5.598%MEDIUM59x

Task detail →

Language Detection

Identifies the primary language of a text snippet. Returns only the two-letter ISO 639-1 code (e.g. 'en', 'es', 'zh'). Used upstream of the translation pipeline.

ModelQuality (% of best)ConfidenceOverpay
NVIDIA Nemotron-3 Nano 30B-A3B 95%MEDIUMbest value
Gemini 3.1 Flash Lite99%RANKED1.4x
DeepSeek V4 Flash98%RANKED1.4x
GPT-5.4 Nano99%RANKED1.5x
GPT-5.4 Mini99%RANKED2.7x
NVIDIA Nemotron-3 Super 120B98%HIGH3.1x
Tencent Hy398%RANKED3.8x
MiniMax M399%RANKED5.8x
DeepSeek V4 Pro98%RANKED6.1x
GPT-5.6 Luna99%RANKED7.2x
NVIDIA Nemotron-3 Ultra 550B99%RANKED14x
Claude Haiku 4.599%RANKED16x
Qwen 3.5 Flash99%RANKED17x
GPT-5.6 Terra99%RANKED17x
Gemini 3.5 Flash99%RANKED35x
Claude Sonnet 5 best100%RANKED35x
GPT-5.599%RANKED37x
GPT-5.6 Sol99%RANKED38x
Qwen 3.7 Plus99%RANKED39x
Gemini 3.1 Pro Preview99%RANKED48x
Qwen 3.6 Plus99%RANKED50x
Claude Sonnet 4.699%RANKED52x
Qwen 3.6 Flash99%RANKED66x
Meta Muse Spark 1.198%HIGH76x
Claude Opus 4.899%RANKED86x
Grok 4.599%RANKED89x
Kimi K2.698%RANKED143x

Task detail →

Subreddit Selection for Research

Picks relevant, active subreddits for researching a subject. Balances large communities (more content) with niche ones (more focused). Filters for accessibility (public, not quarantined) and quality …

ModelQuality (% of best)ConfidenceOverpay
GPT-5.6 Terra 90%MEDIUMbest value
GPT-5.6 Sol best100%MEDIUM3.3x

Task detail →

Topic Grouping and Client Matching

Groups workflow topics (per-article) under broader client topics (persistent categories) — either existing or new. Multiple workflow topics can and should share one client topic when they cover …

ModelQuality (% of best)ConfidenceOverpay
DeepSeek V4 Flash 95%RANKEDbest value
DeepSeek V4 Pro94%RANKED2.4x
GPT-5.6 Luna93%MEDIUM2.7x
Qwen 3.7 Plus95%HIGH4.3x
Qwen 3.6 Plus91%HIGH4.6x
Gemini 3.1 Pro Preview93%RANKED6.2x
Kimi K2.692%HIGH10x
GPT-5.6 Terra98%HIGH11x
Gemini 3.5 Flash99%RANKED12x
Grok 4.598%RANKED14x
Meta Muse Spark 1.196%HIGH14x
GPT-5.6 Sol best100%HIGH16x
GPT-5.594%MEDIUM21x

Task detail →

Engagement Triage

Decide engage/ignore + risk + angle for one social post.

ModelQuality (% of best)ConfidenceOverpay
DeepSeek V4 Flash 92%HIGHbest value
GPT-5.4 Nano94%RANKED1.6x
Gemini 3.1 Flash Lite94%RANKED1.6x
Tencent Hy394%MEDIUM1.8x
GPT-5.4 Mini93%HIGH2.8x
NVIDIA Nemotron-3 Super 120B93%MEDIUM3.8x
MiniMax M398%RANKED4.7x
GPT-5.6 Luna97%HIGH8.1x
Qwen 3.5 Flash94%RANKED9.1x
Claude Haiku 4.598%RANKED11x
GPT-5.6 Terra98%RANKED15x
Qwen 3.7 Plus98%RANKED17x
Gemini 3.5 Flash96%RANKED23x
Gemini 3.1 Pro Preview95%RANKED26x
Claude Sonnet 597%RANKED28x
NVIDIA Nemotron-3 Ultra 550B96%HIGH34x
Claude Sonnet 4.6 best100%RANKED34x
GPT-5.6 Sol93%MEDIUM38x
Qwen 3.6 Plus95%HIGH43x
Grok 4.598%RANKED46x
GPT-5.595%RANKED55x
Meta Muse Spark 1.199%HIGH57x
Claude Opus 4.898%RANKED64x
Kimi K2.697%RANKED77x

Task detail →

X Post Selection

Picks the best N X.com posts to publish within a daily budget. Ranks by engagement potential, topic diversity (avoid bunching), content quality, and timeliness; returns only the chosen post IDs.

ModelQuality (% of best)ConfidenceOverpay
DeepSeek V4 Flash 100%RANKEDbest value
Gemini 3.1 Flash Lite97%HIGH1.6x
GPT-5.4 Mini93%MEDIUM2.1x
Qwen 3.5 Flash97%HIGH2.7x
MiniMax M396%RANKED4.7x
DeepSeek V4 Pro96%HIGH4.9x
GPT-5.6 Luna96%HIGH5.1x
Tencent Hy393%HIGH6x
Qwen 3.6 Plus97%RANKED7.5x
Claude Haiku 4.592%MEDIUM7.6x
GPT-5.6 Terra98%RANKED11x
Gemini 3.1 Pro Preview97%HIGH12x
Kimi K2.699%RANKED17x
Claude Sonnet 595%RANKED20x
Claude Sonnet 4.6100%HIGH23x
Qwen 3.7 Plus95%RANKED24x
GPT-5.5 best100%RANKED25x
GPT-5.6 Sol97%RANKED27x
Qwen 3.6 Flash98%RANKED33x
Gemini 3.5 Flash97%RANKED37x
Grok 4.595%RANKED89x
Meta Muse Spark 1.196%HIGH97x

Task detail →

Confidence — how sure we are about the quality score (more judgments + more agreement = higher confidence): RANKED many independent judges scored this model's outputs and their agreement is very high (most confident) — HIGH many judges have scored it and they mostly agree (well-pinned) — MEDIUM enough judges have weighed in to publish, but they disagree more than we'd like (treat with a small grain of salt). LOW-confidence cells are hidden everywhere on the site. See the methodology for the exact thresholds.