Cost mode:

Cross-check assertions against supplied evidence and apply a fixed rubric or checklist to issue a cited approve/reject verdict on accuracy, support, and wording.

0 capabilities in this category.

Task-by-task breakdown

Confidence — how sure we are about the quality score (more judgments + more agreement = higher confidence): RANKED many independent judges scored this model's outputs and their agreement is very high (most confident) — HIGH many judges have scored it and they mostly agree (well-pinned) — MEDIUM enough judges have weighed in to publish, but they disagree more than we'd like (treat with a small grain of salt). LOW-confidence cells are hidden everywhere on the site. See the methodology for the exact thresholds.