Research review
What the literature can carry. Claim-level sources: evidence ledger. Live instrument: the live AutoQA deployment.
Durable (design)
- Disagreement is multi-causal; forced single labels corrupt indeterminate criteria.
- Raw agreement overstates ability; chance-corrected metrics and cross-benchmark rankings diverge.
- Human label uncertainty breaks naive "match the human" validation.
- Incentives follow measured dimensions (multitask principal-agent).
- Content-only spam defenses lose to capable agents; provenance matters.
- Once consequences attach, employment/selection rules constrain automated tools.
Generation-relative (re-measure on your instrument)
- Hard rubric verification accuracies; flip rates; position bias; preference leakage; long-form instability.
Priors only until remeasured on the exact deployment config and task.
Transfer gap
Almost all published judge numbers are models grading model outputs. AutoQA grades human review writeups. Transfer is open. Praise/positive-claim verification is especially thin - treat as in-house measurement.
Do not train only to predict another reviewer. Do not treat model-reviewer agreement as validation without expert gold and cause tags.
Evidence ledger · Research atlas · Critic and gap-fill · Library