Updated 2026-07-23

Research review

What the literature can carry. Claim-level sources: evidence ledger. Live instrument: the live AutoQA deployment.

Durable (design)

Generation-relative (re-measure on your instrument)

Priors only until remeasured on the exact deployment config and task.

Transfer gap

Almost all published judge numbers are models grading model outputs. AutoQA grades human review writeups. Transfer is open. Praise/positive-claim verification is especially thin - treat as in-house measurement.

Do not train only to predict another reviewer. Do not treat model-reviewer agreement as validation without expert gold and cause tags.

Evidence ledger · Research atlas · Critic and gap-fill · Library