Supporting brief · archive
Measure quality instead of manufacturing consensus
Background only. Current CTA: home. Live instrument: the review program.
Problem
Non-consensus has four causes - unclear instructions, reviewer error, missing expertise, legitimate split. One agreement score mixes them. Driving kappa as the goal confuses error with ambiguity.
Throughput pressure plus "make them agree" is when unmeasured automation becomes a casting vote. That path is refused here.
Proposal shape
Claim-level check of the review against instructions and permitted evidence. Outcomes: supported, contradicted, cannot verify, instructions insufficient (owner, not worker fault). Humans own contract meaning, consequence, and standing audit. Detail: system · KERNEL.
Not claimed
- Universal judge; automated rejection on subjective criteria without measurement.
- Praise verification without in-house gold.
- Removal of expert audit cost.
- Paper numbers substituting for measurement on human writeups.
- Agreement rising only because people imitate the model.
Evidence posture
Architecture case is durable. Specific performance on this workload is open (transfer gap). Research · Evidence ledger.