1 + 1
One check is automatable as a signal only (em dash - never sole proof of AI authorship). One honeypot works after exact sentinels are listed (NAIV-6).
AutoQA on everything is fine. The same unmeasured Pass/Fail power on every check is not - not while the instrument has contradictory preference scales, no authority labels, and checks that should never automatically count against work.
Reviewer disagreement has several causes. Automating a casting vote does not fix them. Long agentic logs (200k+ tokens) make full human re-reads impossible; a model still cannot honestly certify the whole trajectory. The workable path is coverage with limits: AutoQA finds and cites; humans decide on short packets; only measured lanes may auto-act.
Live AutoQA deployment
One check is automatable as a signal only (em dash - never sole proof of AI authorship). One honeypot works after exact sentinels are listed (NAIV-6).
Twenty-seven advise pending validation. Seventeen must not automatically count against work. Twenty-two have no authority until the prompt can be exported.
Contradictory preference scales in the live prompts. One retained case: platform Wrong, human right - an exhibit, not a rate.
In two weeks
Product stays on. We bring back what may advise, escalate, or auto-detect - and the short fix list: one preference scale, enforcement tags, injection hardening, worker-readable names, GRAY exports.