The ask

Keep AutoQA on live work. For two weeks: no new reject, pay, or account automation from AutoQA, and no using it as the casting vote when reviewers disagree. Then we return with what may advise, escalate, or automatically count against work.

Throughput without truth is not quality control.

AutoQA on everything is fine. The same unmeasured Pass/Fail power on every check is not - not while the instrument has contradictory preference scales, no authority labels, and checks that should never automatically count against work.

Reviewer disagreement has several causes. Automating a casting vote does not fix them. Long agentic logs (200k+ tokens) make full human re-reads impossible; a model still cannot honestly certify the whole trajectory. The workable path is coverage with limits: AutoQA finds and cites; humans decide on short packets; only measured lanes may auto-act.

Live AutoQA deployment

Research-aligned architecture. Authority is not ready.

1 + 1

One check is automatable as a signal only (em dash - never sole proof of AI authorship). One honeypot works after exact sentinels are listed (NAIV-6).

27 / 17 / 22

Twenty-seven advise pending validation. Seventeen must not automatically count against work. Twenty-two have no authority until the prompt can be exported.

Lead with facts

Contradictory preference scales in the live prompts. One retained case: platform Wrong, human right - an exhibit, not a rate.

In two weeks

A map, not a kill switch.

Product stays on. We bring back what may advise, escalate, or auto-detect - and the short fix list: one preference scale, enforcement tags, injection hardening, worker-readable names, GRAY exports.