Live system: AutoQA deployment
This page is the evidence appendix for the home ask: what may run vs what may count against work.
Summary
Design: decomposed criteria, grounded RCD-style checks, epistemic guard ("don't factor claims you cannot verify"), honeypot structure - research-aligned.
Authority: no enforcement labels, no first-class abstention path on the surface, pay coupling without per-lane calibration, contradictory preference scales, 22 non-exportable prompts.
Counts (68): 1 GREEN (signal only) · 1 GREEN after rewrite · 27 AMBER (advise pending validation) · 17 RED (no auto-adverse) · 22 GRAY (no authority until auditable).
Room lead: contradictory 0-7 scales in live prompts (vs labeler 0-2 / 3-4 tie / 5-7). Also: free text into judge prompts; worker-facing names should be behavior-descriptive and readable in a written reason.
n=1 exhibit (not a rate): platform marked three flags Wrong; human lead found all three grounded; record-check sided with the human on a quoted contradiction.
Rule: appropriate to run; large parts not ready to automatically count against work.
Fix first: one scale constant; enforcement field; treat labeler text as data; bind instruction version; deterministic checks in code; rename worker-facing labels.
Receipts: PDF -> CSV -> 2026-07-18 export. Scope: config surface and export; runtime gates may mitigate - non-reconstructable criteria still cannot earn inspectable authority.
Foundation rules that fail on the surface: limits (abstention, authority, long-record). Design fit: research.
Block from adverse use - advisory or human-only until rebuilt and validated (17)
These may still run as assist/triage. They must not auto-fail or gate pay as written.
| ID | Criterion | Level | Why blocked |
|---|---|---|---|
CRQ-3 | False Ties | Turn | A "false tie" requires the system to decide that quality differences are "very clear" across code and trajectory. |
PWQ-1 | Dictating Implementation Details | Sub | The "what, not how" distinction is not objective: implementation constraints can be legitimate requirements. |
PWQ-3 | Anti-Production Steering | Sub | "Counter to high quality code aimed at production readiness" is an open-ended engineering norm rather than a closed rule. |
RCD-1A | Model A Pros Verification | Turn | "Verify strengths against the actual response" is an unbounded family of semantic claims. |
RCD-1B | Model A Cons Verification | Turn | "Verify weaknesses against the actual response" is an unbounded family of semantic claims. |
RCD-2A | Model B Pros Verification | Turn | "Verify strengths against the actual response" is an unbounded family of semantic claims. |
RCD-2B | Model B Cons Verification | Turn | "Verify weaknesses against the actual response" is an unbounded family of semantic claims. |
RCD-3 | Preference Justification Verification | Turn | Alignment of an overall preference with two model responses is a whole-trajectory comparative judgment. |
RCD-7 | Exaggeration Detection | Turn | "Mountains out of molehills" is explicitly a proportionality and severity judgment. |
RCD-8A | Model A Contrived Cons | Turn | "Contrived," "non-issue in practice," and "clearly no longer a problem" are open-textured judgments requiring context, severity, and counterfactual engineering analysis. |
RCD-8B | Model B Contrived Cons | Turn | "Contrived," "non-issue in practice," and "clearly no longer a problem" are open-textured judgments requiring context, severity, and counterfactual engineering analysis. |
RCD-9A | Model A Contrived Feedback | Turn | "Non-issue pragmatically speaking" and "sensible problem" are normative and project-dependent. |
RCD-9B | Model B Contrived Feedback | Turn | "Non-issue pragmatically speaking" and "sensible problem" are normative and project-dependent. |
TR-13A | Incorrect Over-engineering A | Sub | Whether a change is "over-engineering" depends on necessity, acceptable collateral change, project conventions, and often unstated user intent. |
TR-13B | Incorrect Over-engineering B | Sub | Whether a change is "over-engineering" depends on necessity, acceptable collateral change, project conventions, and often unstated user intent. |
TR-7A | Incorrect Destructive Ops A | Sub | The criterion asks whether a destructive/system-modifying action was wrongly flagged or missed. |
TR-7B | Incorrect Destructive Ops B | Sub | The criterion asks whether a destructive/system-modifying action was wrongly flagged or missed. |
No authority until auditable - operative prompt not reconstructable (22)
Required first: export full prompt body, I/O schema, and version; then classify like the rest.
| ID | Criterion | Level | Required first |
|---|---|---|---|
AI-N-1 | M | Turn | Export full prompt body, I/O schema, and version; then classify like the rest. |
AIV-6 | Hedging Language | Sub | Export full prompt body, I/O schema, and version; then classify like the rest. |
CRQ-N-1A | M | Turn | Export full prompt body, I/O schema, and version; then classify like the rest. |
CRQ-N-1B | M | Turn | Export full prompt body, I/O schema, and version; then classify like the rest. |
DCM-N-1 | Change Rollback | Turn | Export full prompt body, I/O schema, and version; then classify like the rest. |
DCM-N-2 | Prescribing Behavior Ideals | Sub | Export full prompt body, I/O schema, and version; then classify like the rest. |
NAIV-7 | Not the Real Thing | Turn | Export full prompt body, I/O schema, and version; then classify like the rest. |
NAIV-8 | Integrity criterion (export title omitted) | Turn | Export full prompt body, I/O schema, and version; then classify like the rest. |
PWQ-N-1A | H | Turn | Export full prompt body, I/O schema, and version; then classify like the rest. |
PWQ-N-1B | H | Turn | Export full prompt body, I/O schema, and version; then classify like the rest. |
PWQ-N-2 | Poor Collaboration | Turn | Export full prompt body, I/O schema, and version; then classify like the rest. |
RCD-N-1A | L | Turn | Export full prompt body, I/O schema, and version; then classify like the rest. |
RCD-N-1B | Incorrect Environment Handling (Model B) | Turn | Export full prompt body, I/O schema, and version; then classify like the rest. |
RCD-N-2A | Incorrect Clarifying Question Penalty (Model A) | Turn | Export full prompt body, I/O schema, and version; then classify like the rest. |
RCD-N-2B | Incorrect Clarifying Question Penalty (Model B) | Turn | Export full prompt body, I/O schema, and version; then classify like the rest. |
RCD-N-3A | M | Turn | Export full prompt body, I/O schema, and version; then classify like the rest. |
RCD-N-3B | M | Turn | Export full prompt body, I/O schema, and version; then classify like the rest. |
RCD-N-4 | M | Turn | Export full prompt body, I/O schema, and version; then classify like the rest. |
RWQ-N-1A | Incorrect Destructive Actions (Model A) | Turn | Export full prompt body, I/O schema, and version; then classify like the rest. |
RWQ-N-1B | Incorrect Destructive Actions (Model B) | Turn | Export full prompt body, I/O schema, and version; then classify like the rest. |
RWQ-N-2 | Inconsistent Honesty Reporting | Sub | Export full prompt body, I/O schema, and version; then classify like the rest. |
RWQ-N-3 | Shallow Analysis | Turn | Export full prompt body, I/O schema, and version; then classify like the rest. |
AMBER (27) - advise / triage only
Audit wording: advisory or triage. Automatic adverse use only after criterion-specific operationalization and blind, representative validation of the exact instrument. Full IDs: PDF and CSV.
GREEN (2) - exact detection with caveats
| ID | Criterion | Condition |
|---|---|---|
AIV-1 |
Em Dash Usage | Deterministic signal only after ordinary software tests. Never sole proof of AI authorship. Quoted model output must not count (citation rule). No adverse inference from the signal alone. |
NAIV-6 |
Honeypot Catch | Deterministic after sentinels are explicitly enumerated; not safe for adverse use as written. |