Export 2026-07-18 Verified 2026-07-23

Live system: AutoQA deployment

This page is the evidence appendix for the home ask: what may run vs what may count against work.

Action-list PDF Home ask

Summary

Design: decomposed criteria, grounded RCD-style checks, epistemic guard ("don't factor claims you cannot verify"), honeypot structure - research-aligned.

Authority: no enforcement labels, no first-class abstention path on the surface, pay coupling without per-lane calibration, contradictory preference scales, 22 non-exportable prompts.

Counts (68): 1 GREEN (signal only) · 1 GREEN after rewrite · 27 AMBER (advise pending validation) · 17 RED (no auto-adverse) · 22 GRAY (no authority until auditable).

Room lead: contradictory 0-7 scales in live prompts (vs labeler 0-2 / 3-4 tie / 5-7). Also: free text into judge prompts; worker-facing names should be behavior-descriptive and readable in a written reason.

n=1 exhibit (not a rate): platform marked three flags Wrong; human lead found all three grounded; record-check sided with the human on a quoted contradiction.

Rule: appropriate to run; large parts not ready to automatically count against work.

Fix first: one scale constant; enforcement field; treat labeler text as data; bind instruction version; deterministic checks in code; rename worker-facing labels.

Receipts: PDF -> CSV -> 2026-07-18 export. Scope: config surface and export; runtime gates may mitigate - non-reconstructable criteria still cannot earn inspectable authority.

Foundation rules that fail on the surface: limits (abstention, authority, long-record). Design fit: research.

Block from adverse use - advisory or human-only until rebuilt and validated (17)

These may still run as assist/triage. They must not auto-fail or gate pay as written.

IDCriterionLevelWhy blocked
CRQ-3False TiesTurnA "false tie" requires the system to decide that quality differences are "very clear" across code and trajectory.
PWQ-1Dictating Implementation DetailsSubThe "what, not how" distinction is not objective: implementation constraints can be legitimate requirements.
PWQ-3Anti-Production SteeringSub"Counter to high quality code aimed at production readiness" is an open-ended engineering norm rather than a closed rule.
RCD-1AModel A Pros VerificationTurn"Verify strengths against the actual response" is an unbounded family of semantic claims.
RCD-1BModel A Cons VerificationTurn"Verify weaknesses against the actual response" is an unbounded family of semantic claims.
RCD-2AModel B Pros VerificationTurn"Verify strengths against the actual response" is an unbounded family of semantic claims.
RCD-2BModel B Cons VerificationTurn"Verify weaknesses against the actual response" is an unbounded family of semantic claims.
RCD-3Preference Justification VerificationTurnAlignment of an overall preference with two model responses is a whole-trajectory comparative judgment.
RCD-7Exaggeration DetectionTurn"Mountains out of molehills" is explicitly a proportionality and severity judgment.
RCD-8AModel A Contrived ConsTurn"Contrived," "non-issue in practice," and "clearly no longer a problem" are open-textured judgments requiring context, severity, and counterfactual engineering analysis.
RCD-8BModel B Contrived ConsTurn"Contrived," "non-issue in practice," and "clearly no longer a problem" are open-textured judgments requiring context, severity, and counterfactual engineering analysis.
RCD-9AModel A Contrived FeedbackTurn"Non-issue pragmatically speaking" and "sensible problem" are normative and project-dependent.
RCD-9BModel B Contrived FeedbackTurn"Non-issue pragmatically speaking" and "sensible problem" are normative and project-dependent.
TR-13AIncorrect Over-engineering ASubWhether a change is "over-engineering" depends on necessity, acceptable collateral change, project conventions, and often unstated user intent.
TR-13BIncorrect Over-engineering BSubWhether a change is "over-engineering" depends on necessity, acceptable collateral change, project conventions, and often unstated user intent.
TR-7AIncorrect Destructive Ops ASubThe criterion asks whether a destructive/system-modifying action was wrongly flagged or missed.
TR-7BIncorrect Destructive Ops BSubThe criterion asks whether a destructive/system-modifying action was wrongly flagged or missed.

No authority until auditable - operative prompt not reconstructable (22)

Required first: export full prompt body, I/O schema, and version; then classify like the rest.

IDCriterionLevelRequired first
AI-N-1MTurnExport full prompt body, I/O schema, and version; then classify like the rest.
AIV-6Hedging LanguageSubExport full prompt body, I/O schema, and version; then classify like the rest.
CRQ-N-1AMTurnExport full prompt body, I/O schema, and version; then classify like the rest.
CRQ-N-1BMTurnExport full prompt body, I/O schema, and version; then classify like the rest.
DCM-N-1Change RollbackTurnExport full prompt body, I/O schema, and version; then classify like the rest.
DCM-N-2Prescribing Behavior IdealsSubExport full prompt body, I/O schema, and version; then classify like the rest.
NAIV-7Not the Real ThingTurnExport full prompt body, I/O schema, and version; then classify like the rest.
NAIV-8Integrity criterion (export title omitted)TurnExport full prompt body, I/O schema, and version; then classify like the rest.
PWQ-N-1AHTurnExport full prompt body, I/O schema, and version; then classify like the rest.
PWQ-N-1BHTurnExport full prompt body, I/O schema, and version; then classify like the rest.
PWQ-N-2Poor CollaborationTurnExport full prompt body, I/O schema, and version; then classify like the rest.
RCD-N-1ALTurnExport full prompt body, I/O schema, and version; then classify like the rest.
RCD-N-1BIncorrect Environment Handling (Model B)TurnExport full prompt body, I/O schema, and version; then classify like the rest.
RCD-N-2AIncorrect Clarifying Question Penalty (Model A)TurnExport full prompt body, I/O schema, and version; then classify like the rest.
RCD-N-2BIncorrect Clarifying Question Penalty (Model B)TurnExport full prompt body, I/O schema, and version; then classify like the rest.
RCD-N-3AMTurnExport full prompt body, I/O schema, and version; then classify like the rest.
RCD-N-3BMTurnExport full prompt body, I/O schema, and version; then classify like the rest.
RCD-N-4MTurnExport full prompt body, I/O schema, and version; then classify like the rest.
RWQ-N-1AIncorrect Destructive Actions (Model A)TurnExport full prompt body, I/O schema, and version; then classify like the rest.
RWQ-N-1BIncorrect Destructive Actions (Model B)TurnExport full prompt body, I/O schema, and version; then classify like the rest.
RWQ-N-2Inconsistent Honesty ReportingSubExport full prompt body, I/O schema, and version; then classify like the rest.
RWQ-N-3Shallow AnalysisTurnExport full prompt body, I/O schema, and version; then classify like the rest.

AMBER (27) - advise / triage only

Audit wording: advisory or triage. Automatic adverse use only after criterion-specific operationalization and blind, representative validation of the exact instrument. Full IDs: PDF and CSV.

GREEN (2) - exact detection with caveats

IDCriterionCondition
AIV-1 Em Dash Usage Deterministic signal only after ordinary software tests. Never sole proof of AI authorship. Quoted model output must not count (citation rule). No adverse inference from the signal alone.
NAIV-6 Honeypot Catch Deterministic after sentinels are explicitly enumerated; not safe for adverse use as written.