# AutoQA Foundations - Philosophy & Structure (draft for product-owner review)

**Date:** 2026-07-14
**Status:** Draft v1, synthesized from `00_QUESTIONS.md` and `01_RESEARCH_ATLAS.md`. Every principle carries its evidence anchor; positions the evidence does not settle are marked **[default]** (my recommended posture, reversible) or **[the product owner]** (needs your call).
**Scope:** Project-agnostic foundation for QA of human-annotated AI training data. Implementation infra (the Vercel app) is out of scope here.

---

## 0. The problem, restated precisely

Human attempters produce evaluative work products (annotations, ratings, critiques with rationales) under project instructions. Human reviewers who QA them exhibit unexplained variance in pass/fail rates and reasons. Root causes, ranked: (1) subjective interpretation of evaluation axes, (2) expertise gaps, (3) attention to detail. Attempters are mandatory; the QA layer is what we augment.

The psychometric literature sharpens this: after training, raters' *global* severity differences persist but are a minor variance component - the dominant variance is **rater x item x criterion interaction**: which reviewer, on which submission, on which criterion, this time. That variance cannot be fixed by better rubrics or calibration meetings alone. It can only be attacked by verifying each specific claim on each specific item against evidence - which is exactly what an LLM layer can do at scale and humans at scale cannot. That is the reason this system should exist, stated in the field's own terms. *(Atlas Section 26)*

## 1. The construct: what a verdict means

**Quality = satisfaction of the written project instructions + grounding in the defined evidence record.** Not "would the median reviewer pass this." Human-consensus prediction caps the system's validity at the reviewers' own (unmeasured, evidently low) reliability ceiling; instruction-satisfaction lets the AutoQA be *more* valid than the layer it augments, and makes every verdict traceable to a clause and a span. *(Atlas Section 16-18; peer-reviewed rating-indeterminacy work (NeurIPS 2025) shows forced-choice consensus gold labels can heavily bias judge selection - empirically up to 31% worse - when criteria admit multiple valid readings.)*

Consequences of this choice:
- Gold labels come from **adjudicated expert panels judging against the instructions**, not from single-reviewer or majority crowd labels.
- When the AutoQA and a human reviewer disagree, neither is presumptively right; the arbiter is the instruction text plus the evidence record, adjudicated.
- Where the instructions genuinely underdetermine an item, the correct output is not a verdict but an **instruction-gap report**. Reviewer disagreement on such items is signal about the instructions, not error in the people. *(Diverging Preferences: most annotator disagreement traces to task underspecification and style, not noise or incompetence.)*
- The **ground-truth hierarchy** is explicit: seeded known-verdict items > adjudicated expert panel > criterion author's written amendment > individual human reviewer > AutoQA. Only the first three may feed judge calibration. Overrides by individual humans are logged as claims, not truth.

## 2. The eleven principles

**P1 - Claims, not vibes.** The unit of analysis is the typed claim, not the item. Attempter writeups are decomposed into claims of distinct types, each verified in the lane built for it, each verdict carrying quoted evidence. Item verdicts are transparent, auditable roll-ups. *(Holistic judging of subjective quality measurably fails: near-orthogonal judge/human axes; ~55% on hard per-criterion checks; long-form instability. Decomposed per-criterion reasoning buys +7-12pp and lower cross-judge variance.)*

**P2 - The claim ontology is the type system.** Every extracted statement is triaged before any verification: **(a) grounding claims** (cited evidence exists and says this), **(b) entailment claims** (evidence supports the conclusion drawn), **(c) calibration claims** (stated severity/confidence matches evidence strength), **(d) completeness/absence claims** ("no errors," "fully addresses"), **(e) opinions** - never truth-scored, but checked with real machinery via two reusable passes: *stance-contradiction* (the evidence the attempter cites must not contradict the stated stance) and *stance-calibration* (stance strength must not exceed evidence strength - type-(c) machinery). Opinion verdicts are `consistent` / `inconsistent-with-cited-evidence` (a fail-able finding: the opinion's *warrant* fails grounding, not the opinion itself) / `uncheckable` (never penalized) - **(f) undecidable-given-evidence**. Each type has different machinery and a different reliability ceiling, reported per type.

**Entailment standard [default]:** reasonable-expert support, not strict logical entailment - three-way output per claim (`supported` / `contradicted` / `unsupported`), where `contradicted` is the high-confidence fail trigger and `unsupported` alone triggers an evidence request or escalation rather than an automatic fail. Strict entailment fails nearly all competent human prose; loose support reintroduces the variance we're removing - so the false-flag rate under this standard is a measured ship-gate output (Section 4.3), not an assumption. *(VeriScore/FactBench triage is field-standard; FACTS shows the positive-class asymmetry; AbsenceBench caps type (d).)*

**P3 - Evidence closure is explicit and judge priors are not evidence.** Every project config defines the evidence record: what the attempter saw, what they cited, what's in scope. Verdicts must quote spans from that record to be admissible; "true in the world but unsupported in the record" is `unsupported`, not `pass`. Checks are routed by whether they are **read-checkable or recompute-checkable** - judge confidence is untrustworthy on checks that exceed its own solving ability, where it silently re-solves and trusts itself over the record. **[default - provisional:** the re-solving evidence is a single under-review 2026 study; experiment 3 replicates it internally before this hard-gates routing.**]** Cherry-picking is caught by running both span-closed and context-closed citation checks. The judge gets the same evidence surface the attempter had - architecture beats model. *(Auditing-by-Re-Solving; CiteEval; Gandalf/BVB.)*

**P4 - Positive claims are verified by enumerated class, with honest labels.** "Warranted praise" verification is restricted to checkable praise types compiled per project (accuracy-praise -> grounding check; completeness-praise -> enumerable checklist; clarity-praise -> opinion class). Absence claims get a seeded-defect protocol: unless the judge's planted-flaw recall on this project materially exceeds the attempter population's, its concurrence is labeled **`not-contradicted`, never `verified`**. A vacuousness/eligibility gate precedes grounding checks so hedged empty writeups can't pass by asserting nothing. *(FACTS eligibility gate; AbsenceBench; zero published praise-verification numbers - this is the moat and the first in-house measurement.)*

**P5 - Typed verdicts; abstention and contestation are first-class.** The verdict set is `pass` / `fail` / `unable-to-verify` / `instructions-underdetermine`. High vote-entropy across k samples, cross-family judge disagreement, and clustered reviewer disagreement route items to the contested class - which feeds the instruction-gap channel and never penalizes the attempter. This is how the system enforces consistency where consistency is real, without laundering one arbitrary reading of a contested criterion into false objectivity. *(Toloka's deployed Pass/Fail/Unable-to-verify; CoVal; NUTMEG; rating indeterminacy.)*

**P6 - The human touch is spent where it compounds, and it is blind.** Routing is confidence-calibrated three-lane (auto-verdict / stronger judge / human) with the human budget derived from a target residual error rate, not a fixed per-item quota. The escalated human sees the claim, the criterion clause, and the judge's **collected evidence spans - not its verdict or rationale** (verdict-visible review produces measured over-reliance; evidence-only is the sole format that helps when the AI is right without hurting when it's wrong). Escalation routes the contested *claim*, not the whole item. Every human adjudication writes back - in two tiers that keep the ground-truth hierarchy honest: a single-touch adjudication produces **candidate** artifacts (logged claims, provisional exemplars), and only panel-adjudicated or owner-signed artifacts enter judge calibration and compiled config; the standing audit channel batch-promotes candidates. In enforcement mode the human budget is max(target-residual-error escalations, consequential-fail volume - Section 5); under a pace-bound profile (P11) capacity binds instead, and the residual error rate becomes the reported output rather than the target - never pretend to optimize both. The evidence-only-for-adjudicators result is measured; combining it with full critiques to attempters (P9) is an inferred, untested split **[default]**, A/B'd in experiment 6. Overturn-rate monitoring proves the human lane adds signal; an overturn rate at the judge's error floor means rubber-stamping, and the naive human-verifies-AI design is the modal failure in the literature. *(Trust-or-Escalate/HyPAC; DeepMind complementarity; Vaccaro meta-analysis; Braintrust/LangChain/micro1 write-back practice.)*

**P7 - The QA layer is measured exactly like what it measures.** The judge is enrolled as one more rater facet in the same measurement model as the humans (MFRM-style severity/centrality/fit per rater, human or machine, per project). Scope honestly: MFRM needs 3-5 ratings/output of rater-item linkage, so it runs fully on the dual-review pilot and continuously on the judge facet (which rates everything); production human-reviewer severity is estimable only where deliberate overlap is injected - a low-rate stream of double-reviewed seeded items, a budgeted cost, not a free byproduct. All reporting uses chance-corrected statistics, per-class recall at fixed prevalence, and PPI-corrected population estimates with dual-uncertainty CIs - raw percent-agreement and single-number "accuracy vs humans" are banned artifacts. Feedback passes the same grounding gate it enforces before delivery (fail-closed). Validation stratifies by measured human-agreement level; verdict-agreement alone never certifies a judge, because right-verdict-wrong-rationale poisons feedback and trust (in one benchmark, 24.8% of items got the right verdict with a low-quality critique - RealCritic). **Staging [default]:** v1 ships chance-corrected agreement, PPI-corrected point estimates with CIs, and k-sample entropy routing; MFRM production enrollment, anytime-valid e-processes, and conformal routing are phase-2 refinements - the full stack is the target, not the launch gate. *(MFRM-for-LLM-judges; kappa-deflation cohort; BeyondCorrelation; RealCritic; MetaCritique.)*

**P8 - Adversarial by construction, from day one.** Attempters are paid optimizers; some are agents. Defenses are structural, not detective: stakes-sterile judge prompts (consequence framing measurably softens verdicts with zero CoT trace); comparative/anchored scoring over absolute scales (transfer attacks inflate absolute scores blind); evidence-perturbation probes as a standing citation-theater detector (swap the cited span -> verdict must flip); blind seeded gold injected into live queues (indistinguishability is the binding constraint); **the divergence sensor** - attempter scores rising on the AutoQA while flat on held-out human-graded goldens - as the canonical Goodhart alarm; spam defense that is economics/provenance-based (throttling, work-history consistency, worker-level peer-prediction scoring, randomized deep audits), with content detection demoted to a weak auxiliary signal. Never score anyone on agreement-with-final-verdict - consensus-rewarded auditing trains your best dissenters to conform exactly where independent judgment matters most. *(Evaluation Faking; universal adversarial phrases; TSPA; Westwood PNAS; Community Notes conformity study.)*

**P9 - Feedback is an instrumented intervention, not decoration and not free.** The feedback unit is the 4-tuple: quoted claim, evidence span, criterion clause, concrete fix - task-referenced always, person- or rank-referenced never (a third of feedback interventions reduce performance; high-information task feedback roughly doubles the effect of bare verdicts). Benchmark against exemplar work, never ranked peers. No stylistic guidance by default - feedback is a homogenization pump aimed at the diversity the training data exists to capture, so a population-level output-diversity drift metric ships with the channel. Efficacy is measured per-attempter (repeat-error-rate per criterion, pre/post), expecting modest median effects concentrated in low performers, and separating within-attempter learning from attrition. If high-information feedback doesn't move repeat-error rates in a defined window, conclude decoration and reallocate. *(Kluger & DeNisi; Wisniewski; Cochrane 2025; Dow CSCW 2012; monoculture evidence.)*

**P10 - Everything drifts; every moving part gets its own sensor.** Frozen human-labeled anchor sets interleaved into the live stream attribute drift (judge moved vs population moved) with anytime-valid statistics. Judge versions are pinned per project with golden-set regression gates on upgrade; rubric revision is a versioned loop with explicit retroactive re-scoring policy (adaptation lives in the compilation layer, pinning at the judge layer - that's the resolution of the static-vs-co-evolving fork). Reviewer calibration decays within a session, so reviewer-effect estimates run in rolling windows with hard-item interleaving and order randomization; graduated projects carry randomized, unpredictable audit rates with automatic de-graduation triggers. Queue composition is managed (seeded known-verdict items, easy-case dilution) because prefiltered queues measurably shift reviewer strictness. *(Anchor-set e-process work; DRIFT/OSCE; Pinterest dual monitors; content-moderation practice.)*

**P11 - Quality is a priced dial on disposition, never on perception.** Contractual pace sometimes makes the standard quality threshold unaffordable; the system supports that honestly instead of forcing the loosening to happen off the books. The judge always measures the same way - dial-sterile prompts, same claim verification, same evidence dossier (consequence framing measurably corrupts verdicts with zero CoT trace, and a "be lenient" judge destroys score comparability, calibration, and the divergence sensor simultaneously). Leniency is applied downstream, as thresholds over the continuous scores the pipeline already emits. A project's active **operating profile** (e.g., `strict` / `standard` / `expedited`) moves exactly four things: aggregation cutoffs (how many/how severe findings fail an item), routing thresholds (escalation rate, human budget), verification depth (k-samples, deep-verify sampling fraction, expensive lanes like context-closed citation checks), and enforcement disposition (findings -> fail vs findings -> coaching-only). Two constraints can bind: in error-target mode the residual error rate is the target and human budget is the output; in pace-bound mode (contractual throughput) capacity binds and **the residual error rate becomes the reported output** - either way the number is computed and stamped, never hidden, with CIs that honestly widen when verification depth is thinned. Because measurement persists under every profile, the dial has two independent consumers: the **attempter-disposition dial** (prospective only, logged, announced - the bar people are judged against never moves silently or retroactively) and the **dataset-admission dial** (re-tunable at any time - an expedited batch can be re-filtered to any strictness later from the recorded scores, without re-running QA; the data buyer's strictness is not coupled to the deadline that produced the batch). Floors the dial can never touch: integrity-class verdicts (fabrication, spam, falsehood), seeded-gold injection and audit sampling, the divergence sensor (which conditions on the active profile and reads seeded/audit outcomes, not pass rates), and the Section 5 legal floor. Every delivered batch carries its profile and its corrected residual-error estimate - leniency becomes a contractual artifact instead of quiet degradation. *(Evaluation Faking -> dial-sterile prompts; CriticGPT's FSBS precision/comprehensiveness dial as the precedent for inference-time operating points; Trust-or-Escalate's coverage dial; HyPAC error budgets; corrected-estimator reporting.)*

## 3. The architecture

Eight stages. Stages 0-2 are deterministic-ish and cheap; 3 is where models spend; 4-7 close the loops.

```
project instructions --> [1] COMPILE (workshop loop, versioned config)
                              |
attempter submission --> [0] INTAKE GATE --> [2] DECOMPOSE & TYPE --> [3] VERIFY (lanes)
                                                                          |
                              +-------------------------------------------+
                              v
                         [4] AGGREGATE (typed verdict + evidence dossier)
                              |
                         [5] ROUTE (auto / stronger judge / human, calibrated)
                              |
                         [6] FEEDBACK (4-tuple, self-verified, fail-closed)
                              |
                         [7] META-EVAL & MONITOR (gold, drift, divergence, dashboards)
```

**Stage 0 - Intake & provenance gate.** Worker-level, not item-level: rate/velocity limits, work-history consistency, peer-prediction information scores, session-bounded telemetry (allowlist schema - on-task signals only; see Section 5). Item-level: eligibility/vacuousness check (is this a responsive, contentful submission at all). No LLM-text detection verdicts.

**Stage 1 - Compilation layer (the per-project customization contract).** Project instructions compile into a closed config schema:
- **Axis triage:** every evaluation axis classified into `deterministic` (string/regex/programmatic checks), `judge-groundable` (evidence-anchored verdicts), `judge-assisted` (human decides with judge-collected evidence), `human-only`, or `AutoQA-ineligible` - using per-axis conformal set width and pilot judge-vs-adjudicated agreement as the sorting evidence.
- **Rubric criteria:** <= ~13 active criteria per judge call, each atomic, RIFT-linted; **pitfall/negative criteria human-authored** (the one artifact synthesis reliably can't produce); hard constraints get gating semantics, not soft-criterion listing.
- **Contrastive exemplars:** hard near-miss pass/fail pairs per criterion - discriminative exemplars, not descriptive "what good looks like" prose. These anchor the judge few-shot (the measured-best grounding architecture) and double as reviewer training material.
- **Evidence-closure rules:** what the judge may read, per claim type; citation requirements on load-bearing attempter claims.
- **Severity/weight taxonomy + aggregation formula** (explicit, auditable - accepting the known performance tax vs implicit aggregation).
- **Routing costs and thresholds:** the asymmetric error prices (false-fail vs false-pass), the target residual error rate, calibration set.
- **Operating profiles (the quality/leniency dial, P11):** named threshold sets (`strict` / `standard` / `expedited`) taken from the ship gate's measured operating-point curve, each stamped with its expected fail-class recall and false-pass rate. Profile changes are owner-authorized, logged, prospective-only, and announced to attempters whenever the enforcement bar moves.
Compilation is a **loop, not a compile step**: owner grades real items, criteria drift is expected, every version is kept, and sign-off happens on graded items, not abstract criteria. Compilation fidelity is the binding per-project cost - budget accordingly.

**Stage 2 - Decompose & type.** Claim extraction at adaptive granularity (claim + explicit context annotations, verifier told what's under test; never free-floating atomic rewrites); ontology typing per P2; canonicalization/style-stripping so verbosity and rubric-vocabulary mimicry don't leak into judgment.

**Stage 3 - Verification lanes.**
- *Deterministic lane:* compiled checks run on everything, always.
- *Screening tier:* cheap specialized checkers (MiniCheck-class entailment, per-metric small models) score every claim at ~1/100 frontier cost - triage, never verdicts.
- *Verdict tier:* frontier reasoning judge, per-criterion calls, exemplar-anchored, cross-family from the models attempters plausibly used, stakes-sterile prompts, k-sampled with continuous scores (logit-expectation rather than sampled tokens where available), position/order randomized.
- *Opinion-consistency lane:* stance-contradiction + stance-calibration checks against the attempter's cited evidence, with the P2(e) verdict semantics - runs in the same screening/verdict tiers as other claim types.
- *Contested detector:* vote entropy, paraphrase flips, cross-family disagreement -> `instructions-underdetermine` candidates.

**Stage 4 - Aggregation.** Claim verdicts -> item verdict via the compiled formula; error-tolerant (a ~20% claim-level error rate means naive AND-aggregation is wrong by construction); output is the **verdict + evidence dossier** (every finding with quoted spans and criterion clauses) - the dossier is simultaneously the audit trail, the appeal record, and the feedback source.

**Stage 5 - Routing & the human lane.** Calibrated dual thresholds (HyPAC-style) with the human queue = the abstention set; auto-FAIL held to a stricter bar than auto-PASS (false fails burn trust and pay; consequential fails have a legal floor - Section 5); blind evidence-only adjudication UI; claim-level micro-adjudication; adjudications write back per P6's two-tier rule; overturn telemetry closed-loop. Queue organization (specialist-per-criterion vs generalist-per-item) is a compilation-layer decision informed by the pilot's expertise-share finding - if expertise gaps dominate a project's variance, specialist claim-level routing is the mechanism that closes root cause #2 on the human side. **Cold start** for a zero-label project: the judge runs shadow-mode under full human review until the calibration set reaches the PPI-arithmetic size and thresholds are fitted; **graduation** to confidence-routing happens when calibrated thresholds hit the target residual error on a rolling window - mirrored by P10's de-graduation triggers.

**Stage 6 - Feedback channel.** Per P9. Generated from the dossier (verdict machinery and feedback share one artifact), passed through its own grounding gate, delivered only if it survives.

**Stage 7 - Meta-evaluation & monitoring.** Per-project ship gate before go-live (Section 4); continuous: blind gold injection, frozen anchor set, divergence sensor, per-criterion agreement trends, appeal-overturn bands, population diversity metric, reviewer rolling-window effects, corrected-estimator dashboards. The standing expert-audit channel samples confident passes - **critic-assisted** (a critic-assisted second look finds 4x what an unassisted one does on the "flawless" slice) - and is permanent, because when judge accuracy <= attempter accuracy no statistical cleverness can substitute for gold (the 2x theorem), and false agreement is predicted to rise as model errors correlate.

## 4. The ship gate (per project config, before go-live)

A project's judge config ships only after:
1. **Variance-attribution pilot:** dual-review a common item set; measure the human floor (chance-corrected, per criterion - per-criterion kappa localizes ambiguous criteria); code disagreements against an underspecification-vs-expertise-vs-attention taxonomy. This decides the build emphasis per project and provides the calibration set.
2. **MVVP-style judge validation:** chance-corrected agreement vs adjudicated labels, >=3 replicates, AB/BA position swaps, >=2 gold-set designs, bias audit separate from consistency (a judge can be perfectly repeatable and severely biased).
3. **Perturbation harness:** meaning-preserving paraphrase/reorder/noise on claims AND evidence; evidence-swap probes; verdicts that flip get routed, not shipped.
4. **Seeded-flaw recall:** planted defects (FindTheFlaws-style) measuring fail-class recall at realistic prevalence - the number that actually matters under 90%+ pass rates, where kappa itself misleads (Feinstein-Cicchetti).
5. **Schematic adherence:** do per-criterion scores vary independently and does the rationale's cited schema explain the verdict (halo/factor-collapse check) before per-criterion output is exposed to anyone.
6. **Gold-set arithmetic:** size from the PPI closed form off pilot judge-human correlation per axis (conservative estimates; ~50+ labels per stratum floor; oversample the fail class); low-correlation axes carry human-only until correlation improves.
7. **Beat-the-humans gate:** per axis, judge-vs-adjudicated reliability must exceed the pilot's measured human-reviewer-vs-adjudicated reliability. Axes where the judge doesn't beat the humans it augments are human-only lanes - automating them adds cost and subtracts validity. This is the go/no-go the whole premise rests on, and it is comparative, not absolute.

Beyond go/no-go, the ship gate's data product is the **operating-point curve** - thresholds -> fail-class recall and false-pass rate at pilot prevalence. This is what calibrates P11's dial profiles, so even the `expedited` setting ships with a priced, not guessed, error rate.

## 5. The legal floor (binding if the workforce touches the EU)

**Gate first [the product owner]:** determine whether the attempter population includes EU-resident workers. If yes, this section is a launch constraint; if the pilot is US-only, treat it as forward-looking scope and revisit before any EU expansion. From 2026-12-02 (Platform Work Directive transposition), for attempters in the EU - employee or self-employed, platform established anywhere:
- Decisions of significant detriment (account/pay restriction, termination, or equivalent) **must be taken by a human**; no consent exception. Auto-pass is fine; consequential auto-fail is not. Whether a single-item pay denial qualifies is legally untested - **[default]** treat consequential item fails as requiring the human lane until counsel says otherwise.
- Plain-language explanation of any automated-supported decision, a competent human contact, human review with substantiated written reply within two weeks, rectification/compensation duties. The evidence dossier and 4-tuple feedback ARE this explanation - design them to that bar.
- Telemetry: allowlist only (task-session-bounded behavioral signals); no emotional-state inference, no private-communication monitoring, no off-task collection; DPIA required.
- SCHUFA doctrine: a score with a "determining role" is itself an automated decision, and rubber-stamp review counts as zero-touch - the human lane must have authority, evidence access, and calibrated competence. Enforcement is active.
- Map quality dimensions and verdict evidence onto ISO/IEC 5259 vocabulary now (cheap) rather than retrofit later (expensive).

Convenient truth: the legally mandated design (meaningful human on adverse decisions, explanations, appeals) is the same design the complementarity evidence recommends anyway.

## 6. What this system deliberately does not do

- No per-item LLM-text spam classification as a primary defense (beaten economically and technically).
- No holistic single-call item judging; no per-criterion output without schematic-adherence validation.
- No raw percent-agreement or single-number accuracy anywhere in reporting.
- No truth verdicts on opinion-class claims; no `verified` label on absence claims without seeded-recall evidence.
- No stylistic feedback; no visible attempter leaderboards/rank feedback.
- No consequence language in judge prompts; no trusting CoT rationales as bias evidence.
- No scoring of humans (attempters, reviewers, or auditors) on agreement-with-final-verdict.
- No judge calibration from unadjudicated override streams.
- No leniency via judge prompting - the dial moves thresholds, depth, and disposition, never the judge's perception (dial-sterile is the same discipline as stakes-sterile). No dial setting touches integrity floors, seeded gold, audit sampling, or the legal floor. No silent or retroactive changes to the bar attempters are judged against.
- No claim that the judge validates itself - the expert-audit channel never sunsets.

## 7. First experiments (ordered; each kills or funds a component)

1. **Variance-attribution pilot** on one live project (dual review + disagreement coding). Decides whether rubric compilation or expertise routing is the bigger lever *for our population*, and produces the first calibration set. Everything else calibrates against this.
2. **Positive-claim seeded set:** build warranted/empty-praise pairs per praise type; measure judge precision/recall per type. Nothing published exists - this is both the riskiest premise and the differentiator. Include absence-claim recall (planted defects behind "no errors" writeups).
3. **Transfer check:** take the compiled-config judge and measure per-criterion agreement against adjudicated labels on *human evaluative writeups* (not model outputs) - the number every published anchor is missing. Include a re-solving probe: does our judge's wrong-reference detection collapse on recompute-class checks (internal replication of the under-review 2026 result P3 provisionally leans on)?
4. **Confidence-proxy shootout:** verbalized confidence vs k-sample entropy vs conformal width vs cross-family disagreement, scored on predicting human overturn - this picks the routing statistic. Include the uniform-sampling baseline.
5. **Divergence sensor dry run:** stand up held-out human-graded goldens + AutoQA score tracking from the first live week, so Goodhart is measurable before it's a problem.
6. **Feedback tier A/B:** verdict-only vs 4-tuple high-information feedback, outcome = per-criterion repeat-error rate, attrition separated. Priors: modest median, concentrated in low performers, backfire in a minority. Also A/B the assistance-format split (full critiques to attempters vs evidence-only) - the combination P6/P9 defaults to is untested in any single study.
7. **Agnosticism falsification:** onboard a second project through config alone. Any edit to the core judge prompts falsifies the project-agnostic claim - the fix is moving that surface into the config schema, not patching the core.

## 8. Open decisions **[the product owner]**

1. **Enforcement weight at launch:** advisory/coaching-only vs consequential verdicts. Changes the precision floor, the legal posture (Section 5), and attempter adversarial pressure. My read: launch advisory, attach consequences only per-lane after seeded-recall numbers exist.
2. **Construct authority when project owner intuition contradicts their own compiled rubric:** rubric-as-law with forced amendment **[default]**, or standing owner overrides (which quietly re-import unexplainable variance).
3. **LLM-assistance provenance policy default:** content-validity-only vs provenance-gated (per-project override either way). Interacts with telemetry legality and with the reality that ~1/3 of text workers already use LLMs.
4. **Judge sourcing:** frontier API judges per project (cross-family routing) only, or also an owned fine-tuned screening tier. The evidence supports two-tier; the question is whether to own the bottom tier.
5. **Where the first pilot runs** - which project's data, whether its owner will fund the adjudicated gold set (the one cost that cannot be automated away), and whether that attempter population includes EU-resident workers (switches Section 5 from forward-looking to binding).
6. **Dial governance (P11):** who may change a project's active operating profile (project owner alone, or with client sign-off), and whether expedited batches are disclosed to the data buyer with their stamped residual-error estimate. My read: disclose - a priced error rate converts deadline pressure from hidden liability into a contractual artifact, and it is the posture consistent with the ISO 5259 mapping in Section 5.

---

*Lineage: brainstorm (92 questions, 7 lenses) -> taxonomy (9 themes, 8 tensions) -> 9-domain recency-biased sweep + 27 adversarial verifications + 4 gap-fill literatures -> this synthesis. Full provenance in `atlas/`.*
