Status Proposed; requires scoping
Updated 2026-07-23
The bounded measurement pilot
Optional later path - not the home CTA. Current ask: keep running; define authority.
This page describes a pilot that may return for approval after limits are adopted and a scoping pass returns facts. The open decision today authorizes neither pilot work nor automated enforcement.
Authority boundary: read-only, shadow-mode. No production enforcement, automated rejection, pay or account decisions, or expansion beyond one project. No use of AutoQA as a casting vote to manufacture consensus.
What the pilot must answer
- How reliable is the current human-review baseline by criterion and difficulty?
- Why do reviewers disagree - instruction, error, expertise, or legitimate split - and at what rates?
- Which claims have record-contained truth-makers a configured instrument can check?
- What are claim-extraction and evidence-discovery recall on the exact record format?
- How do false accusations, missed defects, and cost change as automated coverage increases?
- Does evidence-first assistance help adjudicators without creating automation bias?
- Does raising automation rate under a throughput KPI increase false clears or false strikes?
- Which lanes should stop, narrow, stay human-assist only, or return for later limited use?
Published evidence does not answer these for this task and configuration. Measure on untouched, expert-adjudicated data. See research and evidence.
Limits are launch conditions
The pilot may run only inside the capability-limits analysis:
- Result schema keeps supported, contradicted, insufficient, and contract-ambiguous outcomes.
- Intent, deference, steelman, and irreducibly global claims cannot create adverse findings from model output alone.
- Semantic absence and completeness lanes have planted-evidence recall tests.
- Each claim lane names detector, evidence boundary, human owner, and permitted authority.
- Confidence comes from untouched adjudicated data - not vote counts or model self-report.
- Calibration resets after any model, prompt, retrieval, contract, or configuration change.
- Standing blind audit (including auto-accepts and agreements) and injection challenges are in the operating plan.
- Throughput / coverage KPIs cannot disable abstention or force pass/fail.
Phase gates
Phase 1 - Scoping and preflight
Name project and owners. Confirm record boundary, read-only access, legal/security path, budget, limits classification, measurement plan, human-loop staffing, stop conditions.
Gate: approve, revise, or decline recommendation for the bounded pilot.
Phase 2 - Baseline and project contract
Measure current reviewer performance and expert agreement. Tag disagreement causes. Compile instructions into versioned, testable criteria with explicit ambiguity and insufficiency outcomes.
Stop: no stable record boundary, instruction owner, or adjudication capacity.
Phase 3 - Gold data and configured comparison
Development and held-out cases. Measure extraction, discovery, criterion-level judgment, false accusations, misses, abstention, cost for the exact instrument. Test one-interaction human resolution packets.
Stop or narrow: no useful risk-coverage region, no improvement over baseline, unaffordable expert/compute cost, or HITL that never changes outcomes (decoration loop).
Phase 4 - Shadow measurement and report
Score read-only work without changing dispositions. Blind audit and challenge sets. Test evidence-first human assistance. Report lane-level results and disagreement-cause mix.
Return: stop, narrow, human-assist only, or consider later lane-specific proposal.
Required measurements
| Measurement | Why | Decision it informs |
| Expert agreement by criterion | Bounds validation; surfaces unresolved interpretation | Can this criterion support validation or stay judgment-owned? |
| Disagreement cause mix | Stops treating all non-consensus as error | Where to fix instructions vs train vs route experts |
| Human baseline | Separates improvement from model agreement | Does the system add useful signal? |
| Claim-extraction recall | End-to-end recall cannot exceed claims found | Which formats stay in scope? |
| Evidence-discovery recall | Absence and long-record coverage need plants | May a lane assert contradiction or only "evidence found"? |
| False accusations / misses by lane | Aggregate accuracy hides rights-critical error | Authority map per lane |
| Escalation change rate | Human load-bearing vs decoration | Packet design and routing quality |
| Abstention under volume stress | Detect KPI override of honesty | Whether governance holds under pressure |
Decision rule for pilot outcomes
- Stop - no useful risk-coverage region, no operational advantage, unacceptable cost, or a boundary the candidate cannot satisfy.
- Narrow - retain only criteria and record types with bounded truth-makers and measured value.
- Human assist - evidence assembly or contestable feedback; humans keep all consequential authority.
- Later limited use - separate approval naming exact lane, instrument, calibration, false-accusation bound, abstention rule, monitoring.
What the pilot needs from the organization
Named decision rights, one bounded candidate project, expert adjudication capacity, a read-only record path, permission to measure the baseline before comparing automation, and funded audit/contract queues. It does not require production integration during this pilot.
Detail: benchmark protocol and library.