Status Proposed; requires scoping Updated 2026-07-23

The bounded measurement pilot

Optional later path - not the home CTA. Current ask: keep running; define authority.

This page describes a pilot that may return for approval after limits are adopted and a scoping pass returns facts. The open decision today authorizes neither pilot work nor automated enforcement.

Authority boundary: read-only, shadow-mode. No production enforcement, automated rejection, pay or account decisions, or expansion beyond one project. No use of AutoQA as a casting vote to manufacture consensus.

What the pilot must answer

Published evidence does not answer these for this task and configuration. Measure on untouched, expert-adjudicated data. See research and evidence.

Limits are launch conditions

The pilot may run only inside the capability-limits analysis:

Phase gates

Phase 1 - Scoping and preflight

Name project and owners. Confirm record boundary, read-only access, legal/security path, budget, limits classification, measurement plan, human-loop staffing, stop conditions.

Gate: approve, revise, or decline recommendation for the bounded pilot.

Phase 2 - Baseline and project contract

Measure current reviewer performance and expert agreement. Tag disagreement causes. Compile instructions into versioned, testable criteria with explicit ambiguity and insufficiency outcomes.

Stop: no stable record boundary, instruction owner, or adjudication capacity.

Phase 3 - Gold data and configured comparison

Development and held-out cases. Measure extraction, discovery, criterion-level judgment, false accusations, misses, abstention, cost for the exact instrument. Test one-interaction human resolution packets.

Stop or narrow: no useful risk-coverage region, no improvement over baseline, unaffordable expert/compute cost, or HITL that never changes outcomes (decoration loop).

Phase 4 - Shadow measurement and report

Score read-only work without changing dispositions. Blind audit and challenge sets. Test evidence-first human assistance. Report lane-level results and disagreement-cause mix.

Return: stop, narrow, human-assist only, or consider later lane-specific proposal.

Required measurements

MeasurementWhyDecision it informs
Expert agreement by criterionBounds validation; surfaces unresolved interpretationCan this criterion support validation or stay judgment-owned?
Disagreement cause mixStops treating all non-consensus as errorWhere to fix instructions vs train vs route experts
Human baselineSeparates improvement from model agreementDoes the system add useful signal?
Claim-extraction recallEnd-to-end recall cannot exceed claims foundWhich formats stay in scope?
Evidence-discovery recallAbsence and long-record coverage need plantsMay a lane assert contradiction or only "evidence found"?
False accusations / misses by laneAggregate accuracy hides rights-critical errorAuthority map per lane
Escalation change rateHuman load-bearing vs decorationPacket design and routing quality
Abstention under volume stressDetect KPI override of honestyWhether governance holds under pressure

Decision rule for pilot outcomes

What the pilot needs from the organization

Named decision rights, one bounded candidate project, expert adjudication capacity, a read-only record path, permission to measure the baseline before comparing automation, and funded audit/contract queues. It does not require production integration during this pilot.

Detail: benchmark protocol and library.