Status Working design Updated 2026-07-23

The system

The system is defined by a versioned evaluation contract; models are replaceable parts inside it. Every decision must be reconstructible from one auditable unit. No pipeline step recovers truth the observable record does not identify - see limits.

The auditable unit

  1. Criterion C applies.
  2. Attempt span A exhibits behavior B.
  3. Evidence E supports, contradicts, or is insufficient for B.
  4. B does or does not satisfy C.
  5. Project decision policy assigns consequence S - only if authorized for that lane.

Pipeline, in order

  1. Intake and integrity. Provenance, completeness, eligibility. A gap in the record routes to record repair - never held against the worker.
  2. Compile the contract. Instructions become versioned criteria, proof standards, examples, evidence rules - with owner sign-off on graded cases.
  3. Decompose and type. Writeup -> claims typed as grounding, warrant, calibration, completeness, or opinion. Type picks machinery. Uncertifiable checks are marked, not guessed.
  4. Verify in lanes. Deterministic checks always; cheap screeners triage; models judge only with the evidence the worker had; contested items surface instead of being forced.
  5. Adjudicate criteria. Positive evidence, defects, omissions, and open questions recorded separately; rolled up by the project's versioned formula.
  6. Route. Clear cases decide within authority map; borderline get deeper verification; only decision-changing ambiguity reaches a human as one bounded question.
  7. Feedback. Task-referenced only - claim, evidence, clause, fix - and it must pass the same grounding gate it enforces.
  8. Monitor. Blind gold (including auto-accepts and agreements), drift sensors, challenge sets, appeals that write back into the contract. Throughput KPIs must not suppress abstention rates.

Different claims need different proof

Claim typeQuestionTreatment
GroundingDoes the cited evidence say this?Span-level entailment. Fail class measured directly - never assumed.
WarrantDoes the evidence justify this strength of conclusion?Reasoning confined to the record; discovery separate from judgment.
Completeness / absenceIs anything required missing?Explicit checklists; "no errors" is not-contradicted, never verified, until planted-defect recall is measured.
OpinionIs the stance consistent with its cited support?Never truth-scored. Check contradiction and proportion only.
UndecidableDo the inputs decide this at all?Route to owner as instructions problem - no worker penalty. Forced choice on ambiguous items corrupts evaluation.

Disagreement is a first-class output

When two reviewers (or reviewer and AutoQA) split, the system does not average them into a single fake truth. It tags the best-supported cause class:

Using AutoQA solely as "the second reviewer who breaks ties" is an anti-pattern under volume pressure.

Human routing

See guide: human loop. Packet design: criterion, spans, one question, finite options, hide provisional decision for gold and consequence. Escalation rate and escalation-change rate are operating metrics - if humans never change outcomes, the loop is decoration.

Worked case, briefly

Task: does the evidence establish Initiative X as the primary cause of an improvement? Worker correctly says the studies do not establish primacy. Reviewer fails the item: "did not choose yes or no." Project instructions permit inconclusive answers.

System output: work claim supported; review claim contradicted by the instruction clause; no worker fault for following the contract; optional instruction-clarity note if the task note conflicts with the project rule - route to owner, not automatic reject of either party without policy.