The system
The system is defined by a versioned evaluation contract; models are replaceable parts inside it. Every decision must be reconstructible from one auditable unit. No pipeline step recovers truth the observable record does not identify - see limits.
The auditable unit
- Criterion C applies.
- Attempt span A exhibits behavior B.
- Evidence E supports, contradicts, or is insufficient for B.
- B does or does not satisfy C.
- Project decision policy assigns consequence S - only if authorized for that lane.
Pipeline, in order
- Intake and integrity. Provenance, completeness, eligibility. A gap in the record routes to record repair - never held against the worker.
- Compile the contract. Instructions become versioned criteria, proof standards, examples, evidence rules - with owner sign-off on graded cases.
- Decompose and type. Writeup -> claims typed as grounding, warrant, calibration, completeness, or opinion. Type picks machinery. Uncertifiable checks are marked, not guessed.
- Verify in lanes. Deterministic checks always; cheap screeners triage; models judge only with the evidence the worker had; contested items surface instead of being forced.
- Adjudicate criteria. Positive evidence, defects, omissions, and open questions recorded separately; rolled up by the project's versioned formula.
- Route. Clear cases decide within authority map; borderline get deeper verification; only decision-changing ambiguity reaches a human as one bounded question.
- Feedback. Task-referenced only - claim, evidence, clause, fix - and it must pass the same grounding gate it enforces.
- Monitor. Blind gold (including auto-accepts and agreements), drift sensors, challenge sets, appeals that write back into the contract. Throughput KPIs must not suppress abstention rates.
Different claims need different proof
| Claim type | Question | Treatment |
|---|---|---|
| Grounding | Does the cited evidence say this? | Span-level entailment. Fail class measured directly - never assumed. |
| Warrant | Does the evidence justify this strength of conclusion? | Reasoning confined to the record; discovery separate from judgment. |
| Completeness / absence | Is anything required missing? | Explicit checklists; "no errors" is not-contradicted, never verified, until planted-defect recall is measured. |
| Opinion | Is the stance consistent with its cited support? | Never truth-scored. Check contradiction and proportion only. |
| Undecidable | Do the inputs decide this at all? | Route to owner as instructions problem - no worker penalty. Forced choice on ambiguous items corrupts evaluation. |
Disagreement is a first-class output
When two reviewers (or reviewer and AutoQA) split, the system does not average them into a single fake truth. It tags the best-supported cause class:
- Instruction -> contract queue
- Error -> evidence packet; human confirm before adverse action
- Expertise -> specialist queue
- Legitimate split -> record; optional rubric change; not a forced winner
Using AutoQA solely as "the second reviewer who breaks ties" is an anti-pattern under volume pressure.
Human routing
See guide: human loop. Packet design: criterion, spans, one question, finite options, hide provisional decision for gold and consequence. Escalation rate and escalation-change rate are operating metrics - if humans never change outcomes, the loop is decoration.
Worked case, briefly
Task: does the evidence establish Initiative X as the primary cause of an improvement? Worker correctly says the studies do not establish primacy. Reviewer fails the item: "did not choose yes or no." Project instructions permit inconclusive answers.
System output: work claim supported; review claim contradicted by the instruction clause; no worker fault for following the contract; optional instruction-clarity note if the task note conflicts with the project rule - route to owner, not automatic reject of either party without policy.