Reference demonstration
This page and the miniqa bundle show mechanics: versioned contract, claim-level findings, four outcomes, abstention, audit-shaped output. They do not estimate accuracy on real work and do not authorize enforcement.
Not proven here: reviewer benefit, false-accusation rates, transfer from model-output judging to human writeups, fitness under production throughput pressure.
What to inspect
- miniqa 0.2.0 - engine + synthetic demo
- miniqa full - fuller package
- foundation package - schemas, examples, validators
- Evaluation output example
- Human resolution packet example
What "good" looks like in the demo
- Forced binary on an inconclusive case becomes an instruction / policy issue, not silent fail.
- Missing evidence yields cannot-verify, not automatic reject.
- Findings cite spans and criteria a human can contest.
- Work quality and review quality remain separate objects.