Document library

Every source, specification, schema, example, and template.

Full technical archive. CTA: home ยท live map: the review program.

All files are available as rendered in-site documents and as their original raw format. Search covers titles, summaries, tags, and full file contents.

AutoQA Foundation - Research Atlas

Executive synthesis of the recency-biased research sweep, including strongest findings, disagreements, caveats, and gaps.

Core narrativeMD102 lines25.2 KBresearchexecutive synthesisevidence

Raw Questions by Lens

The full 92-question brainstorm across measurement, judge failure modes, rubric operations, human factors, grounding, workflow, and feedback.

Research working papersMD383 lines70.2 KBquestionsbrainstormresearch agenda

Completeness Critic & Gap-Fill Reports

Adversarial audit of the original sweep, plus psychometrics, legal, content moderation, and feedback-science gap fills.

Research working papersMD193 lines64.4 KBcritiquepsychometricslawfeedback

Domain Report - Academic Judges

Academic judge benchmarks, task-dependent capability limits, bias, and model-selection evidence.

Domain researchMD95 lines28.2 KBdomain reportacademic judges

Domain Report - Annotation Quality

Annotation disagreement, rater uncertainty, LLM contamination, label error, and human-assistance effects.

Domain researchMD93 lines31.0 KBdomain reportannotation quality

Benchmark & Calibration Protocol

How to construct adjudicated gold data, calibrate evaluators, run ablations, set launch gates, and monitor drift.

Foundation specificationsMD1064 lines30.7 KBfoundationspecification

Decision Policy Reference

A target-specific rule engine for accept, rework, reject, and escalate decisions, including the bounded human interaction.

Foundation specificationsMD499 lines13.7 KBfoundationspecification

Foundation Package - README

Package overview, core model, invariants, adoption sequence, and conformance levels.

Foundation specificationsMD191 lines7.9 KBfoundationspecification

Worked Example - Benchmark Item

Adjudicated benchmark record for the synthetic case, including acceptable outcomes and ambiguity metadata.

Foundation examplesJSON194 lines6.0 KBexampleworked case

Worked Example - Case Input

A strong attempt incorrectly failed by a human reviewer for not forcing an unsupported binary conclusion.

Foundation examplesJSON108 lines4.1 KBexampleworked case

Worked Example - Evaluation Output

Complete AutoQA result accepting the attempt, returning the annotation for rework, and generating audience-specific feedback.

Foundation examplesJSON1008 lines32.8 KBexampleworked case

Worked Example - Human Resolution

A bounded precedence question that hides the provisional decision and asks only what can change the outcome.

Foundation examplesJSON75 lines2.5 KBexampleworked case

Worked Example - Project Contract

Synthetic closed-world evidence project with positive-proof requirements, anchors, severity, and explicit decision rules.

Foundation examplesJSON656 lines22.0 KBexampleworked case

JSON Schema - Benchmark Item

Runtime schema for adjudicated truth, acceptable actions, ambiguity, challenge attributes, and leakage controls.

Foundation schemasJSON312 lines7.7 KBschemaruntime validation

JSON Schema - Case Input

Runtime schema for task, evidence bundle, attempt, human review, references, and provenance.

Foundation schemasJSON335 lines8.1 KBschemaruntime validation

JSON Schema - Common Definitions

Shared runtime definitions used by the project, case, evaluation, benchmark, and resolution schemas.

Foundation schemasJSON1444 lines32.0 KBschemaruntime validation

JSON Schema - Evaluation Output

Runtime schema for claims, evidence relationships, findings, criterion assessments, decisions, annotation audit, and feedback.

Foundation schemasJSON179 lines4.0 KBschemaruntime validation

JSON Schema - Project Contract

Runtime schema for approved project instructions, criteria, evidence policy, decision policy, and calibration requirements.

Foundation schemasJSON876 lines22.8 KBschemaruntime validation

Template - Criterion Authoring

CSV template for drafting criterion intent, proof standards, failure anchors, evidence requirements, severity, and governance.

Foundation templatesCSV2 lines552 Btemplateauthoring

Template - Gold Set

CSV template for building an adjudicated calibration and test set.

Foundation templatesCSV2 lines593 Btemplateauthoring

Template - Model Bakeoff

CSV template for comparing evaluator configurations by criterion family.

Foundation templatesCSV2 lines582 Btemplateauthoring

TypeScript Domain Model

Strict TypeScript types for project contracts, cases, evaluations, benchmarks, findings, decisions, feedback, and human resolution.

Foundation developer filesTS787 lines20.7 KBtypescriptdeveloper