Status Limits analysis Updated 2026-07-23

What AutoQA cannot do - even if the models get better

Three ceilings: permanent (record and task structure), generation-relative (measure per instrument), authorization (capability is not a license to punish). Full L0-L8 below. CTA: home ยท live map: the review program.

Purpose

Each limit below is stated in strongest defensible form: grounds, what it does not imply, and the signature of operating past it.

The verification model these limits apply to

AutoQA, in any implementation, computes a relation between a reviewer claim q, a governing contract C (the written instructions), and a bounded record R (what the reviewer could see). The possible honest outputs are: supported, contradicted, insufficient (the record does not decide), and contract-ambiguous (the instructions do not decide). Everything below is a statement about what this computation can and cannot establish. A system that emits only pass/fail has not escaped these limits; it has hidden the third and fourth outcomes inside the first two.


L0 - The master limit: record identifiability (permanent)

Statement. Every limit below is an instance of one fact. Let the observable record be X and the correct decision be Y. If two possible worlds produce the identical observable record but have different correct decisions, then no evaluator that sees only the record - any model, any context window, any ensemble, any amount of inference - can be correct in both. The ceiling on any evaluator is the Bayes-optimal accuracy, which reaches 100% only when the correct label is a function of the observable record. Better processing of the record cannot manufacture facts the record does not contain, and cannot make an ambiguous rule produce a unique answer.

Grounds. Immediate from the definitions: identical input forces identical output while the correct outputs differ. (Independently derived, with the same construction, in the companion boundary package's formal standard.)

Does not imply. That record-based verification has no useful operating region. It identifies when the record can decide a bounded claim and when insufficient or contract-ambiguous is the only defensible result.

Violation signature. Any capability argument of the form "the model now supports N tokens" or "we added more compute/votes/agents" offered against a limit that is about what the record determines, not about processing capacity.

L1 - The construct ceiling (permanent)

Statement. A verifier cannot be demonstrated to be better than what the validation design can identify. Pairwise expert agreement is evidence about the reference process, not a numeric cap on model accuracy: two experts who are each 90% accurate agree only about 82% of the time, while a 95%-accurate model would agree with either about 86% of the time - so a genuinely better model shows up above expert pairwise agreement, not below it. The real limit is identification: without a trusted reference, or explicit and stated statistical assumptions, true accuracy is not identified, and accuracy claims beyond what the design identifies are assumptions wearing a number. A stronger reference design (consensus panels, independent adjudication) raises what can be demonstrated - and building it is expert work, which is the gold-data investment a pilot pays for

Grounds. This is the classical attenuation bound from psychometrics: an instrument's correlation with a criterion is capped by the square root of the product of their reliabilities (Spearman, 1904; standard in every measurement text). Generalizability studies of rated work show the dominant variance component is the rater x item x criterion interaction - meaning trained human experts do not converge to a point value on judgment-laden criteria, and training does not remove the interaction [CG-02] [CG-04]. Forced-choice validation against a single human label is provably biased when a criterion admits multiple valid interpretations [AJ-02].

Does not imply. That AutoQA is useless on such criteria - it can still be consistent, evidence-anchored, and cheaper, and it can measure and expose the disagreement structure humans currently hide. It implies only that "the system is right" is an unverifiable claim wherever experts themselves do not agree.

Violation signature. Any accuracy claim on a judgment-laden criterion that exceeds the measured inter-adjudicator agreement for that criterion; any deployment that has never measured that agreement.

L2 - Interpretation is underdetermined by the rule text (permanent)

Statement. Rules can be written to decide every case over a defined domain - exact field checks are precisely that. But criteria written in open-ended language underdetermine a fraction of real cases, and for those the truth of "this violates the instructions" is not a property of the record at all; it is a property of an interpretive decision that has not yet been made. On those cases there is nothing for any verifier, human or machine, to be right about.

Grounds. This is the rule-following / legal-indeterminacy problem, and it is not philosophical decoration: it is why the four-cause decomposition of reviewer disagreement includes "the instructions do not decide the case" as a distinct cause, and why rating indeterminacy work requires response-set validation rather than forced single labels (rating indeterminacy, arXiv:2503.05965; [AJ-02]). The correct output on such a case is contract_ambiguous, routed to the instruction owner as a defect report.

Does not imply. That instructions are hopeless - each routed ambiguity, once resolved by the owner, shrinks the underdetermined set for every future case. The system's proper role here is to find these cases, not to answer them.

Violation signature. A deployment with no underdetermined outcome - where every claim is forced to pass or fail - is converting instruction defects into worker faults at an unknown rate. This is the single most common way to operate past the limits while appearing to operate well.

L3 - Intent and counterfactual truth-makers are not in the record (permanent)

Statement. Criteria whose truth conditions reference mental states ("what the user actually wanted") or counterfactuals ("what a competent engineer would have done," "was there a reasonable non-failure explanation") cannot be verified against the record, because the record does not contain their truth-makers. The best achievable is a consistency check: is the reviewer's interpretation defensible given what the record shows? Consistency is not truth.

Grounds. Structural: a closed-world verifier computes relations between the claim and the record. A claim whose truth depends on facts outside the record (unexpressed intent, counterfactual alternatives) is undecidable from the record by construction. This is why a sound design scores opinions only for record-consistency and never truth [kernel C1 lineage], and why "steelman" checks are the measured weak point of every per-flag protocol - they ask the judge to search a space (reasonable alternative explanations) that no record enumerates.

Does not imply. That such criteria should be dropped - they are often the highest-value ones. It implies their verdicts are judgments, must carry judgment authority (advisory, contestable, human-owned), and can never be fail-capable on model output alone.

Violation signature. Automated adverse findings on intent-, deference-, or steelman-class criteria; any language reporting such findings as "verified."

L4 - Existential and universal claims are asymmetric (permanent structure; measured magnitude)

Statement. "The record contains X" is verifiable by exhibiting X. "The record contains no X" (and its twins: "the reviewer missed nothing," "the work is complete") is a universal claim over the whole record. For a finite, correctly parsed record and a precisely defined property, exhaustive checking can certify absence outright. The limit binds non-exhaustive semantic detectors: their result is only "not found, by a detector whose recall on planted instances is rho." If rho is unmeasured, an absence verdict cannot bound the false-negative risk; it reports only which search was run.

Grounds. The structure is elementary logic. The magnitude is measured on current instruments: grounding judges score ~85 F1 confirming support and ~46 F1 catching its absence (FACTS grounding evaluation, December 2025, re-confirmed at the program's own model comparison) [GF-07]. The asymmetry direction is structural; only its size moves between generations, which is why it is re-measured per instrument.

Does not imply. That completeness checking is impossible - enumerable requirements can be compiled into deterministic checks with recall 1.0 by construction. The limit binds semantic absence detection only.

Violation signature. Any fail-capable "missed issue" or "incomplete work" lane whose detector recall has not been measured with planted evidence - including semantically distant and record-distributed plants, not verbatim needles.

L5 - Local versus global truth-makers: the long-context limit (category: permanent; magnitude: generation-relative)

This is the limit your 500,000-token example lives in, and it is really three limits.

Grounds. This limit combines a measured discovery problem, a structural truth-maker problem, and an identifiability problem. The three parts below state those grounds separately.

(a) Retrieval recall is below 1 and unmeasurable without plants (generation-relative). When a reviewer writes "this change is out of scope" in natural English, the claim's tokens share no surface form with the evidence that grounds or defeats it. The truth-maker could be anywhere in the trajectory - an early user turn, a config file, a convention established in a previous session. There is no locality prior: nothing about the claim bounds where to look. Long-form judging remains measurably unstable for current models on document-level reasoning (LongJudgeBench, June 2026) [AJ-12], and chunked retrieval fixes length but not semantic distance [GF-04]. Agentic processing - externalizing search into files, scripts, and iterative queries - measurably improves long-context task performance over raw attention (arXiv:2603.20432, March 2026), which cuts both ways: it is a real mitigation, and it means the deployed instrument is the agent configuration, not the model, so its discovery recall must be measured as configured. A long window clears exactly one of five gates - input capacity - leaving evidence discovery, semantic binding, policy closure, and authorization unestablished. Effective recall on semantically distant evidence is an empirical property of the exact instrument: assume nothing from the context-window specification; price it with planted evidence. This magnitude will move with model generations - which is why it must be measured per instrument with planted evidence, never assumed from a context-window spec sheet.

(b) Some claims have no span-shaped truth-maker at all (permanent). "Out of scope" is frequently not a property of any quotable span. It is a relation between the diff, the request, the repository's conventions, and intent expressed across turns - a global property assembled from distributed state. A citation-based verifier requires an evidence object of the form "this span grounds this claim." For genuinely global claims that object does not exist as a single span; such claims require assembled, multi-span structured evidence whose assembly must itself be validated. Single-span citation is category-inappropriate for them, not merely hard. The honest moves are: decompose the global claim into local sub-claims where possible; otherwise route to a human or abstain. A quoted span attached to a global claim validates transcription, not the relation (our own prototype's false positive demonstrated exactly this: a verbatim quote grounding a wrong finding).

(c) Verification is entailment over a reconstructed world, and the reconstruction is lossy (permanent). The transcript is a log, not the world: environment state that was never printed, tool effects that were never echoed, and the user's evolving intent are not recoverable from it. Two readers - human or model - can reconstruct different worlds from the same 500k tokens, both consistent with the text. Where the claim's truth depends on the unrecorded part, the record is insufficient in principle, and the correct output is exactly that. For claims about program behavior specifically, a separate universal limit applies: no algorithm decides every nontrivial semantic property of arbitrary programs (Rice's theorem) - so code-semantics conclusions must state their boundary: tested inputs, one environment, one abstraction. Unbounded correctness claims never follow from testing or automated analysis alone; scoped correctness proofs against a formal specification remain possible, at a cost no QA pipeline pays.

Does not imply. That long-record review is out of reach. It implies the design must (1) restructure the record into typed indexes so functional claims become queries [F26-02]; (2) require evidence pointers from claimants where possible, converting search into bounded checking [GF-09]; (3) measure residual discovery recall with plants; and (4) reserve a non-automated lane for irreducibly global claims. Anything else is pretending the limit away.

Violation signature. Fail-capable verdicts on global claims (scope, intent-fit, overall preference) justified by span citations; coverage statements with no named detector and no measured recall; per-record evaluation cost that could not possibly have read the record.

L6 - A pipeline of detectors is bounded by its weakest stage, and self-ensembles cannot bound their own error (permanent)

Statement. End-to-end recall is bounded above by claim-extraction recall: a verifier, however perfect, cannot catch an error in a claim that was never extracted. And repeated samples from the same model are correlated observations - a k-vote majority is a stability heuristic, not k independent trials; it cannot produce a calibrated error probability and can amplify a shared bias. No amount of internal redundancy (more votes, more lenses, more agents from the same family) substitutes for external adjudicated gold.

Grounds. The composition bound is arithmetic. Correlated-ensemble limits are standard statistics; same-family judge weakness on hard cases is measured, and at the rubric level - the granularity AutoQA operates at - frontier judges reached only ~55-56% balanced accuracy on hard verification instances in March 2026 [AJ-04] [F26-06]. Our own prototype's headline number moved when a tied 1-1 vote was resolved by iteration order - a live demonstration that vote counts were never probabilities.

Does not imply. That multiple detectors or repeated samples have no engineering value. They may improve stability or reveal disagreement, but their end-to-end error still requires external adjudicated gold.

Violation signature. Confidence numbers derived from vote counts or model self-report; "we run it three times" offered as calibration; extraction recall never measured as its own detector.

L7 - The authorization limit: capability is not authority (permanent; pure statistics)

Statement. Even inside all capability limits, the right to take or influence adverse action is a statistical property that must be purchased with adjudicated, untouched, per-instrument data - and the price is high. With zero observed false accusations among n adverse predictions, the one-sided exact 95% upper bound is 1 - 0.05^(1/n): a 5% ceiling requires n >= 59; 1% requires n >= 299; 0.1% requires n >= 2,995. (Under the two-sided 95% Wilson convention the 5% figure is ~73; program documents use both - state the convention whenever quoting.) One or two cases certify nothing: at n = 1 the upper bound is roughly 79% under Wilson and 95% under the exact bound. Every change to model, prompt, retrieval, contract, or configuration is a new instrument and resets the count to zero. And the estimate is only valid on the difficulty distribution it was measured on: judge accuracy collapses on precisely the hard-subtle slice where consequences concentrate [AJ-04] [F26-06], so aggregate accuracy is an upper bound that the operating region does not enjoy.

Grounds. The numerical ceilings follow from exact binomial and Wilson intervals. Instrument specificity, untouched evaluation data, and distribution matching are standard conditions for applying those bounds to a deployed system.

The correct frame is a risk-coverage curve, not an accuracy number. A judge that abstains on the hard tail and acts only where calibrated can be legitimate at some coverage level. Leadership demanding full coverage at fixed risk - or quoting one accuracy number with abstentions and hard slices averaged in - is not using a better system; it is spending calibration that was never purchased.

Does not imply. That no automated lane can ever receive authority. It means any authority must be narrow, risk-bounded, instrument-specific, and supported by untouched adjudicated data.

Violation signature. Automated or automation-weighted adverse outcomes with no per-lane, per-instrument calibration on untouched gold; one headline accuracy number; calibration reused across model or prompt versions; no abstention path.

L8 - The instrument degrades under the incentives it creates (permanent under adversarial pressure)

Statement. The judged population adapts. A measured criterion redirects effort toward what is measured [DA-03]; judged text is untrusted input to the judge, and prompt-level framing is not a security boundary - injection resistance is a measured, decaying property, not a design guarantee. A static accuracy estimate therefore cannot be assumed stable once consequences attach to the instrument; stability under optimization pressure is monitored, never presumed.

Grounds. The incentive effect follows from the multitask incentive theorem [DA-03]. The security condition follows from the deployment boundary: judged text is attacker-controlled input, while resistance is a measured property of one configured instrument and challenge set rather than a guarantee created by prompt wording.

Does not imply. That blind audit and challenge testing eliminate adaptation or attack. They make degradation observable and force the instrument to earn its operating region again.

Violation signature. No standing blind gold stream; no injection challenge set run per instrument version; calibration measured once and trusted indefinitely.


The worked example: a scoping claim against a 500k-token transcript

A reviewer flags a code change as out of scope, justifying it in natural English. What can any AutoQA legitimately do with this?

  1. Extract the claim and its cited locus - subject to L6: extraction recall is now part of the error budget.
  2. Classify the truth-maker. If the claim is local ("the diff adds a config system the user never asked for") it may reduce to checkable sub-claims: the request text, the diff content, a convention document. If it is global ("this expansion does not serve the user's intent across the session") - L5(b) and L3 apply: no span verifies it, and part of its truth-maker (intent) is not in the record at all.
  3. Search for grounding or defeating evidence - subject to L5(a): recall over 500k semantically-distant tokens is an unmeasured detector until planted-evidence tests price it, and "found nothing" means insufficient, never false (L4).
  4. Judge the interpretive component - "does the scoping rule cover this?" - subject to L2: on a real fraction of cases the rule does not decide, and the honest output is an instruction-defect report, not a verdict on the reviewer.
  5. Aggregate and act - subject to L1 (validation bounded by the reliability of the standard on scoping judgments), L6 (votes are not confidence), and L7 (no adverse authority without per-lane calibration that almost certainly does not exist for this lane).

The conclusion is not that the system can do nothing. On this one claim it can legitimately: verify the checkable sub-claims with quoted spans; assemble the evidence trail a human adjudicator needs; flag the rule gap if the rule is the problem; and abstain with a stated reason otherwise. What it cannot legitimately do - at any model generation for steps 2's global case and L2/L3's interpretive core, and at the current generation for step 3's recall - is return an autonomous verdict on the reviewer.

What remains inside the limits

The limits above leave a large, valuable, defensible operating region:

How to tell a deployment is past the limits - the checklist

A deployment is operating beyond AutoQA's theoretical limits if any of the following is true:

  1. Judgment-laden criteria carry accuracy claims above measured expert agreement - or expert agreement was never measured (L1).
  2. Every claim is forced to pass/fail; there is no underdetermined or insufficient outcome in the schema (L2, L4).
  3. Intent, deference, or steelman findings are fail-capable on model output (L3).
  4. "Missed issue" or completeness verdicts ship without planted-evidence recall measurement (L4, L5a).
  5. Global claims (scope, overall preference, intent-fit) are auto-judged with span citations as justification (L5b).
  6. Confidence is derived from vote counts or self-report rather than untouched adjudicated gold (L6).
  7. Adverse authority exists without per-lane, per-instrument calibration and a stated false-accusation bound - or one accuracy headline covers all lanes and difficulties (L7).
  8. There is no standing blind audit and no injection challenge set (L8).

Each item is checkable against a deployment's configuration and records in an afternoon. That is what makes the limits presentable beyond dispute: they are not opinions about model quality; they are properties of the task, the statistics, and the deployment's own paperwork.

Anticipated objections

"Long-context models are getting better." True, and irrelevant to L1, L2, L3, L5(b), L6, L7, L8, which are not capability limits. It moves L5(a) and L4's magnitude only - and the response is already in the design: measure recall per instrument with plants. A bigger window changes the price of the measurement, not the need for it.

"We'll add more votes / more agents." L6: correlated observations. More samples from the same family narrow variance around the same bias. Only external gold bounds error.

"The model explains its reasoning and quotes the record." Our own prototype produced a false positive with a verbatim, script-validated quote. Quotes prove transcription; explanation is generation, not evidence (L5b).

"Humans are unreliable too." Correct - that is L1's foundation, not a rebuttal. The human ceiling caps what any judge can be proven to do. The question that remains is who holds authority, who is accountable, and whether disagreement is decomposed or erased - questions of governance, which no capability improvement answers.

"The pilot will prove it works." A pilot can establish feasibility on the lanes inside the limits and price the lanes near them. It cannot establish anything about the lanes the limits close - no experiment verifies claims whose truth-makers are absent from the record.

Companion artifacts

An independently produced boundary package dated 2026-07-19 reaches the same determination by the identifiability route. It supplies a formal boundary standard, an automatic-authority policy, a pre-registered challenge protocol, and a leadership briefing deck. The package is archive material, not a current accuracy estimate. Its long-context section uses 2024-generation measurements as historical motivation; those figures are not used in site display prose and must not set a current threshold.

Sources

Ledger-verified: [CG-02] [CG-04] rater-variance G-theory; [AJ-02] forced-choice bias under indeterminacy; [AJ-03] chance-corrected vs raw agreement; [AJ-04] [F26-06] hard-slice judge collapse; [GF-07] grounding asymmetry (Dec 2025); [AJ-12] long-form judge instability; [GF-04] decomposition/decontextualization tension; [F26-02] typed-index retrieval; [GF-09] citation-bounded checking; [CM-05] [CM-03] assisted- review gains; [DA-03] multitask incentive theorem; [CG-07] the legal floor. Named results: Spearman attenuation (1904); Wilson score interval; exact binomial bound; coding-agents long-context processing (arXiv:2603.20432, 2026); Rice (1953); rating indeterminacy (arXiv:2503.05965); prompt-injection against judges (arXiv:2403.17710). The prototype evidence (false positive with verified quote; tie nondeterminism) is in the working proof record.

Next action: apply the violation checklist to the proposed pilot, then record only the capped scoping decision.

Apply the limits to the pilot | Record the scoping decision