{
 "summary": {
  "generated": "2026-07-15",
  "n_claims": 129,
  "n_corpus_claims": 132,
  "n_da_admitted": 3,
  "da_candidates_not_admitted": [
   {
    "id": "DA-04",
    "title": "ANSI/ASQ Z1.4 - Sampling Procedures and Tables for Inspection by Attributes",
    "reason": "not_retrievable",
    "notes": "Both primary and fallback URLs failed retrieval with HTTP 403 (ASQ bot-block). Per instructions, anchors failing both URLs are not_retrievable and no substitute sources were used. No verdict on the claim's substance is possible from retrieved text."
   }
  ],
  "n_sources": 124,
  "n_load_bearing": 36,
  "n_used_on_site": 65,
  "external_verdicts": {
   "confirmed": 83,
   "partially_confirmed": 45,
   "not_retrievable": 7
  },
  "in_corpus_verdicts": {
   "confirmed": 20,
   "partially_confirmed": 7
  },
  "decision_grades": {
   "C": 16,
   "A": 79,
   "B": 40
  },
  "measured_on": {
   "model-outputs": 49,
   "structural": 46,
   "adjacent-domain": 14,
   "human-work": 26
  },
  "deployment_tier_measurements": {
   "count": 0,
   "ids": [],
   "definition": "claims whose measurement includes the deployment tier this program would run on (Claude Fable/Opus 4.8 class, GPT-5.5/5.6 class)"
  },
  "missing_batches": [],
  "method": {
   "layer1": "In-corpus adversarial spot-checks performed 2026-07-14 during the research sweep (27 load-bearing claims; transcripts preserved per claim).",
   "layer2": "Independent re-verification 2026-07-15: every claim checked against ONLY the source URL named in the corpus (plus served redirects and same-document arXiv variants). Verdicts derive exclusively from retrieved text; numbers not visible in retrievable text are listed per claim rather than assumed.",
   "constraint": "No substitute sources, no search engines, no background knowledge. Abstract-level retrieval cannot audit full-PDF-only numbers; those remain attributed to the corpus's own layer-1 transcripts where they exist."
  },
  "corpus_meta": {
   "files": {
    "domain-academic-judges.md": {
     "domain": "academic-judges",
     "n_claims": 12,
     "disagreements": [
      "Dedicated trained judges vs frontier generalists: JudgeBench (ICLR 2025) shows fine-tuned judges and reward models plateau near random-to-64% on hard reasoning-based judging while frontier reasoning models lead (o1-preview 75.4%); but J1 (Meta, 2025-05), CompassJudger-2 (2025-07), and Skywork-Reward-V2 (2025-07) each claim small trained models beat o1-mini/o3-class models on judge/reward benchmarks; and Lambert (RewardBench 2 author, 2025-06) argues scalar reward models STILL beat generative LLM judges at output ranking. The conflict is largely benchmark-dependent (preference-style vs verifiable-reasoning tasks) - no source resolves which wins for criteria-based QA of human annotations, which resembles neither exactly.",
      "Judge-human agreement levels: practitioner sources claim ~85% agreement 'higher than human-human agreement' (Confident AI, 2026) and '>80%, matching human-to-human levels' (Galileo); academic chance-corrected measurements show best-model kappa of 0.28 +/- 0.32 across 20 tasks (JUDGE-BENCH, ACL 2025), universal 33-41pp kappa deflation vs raw agreement (arXiv 2606.19544, 2026-06), and near-orthogonal judge-vs-human evaluation axes on subjective rubrics (arXiv 2606.03043, 2026-06). Much of the gap is percent-agreement vs chance-corrected statistics plus task subjectivity - the practitioner numbers are not wrong so much as unfalsifiably framed.",
      "Verbosity bias magnitude: the 2606.19544 large-scale audit found verbosity bias small (<0.011) across 21 judges under a single controlled pairwise rubric, whereas the broader literature (Zheng 2023 MT-Bench, Gu et al. survey 2026, arXiv 2510.12462) treats verbosity bias as a major systematic failure. Possible reconciliation: bias magnitude depends heavily on rubric structure and elicitation format, meaning it is a controllable design variable rather than a constant.",
      "Run-to-run consistency: The Coin Flip Judge (2026-04) reports 13.6% average pairwise flip rate and recommends ~11-vote majorities (mini-tier OpenAI judges only), while Reliability without Validity (2026-06) measured test-retest reliability >0.95 for some production judges - consistency is highly judge-specific, and the latter paper explicitly warns high consistency coexists with severe position bias (consistency is not validity)."
     ],
     "gaps": [
      "No benchmark meta-evaluates judges on the AutoQA task itself - verifying HUMAN-written evaluative claims (grounding of praise, evidence-conclusion entailment in critiques, severity calibration). Closest proxies are rubric-verification (RubricEval), factuality judging (RewardBench 2 Factuality), and critique evaluation inside J1/CompassJudger training; the positive-claim/'warranted vs empty praise' verification question (design Q3) has no published precision/recall numbers anywhere I could find.",
      "Could not obtain a verified July-2026 leaderboard snapshot for the newest frontier models (GPT-5.x, Claude Opus 4.x, Gemini 3) on JudgeBench or RewardBench 2 - both official leaderboards render via dynamic JS (HuggingFace Spaces) and did not yield data to fetches; the largest systematic evaluation covering 'the April 2026 frontier' (arXiv 2606.19544) reports cohort-level findings but does not name per-model winners in accessible text. Best-available anchor numbers are therefore early-2026 or older.",
      "No single study reports judge-vs-adjudicated-expert agreement broken out per evaluation-axis type (factual grounding vs instruction compliance vs tone vs holistic) - the per-axis agreement ceiling matrix needed for design Q5's lane assignment must be assembled from fragments (RewardBench 2 subsets, JUDGE-BENCH task variance, RubricEval categories) or measured in-house.",
      "Human inter-reviewer reliability floors (Krippendorff's alpha) for annotation-QA reviewers specifically were out of this domain's scope and not covered by the judge meta-evaluation literature; the rating-indeterminacy framework (NeurIPS 2025) is the right machinery but its 11 tasks are generic rating tasks, not training-data QA.",
      "Atla (Selene) company status as of mid-2026 unverified - Selene Mini (8B, open weights, Jan 2025) claims remain reproducible from the model card, but I found no 2026 successor release, so treat Selene as a frozen artifact rather than a maintained product line."
     ],
     "weak_spots": [],
     "n_verifications": 3,
     "unmatched_verifications": []
    },
    "domain-annotation-quality.md": {
     "domain": "annotation-quality",
     "n_claims": 12,
     "disagreements": [
      "Detectability of LLM-generated submissions: Veselovsky et al. (2023/2025) achieved usable classifier-based detection on a long-form text task, and TU Delft's prompt-injection work (TOCHI 2025) claims in-instrument countermeasures work; Westwood (PNAS 2025) and the community survey (arXiv 2606.04924, June 2026) show detection and quality checks are comprehensively evaded by agents and degrade under paraphrase/out-of-domain, with mixed human-AI authorship hardest. The recent evidence favors Westwood, but the sources genuinely conflict on whether content-side detection retains any value as a secondary signal.",
      "Nature of disagreement: the aggregation tradition (Dawid-Skene, MACE, majority vote) models disagreement as noise around one latent truth to be recovered; the perspectivist tradition (Aroyo & Welty, DICES, LeWiDi-2025, Diverging Preferences) treats much of it as legitimate signal to be preserved. Neither side claims universality, but they prescribe opposite data models (adjudicated gold vs. per-rater distributions) - the design must pick per-axis, and no source offers a validated automatic classifier for which regime an axis falls into.",
      "Agreement statistics under skew: Krippendorff defends alpha as the general-purpose reliability standard; Gwet and successors show alpha/kappa collapse under extreme prevalence and push AC1/AC2; Elangovan et al. (ICLR 2025) argue all human-reliability coefficients are misapplied to machine-vs-human comparison in the first place. There is no consensus single statistic - only convergence that one-number reporting is wrong.",
      "Mitigation of crowdworker LLM use: Veselovsky et al. measured that request+hurdle interventions halve usage (a meaningful effect worth deploying); Westwood concludes no instrument-level mitigation withstands agent-based completion ('no magical fix'). These are consistent about direction (partial reduction) but conflict sharply on whether prevention is worth its measured quality cost."
     ],
     "gaps": [
      "No published measurement of the false-agreement rate (judge and attempter agreeing while both wrong per expert panel) in an annotation-QA setting; the nearest proxy is Northcutt's 3.3%+ gold-label error floor. This number would have to be produced in-house via an independent expert-audit channel.",
      "No public 2025-2026 prevalence data for LLM use specifically among paid annotators/reviewers on managed AI-training-data platforms (Scale, Surge, Outlier, Prolific-managed RLHF work); all published prevalence numbers come from MTurk/Prolific survey and summarization tasks, which may under- or over-state prevalence for critique/rating work.",
      "Item-level (as opposed to worker-level) detection of LLM-generated annotations for short/low-dimensional outputs remains unsolved; the NeurIPS 2025 peer-prediction mechanism scores workers, not items, and requires overlapping assignments across workers.",
      "No annotation-science source found that gives sample-size arithmetic for certifying a judge against a gold set under low failure base rates (per-project vs pooled hierarchical validation); this appears only in general eval-statistics literature outside this domain's charter.",
      "No longitudinal studies found measuring whether QA feedback changes annotator behavior over time (repeat-error rates per criterion); the feedback-efficacy question (taxonomy Q7) has essentially no direct evidence in the annotation-quality literature.",
      "Could not open the full text of the Diverging Preferences taxonomy to extract exact percentage splits across its 10 disagreement categories (abstract confirms only that underspecification + style form the majority); exact fractions would sharpen the variance-attribution prior."
     ],
     "weak_spots": [],
     "n_verifications": 3,
     "unmatched_verifications": []
    },
    "domain-critique-models.md": {
     "domain": "critique-models",
     "n_claims": 12,
     "disagreements": [
      "Does adding a human to an AI judge help at all? Vaccaro et al. (Nature Human Behaviour 2024) find human-AI combos on average significantly WORSE than the best solo party on decision tasks, while CriticGPT (2024) finds Human+AI teams beyond the model-only Pareto frontier and DeepMind (2025) gets 91.3% hybrid vs 87.7% AI-alone. Reconciliation offered by DeepMind: complementarity only appears with confidence-based routing to a slice where humans genuinely beat the AI and with non-leading assistance formats - but the DeepMind-adjacent follow-up ('Toward Human-AI Complementarity Across Diverse Tasks') found baseline hybridization added only +0.4pp when model confidence failed to identify the complementarity region. The one-touch design's value is conditional, not guaranteed.",
      "What should the human be shown? DeepMind (2025) finds showing the AI's judgment/reasoning/confidence causes over-reliance and that evidence-only assistance is the only safe format, whereas CriticGPT's protocol and the ICLR 2025 feedback agent delivered full natural-language critiques and reported net gains. Task framing differs (assisting a second-party judge vs coaching the first-party author), which suggests the AutoQA should show full critiques to ATTEMPTERS but evidence-only to human ADJUDICATORS - no single source tests both arms.",
      "Are LLM critics better or worse than humans? CriticGPT: model critiques preferred 63% and catch more bugs than paid contractors. MetaCritique: human critiques have much higher factual precision (87.6% vs 71.9% of atomic units factual) while LLM critiques have higher coverage. Both true simultaneously - the disagreement is about which axis (precision vs comprehensiveness) a QA system should privilege, and sources implicitly pick opposite defaults.",
      "Can critics increase comprehensiveness of review? CriticGPT/Saunders show critics make human review MORE comprehensive (catching missed flaws), but AbsenceBench (2025) shows LLMs are structurally near-blind to omissions (56.9 F1 drop). Reconciliation: the critic gains are demonstrated on commission-type errors (inserted/tampered bugs, incorrect statements); no source demonstrates critic gains on omission-type attempter failures, and AbsenceBench predicts they won't appear without checklist/placeholder scaffolding.",
      "What metric certifies a judge? Open-loop verdict-agreement benchmarks (CriticBench 2024, CriticEval 2024, MT-Bench-style agreement) vs RealCritic's closed-loop position that verdict accuracy without correction-effectiveness is 'superficial success' (24.8% right-verdict-wrong-critique), reinforced by RM-NLHF (ICML 2026) and by 'Reliability without Validity' (2026) showing judge rankings flip up to 15 places depending on which benchmark style you trust."
     ],
     "gaps": [
      "No public post-deployment report from OpenAI on CriticGPT's integration into the production RLHF labeling pipeline (announced intent, June 2024) - no follow-up metrics, and no officially branded successor found through July 2026; the lineage continued in academic work (CTRL, DeepCritic, Critique-RL, MultiCritique) rather than vendor deployment reports.",
      "No published RCT of AI critique/feedback specifically on human DATA-ANNOTATION QA work. Closest analogs are peer review (ICLR 2025) and fact-verification rater assistance (DeepMind 2025). No longitudinal data anywhere on repeat-error-rate reduction per annotator under sustained AI feedback - Q7's efficacy metric has no published precedent to calibrate against.",
      "No measured multi-cycle adversarial study of annotators gaming a deployed critic (Q8's decision-boundary leakage): the NeurIPS'24 checklist gaming is single-shot and author-side; citation-theater/evidence-swap perturbation tests on critics were not found.",
      "Toloka's deployment report - the single closest production analog - publishes its benchmark design (300+ submissions, 20+ projects) but its quantitative precision/recall table is embedded in an image and was not extractable; no independent audit of any commercial annotation-QA LLM system was found.",
      "False-agreement rate (judge and attempter both wrong vs expert panel) is unmeasured in the literature; the CriticGPT 24%-vs-6% flawless-data result is the nearest proxy and is one study, one domain (mostly), one vendor.",
      "Positive-claim/praise verification (is 'accurate/complete/clear' praise warranted?) has no dedicated benchmark or precision/recall measurements - MetaCritique-style AIU machinery has only been applied to flaw-finding critiques. This is a genuine hole given the mission's requirement to evaluate positive statements.",
      "Semantic Scholar returned HTTP 429 during literature-graph search (OpenAlex/Crossref coverage only for that call); one WebSearch call was flagged by a per-call safeguard and was successfully re-run via Exa - no coverage lost, noted for provenance."
     ],
     "weak_spots": [],
     "n_verifications": 3,
     "unmatched_verifications": []
    },
    "domain-grounding-faithfulness.md": {
     "domain": "grounding-faithfulness",
     "n_claims": 12,
     "disagreements": [
      "HHEM's fitness as an automated grounding judge: Vectara's own next-generation hallucination leaderboard (blog, 2025-11-19) still uses commercial HHEM as the sole scorer and reports rates like Gemini-2.5-flash-lite 3.3% vs Gemini-3-pro 13.6%, while Vectara's own research paper (FaithJudge, EMNLP Industry 2025) shows HHEM-2.1-Open at only 66.7% balanced accuracy on FaithBench - near-random on hard hallucinations - and builds an LLM-judge replacement because of 'limitations observed in current hallucination detection methods'. Detector choice materially changes model rankings; automated hallucination-rate leaderboards from different detectors disagree and cannot be treated as interchangeable ground truth.",
      "Atomic decomposition: the FActScore/SAFE/RAGAS lineage (2023-2024) treats finer-grained atomic claims as strictly better for verification, but Decomposition Dilemmas (NAACL 2025) shows decomposition DEGRADES strong verifiers (80.0 -> 71.1 bacc on WiCE with MiniCheck) and DnDScore (EMNLP 2025) shows atomization and decontextualization actively conflict. The field has not converged: long-form factuality pipelines still default to atomic claims while the 2025 analysis papers say granularity must be adaptive.",
      "Specialized small detectors vs frontier LLM judges: LLM-AggreFact (through 2025) and FaithLens (ACL 2026) show small specialized models beating frontier LLMs (Bespoke-MiniCheck-7B > Claude-3.5/GPT-4o; FaithLens-8B > GPT-5.2/o3), but FaithJudge (2025) shows frontier reasoning judges WITH human-annotated exemplars beating all specialized detectors on adversarial sets (84.0 vs 71.2 bacc). Both are right in their regime: fixed checkers win zero-shot on broad distributions; exemplar-anchored frontier judges win on hard, domain-specific cases - which is an argument for a two-tier architecture rather than either alone.",
      "Citation-support standard: ALCE-style span-closed NLI citation precision/recall remains the de-facto automated standard, but CiteEval (ACL 2025) calls NLI-vs-cited-spans 'a suboptimal proxy' and CiteGuard (ACL 2026) additionally doubts LLM-as-judge reliability for citation evaluation - i.e., the two dominant automated approaches each have a 2025-2026 paper attacking them, with no settled successor."
     ],
     "gaps": [
      "No published measurements of run-to-run verdict flip rates (self-consistency) for grounding checkers or LLM faithfulness judges on identical inputs - the taxonomy's 'flip rate: noise vs underspecification signal' question has no literature anchor found; specialized classifiers (MiniCheck/HHEM) are deterministic by construction but that property is unstated in evaluations.",
      "Positive-claim verification of EVALUATIVE statements (warranted vs empty praise, 'stated severity matches evidence strength' calibration claims) is essentially unstudied: the entire 2024-2026 literature verifies factual/descriptive claims against documents; no benchmark tests whether a checker can verify claims like 'this response is complete' or 'this citation is the strongest available'. The AutoQA's core use case - verifying human evaluative writeups about AI outputs - is a domain transfer no found benchmark covers.",
      "No dedicated 2025-2026 study found on evidence-swap/citation-theater probes for verifiers (does the verdict flip when cited evidence is replaced with equally-plausible irrelevant text?). FactEval perturbs claims, not evidence; FACTS documents gaming by response vagueness, not by decorative citation. This test would need to be built in-house.",
      "Aggregation from claim-level verdicts to item-level pass/fail with severity weighting is under-researched: FActScore-style %-supported and FACTS' all-or-nothing eligibility+grounding are the only documented rules; RAGTruth/FaithBench severity taxonomies (benign/questionable/unwanted, evident/subtle) exist as LABELS but FaithJudge found judges classify the middle severity classes unreliably and retreated to binary - no validated severity-weighted aggregation scheme found.",
      "Judge-vs-adjudicated-human agreement ceilings per axis type are only available for grounding (~65 macro-F1 in FACTS on held-out human labels; 84 bacc for FaithJudge on FaithBench) - no comparable 2025-2026 numbers found for instruction-compliance or tone/style axes from the faithfulness literature, so the lane-assignment decision (judge-autonomous vs assisted vs human-only per axis) lacks published anchors outside factual grounding."
     ],
     "weak_spots": [],
     "n_verifications": 3,
     "unmatched_verifications": []
    },
    "domain-hybrid-statistical.md": {
     "domain": "hybrid-statistical",
     "n_claims": 12,
     "disagreements": [
      "Uncertainty-routed vs uniform human sampling: Zrnic & Candes (ICML 2024) and Gligoric et al. (NAACL 2025) report provably valid inference with >25% fewer human labels when the model's uncertainty guides which items humans label; Sfyraki & Wang (ICLR 2026, arXiv:2604.18569) find in the sequential regime that the uncertainty component contributes little and near-constant query probabilities at the budget ceiling give the tightest intervals. Both sides are peer-reviewed; the honest read is that active-sampling gains are regime-dependent (batch design + informative confidence signals vs sequential streaming), so routing policy must be validated per project, not assumed.",
      "Explicit judge-error modeling vs black-box residual calibration: Chen et al. (arXiv:2601.05420) show Rogan-Gladen-style TPR/FPR-correction estimators yield intervals reportedly 3-15x wider than PPI++/EIF for score ESTIMATION; Feng et al. (ICLR 2026, arXiv:2601.20913) argue the opposite direction for CERTIFICATION - explicitly modeling judge TPR/FPR yields better-powered hypothesis tests than treating PPI as a black box. Partly reconcilable (different targets: estimation vs testing), but they issue opposing default recommendations and the synthesis should pick per use-case: PPI++/EIF for dashboards and rates, TPR/FPR-corrected tests for threshold certification.",
      "What the human labels are for - target vs corrector: Trust or Escalate guarantees agreement with human majority preference (capping the system's validity at human panel reliability, i.e., consistency-of-application); the PPI/certification line treats human labels as gold-standard truth for bias correction, and both 'Noisy but Valid' and 'How to Correctly Report' prove regimes where corrected judge-based evaluation is strictly MORE reliable/powerful than human-only evaluation, while Kim (arXiv:2605.16354) explicitly warns her framework breaks when human raters are themselves unreliable. No paper in this literature resolves what to do when the gold itself is noisy - the taxonomy Q1 construct choice (instruction-satisfaction vs consensus-prediction) is genuinely open in the statistics literature."
     ],
     "gaps": [
      "No 2025-2026 paper found that applies PPI, selective escalation, or conformal routing specifically to QA of HUMAN-produced annotations (reviewing human attempters); every application evaluates model outputs with humans as gold. Transfer is plausible but unvalidated, especially under adversarial attempter adaptation, which violates the i.i.d. assumptions most PPI variants rely on (only arXiv:2511.21140 claims calibration-to-test shift robustness).",
      "All PPI-family guarantees are population-level (rates, means, rankings) and selective-evaluation guarantees are marginal, not conditional: a system can meet its aggregate agreement/coverage target while concentrating its failures on specific attempters, criteria, or hard-item subgroups. Conditionally valid (per-subgroup) escalation guarantees for LLM judges remain an open problem in this literature.",
      "The specific fork 'one shallow human touch per item vs zero-touch plus randomized deep audits' has no direct head-to-head study measuring total error caught per human-hour; the statistical literature implicitly favors the zero-touch+audit side (valid conclusions with human labels on a small fraction of items) but only for measurement/certification objectives, not for per-item verdict correction or attempter coaching.",
      "Calibration economics for rubric-based QA: Trust or Escalate's Simulated Annotators and fixed-sequence testing were validated on pairwise preference tasks; required calibration-set sizes and achievable coverage for per-criterion rubric verdicts (the AutoQA's actual workload) are not established anywhere I found.",
      "Could not open the Trust or Escalate OpenReview forum (CAPTCHA); handling of abstained items and calibration-set-size details confirmed only from abstract-level and secondary sources, not reviewer discussion. The 3-15x interval-width figure in Chen et al. was corroborated from search-indexed paper text, not independently recomputed.",
      "No published head-to-head comparison of conformal-width routing vs Simulated-Annotators-style confidence routing vs verbalized-confidence routing for deciding the human-escalation queue - the three candidate mechanisms have never been benchmarked against each other."
     ],
     "weak_spots": [],
     "n_verifications": 3,
     "unmatched_verifications": []
    },
    "domain-rubrics-recent.md": {
     "domain": "rubrics-recent",
     "n_claims": 12,
     "disagreements": [
      "LLM-generated vs expert rubrics: Scale's RaR (Jul 2025) reports reference-grounded synthetic rubrics match human-authored ones (0.359 vs 0.348 as RL rewards), while RubricBench (Mar 2026) measures a stable ~26-27pt preference-accuracy gap that test-time compute cannot close, and SVR (Jun 2026) claims the gap closes (24.1 -> 0.3) only by mining discriminative rubrics from preference pairs. The likely reconciling variable is grounding: RaR's generator saw expert reference answers, RubricBench's saw only the instruction - but no paper tests all conditions head-to-head, so 'can compilation skip expert grounding' remains unsettled; the safe design assumes it cannot.",
      "Aggregation rule: RaR found implicit aggregation (judge internally weighs all criteria into one score) consistently beats explicit per-criterion binary + weighted-sum for reward quality (up to 28% relative), yet the flagship production benchmarks (HealthBench, PaperBench) and OpenAI's grader doctrine use explicit independent per-criterion judgments with arithmetic roll-ups for auditability. Performance and auditability pull in opposite directions; no source resolves the tradeoff for a QA setting where verdicts must be explained.",
      "Decomposition granularity: PaperBench succeeds with 8,316 atomic binary leaves (judge F1 0.83, each leaf graded separately), while RubricBench finds checklists of 13+ items degrade judgment via 'attention displacement', and RIFT lists both Non-Atomic AND Redundant as failure modes. The apparent resolution - atomicity helps when each criterion gets its own judge call, hurts when one call processes a long checklist - is inferred, not directly tested anywhere.",
      "Static vs dynamic rubrics: production eval practice (HealthBench, PaperBench, OpenAI graders) pins rubrics and regression-tests judges against them, treating rubric stability as integrity; the 2026 RL literature (Rubric-ARM, EvoRubrics, ARBOR, AMARIS) argues static rubrics are inherently exploitable specifications that must co-evolve with the optimizing population. For an AutoQA facing adaptive paid attempters, these prescriptions conflict: pin-and-regression-gate vs continuously-adapt.",
      "Judge-model reasoning strength: HealthBench found the non-reasoning GPT-4.1 outperformed reasoning models o3/o4-mini as a rubric grader (MF1 0.709 vs 0.681/0.692), while PaperBench selected reasoning model o3-mini (high effort) and found o1 best (F1 0.84). Whether rubric grading benefits from reasoning-class judges appears task-dependent and both labs partially attribute results to prompt-tuning conducted with the winning model."
     ],
     "gaps": [
      "No published 2025-2026 study applies rubric-based LLM autograding to QA of HUMAN annotators' work product (annotations/critiques/ratings). All quantitative evidence in this domain grades model outputs or RL rollouts; transfer to human-attempter QA - including how paid humans adapt adversarially versus how policies reward-hack - is an assumption, not a finding.",
      "No major training-data vendor (Scale, Surge, Mercor) has published a complete rubric-writing methodology guide. The closest public artifacts are RaR's four design principles, OpenAI's grader-doc guidance, HealthBench's physician criteria process, and RIFT's failure taxonomy; Appen markets 'rubric design' services but publishes no method detail. Vendor-internal rubric-writing guides likely exist but are not public.",
      "Criteria drift has no 2026 quantitative follow-up: EvalGen's finding is qualitative (small-n user study, 2024). No source measures drift rates, drift half-life, or drift-triggered revision thresholds in a production rubric pipeline - the exact parameters design question 9 needs.",
      "The 2026 rubric co-evolution cluster (EvoRubrics, ARBOR, AMARIS) was verified only at abstract/title level; their specific numbers were not confirmed by opening full texts. The trend claim rests on Rubric-ARM (opened) plus five-plus corroborating 2026 titles.",
      "OpenAI's graders documentation page carries a deprecation notice for the graders/evals workflows; what replaces them in OpenAI's mid-2026 stack was not tracked down, so the grader taxonomy finding may describe a sunsetting API surface (the design doctrine it encodes remains valid evidence of practice).",
      "No source directly measures the minimum number of adjudicated exemplars per criterion before judge-human agreement plateaus (the 5 vs 50 vs 500 question); Autorubric's '5-shot calibration -> 80%' on one chemistry dataset is the only datapoint found, and SVR implies exemplar QUALITY (hard contrastive pairs) dominates count."
     ],
     "weak_spots": [],
     "n_verifications": 3,
     "unmatched_verifications": []
    },
    "domain-industry-practice.md": {
     "domain": "industry-practice",
     "n_claims": 12,
     "disagreements": [
      "Judge-vs-human ceiling: LangChain marketing materials claim LLM judges reach ~85% alignment with humans, exceeding ~81% human-human agreement (an MT-Bench-style chat-preference number), while Handshake's BVB shows the best 2026 verifiers at F1 0.633-0.664 against practicing-expert labels (humans at 89.5% agreement) and Surge's AdvancedIF verifier at F1 0.728. The 'judges already beat humans' claim holds only for casual preference tasks, not expert-graded artifact-heavy work - the ceiling is domain-dependent by ~30 F1 points.",
      "Agreement statistic: a major annotation platform publishes chance-corrected Krippendorff's alpha thresholds as its quality standard, while LangSmith/Braintrust standardize on raw percent agreement over tiny reference sets. The industry has not converged, and the eval-tooling side uses exactly the statistic our taxonomy proposes banning.",
      "Is human review the gold standard? Handshake treats dual-coded, adjudicated practicing-expert labels as ground truth for meta-evaluation; Braintrust explicitly warns human review is 'not automatically a gold standard' (untrained reviewers, vague rubrics, fatigue can underperform a decent judge). Reconcilable - adjudicated expert panels vs ad-hoc reviewers - but the sources take opposite rhetorical stances.",
      "Vendor QA claims vs observed integrity: major annotation vendors market rigorous multi-layer QA, SLAs, and 3% acceptance rates, while internal Scale documents (via Inc./Futurism) show spam paid at scale for 11 months, ZeroGPT as a stopgap, and country-level bans; Scale's spokesperson disputes the reporting ('filled with so many inaccuracies'). Marketing descriptions of QA pyramids should not be taken as evidence they perform as described.",
      "Content-based vs economics-based spam defense: Patronus/Galileo/detection vendors implicitly sell content-based classification as sufficient, while the Scale/Outlier record and the annotator-account black market (AlgorithmWatch) show content detection losing to spam economics at scale; Handshake's Cleanlab purchase (statistical label auditing) is a third position - detect via label statistics rather than either content classifiers or human re-review."
     ],
     "gaps": [
      "No data vendor publishes actual achieved inter-reviewer reliability numbers (e.g., Krippendorff alpha values on common item sets) - only target thresholds. The one hard human-agreement number found (Handshake's 89.5%, dual-coded adjudicated subset) is from a benchmark paper, not production QA reporting.",
      "No public data anywhere on false-agreement rates (judge and attempter both wrong vs expert ground truth), on per-project gold-set sizes/refresh rates at data vendors, or on judge-model version pinning / golden-set regression gates at data vendors (taxonomy Q1, Q9). Eval-tooling vendors support regression testing generically but no vendor documents pinning policy for annotation-QA judges.",
      "No published evidence on whether QA feedback changes attempter behavior longitudinally (repeat-error rates per criterion, revision acceptance) - no vendor reports feedback-efficacy metrics, and no controlled feedback-vs-verdict-only comparison was found (taxonomy Q7).",
      "a large data vendor's current (2025-2026) production auto-QA internals are undocumented publicly except via leaks and contributor folklore; Turing publishes essentially nothing concrete; Invisible's public trace is limited to a job post showing AI QA Trainers build rubrics/pass-fail criteria and regression suites. Mercor's QA is visible only through reviewer-hiring materials emphasizing consistency screening.",
      "Could not verify a major annotation platform data-factory blog's exact publication date (undated on page; content verified live 2026-07-14). Sacra figures for Surge/micro1 are analyst estimates, not audited. The Inc. internal documents were verified only via Futurism's direct quotations (Inc. itself 403s).",
      "No vendor was found admitting or measuring feedback-loop gaming (attempters overfitting the QA judge's decision boundary) or monoculture drift from judge feedback - taxonomy Q8's core dynamic is unstudied in public industry sources."
     ],
     "weak_spots": [],
     "n_verifications": 3,
     "unmatched_verifications": []
    },
    "domain-contrarian.md": {
     "domain": "contrarian",
     "n_claims": 12,
     "disagreements": [
      "Can LLM judges match human evaluators at all? Counter-evidence exists: some studies report GPT-4-class judges agreeing with human CONSENSUS at rates exceeding individual human raters (Chatbot Arena-style preference tasks), and one 2026 study reports inter-judge Krippendorff's alpha of 0.867 above the 0.80 threshold. The contrarian corpus (kappa deflation, 64-68% SME agreement in expert domains, schema incoherence) does not refute this; the reconciliation is regime-dependence - judges can beat the mean crowd rater on generic preference tasks while falling below adjudicated expert panels on expertise-heavy and subjective axes. Design consequence: per-axis ceilings, not a global verdict on judge viability.",
      "Human-in-the-loop value: Vaccaro et al. find human-AI combinations LOSE on decision tasks when AI outperforms the human, but the same meta-analysis finds GAINS when the human outperforms the AI - so the one-touch design is not universally bad; it is bad specifically where the judge is stronger than the reviewer, and potentially positive on expert axes where humans still beat judges (per the IUI 2025 expert-domain results). The two contrarian findings point in opposite directions depending on axis type.",
      "Claim decomposition: the decompose-then-verify literature is internally split. Proponents (FactScore lineage, LREC 2026 Japanese decomposition dataset) find atomic claims improve explainability and reduce annotator variability; critics ('Does Claim Decomposition Boost or Burden Fact-checking Performance?' OpenReview; Wanner et al. 2024; the ACL 2025 dynamic-decomposition paper's own literature review) find decomposition does not consistently improve verification across input lengths and verifier strengths, and that atomicity choice itself changes verdicts (accuracy swings of ~0.12 from decomposition policy alone). Neither side has settled the granularity question the AutoQA design poses.",
      "Determinism as a fix for judge inconsistency: common practitioner guidance says run judges at temperature 0 for reproducibility; Rating Roulette (EMNLP 2025) shows disabling sampling measurably REDUCES agreement with human judgment - reproducibility and validity trade off rather than align.",
      "Bias against AI-generated content: earlier work suggested LLM judges penalize (or favor) AI-generated text; Balog, Metzler & Qin (SIGIR 2025) report finding no evidence of bias against AI-generated content in their preliminary study, while confirming a significant bias TOWARD LLM-based rankers' outputs. The direction and existence of AI-content bias is unsettled; the similarity-based affinity bias (Great Models Think Alike) is the better-supported formulation."
     ],
     "gaps": [
      "No published longitudinal study of judge-feedback-induced homogenization of a paid annotator population (the monoculture finding rests on analogical evidence from ideation/creative-writing settings, not annotation QA feedback loops).",
      "No published measurement of the false-agreement rate (AutoQA judge and attempter agreeing while both wrong per expert ground truth) in a production annotation-QA setting; Great Models Think Alike predicts it rises with model similarity but does not measure it in this configuration.",
      "Found no practitioner postmortem of an LLM-based AutoQA layer for human annotation being Goodharted by attempters over feedback cycles - the a large data vendor postmortem is from the pre-LLM-judge QA era, and adversarial-judge evidence comes from benchmark/exam settings; the specific 'scores rise on AutoQA, flat on held-out human-graded goldens' divergence experiment appears not to have been run or published anywhere.",
      "Little direct evidence on positive-claim (praise) verification: no seeded-set study measuring whether judges pass plausible-but-ungrounded praise at higher rates than they fail grounded praise, per praise type - the leniency/sycophancy literature is about response evaluation generally, not verification of an evaluator's positive statements.",
      "Rating Roulette's per-benchmark intra-rater ICC/Krippendorff numbers were not extracted (only pages 1-3 read: abstract-level claims plus the SummEval human-kappa figures verified); the synthesis engine should not cite specific LLM intra-rater coefficients from this report.",
      "The 'Reliability without Validity' (2026) preprint is 0-citation and un-peer-reviewed as of 2026-07-14; its headline numbers are consistent with the older peer-reviewed literature (kappa deflation, position bias) but the specific 33-41pp figure rests on one team's protocol."
     ],
     "weak_spots": [],
     "n_verifications": 3,
     "unmatched_verifications": []
    },
    "domain-frontier-2026.md": {
     "domain": "frontier-2026",
     "n_claims": 12,
     "disagreements": [
      "Decomposition granularity: RubricEval (Mar 2026) finds rubric-level evaluation OUTPERFORMS finer checklist-level decomposition, while LLM-as-a-Verifier (Jul 2026) finds criteria decomposition monotonically IMPROVES verification accuracy. These may be reconcilable (structured-criteria-with-context vs flat atomic checklists) but as stated the 2026 evidence points both ways on how far to decompose.",
      "Vendor optimism vs academic meta-evals: practitioner sources claim judges now agree with humans ~85%, 'higher than two humans agree with each other' (Confident AI, 2026) and treat LLM-judge-first as safe default; academic 2026 meta-evals show GPT-4o at 55.97% on RubricEval-Hard per-criterion judgments and near-total unreliability on fluency/consistency axes (conformal-sets paper). The gap is largely about task granularity and difficulty slice, and the vendor number is unaudited.",
      "Is disagreement noise or signal: the dominant production calibration loop (LangChain/FutureAGI) treats judge-human and human-human disagreement as error to be driven down via corrections and kappa targets, while OpenAI's CoVal explicitly argues that for contested items there is 'no single stable target to predict' and any single aggregated score encodes an arbitrary compromise -- directly opposing answers to taxonomy Q2's consistency-vs-laundering question.",
      "Trusting judge uncertainty for routing: HyPAC's entire mechanism routes on calibrated model uncertainty with PAC guarantees, but Auditing-by-Re-Solving shows judge self-assessment collapses (88% false positives on clean items) precisely when the check exceeds the judge's own solving capability -- so uncertainty-routing is sound only within the judge's competence envelope, a boundary HyPAC's guarantees assume away via its calibration distribution."
     ],
     "gaps": [
      "No dedicated new open judge/evaluator MODEL release found in Jan-Jul 2026 (no Prometheus/Selene/Glider successor surfaced in multiple date-filtered searches); 2026 activity concentrated in verifier architectures, meta-eval benchmarks, and statistical frameworks instead. Absence-of-evidence: a release could have been missed.",
      "No 2026 source measures attempter-side adaptation to an AI QA layer (feedback leakage, score divergence on held-out golden items, monoculture drift) in a production annotation workforce -- Evaluation Faking covers judge-side gaming only. Taxonomy Q8's core empirical questions remain unanswered in the open literature.",
      "No public 2026 vendor product specifically markets claim-level AutoQA of HUMAN annotations/critiques; annotation platforms advertise generic 'AI-assisted QA' and 'model-in-the-loop' (Encord, Kili, major annotation platforms), and Handshake's Gandalf verifies agent rollouts, not human reviewer writeups. The exact product this project is building appears unoccupied in public 2026 positioning.",
      "No 2026 empirical estimate of the false-agreement rate (judge and attempter agreeing but both wrong vs expert ground truth) on annotation-QA-shaped tasks; FindTheFlaws enables the measurement but the number has not been published.",
      "No 2026 study on longitudinal feedback efficacy for human annotators (repeat-error-rate reduction from AI-generated feedback) -- taxonomy Q7's decoration-vs-value question is empirically open.",
      "Could not open OpenReview pages via plain fetch (Cloudflare); Auditing-by-Re-Solving and Evaluation Faking were verified via crawled abstracts and metadata, but full experimental appendices were not inspected. Auditing-by-Re-Solving is an ACL ARR May 2026 submission (under review, not yet accepted)."
     ],
     "weak_spots": [],
     "n_verifications": 3,
     "unmatched_verifications": []
    },
    "critic-and-gapfill.md": {
     "domain": "critic-and-gapfill",
     "n_claims": 24,
     "disagreements": [],
     "gaps": [],
     "weak_spots": [
      "The single most load-bearing transfer assumption is unverified everywhere: every quantitative judge result in the sweep is measured on grading MODEL outputs, while the AutoQA grades HUMAN evaluative writeups (grounding of praise, severity calibration, evidence-conclusion entailment). All nine domains flag this and then the numbers get cited anyway - no domain even found a partial-transfer experiment, so the anchor numbers (77.4 bacc grounding, 55-56% rubric-hard, kappa 0.28) may be systematically off in either direction for the actual workload.",
      "The headline meta-evaluation figures (33-41pp kappa deflation, 14-15 position ranking flips, consistency-bias paradox) all trace to one un-peer-reviewed, 0-citation June 2026 preprint ('Reliability without Validity', arXiv 2606.19544), cited independently by four domains as if it were four sources. Cross-domain repetition is masquerading as corroboration.",
      "Much of the frontier-2026 and rubrics-recent evidence was verified only at abstract/title level (rubric co-evolution cluster, Auditing-by-Re-Solving, Evaluation Faking, Trust or Escalate details, Toloka's precision/recall table locked in an image) - several architecturally decisive claims (agent-in-environment verifiers beat text-only 10x cheaper; consequence-framing softens verdicts with zero CoT trace) rest on numbers nobody in the sweep recomputed or read in full.",
      "The spam-economics narrative (Scale/Bulba, paid gibberish for 11 months, account black market) is sourced via Futurism quoting Inc. documents the researchers could not access (403), is disputed on record by Scale, and dates to the pre-LLM-judge QA era - yet it anchors the taxonomy Q8 conclusion that content-based detection has 'already lost'. Westwood's agent study is the solid leg; the Scale story is journalism-of-journalism.",
      "CriticGPT's 24%-vs-6% flawless-slice result - the best evidence for the whole human+critic premise - is one study, one vendor, mostly one domain (code), with no post-deployment report and no replication in 24+ months. The critique-models domain flags this but the claim still functions as the design's cornerstone.",
      "Several 'disagreement resolutions' offered across domains are inferred reconciliations, not tested ones: verbosity bias as 'a controllable design variable', RaR-vs-RubricBench reconciled by 'grounding' as the variable, atomicity 'helps with per-criterion calls, hurts in long checklists' (explicitly labeled 'inferred, not directly tested'), and 'show critiques to attempters but evidence-only to adjudicators'. If the synthesis treats these as findings, the design inherits speculation as doctrine.",
      "The premise metric is missing on both sides: no measured Krippendorff's alpha exists for the actual human reviewer population this system replaces (vendors publish only thresholds), and none was gathered in-house per the mission context. The system is being designed against 'unexplainable variance' that has never been quantified - the variance-attribution pilot in taxonomy Q1 is not optional and no sweep evidence substitutes for it.",
      "Judge-model evidence is stale relative to the decision date: no verified July-2026 leaderboard for current frontier models (both major leaderboards unfetchable), so any per-model judge-selection guidance in the synthesis is anchored to early-2026-or-older cohorts.",
      "Positive-claim/praise verification - an explicit mission requirement - has literally zero direct evidence in any domain (all six domain gap-lists say so). The synthesis has no precision/recall prior at all for the 'is this praise warranted' machinery and should present it as a from-scratch in-house measurement, not extrapolate from flaw-detection numbers, especially given AbsenceBench predicts critics are near-blind to omissions (empty praise is an omission-shaped failure).",
      "Non-US/non-English practice was never searched: the Chinese data-labeling industry (Baidu/Alibaba crowdsourcing QA, government-backed labeling bases) and Japanese/Korean/Indian vendor practice are absent, so 'industry practice' claims generalize from a US/EU vendor sample."
     ],
     "n_verifications": 0,
     "unmatched_verifications": []
    }
   }
  }
 },
 "claims": [
  {
   "id": "AJ-01",
   "domain": "academic-judges",
   "area": null,
   "claim": "On hard, objectively-verifiable judging tasks (JudgeBench, ICLR 2025), frontier reasoning models dominate dedicated judge models: o1-preview scored 75.4% overall while GPT-4o scored 50.9-56.6% (near random), the best reward model (Skywork-Reward-Gemma-2-27B) hit 64.3%, and fine-tuned judges like PandaLM fell at or below random.",
   "load_bearing": true,
   "evidence": "JudgeBench uses 350 rigorously verified response pairs across knowledge/reasoning/math/coding where correctness is checkable. Verified numbers: GPT-4o vanilla 50.9% overall, Arena-Hard pipeline 56.6%, o1-preview 75.4% (85.7% math/coding), Skywork-Reward-27B 64.3%, reward models 60-64%, ChatEval multi-agent ~34%. Fine-tuned preference judges plateau near random on knowledge/reasoning.",
   "implication": "The AutoQA core judge should be a frontier reasoning model, not an off-the-shelf fine-tuned judge or scalar reward model, because annotation-QA requires reasoning about correctness, not preference matching. Reasoning-model headroom (75% vs 51% for the same lab's non-reasoning model) is the single biggest capability lever.",
   "source": {
    "raw": "JudgeBench: A Benchmark for Evaluating LLM-Based Judges (ICLR 2025) + EmergentMind topic synthesis (updated Jan 2026) | https://arxiv.org/abs/2410.12784 | 2024-10 (ICLR 2025; synthesis updated 2026-01) | academic",
    "title": "JudgeBench: A Benchmark for Evaluating LLM-Based Judges (ICLR 2025) + EmergentMind topic synthesis (updated Jan 2026)",
    "url": "https://arxiv.org/abs/2410.12784",
    "date": "2024-10 (ICLR 2025; synthesis updated 2026-01)",
    "type": "academic"
   },
   "venue": {
    "peer_reviewed": true,
    "marker": "ICLR"
   },
   "same_source_claims": [],
   "in_corpus_verification": {
    "verdict": "confirmed",
    "transcript": "Every quantitative claim checks against the paper's own tables (arXiv:2410.12784 v2, ICLR 2025 camera-ready), read directly from the PDF: GPT-4o vanilla 50.86% / Arena-Hard 56.57% overall (Table 1), o1-preview 75.43% overall with 85.71% on both math and coding (Table 2), Skywork-Reward-Gemma-2-27B 64.29% as best reward model (Table 3, full range 59.43-64.29, paper says \"approximately 59% to 64%\" so the claim's \"60-64%\" is a minor round-up at the low end), PandaLM 13.14% (far below random; paper: all fine-tuned judges except Skywork significantly below random), ChatEval 34.00%. 350 verified pairs across knowledge (154)/reasoning (98)/math (56)/coding (42) confirmed - note this is the GPT-4o-generated split; a separate 270-pair Claude-3.5-Sonnet split exists. Dates confirmed: arXiv v1 2024-10-16, ICLR 2025, EmergentMind synthesis updated 2026-01-09. Two contextual nuances, neither refuting: (1) in the v2 tables o1-preview is not the single best judge - o3-mini (high) scores 80.86% and DeepSeek-R1 73.14%, which strengthens rather than weakens the \"reasoning models dominate\" thesis; (2) 2025-2026 follow-on work (J1, RM-R1, meta-judging, RRD) has pushed JudgeBench accuracies to 77-81%+, so \"GPT-4o near random\" describes the late-2024 snapshot of non-reasoning judges, not the current frontier. The claim as stated about the paper's findings is accurate.",
    "corrected": null
   },
   "external_verification": {
    "verdict": "confirmed",
    "checked": [
     "o1-preview Overall 75.43 (Table 2, Arena-Hard Judge with o1-preview backbone)",
     "o1-preview Math 85.71 and Coding 85.71 (Table 2) = 85.7% math/coding",
     "Vanilla (GPT-4o) Overall 50.86; Arena-Hard Judge (GPT-4o) Overall 56.57 (Table 1/2) = 50.9-56.6%",
     "Skywork-Reward-Gemma-2-27B Overall 64.29, highest reward model (Table 3)",
     "PandaLM Overall 13.14 (Table 1) = well below random",
     "ChatEval Overall 34.00 (Table 1)",
     "350 questions: 154 Knowledge, 98 Reasoning, 56 Mathematics, 42 Coding (p.6)",
     "p.8: GPT-4o 'achieving accuracy no better than random guessing when using the vanilla prompt'"
    ],
    "not_visible": [],
    "quote": "our dataset consists of a total of 350 questions: 154 in Knowledge, 98 in Reasoning, 56 in Mathematics, and 42 in Coding.",
    "notes": "All headline numbers in the claim match the tables exactly. Minor: the paper's reward-model overall range is stated as 'approximately 59% to 64%' (lowest = 59.43), so the evidence field's '60-64%' is a slight low-end rounding; does not affect the claim body. arXiv HTML full text 404'd for both v1/v2, so numbers were read from the PDF (allowed variant b).",
    "checked_at": "2026-07-15",
    "retrieval": "pdf"
   },
   "source_retrieval_meta": {
    "retrieval": "pdf",
    "fetched": [
     "https://arxiv.org/abs/2410.12784",
     "https://arxiv.org/pdf/2410.12784v2"
    ],
    "resolved_title": "JudgeBench: A Benchmark for Evaluating LLM-based Judges",
    "resolved_date": "arXiv v1 2024-10-16, v2 2025-04-05; Published as a conference paper at ICLR 2025",
    "title_match": true
   },
   "used_on": [
    {
     "page": "decisions.html",
     "anchor": "d4",
     "label": "Decisions - judge sourcing"
    }
   ],
   "decision_use": {
    "grade": "C",
    "basis": "superseded-gen capability",
    "measured_on": "model-outputs",
    "note": "Late-2024 snapshot (o1-preview vs GPT-4o, PandaLM). Direction (reasoning models dominate trained judges) has held through 2026 follow-ons; the specific accuracies are two generations stale. Re-rank at the deployment-tier bakeoff.",
    "models_measured": [
     "o1-preview",
     "GPT-4o",
     "Skywork-Reward-Gemma-2-27B",
     "PandaLM",
     "Skywork-Reward-27B"
    ]
   }
  },
  {
   "id": "AJ-02",
   "domain": "academic-judges",
   "area": null,
   "claim": "Standard LLM-judge validation via forced-choice human gold labels is provably biased when rating criteria admit multiple valid interpretations: across 11 real-world rating tasks and 9 commercial LLMs, forced-choice validation selected judge systems performing up to 31% worse than validation using multi-label 'response set' ratings that model indeterminacy.",
   "load_bearing": true,
   "evidence": "Guerdan, Barocas, Holstein, Wallach, Wu, Chouldechova (NeurIPS 2025 poster, abstract opened and confirmed). They formalize 'rating indeterminacy' - items where multiple ratings are reasonable - and show humans and LLMs resolve forced choices differently, so aggregated gold labels systematically mis-rank judges. They provide concrete recommendations for elicitation and aggregation.",
   "implication": "Directly answers design Q1/Q2: the ground-truth construct cannot be a single forced pass/fail human label. The gold set must elicit response-set labels ('which verdicts are defensible?'), and the verdict ontology needs a first-class 'instructions underdetermine this case' class. Validation against collapsed consensus labels will select the wrong judge configuration.",
   "source": {
    "raw": "Validating LLM-as-a-Judge Systems under Rating Indeterminacy (NeurIPS 2025) | https://neurips.cc/virtual/2025/poster/117308 | 2025-09-19 | academic",
    "title": "Validating LLM-as-a-Judge Systems under Rating Indeterminacy (NeurIPS 2025)",
    "url": "https://neurips.cc/virtual/2025/poster/117308",
    "date": "2025-09-19",
    "type": "academic"
   },
   "venue": {
    "peer_reviewed": true,
    "marker": "NeurIPS"
   },
   "same_source_claims": [],
   "in_corpus_verification": {
    "verdict": "partially_confirmed",
    "transcript": "Core claim confirmed against the actual NeurIPS poster page (https://neurips.cc/virtual/2025/poster/117308): authors, rating-indeterminacy framing, forced-choice vs multi-label response-set elicitation, 11 tasks, 9 commercial LLMs, and the \"as much as 31% worse\" figure all match verbatim. Two corrections: (1) \"provably biased\" overstates the paper's own language - it says differing resolution of indeterminacy \"can heavily bias\" validation, and the 31% gap is empirical, not a theorem (the theory links performance measures/elicitation schemes); (2) the 2025-09-19 date is not shown on the poster page - plausibly the acceptance date, but unverified; the paper dates to arXiv March 2025 and NeurIPS Dec 2025, with an ML@CMU blog post 2025-12-09. Search for later work (as of July 2026) found follow-on papers (IRT-based judge reliability, arXiv 2602.00521; grading-scale alignment, arXiv 2601.03444) that build on, not refute or supersede, this result. Code: https://github.com/lguerdan/indeterminacy.",
    "corrected": "Guerdan, Barocas, Holstein, Wallach, Wu & Chouldechova (NeurIPS 2025; arXiv:2503.05965, first posted March 2025) show that standard LLM-judge validation via forced-choice human gold labels can be heavily biased when rating criteria admit multiple valid interpretations (\"rating indeterminacy\"): across 11 real-world rating tasks and 9 commercial LLMs, forced-choice-based validation selected judge systems performing as much as 31% worse than those selected by their framework using multi-label \"response set\" ratings. The bias mechanism is supported by a theoretical framework, but the 31% suboptimality figure is an empirical finding, not a proof."
   },
   "external_verification": {
    "verdict": "confirmed",
    "checked": [
     "11 real-world rating tasks and 9 commercial LLMs",
     "forced-choice validation selects judge systems performing as much as 31% worse than the paper's response-set approach",
     "rating indeterminacy defined as criteria admitting multiple valid interpretations",
     "forced-choice differences 'can heavily bias LLM-as-a-judge validation'",
     "solution uses multi-label 'response set' ratings"
    ],
    "not_visible": [
     "exact source date 2025-09-19 (page shows only 'NeurIPS 2025 Poster')"
    ],
    "quote": "The experiments involve \"11 real-world rating tasks and 9 commercial LLMs\" ... forced-choice validation selects judge systems \"performing as much as 31% worse than judge systems selected by our approach.\"",
    "notes": "Core assertion and both headline numbers (11 tasks, 9 LLMs, up-to-31% gap) match. Claim's '31% worse than multi-label response-set ratings' is consistent with retrieved '31% worse than judge systems selected by our approach' (their approach = multi-label response set). Authors match: Guerdan, Barocas, Holstein, Wallach, Wu, Chouldechova.",
    "checked_at": "2026-07-15",
    "retrieval": "abstract"
   },
   "source_retrieval_meta": {
    "retrieval": "abstract",
    "fetched": [
     "https://neurips.cc/virtual/2025/poster/117308"
    ],
    "resolved_title": "Validating LLM-as-a-Judge Systems under Rating Indeterminacy",
    "resolved_date": "NeurIPS 2025 (no specific day shown on page)",
    "title_match": true
   },
   "used_on": [
    {
     "page": "index.html",
     "anchor": "decision",
     "label": "Overview - the five authorizations"
    },
    {
     "page": "architecture.html",
     "anchor": "ontology",
     "label": "System - claim types"
    },
    {
     "page": "pilot.html",
     "anchor": "main",
     "label": "Pilot - gates and objectives"
    },
    {
     "page": "decisions.html",
     "anchor": "a1",
     "label": "Decisions - the standard"
    }
   ],
   "decision_use": {
    "grade": "A",
    "basis": "validation-design result",
    "measured_on": "structural",
    "note": "A result about how gold labels are elicited and aggregated, not about model skill. Applies to validating Fable/GPT-5.5-class judges exactly as it did to GPT-4o-class.",
    "models_measured": []
   }
  },
  {
   "id": "AJ-03",
   "domain": "academic-judges",
   "area": null,
   "claim": "The largest systematic judge meta-evaluation to date (21 judges, 9 providers, ~541k judgments, including April-2026 frontier models) found raw exact-match agreement universally overstates judge ability - chance-corrected Cohen's kappa is 33-41 percentage points lower on MT-Bench - and judge rankings shift by up to 14 positions depending on which benchmark you use.",
   "load_bearing": true,
   "evidence": "Norman, Rivera, Hughes (arXiv 2606.19544, abstract opened and confirmed). Four cohort-wide findings: universal kappa deflation; benchmark-dependent rankings (up to 14 position shifts across MT-Bench/JudgeBench/RewardBench); a 'consistency-bias paradox' where test-retest reliability >0.95 coexists with position bias >0.10 in two production judges; verbosity bias small (<0.011) under a single controlled pairwise rubric. They distill a Minimum Viable Validation Protocol.",
   "implication": "Answers design Q1's metric question: ban raw percent-agreement and single-number accuracy from all AutoQA reporting; standardize on chance-corrected statistics per benchmark. Also answers Q5: a judge that is perfectly self-consistent can still be systematically biased - consistency monitoring and bias audits are separate mandatory sensors, and per-project validation must use multiple gold-set designs since rankings do not transfer.",
   "source": {
    "raw": "Reliability without Validity: A Systematic, Large-Scale Evaluation of LLM-as-a-Judge Models Across Agreement, Consistency, and Bias | https://arxiv.org/abs/2606.19544 | 2026-06-17 | academic",
    "title": "Reliability without Validity: A Systematic, Large-Scale Evaluation of LLM-as-a-Judge Models Across Agreement, Consistency, and Bias",
    "url": "https://arxiv.org/abs/2606.19544",
    "date": "2026-06-17",
    "type": "academic"
   },
   "venue": {
    "peer_reviewed": true,
    "marker": "Findings"
   },
   "same_source_claims": [
    "CM-09",
    "CT-01"
   ],
   "in_corpus_verification": {
    "verdict": "confirmed",
    "transcript": "Fetched https://arxiv.org/abs/2606.19544 directly. Every quantitative element checks out against the abstract: 21 judges from 9 providers, 118 runs / ~541,000 judgments across MT-Bench/JudgeBench/RewardBench; kappa deflation of 33-41 percentage points on MT-Bench across all 21 models; ranking shifts up to 14 positions across benchmarks; consistency-bias paradox (test-retest reliability >0.95 with position bias >0.10 in two production judges); verbosity bias <0.011; Minimum Viable Validation Protocol; April 2026 frontier models included. Submission date June 17, 2026 matches; authors Norman, Rivera, Hughes match. \"Largest to date\" is the authors' own framing, but two independent searches (July 14, 2026) found no larger or superseding meta-evaluation; follow-on work (AURA arXiv 2606.19714, Apple correlated-panels paper, \"Below the Reliability Floor\" on OpenReview) complements rather than supersedes it. Caveats: preprint with 0 citations; authors themselves limit dataset authority to the March-April 2026 measurement window due to hosted-endpoint drift.",
    "corrected": null
   },
   "external_verification": {
    "verdict": "confirmed",
    "checked": [
     "21 judges from nine providers",
     "approximately 541,000 individual judgments",
     "118 runs",
     "kappa deflation ... universal (33--41 pp on MT-Bench)",
     "judge rankings shift by up to 14 positions across benchmarks",
     "high test--retest reliability (>0.95) with severe position bias (>0.10) in two production-deployed judges",
     "verbosity bias is small (<0.011)",
     "consistency--bias paradox",
     "Minimum Viable Validation Protocol",
     "April 2026 frontier"
    ],
    "not_visible": [],
    "quote": "kappa deflation between exact match and Cohen's kappa is universal (33--41 pp on MT-Bench) ... judge rankings shift by up to 14 positions across benchmarks",
    "notes": "All four cohort-wide findings and every headline number verified in the abstract; verbosity bias <0.011, MVVP, and consistency-bias paradox all present verbatim.",
    "checked_at": "2026-07-15",
    "retrieval": "fulltext"
   },
   "source_retrieval_meta": {
    "retrieval": "fulltext",
    "fetched": [
     "https://arxiv.org/abs/2606.19544",
     "https://arxiv.org/html/2606.19544v1"
    ],
    "resolved_title": "Reliability without Validity: A Systematic, Large-Scale Evaluation of LLM-as-a-Judge Models Across Agreement, Consistency, and Bias",
    "resolved_date": "2026-06-17",
    "title_match": true
   },
   "used_on": [
    {
     "page": "index.html",
     "anchor": "evidence",
     "label": "Overview - what the evidence supports"
    },
    {
     "page": "brief.html",
     "anchor": "measurement",
     "label": "Summary - measurement point"
    },
    {
     "page": "research.html",
     "anchor": "research-cards",
     "label": "Research - findings by decision weight"
    },
    {
     "page": "pilot.html",
     "anchor": "main",
     "label": "Pilot - gates and objectives"
    },
    {
     "page": "decisions.html",
     "anchor": "invariants",
     "label": "Decisions - settled constraints"
    }
   ],
   "decision_use": {
    "grade": "B",
    "basis": "current-gen capability",
    "measured_on": "model-outputs",
    "note": "Includes April-2026 frontier models; the kappa-deflation statistics are properties of agreement measurement and transfer; per-judge rankings do not. Deployment tier and human-writeup QA untested.",
    "models_measured": []
   }
  },
  {
   "id": "AJ-04",
   "domain": "academic-judges",
   "area": null,
   "claim": "At the rubric level - the exact granularity AutoQA would operate at - even frontier judges achieve only ~55-56% balanced accuracy on hard rubric-verification instances (GPT-4o 55.97%, Claude-Sonnet-4.5 55.65%), but rubric-level evaluation with explicit chain-of-thought reasoning beats checklist-level evaluation by 7-12 percentage points and reduces cross-judge variance.",
   "load_bearing": true,
   "evidence": "RubricEval (arXiv 2603.25133, full alphaXiv overview read). Binary satisfied/not-satisfied verification of atomic rubrics against responses; EASY subset ~90% accuracy, HARD subset ~55%. Compositional (conditional-logic) instructions, Role Persona, and Format Structure rubrics hardest. Their multi-model arbitration labeling pipeline (RAF) reached 85.0% agreement with humans, Cohen's kappa 0.702. Reasoning-intensive models (o3) notably outperform standard instruct models.",
   "implication": "Answers design Q3 granularity: decompose to atomic criteria and require per-criterion explicit reasoning (worth +7-12pp), but budget for the cost multiplier and expect a hard residual class (~45% error on contested rubrics) that must route to humans. Easy/hard stratification should be built in: judge consensus at the coarse pass identifies which claims need escalation - the RAF arbitration cascade is a directly reusable architecture for gold-label production.",
   "source": {
    "raw": "RubricEval: A Rubric-Level Meta-Evaluation Benchmark for LLM Judges in Instruction Following | https://www.alphaxiv.org/overview/2603.25133 | 2026-03-26 | academic",
    "title": "RubricEval: A Rubric-Level Meta-Evaluation Benchmark for LLM Judges in Instruction Following",
    "url": "https://www.alphaxiv.org/overview/2603.25133",
    "date": "2026-03-26",
    "type": "academic"
   },
   "venue": {
    "peer_reviewed": false,
    "marker": "preprint"
   },
   "same_source_claims": [
    "F26-06"
   ],
   "in_corpus_verification": null,
   "external_verification": {
    "verdict": "confirmed",
    "checked": [
     "GPT-4o 55.97% balanced accuracy on Hard subset",
     "Claude-Sonnet-4.5 55.65% on Hard subset",
     "rubric-level with reasoning beats checklist-level by a gap of 7-12 points",
     "combining rubric+checklist reduces inter-judge (cross-judge) variance",
     "Easy subset ~90% (Qwen3-235B / gpt-oss-120b ~90%, Table 2: 89.87 / 89.55)",
     "RAF human agreement 85.0% accuracy, Cohen's kappa 0.702",
     "3,486 quality-controlled instances with Easy/Hard subsets"
    ],
    "not_visible": [],
    "quote": "GPT-4o achieves merely 55.97% balanced accuracy, and Claude-Sonnet-4.5 reaches 55.65% ... while checklist-level achieves only 69.90% and 70.44%--a gap of 7-12 points ... Human-RAF agreement reaches 85.0% accuracy with Cohen's kappa = 0.702",
    "notes": "Given alphaXiv URL returned HTTP 403 to WebFetch; curl returned only the SPA shell (page title and meta description confirm the exact title/authors). Verified against the allowed arXiv variant (abs + HTML full text) since the URL contains arXiv ID 2603.25133. All core numbers confirmed verbatim in the HTML full text. Two evidence-field nuances: (a) the paper does NOT explicitly assert reasoning models like o3 outperform instruct models as a class, though o3 leads the Hard split at 84.81 BAcc and is the best single judge at 79.9% on disputed cases; (b) evidence equates 'Compositional' with 'conditional-logic', but the paper shows Compositional *instructions* are hardest while the Conditional Logic *rubric type* actually scores relatively well (Table 4). Neither touches the core claim.",
    "checked_at": "2026-07-15",
    "retrieval": "fulltext"
   },
   "source_retrieval_meta": {
    "retrieval": "fulltext",
    "fetched": [
     "https://www.alphaxiv.org/overview/2603.25133",
     "https://arxiv.org/abs/2603.25133",
     "https://arxiv.org/html/2603.25133v1"
    ],
    "resolved_title": "RubricEval: A Rubric-Level Meta-Evaluation Benchmark for LLM Judges in Instruction Following",
    "resolved_date": "2026-03-26",
    "title_match": true
   },
   "used_on": [
    {
     "page": "index.html",
     "anchor": "evidence",
     "label": "Overview - what the evidence supports"
    },
    {
     "page": "research.html",
     "anchor": "research-cards",
     "label": "Research - findings by decision weight"
    }
   ],
   "decision_use": {
    "grade": "B",
    "basis": "current-gen capability",
    "measured_on": "model-outputs",
    "note": "March-2026 on GPT-4o / Claude-Sonnet-4.5 / o3. One tier behind deployment; hard-subset ~55% is a prior, not a ceiling, for Fable/GPT-5.5-class. The +7-12pp per-criterion-reasoning effect is the durable part.",
    "models_measured": [
     "GPT-4o",
     "Claude-Sonnet-4.5",
     "o3"
    ]
   }
  },
  {
   "id": "AJ-05",
   "domain": "academic-judges",
   "area": null,
   "claim": "On subjective rubrics, LLM judges' evaluation axis is nearly orthogonal to the human axis (87-89 degrees vs 78-81 degrees human-to-human), judges use only 0.3-0.5x the human score spread, and inter-LLM agreement (r~0.35) exceeds LLM-human agreement (r~0.27-0.32) - while on a rubric with a verifiable factual answer the same judges fall back into the human range (58.5 degrees, r=0.519).",
   "load_bearing": false,
   "evidence": "Mukherjee, Hamna, Bali, Sitaram (arXiv 2606.03043, abstract opened). 41 LLM judges, 4 datasets, 8 Indic languages, geometric analysis with bootstrap CIs. Fine-tuning/preference optimization recovers score spread (0.32 to 1.08) but barely moves the axis; only post-hoc calibration on a small human-anchored set improved all rubrics. Scope caveat: Indic-language community-health data, but the verifiable-vs-subjective contrast is the mechanism of interest.",
   "implication": "Multi-judge ensembles and inter-judge agreement are NOT validity evidence on subjective axes - consensus can reflect a shared collapsed subspace. The objective-subjective boundary (design Q2) is empirically real and large: verifiable criteria are judge-autonomous territory; subjective criteria need human-anchored calibration sets, not more judges.",
   "source": {
    "raw": "The Geometry of LLM-as-Judge: Why Inter-LLM Consensus Is Not Human Alignment | https://arxiv.org/abs/2606.03043 | 2026-06-02 | academic",
    "title": "The Geometry of LLM-as-Judge: Why Inter-LLM Consensus Is Not Human Alignment",
    "url": "https://arxiv.org/abs/2606.03043",
    "date": "2026-06-02",
    "type": "academic"
   },
   "venue": {
    "peer_reviewed": false,
    "marker": "preprint"
   },
   "same_source_claims": [],
   "in_corpus_verification": null,
   "external_verification": {
    "verdict": "confirmed",
    "checked": [
     "41 LLM judges, 4 community-built Indic datasets, 8 Indic languages",
     "judge/human score-spread ratio sigma_J/sigma_H approx 0.3-0.5",
     "principal angle to human axis 87-89 deg vs human-human 78-81 deg",
     "inter-LLM r_LL approx 0.35 vs LLM-human r_LH approx 0.27-0.32",
     "verifiable factual rubric: axis 58.5 deg, r_LH = 0.519",
     "fine-tuning recovers spread 0.32 -> 1.08 but axis stays 87-88 deg",
     "geometry measured with bootstrap confidence intervals; post-hoc calibration on a small human-anchored set improves all rubrics"
    ],
    "not_visible": [],
    "quote": "87 deg--89 deg versus 78 deg--81 deg ... r_LL approx 0.35 versus r_LH approx 0.27--0.32 ... axis 58.5 deg; r_LH = 0.519 ... sigma_J / sigma_H approx 0.3--0.5",
    "notes": "Every headline number in the claim is present verbatim in the abstract, including the verifiable-vs-subjective contrast (58.5 deg / r=0.519) and the fine-tuning spread recovery (0.32->1.08) with a near-stationary axis. Scope caveat (Indic community-health data) also matches.",
    "checked_at": "2026-07-15",
    "retrieval": "abstract"
   },
   "source_retrieval_meta": {
    "retrieval": "abstract",
    "fetched": [
     "https://arxiv.org/abs/2606.03043"
    ],
    "resolved_title": "The Geometry of LLM-as-Judge: Why Inter-LLM Consensus Is Not Human Alignment",
    "resolved_date": "2026-06-02",
    "title_match": true
   },
   "used_on": [],
   "decision_use": {
    "grade": "B",
    "basis": "current-gen capability",
    "measured_on": "model-outputs",
    "note": "June-2026, 41 judges. Subjective-vs-verifiable axis geometry is the mechanism; Indic community-health scope noted in corpus. Consensus-is-not-validity conclusion is method-level and durable.",
    "models_measured": []
   }
  },
  {
   "id": "AJ-06",
   "domain": "academic-judges",
   "area": null,
   "claim": "Single-trial LLM judging is measurably noisy: across 29 tasks with 50 repeated trials, pairwise preferences flipped on average 13.6% of the time (28% of questions exceeded 20% flip rate), semantically equivalent prompt templates changed majority outcomes in 25% of tested cases, and ~11 repeated trials were needed for majority vote to recover the reference verdict with 95% probability.",
   "load_bearing": false,
   "evidence": "Yagubyan (arXiv 2606.13685, abstract opened). GPT-4o-mini and GPT-4.1-mini judges; also found significant first-position bias (72% A-majority, p=0.024), cross-judge agreement only 76% (kappa 0.51), and a pairwise-pointwise gap where judges pick winners even when their own scalar scores show no meaningful difference. Caveat: both judges from one provider, mini-tier models; the Reliability-without-Validity cohort shows some production judges reach test-retest >0.95.",
   "implication": "Answers design Q5's flip-rate fork with 'both': k-sample voting (k around 5-11) plus position randomization is baseline hygiene, AND per-item flip-rate is a cheap, free signal of criterion underspecification - items with high vote entropy are exactly the 'instructions underdetermine this' class. A perturbation-robustness harness is justified as a ship gate.",
   "source": {
    "raw": "The Coin Flip Judge? Reliability and Bias in LLM-as-a-Judge Evaluation | https://arxiv.org/abs/2606.13685 | 2026-04-23 | academic",
    "title": "The Coin Flip Judge? Reliability and Bias in LLM-as-a-Judge Evaluation",
    "url": "https://arxiv.org/abs/2606.13685",
    "date": "2026-04-23",
    "type": "academic"
   },
   "venue": {
    "peer_reviewed": false,
    "marker": "preprint"
   },
   "same_source_claims": [],
   "in_corpus_verification": null,
   "external_verification": {
    "verdict": "confirmed",
    "checked": [
     "29 tasks spanning 10 categories",
     "50 pairwise trials (and 50 pointwise) per question",
     "pairwise preferences flip on average 13.6%",
     "28% of questions exceed a 20% flip rate (one hits 56%)",
     "equivalent templates change majority outcomes in 25% of tested cases",
     "11 repeated trials needed for majority vote to match 50-trial reference at 95%",
     "first-position bias 72% A-majority, p=0.024 (GPT-4o-mini)",
     "cross-judge agreement 76%, kappa 0.51",
     "judges: GPT-4o-mini and GPT-4.1-mini",
     "pairwise-pointwise gap: small non-significant scalar gaps despite winner picks"
    ],
    "not_visible": [
     "the claim's own caveat that a 'Reliability-without-Validity cohort' shows some production judges reach test-retest >0.95 (external context, not from this paper)"
    ],
    "quote": "pairwise preferences flip on average 13.6% of the time ... 28% of questions exceeding a 20% flip rate ... two OpenAI judge models (GPT-4o-mini and GPT-4.1-mini) ... a significant first-position bias (72% A-majority, p = 0.024)",
    "notes": "Every headline number in the claim matches the abstract verbatim. Single-provider mini-tier caveat is stated by the paper (recommends cross-provider replication as future work).",
    "checked_at": "2026-07-15",
    "retrieval": "abstract"
   },
   "source_retrieval_meta": {
    "retrieval": "abstract",
    "fetched": [
     "https://arxiv.org/abs/2606.13685"
    ],
    "resolved_title": "The Coin Flip Judge? Reliability and Bias in LLM-as-a-Judge Evaluation",
    "resolved_date": "2026-04-23",
    "title_match": true
   },
   "used_on": [],
   "decision_use": {
    "grade": "C",
    "basis": "superseded-gen capability",
    "measured_on": "model-outputs",
    "note": "Measured on mini-tier judges (GPT-4o-mini / GPT-4.1-mini). Flip rates for deployment-tier reasoners are unknown and plausibly far lower; k-sampling remains cheap hygiene either way.",
    "models_measured": [
     "GPT-4o-mini",
     "GPT-4.1-mini"
    ]
   }
  },
  {
   "id": "AJ-07",
   "domain": "academic-judges",
   "area": null,
   "claim": "The broadest LLM-vs-human-judge study on existing human-annotated datasets (JUDGE-BENCH, 20 datasets) found the best model (GPT-4o) reached only kappa = 0.28 +/- 0.32 on categorical judgments and Spearman rho = 0.50 +/- 0.21 on graded ones, with enormous task-to-task variance - and LLMs agreed MORE with non-expert annotators than with experts.",
   "load_bearing": false,
   "evidence": "Bavaresco et al., 'LLMs instead of Human Judges?' (ACL 2025). Structured/rule-like attributes correlate reliably; subjective attributes (engagingness) poorly; safety judgments produced negative correlations due to guardrail refusals. Authors conclude task-specific human validation is essential before deploying LLM judges. Read via detailed paper-note secondary source; consistent with the ACL publication.",
   "implication": "Chance-corrected agreement with humans on subjective axes is far lower than the practitioner '85% agreement' narrative. The agree-more-with-non-experts finding is a red flag for the expertise-gap root cause: an uncalibrated judge may replicate crowd-level judgment, not expert judgment - expert-adjudicated (not crowd-consensus) gold labels are required for validation.",
   "source": {
    "raw": "LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks (JUDGE-BENCH) | https://en.papernotes.org/ACL2025/llm_nlp/llm_vs_human_judges_study/ | 2025-07 (ACL 2025) | academic",
    "title": "LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks (JUDGE-BENCH)",
    "url": "https://en.papernotes.org/ACL2025/llm_nlp/llm_vs_human_judges_study/",
    "date": "2025-07 (ACL 2025)",
    "type": "academic"
   },
   "venue": {
    "peer_reviewed": true,
    "marker": "ACL"
   },
   "same_source_claims": [],
   "in_corpus_verification": null,
   "external_verification": {
    "verdict": "confirmed",
    "checked": [
     "Judge-Bench with 20 datasets (70k+ instances), 11 LLM judges",
     "GPT-4o overall leader",
     "GPT-4o Categorical Avg kappa = 0.28 +/- 0.32 (Table 1)",
     "GPT-4o Graded Avg rho = 0.50 +/- 0.21 (Table 1)",
     "higher correlation with non-expert than expert evaluators",
     "every model struggles on engagingness",
     "safety guardrail refusals -> negative correlations",
     "task-specific human validation remains critical"
    ],
    "not_visible": [
     "author name 'Bavaresco' (not present on this secondary paper-note page)"
    ],
    "quote": "Table 1 lists GPT-4o Categorical Avg kappa as 0.28 +/- 0.32 ... GPT-4o Graded Avg rho as 0.50 +/- 0.21. 'Models exhibit higher correlation with non-expert evaluators.' 'Achieves negative correlations on DICES and Medical-safety due to overactive guardrails.'",
    "notes": "Secondary source (a paper note dated 2026-05-08 discussing the ACL 2025 paper, arXiv 2406.18403 cited on the page); the claim self-identifies as read via a paper note. All headline numbers (kappa 0.28+/-0.32, rho 0.50+/-0.21) and the non-expert>expert and safety-refusal findings match verbatim. source_date '2025-07 (ACL 2025)' is the paper's venue; page metadata date is 2026-05-08.",
    "checked_at": "2026-07-15",
    "retrieval": "fulltext"
   },
   "source_retrieval_meta": {
    "retrieval": "fulltext",
    "fetched": [
     "https://en.papernotes.org/ACL2025/llm_nlp/llm_vs_human_judges_study/"
    ],
    "resolved_title": "LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks (ACL 2025 paper note)",
    "resolved_date": "2026-05-08",
    "title_match": true
   },
   "used_on": [],
   "decision_use": {
    "grade": "C",
    "basis": "superseded-gen capability",
    "measured_on": "model-outputs",
    "note": "GPT-4o-era kappa 0.28+/-0.32. Use as the canonical warning against percent-agreement narratives, not as an estimate of current judge-human agreement.",
    "models_measured": [
     "GPT-4o"
    ]
   }
  },
  {
   "id": "AJ-08",
   "domain": "academic-judges",
   "area": null,
   "claim": "Preference leakage is a measured, large contamination bias: a judge favors outputs of models trained on data from a related generator by up to 28.7 percentage points (Preference Leakage Score), with a relatedness gradient - same model 23.6% average, fine-tuned descendants 19-22%, same family 2.8-8.9% - and judges cannot detect their own students (near-chance) while a BERT classifier can (82.4%).",
   "load_bearing": false,
   "evidence": "Li et al., ICLR 2026 (arXiv 2502.01534). SFT transmits the bias strongly (23.6% avg PLS) vs DPO (5.2%); effect scales roughly linearly with synthetic-data mixing ratio with no threshold; style/format removal cuts PLS sharply; only contextual calibration mitigated it (17.8 to 7.3); prompting-based mitigation failed.",
   "implication": "Answers design Q5 self-preference: when attempters critique (or produce with the help of) outputs from model family X, the judge should not be family X or its distillation lineage - and since lineage is often undisclosed, cross-family judge routing plus style-stripping/canonicalization is the practical defense. Prompt-level debiasing will not fix this.",
   "source": {
    "raw": "Preference Leakage: A Contamination Problem in LLM-as-a-judge (ICLR 2026) | https://arxiv.org/abs/2502.01534 | 2026-02 (ICLR 2026; arXiv 2025-02) | academic",
    "title": "Preference Leakage: A Contamination Problem in LLM-as-a-judge (ICLR 2026)",
    "url": "https://arxiv.org/abs/2502.01534",
    "date": "2026-02 (ICLR 2026; arXiv 2025-02)",
    "type": "academic"
   },
   "venue": {
    "peer_reviewed": true,
    "marker": "ICLR"
   },
   "same_source_claims": [],
   "in_corpus_verification": null,
   "external_verification": {
    "verdict": "confirmed",
    "checked": [
     "up to 28.7% preference leakage score (Table 1/2)",
     "same model 23.6% average PLS (Table 2)",
     "inheritance/fine-tuned 19.3% and 22.3% (in 19-22% band)",
     "same family 8.9% same-series vs 2.8% different-series",
     "SFT 23.6% avg vs DPO 5.2%",
     "judge self-detection near random (e.g. 41.0%, 52.0%)",
     "BERT classifier 82.4%",
     "leakage proportional to synthetic-data amount, no clear threshold",
     "style/format removal is largest decrease (e.g. 17.5% to 9.0%)",
     "contextual calibration Error Bias 17.8 to 7.3",
     "prompting mitigation failed (17.8 to 18.3, slightly worse)",
     "accepted ICLR 2026 (abstract page)"
    ],
    "not_visible": [],
    "quote": "contextual calibration with an additional held-out set for bias adjustment is the most effective, reducing Error Bias from 17.8 to 7.3.",
    "notes": "Abstract page confirms core + ICLR 2026 acceptance but shows no numbers; all 11 headline figures verified verbatim in the arxiv.org/html full text. Authors: Dawei Li et al. (Huan Liu group).",
    "checked_at": "2026-07-15",
    "retrieval": "fulltext"
   },
   "source_retrieval_meta": {
    "retrieval": "fulltext",
    "fetched": [
     "https://arxiv.org/abs/2502.01534",
     "https://arxiv.org/html/2502.01534"
    ],
    "resolved_title": "Preference Leakage: A Contamination Problem in LLM-as-a-judge",
    "resolved_date": "arXiv v1 2025-02-03; v3 2026-03-04; accepted ICLR 2026",
    "title_match": true
   },
   "used_on": [
    {
     "page": "pilot.html",
     "anchor": "main",
     "label": "Pilot - gates and objectives"
    }
   ],
   "decision_use": {
    "grade": "B",
    "basis": "current-gen mechanism",
    "measured_on": "model-outputs",
    "note": "ICLR 2026. Lineage-based contamination is structural to training pipelines and likely persists at the deployment tier; exact PLS magnitudes are tier-bound. Cross-family routing stays the defense.",
    "models_measured": []
   }
  },
  {
   "id": "AJ-09",
   "domain": "academic-judges",
   "area": null,
   "claim": "Dedicated trained judges advanced substantially in 2025: Meta's J1 (RL-trained thinking-judge, May 2025) at 32B outperforms o1-mini, o3, and 671B DeepSeek-R1 on some judge benchmarks; CompassJudger-2-7B (July 2025) matches far larger generalists on judge benchmarks; Skywork-Reward-V2-Llama-3.1-8B (July 2025) topped seven reward benchmarks including JudgeBench.",
   "load_bearing": false,
   "evidence": "J1 arXiv 2505.10320 (verified): unified verifiable-reward RL over 22K synthetic pairs, mitigates position bias by construction, gains from test-time self-consistency. CompassJudger-2 arXiv 2507.09104: verifiable-reward + margin-loss training; 32B scores 80.9% accuracy on JudgerBenchV2. Skywork-Reward-V2 arXiv 2507.01352: 26M human-AI-curated preference pairs; 8B model led RewardBench v1/v2, PPE, RM-Bench, JudgeBench among reward models.",
   "implication": "Small trained judges are now viable for high-volume, well-specified sub-checks (cost tier), but their wins are on preference-style benchmarks; for reasoning-heavy annotation QA the JudgeBench evidence still favors frontier reasoning models. A two-tier design (cheap trained judge for triage, frontier reasoner for verdicts) is supported; the J1 training recipe (verifiable rewards, position-bias-symmetric training) is reusable if a custom judge is ever trained.",
   "source": {
    "raw": "J1: Incentivizing Thinking in LLM-as-a-Judge via Reinforcement Learning (with CompassJudger-2, Skywork-Reward-V2) | https://arxiv.org/abs/2505.10320 | 2025-05 to 2025-07 | academic",
    "title": "J1: Incentivizing Thinking in LLM-as-a-Judge via Reinforcement Learning (with CompassJudger-2, Skywork-Reward-V2)",
    "url": "https://arxiv.org/abs/2505.10320",
    "date": "2025-05 to 2025-07",
    "type": "academic"
   },
   "venue": {
    "peer_reviewed": false,
    "marker": "preprint"
   },
   "same_source_claims": [],
   "in_corpus_verification": null,
   "external_verification": {
    "verdict": "partially_confirmed",
    "checked": [
     "J1-Qwen-32B outperforms o1-mini, o3, and a much larger 671B DeepSeek-R1 on some benchmarks",
     "RL framework teaching LLM judges to think before deciding (thinking-judge)",
     "unified format with verifiable rewards; reduces positional bias",
     "trained at scales of 8B, 32B, and 70B",
     "trained only on synthetic data; v1 submitted May 15 2025"
    ],
    "not_visible": [
     "'22K synthetic pairs' (abstract says only 'synthetic data', no count)",
     "'test-time self-consistency' (abstract mentions iterative self-correction, not self-consistency)",
     "attribution to 'Meta' (authors are Meta FAIR but abstract does not state Meta)",
     "CompassJudger-2-7B claim and 80.9% JudgerBenchV2 (cites arXiv 2507.09104, a different URL not fetched)",
     "Skywork-Reward-V2-Llama-3.1-8B, 26M pairs, topping seven benchmarks incl JudgeBench (cites arXiv 2507.01352, a different URL not fetched)"
    ],
    "quote": "J1-Qwen-32B, our multitasked pointwise and pairwise judge also outperforms o1-mini, o3, and a much larger 671B DeepSeek-R1",
    "notes": "The J1-specific core assertion (32B beats o1-mini/o3/671B DeepSeek-R1, RL thinking-judge, verifiable rewards, position-bias mitigation) is confirmed. The claim bundles two other papers (CompassJudger-2, Skywork-Reward-V2) with their own arXiv IDs that the hard constraints forbid fetching; those portions are not verifiable from this source.",
    "checked_at": "2026-07-15",
    "retrieval": "abstract"
   },
   "source_retrieval_meta": {
    "retrieval": "abstract",
    "fetched": [
     "https://arxiv.org/abs/2505.10320"
    ],
    "resolved_title": "J1: Incentivizing Thinking in LLM-as-a-Judge via Reinforcement Learning",
    "resolved_date": "2025-05-15 (v1); v3 2025-10-13",
    "title_match": true
   },
   "used_on": [],
   "decision_use": {
    "grade": "C",
    "basis": "superseded-gen capability",
    "measured_on": "model-outputs",
    "note": "Mid-2025 trained-judge landscape (J1, CompassJudger-2, Skywork-V2). This market moves quarterly; treat as a map of the design space, not a current ranking.",
    "models_measured": [
     "o1-mini",
     "o3",
     "DeepSeek-R1",
     "Skywork-Reward-V2-Llama-3.1-8B",
     "Skywork-Reward-V2"
    ]
   }
  },
  {
   "id": "AJ-10",
   "domain": "academic-judges",
   "area": null,
   "claim": "RewardBench 2 (Ai2, June 2025) made reward-model evaluation ~20 points harder via best-of-4 format (random = 25%) and unseen human prompts; top models score below 40% on Precise Instruction Following, yet benchmark scores correlate 0.87 (Pearson) with downstream best-of-N performance across 113 reward models.",
   "load_bearing": false,
   "evidence": "arXiv 2506.01937 (verified via abstract, Ai2 dataset page, and Lambert's analysis post). 1,865 cases across Factuality, Focus, Math, Precise IF, Safety, Ties. Lambert (author) notes generative LLM-as-judge approaches remain 'weaker than expected relative to standard reward models' at output ranking, and that frontier models still fail trivial rankings ('name a color in the rainbow').",
   "implication": "Precise-instruction-following is the weakest judged capability (<40% at best-of-4) - exactly the 'is this annotation compliant with written project instructions?' skill AutoQA needs. Do not assume instruction-compliance checking is a solved sub-problem; it needs its own gold set, and deterministic/programmatic checks should absorb whatever can be compiled to rules.",
   "source": {
    "raw": "RewardBench 2: Advancing Reward Model Evaluation | https://arxiv.org/abs/2506.01937 | 2025-06-02 | academic",
    "title": "RewardBench 2: Advancing Reward Model Evaluation",
    "url": "https://arxiv.org/abs/2506.01937",
    "date": "2025-06-02",
    "type": "academic"
   },
   "venue": {
    "peer_reviewed": false,
    "marker": "preprint"
   },
   "same_source_claims": [],
   "in_corpus_verification": null,
   "external_verification": {
    "verdict": "confirmed",
    "checked": [
     "~20 points harder: 'models score about 20 points on average lower on RewardBench 2 compared to the first RewardBench'",
     "best-of-4 format with random baseline 25%",
     "1,865 prompts",
     "leading models below 40% on Precise Instruction Following",
     "Pearson 0.87 correlation with downstream BoN sampling",
     "113 reward models evaluated on BoN",
     "six domains: factuality, precise instruction following, math, safety, focus, ties",
     "sources new human prompts instead of existing prompts from downstream evaluations"
    ],
    "not_visible": [
     "Lambert blog quotes 'weaker than expected relative to standard reward models' and 'name a color in the rainbow' (attributed in evidence to an external Lambert analysis post, not the arXiv text)"
    ],
    "quote": "There is only one correct chosen response, meaning the random baseline is 25% accuracy ... average score on downstream tasks with BoN sampling has a high Pearson correlation of 0.87 ... We evaluated 113 RMs",
    "notes": "Abstract page gave only the ~20-point figure and 'new human prompts'; all specific numbers (best-of-4, 25%, 1,865, <40% Precise IF, 0.87, 113, six domains) were verified verbatim from the arXiv HTML full text (v2). Every headline number in the claim matches.",
    "checked_at": "2026-07-15",
    "retrieval": "fulltext"
   },
   "source_retrieval_meta": {
    "retrieval": "fulltext",
    "fetched": [
     "https://arxiv.org/abs/2506.01937",
     "https://arxiv.org/pdf/2506.01937",
     "https://arxiv.org/html/2506.01937v2"
    ],
    "resolved_title": "RewardBench 2: Advancing Reward Model Evaluation",
    "resolved_date": "2025-06-02 (v1); v2 2026-04-23; ICLR 2026",
    "title_match": true
   },
   "used_on": [],
   "decision_use": {
    "grade": "C",
    "basis": "superseded-gen capability",
    "measured_on": "model-outputs",
    "note": "June-2025; top models <40% on Precise Instruction Following predates GPT-5-class. The warning (instruction-compliance checking is not solved) held then; current magnitude unknown - exactly why it needs its own gold set.",
    "models_measured": []
   }
  },
  {
   "id": "AJ-11",
   "domain": "academic-judges",
   "area": null,
   "claim": "In expert-knowledge domains, LLM judges agreed with subject-matter experts only 68% (dietetics, registered dietitians) and 64% (mental health, clinical psychologists) on overall pairwise preference, with agreement varying further across domain-specific aspect questions.",
   "load_bearing": false,
   "evidence": "Szymanski et al. (arXiv 2410.20266; ACM IUI 2025, ~220 citations). Mixed-methods pairwise comparison study; authors conclude LLMs alone lack the depth for complex knowledge-specific tasks and experts must stay in the evaluation loop. Full abstract text opened and confirmed.",
   "implication": "For projects requiring genuine domain expertise (root cause 2), expect judge-expert agreement in the 60s on holistic preference - below any reasonable autonomy bar. Expertise-heavy axes belong in the judge-assisted or human-only lane, and the axis-triage classifier (design Q2) should treat 'requires specialized professional knowledge' as an explicit routing feature.",
   "source": {
    "raw": "Limitations of the LLM-as-a-Judge Approach for Evaluating LLM Outputs in Expert Knowledge Tasks | https://arxiv.org/abs/2410.20266 | 2024-10 (IUI 2025) | academic",
    "title": "Limitations of the LLM-as-a-Judge Approach for Evaluating LLM Outputs in Expert Knowledge Tasks",
    "url": "https://arxiv.org/abs/2410.20266",
    "date": "2024-10 (IUI 2025)",
    "type": "academic"
   },
   "venue": {
    "peer_reviewed": true,
    "marker": "IUI"
   },
   "same_source_claims": [],
   "in_corpus_verification": null,
   "external_verification": {
    "verdict": "confirmed",
    "checked": [
     "68% SME-LLM agreement in dietetics (registered dietitians)",
     "64% SME-LLM agreement in mental health (clinical psychologists)",
     "overall pairwise preference comparison; agreement varied across domain-specific aspect questions",
     "two fields: dietetics with registered dietitian experts, mental health with clinical psychologist experts",
     "conclusion: keep human experts in the evaluation loop; LLMs alone lack depth for complex knowledge-specific tasks"
    ],
    "not_visible": [],
    "quote": "SMEs agreed with LLM judges 68% of the time in the dietetics domain and 64% in mental health",
    "notes": "Abstract page fully supports the claim's core assertion and both headline percentages. Title and Oct-2024 date match; IUI 2025 venue appears in the Comments field.",
    "checked_at": "2026-07-15",
    "retrieval": "abstract"
   },
   "source_retrieval_meta": {
    "retrieval": "abstract",
    "fetched": [
     "https://arxiv.org/abs/2410.20266"
    ],
    "resolved_title": "Limitations of the LLM-as-a-Judge Approach for Evaluating LLM Outputs in Expert Knowledge Tasks",
    "resolved_date": "2024-10-26 (submitted; Comments cite ACM IUI 2025)",
    "title_match": true
   },
   "used_on": [
    {
     "page": "research.html",
     "anchor": "research-cards",
     "label": "Research - findings by decision weight"
    },
    {
     "page": "pilot.html",
     "anchor": "main",
     "label": "Pilot - gates and objectives"
    }
   ],
   "decision_use": {
    "grade": "C",
    "basis": "superseded-gen capability",
    "measured_on": "model-outputs",
    "note": "2024 study (IUI 2025). The 64-68% expert-agreement figures predate reasoning-model maturity. Mechanism (expertise-heavy axes are the weak lane) is plausible but the deployment-tier boundary is E1/E3's to measure - in either direction.",
    "models_measured": []
   }
  },
  {
   "id": "AJ-12",
   "domain": "academic-judges",
   "area": null,
   "claim": "LLM judges are measurably less reliable on long-form outputs: LongJudgeBench (June 2026) finds current judges unstable across real-world long-form scenarios requiring document-level assessment of organization, coverage, and cross-section consistency, and rubrics or references help but are 'not always sufficient'.",
   "load_bearing": false,
   "evidence": "Chen et al., arXiv 2606.01629 (abstract opened and confirmed). First meta-evaluation benchmark targeting long-form judging specifically; systematically evaluates judges across base models and judging protocols; code public. Complements the short-form-centric older benchmarks (MT-Bench, LLMBar, JudgeBench).",
   "implication": "Attempter writeups/critiques are long-form documents; holistic single-pass judging of them inherits this instability. Supports the decomposition-first architecture (per-claim verification then aggregation) over holistic item scoring, and flags 'global coherence of the writeup' as a check that per-claim decomposition alone will miss (checklist myopia risk from design Q3 is real in both directions).",
   "source": {
    "raw": "Benchmarking LLM-as-a-Judge for Long-Form Output Evaluation (LongJudgeBench) | https://arxiv.org/abs/2606.01629 | 2026-06-01 | academic",
    "title": "Benchmarking LLM-as-a-Judge for Long-Form Output Evaluation (LongJudgeBench)",
    "url": "https://arxiv.org/abs/2606.01629",
    "date": "2026-06-01",
    "type": "academic"
   },
   "venue": {
    "peer_reviewed": false,
    "marker": "preprint"
   },
   "same_source_claims": [],
   "in_corpus_verification": null,
   "external_verification": {
    "verdict": "confirmed",
    "checked": [
     "judges unstable across long-form scenarios: 'a substantial reliability gap: current LLM judges remain unstable across scenarios'",
     "document-level assessment of 'overall organization, task-relevant coverage and depth, cross-section consistency'",
     "rubrics/references 'are helpful but not always sufficient'",
     "code public at github.com/cjj826/LongJudgeBench",
     "authors Chen et al.; submitted June 1 2026 (matches source_date 2026-06-01)"
    ],
    "not_visible": [
     "evidence-field characterization 'First meta-evaluation benchmark targeting long-form judging specifically' -- abstract calls it 'a comprehensive benchmark' and does not explicitly claim to be first"
    ],
    "quote": "current LLM judges remain unstable across scenarios ... more complex document-level assessments of overall organization, task-relevant coverage and depth, cross-section consistency ... helpful but not always sufficient",
    "notes": "Claim text fully supported by the abstract. The only unsupported specific is the 'first benchmark' framing, which appears in the evidence field, not the claim; abstract says 'comprehensive benchmark' without a primacy claim.",
    "checked_at": "2026-07-15",
    "retrieval": "abstract"
   },
   "source_retrieval_meta": {
    "retrieval": "abstract",
    "fetched": [
     "https://arxiv.org/abs/2606.01629"
    ],
    "resolved_title": "Benchmarking LLM-as-a-Judge for Long-Form Output Evaluation",
    "resolved_date": "2026-06-01 (v1); 2026-06-02 (v2)",
    "title_match": true
   },
   "used_on": [
    {
     "page": "research.html",
     "anchor": "research-cards",
     "label": "Research - findings by decision weight"
    },
    {
     "page": "architecture.html",
     "anchor": "ontology",
     "label": "System - claim types"
    },
    {
     "page": "pilot.html",
     "anchor": "main",
     "label": "Pilot - gates and objectives"
    }
   ],
   "decision_use": {
    "grade": "B",
    "basis": "current-gen capability",
    "measured_on": "model-outputs",
    "note": "June-2026 benchmark. Long-form instability at the current tier; deployment-tier behavior untested.",
    "models_measured": []
   }
  },
  {
   "id": "AQ-01",
   "domain": "annotation-quality",
   "area": null,
   "claim": "Roughly one-third of crowdworkers used LLMs on an LLM-advantaged text-production task (33-35% on Prolific, July 2023; ~34% Prolific self-report in a 2025 follow-up by Zhang et al.), and the best tested mitigations (direct request + copy-paste disable or image-only presentation) only cut usage roughly in half (27.6% to ~15.9%), never to zero, while also degrading response quality (direct requests reduced keyword retention 6.2%).",
   "load_bearing": true,
   "evidence": "Veselovsky et al. ran two preregistered Prolific studies (n=168 and n=720, medical-abstract summarization) with a calibrated classifier plus self-report and keystroke heuristics; prevalence 33.3-35.4% baseline (95% CIs ~[26%,43%]); a 3x3 factorial of requests x hurdles halved but did not eliminate use. The June 2026 community survey (arXiv 2606.04924) corroborates: 30-40% estimated reliance and 34% Prolific self-report, and notes mitigations 'reduce rather than eliminate' LLM-assisted responses.",
   "implication": "AutoQA must assume a contamination base rate of ~15-35% on any text-heavy attempter task even after friction measures; per-item content detection and 'please don't' policies cannot be the primary defense, and mitigation friction has a measurable quality cost that the QA layer must not misattribute to attempter skill.",
   "source": {
    "raw": "Prevalence and Prevention of Large Language Model Use in Crowd Work (CACM; arXiv 2310.15683) | https://arxiv.org/abs/2310.15683 | 2025-02-18 | academic",
    "title": "Prevalence and Prevention of Large Language Model Use in Crowd Work (CACM; arXiv 2310.15683)",
    "url": "https://arxiv.org/abs/2310.15683",
    "date": "2025-02-18",
    "type": "academic"
   },
   "venue": {
    "peer_reviewed": false,
    "marker": "preprint"
   },
   "same_source_claims": [],
   "in_corpus_verification": {
    "verdict": "partially_confirmed",
    "transcript": "Core claim VERIFIED against the full PDF of arXiv 2310.15683 (Veselovsky, Horta Ribeiro, Cozzolino, Gordon, Rothschild, West): Study 1, n=168 Prolific workers, 3 July 2023, medical-abstract summarization; prevalence 33.3% [25.9,40.1] (classify-and-count), 35.2% [29.8,40.6] (probabilistic), 35.4% [27.8,43.0] (corrected) - matches \"33.3-35.4%, CIs ~[26%,43%]\". Study 2, n=720, 23 July 2023, 3x3 factorial (request: none/indirect/direct x hurdle: none/image/no-Ctrl-C+V); classifier estimate dropped 27.6% -> 15.9% (direct+image) and 15.8% (direct+Ctrl C+V) - matches \"roughly halved, not eliminated\" (paper: \"reduced LLM use by nearly 50%, it could not fully prevent it\"). Direct request decreased keyword retention by 6.2% (p=0.009) - exact match. Classifier was finetuned e5-base-v2, calibrated (Card & Smith), with self-report and high-precision heuristics (completion time + paste artifacts) - matches. The June 2026 community survey arXiv 2606.04924 (Velutharambath et al., 155 researchers, survey run Aug 2025-Mar 2026) does cite \"Veselovsky et al. estimate 30-40%\", cites \"Zhang et al. (2025): 34% of Prolific participants self-report using LLMs for open-ended questions\", and states mitigations \"reduce rather than eliminate LLM-assisted responses\" - so the follow-up attribution to Zhang et al. 2025 is consistent with that survey's citations (note: I verified the citation exists in the survey, not the underlying Zhang et al. paper itself). ISSUES: (1) Date \"2025-02-18\" is wrong for the arXiv link given - arXiv has only v1, dated 24 Oct 2023; the published version is Communications of the ACM 68(3):42-47, March 2025 issue (DOI 10.1145/3685527). (2) \"Preregistered\" is not supported by the paper text - no preregistration is mentioned in the arXiv PDF. (3) Minor nuance: \"never to zero\" - by the classifier measure true, but the high-precision heuristics table actually shows 0.0% in the direct+Ctrl C+V cell; the paper's own framing (\"cannot fully prevent\") still supports the claim's spirit. (4) 44% of surveyed researchers in the 2026 survey observed LLM use in their data - newer corroborating, not superseding, evidence; no later work found that overturns the prevalence estimates. Sources: https://arxiv.org/abs/2310.15683, https://arxiv.org/abs/2606.04924, https://dl.acm.org/doi/10.1145/3685527.",
    "corrected": "Roughly one-third of Prolific crowdworkers used LLMs on an LLM-advantaged text-production task (33.3-35.4% across three estimators, July 2023, n=168; Veselovsky et al., arXiv 2310.15683 v1 posted 24 Oct 2023, published in Communications of the ACM 68(3), March 2025). In a second study (n=720, 3x3 factorial), the best mitigations (direct request + image presentation or + copy-paste disable) roughly halved classifier-estimated usage from 27.6% to 15.9%/15.8% without eliminating it, and directly requesting non-use reduced keyword retention by 6.2% (p=0.009). The studies were not described as preregistered. A June 2026 community survey of 155 researchers (arXiv 2606.04924, Velutharambath et al.) corroborates, citing Veselovsky et al.'s 30-40% estimate and Zhang et al. (2025)'s 34% Prolific self-report, and finding 44% of researchers observed LLM use in their crowdsourced free-text data; it echoes that mitigations reduce rather than eliminate LLM-assisted responses."
   },
   "external_verification": {
    "verdict": "partially_confirmed",
    "checked": [
     "Study #1: n=168 workers on Prolific, 3 July 2023, summarizing abstracts",
     "Prevalence estimates 33.3% [25.9%,40.1%], 35.2% [29.8%,40.6%], 35.4% [27.8%,43.0%]",
     "Study #2: n=720 users, 23 July 2023, 3x3 factorial (request x hurdle)",
     "Table 1a baseline 27.6% dropping to 15.9% (direct request + image, preventing copy-pasting)",
     "'almost halved, dropping from 27.6% to 15.9%'",
     "abstract: mitigations 'significantly reduce, but not eliminate' LLM use",
     "directly requesting workers not to use LLMs decreased keyword retention by 6.2% (p=0.009)",
     "task = summarize medical paper abstracts (NEJM), per Materials 4.1 and Appendix A"
    ],
    "not_visible": [
     "~34% Prolific self-report in a 2025 follow-up by Zhang et al. (external source, not this paper)",
     "June 2026 community survey arXiv 2606.04924 corroboration (external source)",
     "the word 'preregistered'"
    ],
    "quote": "almost halved, dropping from 27.6% to 15.9%",
    "notes": "The Veselovsky-attributable core and every headline number (33-35% baseline, 27.6%->15.9%, 6.2% keyword retention, n=168/n=720, medical abstracts, halve-not-eliminate) are fully confirmed by the body. Note the ABSTRACT rounds prevalence to 'around 30%', but Study #1 body reports 33.3-35.4%, so the claim's '33-35%' is supported by the body. Downgraded to partially_confirmed only because the claim additionally cites two external sources (Zhang et al. 2025; arXiv 2606.04924) not verifiable from this document. Also: source_date 2025-02-18 refers to the CACM publication; the fetched arXiv preprint is dated 2023-10-24.",
    "checked_at": "2026-07-15",
    "retrieval": "pdf"
   },
   "source_retrieval_meta": {
    "retrieval": "pdf",
    "fetched": [
     "https://arxiv.org/abs/2310.15683",
     "https://arxiv.org/pdf/2310.15683"
    ],
    "resolved_title": "Prevalence and prevention of large language model use in crowd work",
    "resolved_date": "arXiv:2310.15683v1, 24 Oct 2023",
    "title_match": true
   },
   "used_on": [
    {
     "page": "decisions.html",
     "anchor": "d3",
     "label": "Decisions - AI-assistance policy"
    }
   ],
   "decision_use": {
    "grade": "A",
    "basis": "behavioral floor",
    "measured_on": "adjacent-domain",
    "note": "Crowdworker LLM-use base rate (2023, replicated 2025). A floor that rises as tools improve; the decision input (assume pervasive assistance) only strengthens.",
    "models_measured": []
   }
  },
  {
   "id": "AQ-02",
   "domain": "annotation-quality",
   "area": null,
   "claim": "An autonomous LLM agent passed 99.8% of standard attention/quality checks (6,000 trials, 20 check types), evaded behavioral detectors (simulated reading times, mouse movement, typo-laden keystrokes), refused AI-revealing 'reverse shibboleth' tasks 97.7% of the time, and cost ~$0.05 per completed survey - versus $1.50 paid - making content-based spam detection economically and technically beaten as of late 2025.",
   "load_bearing": true,
   "evidence": "Westwood (PNAS, Nov 20 2025; 43,800 evaluations across 9 LLMs) shows instruction-following checks, logic puzzles, reCAPTCHA, and rare-event probes all fail; persona calibration even masks superhuman-accuracy tells. Recommended defenses are entirely provenance/economics-based: identity revalidation, per-worker throttling, completion-history transparency, pass/fail history, location/VPN checks - 'there is no magical fix.' The 2026 crowdsourcing survey (arXiv 2606.04924) independently confirms detectors degrade under paraphrase and out-of-domain, and mixed human-AI authorship is harder still.",
   "implication": "Directly answers taxonomy Q8: wholesale-LLM-submission defense must be economics/provenance-based (rate limits, work-history consistency, randomized deep audits, process telemetry), not per-item text classification; any AutoQA spam layer built on stylistic detection will be silently beaten within the cost asymmetry.",
   "source": {
    "raw": "The potential existential threat of large language models to online survey research (PNAS 122(47)) | https://www.pnas.org/doi/10.1073/pnas.2518075122 | 2025-11-20 | academic",
    "title": "The potential existential threat of large language models to online survey research (PNAS 122(47))",
    "url": "https://www.pnas.org/doi/10.1073/pnas.2518075122",
    "date": "2025-11-20",
    "type": "academic"
   },
   "venue": {
    "peer_reviewed": true,
    "marker": "PNAS"
   },
   "same_source_claims": [],
   "in_corpus_verification": {
    "verdict": "confirmed",
    "transcript": "Every quantitative element of the CLAIM verifies against the primary source (Westwood, \"The potential existential threat of large language models to online survey research,\" PNAS 122(47), published Nov 20, 2025; PubMed 41264250, cited 81x by Jul 2026; won a 2025 Cozzarelli Prize). Extracted paper text confirms verbatim: (1) 99.8% pass rate across 6,000 trials (20 ACQ types x 300 repetitions), only 10 errors, perfect on 18/20 check types; (2) execution layer \"simulates realistic reading times calibrated to the persona's education level, generates human-like mouse movements, and types open-ended responses keystroke-by-keystroke, complete with plausible typos and corrections\"; (3) strategically declines 97.7% of \"reverse shibboleth\" tasks; (4) ~$0.05 per survey with commercial models vs $1.50 payment (>96.8% margin), marginal cost near zero with local open-weight models; (5) 43,800 total evaluations, 139 questions, 6,700 trials, with a key subset validated across nine LLMs (DeepSeek R1, Mistral Large, Claude 3.7 Sonnet, Grok 3, Gemini 2.5, Qwen Plus, Llama 4 Maverick, GPT-4.1, o4-mini); (6) 0% error rate on impossible-event/rare-event probes. The paper's abstract and press framing support the interpretive conclusion that coherence-/content-based detection is no longer tenable and defenses must be provenance/economic. Supersession check: no 2026 work refutes it - follow-ups (Westwood & Frederick 2026 Prolific audit; a 2026 SAGE AMPPS review of AI-mediated contamination) build on it and treat detection as an unsolved provenance problem. Two minor caveats, both in the EVIDENCE SUMMARY rather than the claim: (a) arXiv 2606.04924 (\"Can Crowdsourcing Survive the LLM Era?\", Jun 2026) is a community survey of 155 researchers' practices/experiences, not a technical detector benchmark; its abstract does not state the paraphrase/out-of-domain/mixed-authorship degradation findings attributed to it (those likely come from a different detection-benchmark paper). (b) On reCAPTCHA the paper says the system \"is designed to accommodate tools for bypassing\" reCAPTCHA - an architectural capability, slightly weaker than \"reCAPTCHA... fail[s]\" as an empirically demonstrated result. Neither caveat affects the claim text itself, which is fully accurate.",
    "corrected": null
   },
   "external_verification": {
    "verdict": "not_retrievable",
    "checked": [],
    "not_visible": [
     "99.8% pass rate on attention/quality checks",
     "6,000 trials / 20 check types",
     "43,800 evaluations across 9 LLMs",
     "97.7% reverse-shibboleth refusal",
     "$0.05 per survey vs $1.50 paid",
     "'there is no magical fix' and provenance-based defenses"
    ],
    "quote": "",
    "notes": "Page is protected by a Cloudflare JS challenge ('Just a moment...'). WebFetch returned HTTP 403 Forbidden; curl -sL with two different browser user-agents both returned the Cloudflare interstitial (HTTP 403, no redirect). Constraints forbid substitute sources or other pages on pnas.org, so none of the claimed figures could be verified.",
    "checked_at": "2026-07-15",
    "retrieval": "failed"
   },
   "source_retrieval_meta": {
    "retrieval": "failed",
    "fetched": [
     "https://www.pnas.org/doi/10.1073/pnas.2518075122"
    ],
    "resolved_title": null,
    "resolved_date": null,
    "title_match": false
   },
   "used_on": [
    {
     "page": "research.html",
     "anchor": "research-cards",
     "label": "Research - findings by decision weight"
    },
    {
     "page": "pilot.html",
     "anchor": "main",
     "label": "Pilot - gates and objectives"
    },
    {
     "page": "decisions.html",
     "anchor": "d3",
     "label": "Decisions - AI-assistance policy"
    },
    {
     "page": "decisions.html",
     "anchor": "invariants",
     "label": "Decisions - settled constraints"
    }
   ],
   "decision_use": {
    "grade": "A",
    "basis": "adversarial floor",
    "measured_on": "adjacent-domain",
    "note": "Survey-agent result. Attack capability is a lower bound that strengthens with every model generation; the economic asymmetry ($0.05 vs $1.50) widens. Content detection stays beaten.",
    "models_measured": []
   }
  },
  {
   "id": "AQ-03",
   "domain": "annotation-quality",
   "area": null,
   "claim": "Correlation-based meta-evaluation of automatic judges against human labels is systematically distorted by human label uncertainty: on high-disagreement items a machine judge can superficially match or beat human-human correlation, while on high-agreement strata machine-human correlation drops well below the human-human baseline - so a single aggregate 'agreement with humans' number is invalid as a validation target.",
   "load_bearing": true,
   "evidence": "Elangovan et al. (Amazon), ICLR 2025 (arXiv 2410.03775, opened and verified): Krippendorff's alpha / kappa were designed for human-human reliability and their assumptions break for machine labels; the paper's fix is (1) stratify meta-evaluation by human-label uncertainty, (2) use binned Jensen-Shannon divergence for perception-type (inherently subjective) tasks, (3) report 'perception charts' instead of one correlation. Code open-sourced (amazon-science/BeyondCorrelation).",
   "implication": "Answers taxonomy Q1 directly: meta-evaluation must stratify gold/validation items by measured human agreement level and report per-stratum judge performance; 'accuracy vs. humans' as a single number should be banned from reporting, and judge validity claims are only meaningful on the high-human-agreement stratum.",
   "source": {
    "raw": "Beyond correlation: The impact of human uncertainty in measuring the effectiveness of automatic evaluation and LLM-as-a-judge (ICLR 2025) | https://arxiv.org/abs/2410.03775 | 2025-05-01 | academic",
    "title": "Beyond correlation: The impact of human uncertainty in measuring the effectiveness of automatic evaluation and LLM-as-a-judge (ICLR 2025)",
    "url": "https://arxiv.org/abs/2410.03775",
    "date": "2025-05-01",
    "type": "academic"
   },
   "venue": {
    "peer_reviewed": true,
    "marker": "ICLR"
   },
   "same_source_claims": [],
   "in_corpus_verification": {
    "verdict": "partially_confirmed",
    "transcript": "Substance CONFIRMED on every checked point against the arXiv record (https://arxiv.org/abs/2410.03775): (1) abstract states machine labels \"may superficially appear to have similar or better correlation with the human majority\" when human uncertainty is high, and machine-human correlation falls \"well below\" human-human correlation as human label consistency increases; (2) it flags Krippendorff's alpha / Randolph's kappa as designed for human-human reliability with assumptions inapplicable to machine labels; (3) proposed fixes match exactly - stratification by human label uncertainty, binned Jensen-Shannon divergence for perception-type tasks, and 'perception charts'; (4) venue is ICLR 2025 per the arXiv comments field; (5) code is open-sourced at github.com/amazon-science/BeyondCorrelation. The only inaccuracy is the SOURCE date: the claim says 2025-05-01, but the latest arXiv version is v3 dated 2025-01-27 (original submission 2024-10-03) - no May 2025 version exists on the record. Supersession check: a follow-up search surfaced related later work - \"Validating LLM-as-a-Judge Systems under Rating Indeterminacy\" (arXiv 2503.05965), \"Hidden Measurement Error in LLM Pipelines Distorts Annotation, Evaluation, and Benchmarking\" (arXiv 2604.11581), and \"CARE: Confounder-Aware Aggregation for Reliable LLM Evaluation\" (arXiv 2603.00039, Feb 2026). These extend and reinforce the human-uncertainty/measurement-error critique rather than refute it; nothing found supersedes or contradicts the core finding. Verdict is partially_confirmed solely because of the incorrect source date; the claim itself stands.",
    "corrected": "Correlation-based meta-evaluation of automatic judges against human labels is systematically distorted by human label uncertainty: on high-disagreement items a machine judge can superficially match or beat human-human correlation, while on high-agreement strata machine-human correlation drops well below the human-human baseline - so a single aggregate 'agreement with humans' number is invalid as a validation target. Source: Elangovan et al., \"Beyond correlation...\" (Amazon Science), accepted at ICLR 2025, arXiv 2410.03775 (v1 2024-10-03, latest v3 2025-01-27), code at github.com/amazon-science/BeyondCorrelation."
   },
   "external_verification": {
    "verdict": "confirmed",
    "checked": [
     "a single aggregate correlation score can obscure fundamental differences between human labels and those from automatic evaluation",
     "high uncertain-label proportion: machine labels may superficially appear to have similar or better correlation with the human majority label (an illusion)",
     "as consistent-label samples increase, correlation between machine and human labels fall well below HH correlation",
     "propose stratifying data by human label uncertainty",
     "binned Jensen-Shannon Divergence for perception",
     "perception charts",
     "code open-sourced at github.com/amazon-science/BeyondCorrelation",
     "Accepted at ICLR 2025"
    ],
    "not_visible": [
     "explicit 'Amazon' author affiliation -- page lists no affiliations; Amazon is only implied by the 'amazon-science' GitHub org name"
    ],
    "quote": "we propose ... a new metric - binned Jensen-Shannon Divergence for perception ... perception charts, to contextualize correlation measures appropriately",
    "notes": "Core claim and all three prescribed fixes (stratify by uncertainty, binned JS divergence, perception charts) confirmed verbatim. Only the 'Amazon' affiliation label is not explicitly on the page.",
    "checked_at": "2026-07-15",
    "retrieval": "abstract"
   },
   "source_retrieval_meta": {
    "retrieval": "abstract",
    "fetched": [
     "https://arxiv.org/abs/2410.03775"
    ],
    "resolved_title": "Beyond correlation: The Impact of Human Uncertainty in Measuring the Effectiveness of Automatic Evaluation and LLM-as-a-Judge",
    "resolved_date": "2024-10-03 (v1); v3 2025-01-27; Accepted at ICLR 2025",
    "title_match": true
   },
   "used_on": [
    {
     "page": "decisions.html",
     "anchor": "a1",
     "label": "Decisions - the standard"
    }
   ],
   "decision_use": {
    "grade": "A",
    "basis": "validation-design result",
    "measured_on": "structural",
    "note": "How to validate judges under human-label uncertainty - statistics of the gold set, not model capability.",
    "models_measured": []
   }
  },
  {
   "id": "AQ-04",
   "domain": "annotation-quality",
   "area": null,
   "claim": "In human preference/rating data, the majority of annotator disagreements are attributable to task underspecification and response-style preferences - not random noise and not primarily expertise gaps - and standard aggregation (Bradley-Terry reward modeling) plus LLM-as-judge evaluation both fail to account for this divergence.",
   "load_bearing": true,
   "evidence": "Zhang et al., ICML 2025 (arXiv 2410.14632, opened and verified): built a 10-category, 4-class taxonomy of disagreement sources on preference datasets; abstract states 'the majority of disagreements are due to factors such as task underspecification or response style,' explicitly challenging the noise assumption; they show LLM evaluations are heavily swayed by divisive style features and develop divergence-identification methods. Converges with Jiang & de Marneffe (TACL 2022) who found the same structure in NLI.",
   "implication": "Supports root-cause-1 primacy and taxonomy Q2: the verdict ontology needs a first-class 'instructions underdetermine this case' class detected via measured divergence, and a cheap variance-attribution pilot is empirically justified because published decompositions find underspecification (fixable by rubric compilation) dominates over noise (fixable by averaging).",
   "source": {
    "raw": "Diverging Preferences: When do Annotators Disagree and do Models Know? (ICML 2025) | https://arxiv.org/abs/2410.14632 | 2025-07-15 | academic",
    "title": "Diverging Preferences: When do Annotators Disagree and do Models Know? (ICML 2025)",
    "url": "https://arxiv.org/abs/2410.14632",
    "date": "2025-07-15",
    "type": "academic"
   },
   "venue": {
    "peer_reviewed": true,
    "marker": "ACL"
   },
   "same_source_claims": [],
   "in_corpus_verification": null,
   "external_verification": {
    "verdict": "confirmed",
    "checked": [
     "majority of disagreements due to task underspecification or response style (not simple noise)",
     "standard reward modeling (Bradley-Terry) and LLM-as-Judge fail to account for divergence",
     "develops methods for identifying diverging preferences",
     "ten-category taxonomy across four high-level classes",
     "ICML 2025 venue note"
    ],
    "not_visible": [
     "explicit statement that expertise gaps are NOT a main cause (excluded only by implication, since the majority is attributed to underspecification/style)",
     "Jiang & de Marneffe TACL 2022 corroboration (a separate citation, not part of this source)"
    ],
    "quote": "the majority of disagreements are due to factors such as task underspecification or response style ... standard reward modeling (e.g., Bradley-Terry) and LLM-as-Judge evaluation methods fail to account for divergence between annotators.",
    "notes": "Core assertion and all listed mechanisms confirmed from the abstract. The 'not primarily expertise gaps' element is supported by implication rather than an explicit exclusionary sentence.",
    "checked_at": "2026-07-15",
    "retrieval": "abstract"
   },
   "source_retrieval_meta": {
    "retrieval": "abstract",
    "fetched": [
     "https://arxiv.org/abs/2410.14632"
    ],
    "resolved_title": "Diverging Preferences: When do Annotators Disagree and do Models Know?",
    "resolved_date": "2024-10-18 (v1); v3 2026-03-03; venue ICML 2025",
    "title_match": true
   },
   "used_on": [
    {
     "page": "decisions.html",
     "anchor": "a1",
     "label": "Decisions - the standard"
    }
   ],
   "decision_use": {
    "grade": "A",
    "basis": "human science",
    "measured_on": "human-work",
    "note": "Attribution of annotator disagreement to underspecification/style. Human raters have not changed.",
    "models_measured": []
   }
  },
  {
   "id": "AQ-05",
   "domain": "annotation-quality",
   "area": null,
   "claim": "Even on a nominally objective task (NLI), annotator disagreement decomposes into a 10-category taxonomy across three high-level sources - uncertainty in sentence meaning, underspecification in guidelines, and annotator behavior - proving that a fraction of inter-reviewer variance is item-intrinsic and survives perfect rubric operationalization.",
   "load_bearing": false,
   "evidence": "Jiang & de Marneffe, TACL 2022, developed and hand-applied the taxonomy to high-disagreement NLI items; disagreement persisted among trained annotators on items with genuine interpretive ambiguity. D3CODE (EMNLP 2024; 4.5K sentences, 4K+ annotators, 21 countries) adds that residual variance is systematically structured by annotators' moral values and region, not random.",
   "implication": "The target metric cannot be raw agreement: some disagreement is item-intrinsic signal. The variance-attribution pilot should code disagreements against a taxonomy like this to split rubric-fixable variance (underspecification) from irreducible variance (ambiguity/values) before setting build order.",
   "source": {
    "raw": "Investigating Reasons for Disagreement in Natural Language Inference (TACL 10) | https://aclanthology.org/2022.tacl-1.78/ | 2022-12-01 | academic",
    "title": "Investigating Reasons for Disagreement in Natural Language Inference (TACL 10)",
    "url": "https://aclanthology.org/2022.tacl-1.78/",
    "date": "2022-12-01",
    "type": "academic"
   },
   "venue": {
    "peer_reviewed": true,
    "marker": "ACL"
   },
   "same_source_claims": [],
   "in_corpus_verification": null,
   "external_verification": {
    "verdict": "partially_confirmed",
    "checked": [
     "taxonomy of disagreement sources with 10 categories spanning 3 high-level classes",
     "one class is uncertainty in the sentence meaning",
     "task is NLI (natural language inference) annotation disagreement"
    ],
    "not_visible": [
     "'underspecification in guidelines' as a named high-level class (abstract instead names 'annotator biases and task artifacts')",
     "'survives perfect rubric operationalization' framing",
     "disagreement persisting specifically among trained annotators",
     "D3CODE (EMNLP 2024; 4.5K sentences, 4K+ annotators, 21 countries) - a separate source not on this page"
    ],
    "quote": "We developed a taxonomy of disagreement sources with 10 categories spanning 3 high-level classes. We found that some disagreements are due to uncertainty in the sentence meaning, others to annotator biases and task artifacts",
    "notes": "Core (10-category / 3-class taxonomy of intrinsic NLI disagreement) confirmed verbatim. Two of the claim's three named high-level sources map to the abstract ('uncertainty in sentence meaning'; 'annotator behavior' ~ 'annotator biases and task artifacts'); the third, 'underspecification in guidelines', is not named in the abstract. The D3CODE material cited in the evidence is a different paper, unverifiable from this URL. Only the abstract page is accessible (PDF/other same-site pages disallowed).",
    "checked_at": "2026-07-15",
    "retrieval": "abstract"
   },
   "source_retrieval_meta": {
    "retrieval": "abstract",
    "fetched": [
     "https://aclanthology.org/2022.tacl-1.78/"
    ],
    "resolved_title": "Investigating Reasons for Disagreement in Natural Language Inference",
    "resolved_date": "TACL Volume 10 (2022), pp. 1357-1374",
    "title_match": true
   },
   "used_on": [],
   "decision_use": {
    "grade": "A",
    "basis": "human science",
    "measured_on": "human-work",
    "note": "2022 NLI disagreement taxonomy. Human disagreement structure, not model capability.",
    "models_measured": []
   }
  },
  {
   "id": "AQ-06",
   "domain": "annotation-quality",
   "area": null,
   "claim": "Chance-corrected agreement coefficients (Cohen's kappa, and analogously Krippendorff's alpha) can be near zero despite very high raw agreement when class prevalence is extreme (the Feinstein-Cicchetti paradoxes), so under the 90%+ pass rates typical of QA pipelines, single-coefficient reliability reporting is misleading in the opposite direction from raw percent agreement.",
   "load_bearing": false,
   "evidence": "Feinstein & Cicchetti (J Clin Epidemiol, 1990) formalized the two kappa paradoxes (high agreement/low kappa under imbalance; asymmetric marginals inflating kappa); the effect is prevalence-driven, extensively replicated (Quarfoot & Levine, Am. Statistician 2016; Gwet's AC1 literature). This is exactly the regime of QA pass/fail data with low failure base rates.",
   "implication": "Answers taxonomy Q1 stats question: standardize on reporting a chance-corrected coefficient PLUS per-class recall at fixed prevalence PLUS CIs; ban both raw percent-agreement AND any single coefficient in isolation, because at extreme pass rates alpha/kappa collapse even for a well-functioning judge.",
   "source": {
    "raw": "High agreement but low kappa: I. The problems of two paradoxes (Feinstein & Cicchetti) | https://doi.org/10.1016/0895-4356(90)90158-L | 1990-01-01 | academic",
    "title": "High agreement but low kappa: I. The problems of two paradoxes (Feinstein & Cicchetti)",
    "url": "https://doi.org/10.1016/0895-4356(90)90158-L",
    "date": "1990-01-01",
    "type": "academic"
   },
   "venue": {
    "peer_reviewed": null,
    "marker": "academic"
   },
   "same_source_claims": [],
   "in_corpus_verification": null,
   "external_verification": {
    "verdict": "partially_confirmed",
    "checked": [
     "resolved article title confirms the core thesis: high agreement, low kappa, two paradoxes"
    ],
    "not_visible": [
     "abstract-level mechanism: kappa near zero under extreme prevalence (prevalence-driven)",
     "second paradox: asymmetric marginal distributions inflating/affecting kappa",
     "the 'opposite direction from raw percent agreement' framing tied to QA pass rates",
     "external replications cited in the claim's evidence (Quarfoot & Levine 2016; Gwet AC1 literature) - not this source"
    ],
    "quote": "articleName : 'High agreement but low Kappa: I. the problems of two paradoxes'",
    "notes": "Server redirect chain doi.org -> linkinghub.elsevier.com -> jclinepi.com. Abstract body is behind a Cloudflare JS challenge (curl) and the Elsevier landing page returns only 'Redirecting'; a full WebFetch of the DOI/landing was flagged by a stochastic model safeguard. Title confirmed via the linkinghub siteCatalyst metadata, which locks the core high-agreement/low-kappa/two-paradoxes assertion; the specific paradox mechanisms could not be retrieved from the abstract text.",
    "checked_at": "2026-07-15",
    "retrieval": "landing"
   },
   "source_retrieval_meta": {
    "retrieval": "landing",
    "fetched": [
     "https://doi.org/10.1016/0895-4356(90)90158-L",
     "https://linkinghub.elsevier.com/retrieve/pii/089543569090158L",
     "http://jclinepi.com/retrieve/pii/089543569090158L"
    ],
    "resolved_title": "High agreement but low Kappa: I. the problems of two paradoxes",
    "resolved_date": "1990 (J Clin Epidemiol; exact date not shown)",
    "title_match": true
   },
   "used_on": [
    {
     "page": "decisions.html",
     "anchor": "c4",
     "label": "Decisions - default: statistics staging"
    },
    {
     "page": "decisions.html",
     "anchor": "invariants",
     "label": "Decisions - settled constraints"
    }
   ],
   "decision_use": {
    "grade": "A",
    "basis": "statistical fact",
    "measured_on": "structural",
    "note": "Kappa paradox under extreme base rates. Mathematics.",
    "models_measured": []
   }
  },
  {
   "id": "AQ-07",
   "domain": "annotation-quality",
   "area": null,
   "claim": "Classical label-aggregation models (Dawid-Skene 1979, MACE 2013 with its explicit spammer latent variable, GLAD) all assume annotators are conditionally independent given the true label - an assumption measurably violated when the 'annotators' are LLM judges sharing data/architectures/prompts, causing miscalibrated posteriors and confidently wrong aggregate verdicts; Feb 2026 work replaces this with dependence-aware Ising-model aggregation.",
   "load_bearing": false,
   "evidence": "Balasubramanian, Podkopaev & Kasiviswanathan (Amazon/UC Davis, arXiv 2601.22336, Feb 2 2026) state the CI assumption 'is often violated by LLM judges due to shared data, architectures, prompts, and failure modes' and that ignoring dependencies 'can yield miscalibrated posteriors and even confidently incorrect predictions'; they build aggregation via Ising models. MACE (Hovy et al., NAACL 2013) remains the canonical spam-aware human aggregator, with production implementations in Toloka's crowd-kit.",
   "implication": "If AutoQA uses k-sample voting or multi-judge ensembles, naive Dawid-Skene/majority aggregation will overstate confidence because same-family judge samples share failure modes; confidence estimates must model correlated errors (or use genuinely different model families), and MACE-style competence/spam latent variables remain the right tool for the human-reviewer side.",
   "source": {
    "raw": "Dependence-Aware Label Aggregation for LLM-as-a-Judge via Ising Models (arXiv 2601.22336) | https://arxiv.org/pdf/2601.22336 | 2026-02-02 | academic",
    "title": "Dependence-Aware Label Aggregation for LLM-as-a-Judge via Ising Models (arXiv 2601.22336)",
    "url": "https://arxiv.org/pdf/2601.22336",
    "date": "2026-02-02",
    "type": "academic"
   },
   "venue": {
    "peer_reviewed": true,
    "marker": "ACL"
   },
   "same_source_claims": [],
   "in_corpus_verification": null,
   "external_verification": {
    "verdict": "partially_confirmed",
    "checked": [
     "classical methods assume annotators are conditionally independent",
     "assumption often violated by LLM judges due to shared data, architectures, prompts, and failure modes",
     "ignoring dependencies yields miscalibrated posteriors and even confidently incorrect predictions",
     "dependence-aware aggregation via Ising graphical models",
     "Dawid-Skene named"
    ],
    "not_visible": [
     "MACE (2013)",
     "'spammer' latent variable",
     "GLAD",
     "Dawid-Skene '1979' date",
     "Toloka crowd-kit production implementations"
    ],
    "quote": "Most classical methods, e.g., Dawid-Skene or (weighted) majority voting, assume annotators are conditionally independent ... Ignoring such dependencies can yield miscalibrated posteriors and even confidently incorrect predictions.",
    "notes": "Both direct quotes the claim attributes to the paper are confirmed verbatim, and the Ising-model core is confirmed. Retrieved abstract lists CI examples as 'Dawid-Skene or (weighted) majority voting' only; MACE, GLAD, the spammer latent variable, and the crowd-kit implementation note are the claim author's additions and are not visible in retrieved text. Submitted 2026-01-29 (source_date 2026-02-02).",
    "checked_at": "2026-07-15",
    "retrieval": "abstract"
   },
   "source_retrieval_meta": {
    "retrieval": "abstract",
    "fetched": [
     "https://arxiv.org/abs/2601.22336"
    ],
    "resolved_title": "Dependence-Aware Label Aggregation for LLM-as-a-Judge via Ising Models",
    "resolved_date": "2026-01-29",
    "title_match": true
   },
   "used_on": [],
   "decision_use": {
    "grade": "A",
    "basis": "statistical method",
    "measured_on": "structural",
    "note": "Dawid-Skene/MACE-class aggregation. Method, not capability.",
    "models_measured": []
   }
  },
  {
   "id": "AQ-08",
   "domain": "annotation-quality",
   "area": null,
   "claim": "LLM contamination in low-dimensional annotation data (multiple-choice labels, ratings) can be evaluated at the worker level without any ground truth, using peer-prediction scores that condition on LLM-generated reference labels - a training-free mechanism with theoretical guarantees under an LLM-collusion model, published at NeurIPS 2025 - because text-based detectors are inapplicable to short label outputs.",
   "load_bearing": false,
   "evidence": "Zhang, Pang, Zhu & Liu (NeurIPS 2025; arXiv 2506.06991, opened and verified): existing detectors need high-dimensional text; their mechanism scores the informativeness of worker answers via inter-worker correlation conditioned on a subset of LLM labels, explicitly modeling collusion (multiple workers pasting from the same LLM), and empirically detects low-effort cheating on real crowdsourcing datasets. Scoring is per-worker, not per-item.",
   "implication": "For rating/label-type attempter outputs where stylistic detection is impossible, the viable contamination sensor is worker-level statistical scoring (cross-worker correlation vs. an LLM-conditioned baseline) run continuously - reinforcing that provenance/behavioral defense operates at the annotator level, not the item level.",
   "source": {
    "raw": "Evaluating LLM-Contaminated Crowdsourcing Data Without Ground Truth (NeurIPS 2025) | https://arxiv.org/abs/2506.06991 | 2025-12-01 | academic",
    "title": "Evaluating LLM-Contaminated Crowdsourcing Data Without Ground Truth (NeurIPS 2025)",
    "url": "https://arxiv.org/abs/2506.06991",
    "date": "2025-12-01",
    "type": "academic"
   },
   "venue": {
    "peer_reviewed": true,
    "marker": "NeurIPS"
   },
   "same_source_claims": [],
   "in_corpus_verification": null,
   "external_verification": {
    "verdict": "partially_confirmed",
    "checked": [
     "evaluates contaminated data without ground truth via peer prediction",
     "quantifies correlations between worker answers conditioning on (a subset of) LLM-generated labels",
     "training-free scoring mechanism with theoretical guarantees under a crowdsourcing model that accounts for LLM collusion",
     "existing text-based detectors rely on high-dimensional text, unsuitable for annotation tasks",
     "multiple-choice labeling named as target task",
     "empirically robust in detecting low-effort cheating on real-world crowdsourcing datasets",
     "authors Yichi Zhang, Jinlong Pang, Zhaowei Zhu, Yang Liu (matches Zhang, Pang, Zhu & Liu)"
    ],
    "not_visible": [
     "NeurIPS 2025 venue (no venue stated anywhere on the abstract page)",
     "'ratings' as a target task type (only multiple-choice explicitly named)",
     "explicit 'per-worker not per-item' framing (implied by inter-worker correlation, not stated)",
     "explicit 'multiple workers pasting from the same LLM' scenario",
     "source_date 2025-12-01 (page shows Jun/Nov 2025 only)"
    ],
    "quote": "a training-free scoring mechanism with theoretical guarantees under a crowdsourcing model that accounts for LLM collusion",
    "notes": "Core method claims fully match the abstract. The asserted NeurIPS 2025 publication and the 2025-12-01 date are not visible on the page; abstract dates are v1 2025-06-08, v2 2025-11-06.",
    "checked_at": "2026-07-15",
    "retrieval": "abstract"
   },
   "source_retrieval_meta": {
    "retrieval": "abstract",
    "fetched": [
     "https://arxiv.org/abs/2506.06991"
    ],
    "resolved_title": "Evaluating LLM-Contaminated Crowdsourcing Data Without Ground Truth",
    "resolved_date": "arXiv v1 2025-06-08; v2 2025-11-06",
    "title_match": true
   },
   "used_on": [
    {
     "page": "decisions.html",
     "anchor": "d3",
     "label": "Decisions - AI-assistance policy"
    }
   ],
   "decision_use": {
    "grade": "A",
    "basis": "statistical method",
    "measured_on": "structural",
    "note": "Worker-level contamination auditing without per-item detection. Method survives model improvement on both sides.",
    "models_measured": []
   }
  },
  {
   "id": "AQ-09",
   "domain": "annotation-quality",
   "area": null,
   "claim": "The NLP field has institutionalized disagreement-as-signal: the third Learning-With-Disagreements shared task (LeWiDi-2025, EMNLP 2025) standardizes dual evaluation - soft-label (predict the population label distribution, Wasserstein/Manhattan distance) and perspectivist (recover individual annotators' labels) - across four datasets that all ship per-annotator labels rather than adjudicated gold.",
   "load_bearing": false,
   "evidence": "Leonardelli et al. (arXiv 2510.08460, Oct 2025; le-wi-di.github.io) describe the two complementary paradigms and public leaderboard; lineage runs from Aroyo & Welty's 'truth is a lie' perspectivism through DICES (NeurIPS 2023 safety dataset with up to 100+ raters/item) and jury learning (Gordon et al., CHI 2022, which composes verdicts from explicitly selected rater subpopulations); Frenda et al.'s survey (LREV 2024) documents the dataset ecosystem.",
   "implication": "The AutoQA data model should preserve rater identity and full label distributions (never only adjudicated verdicts), and for judgment-laden axes the judge's target should be the distribution/jury composition, not a forced scalar - this is now the field-standard evaluation contract, giving ready-made metrics (Wasserstein/Manhattan soft-label distance) for 'consistency-of-application' monitoring.",
   "source": {
    "raw": "LeWiDi-2025 at NLPerspectives: Third Edition of the Learning with Disagreements Shared Task | https://arxiv.org/pdf/2510.08460 | 2025-10-09 | academic",
    "title": "LeWiDi-2025 at NLPerspectives: Third Edition of the Learning with Disagreements Shared Task",
    "url": "https://arxiv.org/pdf/2510.08460",
    "date": "2025-10-09",
    "type": "academic"
   },
   "venue": {
    "peer_reviewed": true,
    "marker": "NeurIPS"
   },
   "same_source_claims": [],
   "in_corpus_verification": null,
   "external_verification": {
    "verdict": "partially_confirmed",
    "checked": [
     "Third edition of the Learning With Disagreements (LEWIDI) shared task",
     "NLPerspectives workshop at EMNLP 2025",
     "soft-label approach: models predict population-level distributions of judgments",
     "perspectivist approach: models predict interpretations of individual annotators",
     "four datasets (paraphrase, irony, sarcasm, NLI)",
     "went beyond standard metrics such as cross-entropy with new metrics for both paradigms"
    ],
    "not_visible": [
     "'Wasserstein/Manhattan distance' as the soft-label/perspectivist metrics (not named in abstract)",
     "'public leaderboard' (not in abstract; a Codabench competition link appears in the PDF structure but term not confirmed)",
     "'per-annotator labels rather than adjudicated gold' (implied by perspectivist paradigm, not stated explicitly)",
     "lineage claims: DICES (NeurIPS 2023), jury learning (Gordon et al. CHI 2022), Frenda et al. survey (LREV 2024) - cite other works not fetchable here; Aroyo & Welty citation anchor was present in the PDF"
    ],
    "quote": "the soft-label approach, in which models predict population-level distributions of judgments ... the perspectivist approach, in which models predict the interpretations of individual annotators",
    "notes": "Given PDF was FlateDecode-compressed (no readable abstract); fetched the same-arXiv-ID abs page (allowed variant) for the abstract. Core dual-evaluation design confirmed; specific distance metrics and leaderboard not visible.",
    "checked_at": "2026-07-15",
    "retrieval": "abstract"
   },
   "source_retrieval_meta": {
    "retrieval": "abstract",
    "fetched": [
     "https://arxiv.org/pdf/2510.08460",
     "https://arxiv.org/abs/2510.08460"
    ],
    "resolved_title": "LeWiDi-2025 at NLPerspectives: Third Edition of the Learning with Disagreements Shared Task",
    "resolved_date": "2025-10-09 (v1); v2 2026-02-06",
    "title_match": true
   },
   "used_on": [],
   "decision_use": {
    "grade": "A",
    "basis": "field practice",
    "measured_on": "structural",
    "note": "Disagreement-as-signal institutionalized (LeWiDi-2025). Practice fact.",
    "models_measured": []
   }
  },
  {
   "id": "AQ-10",
   "domain": "annotation-quality",
   "area": null,
   "claim": "Gold/test labels themselves are wrong at material rates - a lower-bound average of 3.3% label errors across 10 canonical benchmarks (at least 6% in the ImageNet validation set), validated by human review of confident-learning-flagged candidates (51% flag precision) - and error rates this size are enough to flip model rankings.",
   "load_bearing": false,
   "evidence": "Northcutt, Athalye & Mueller (NeurIPS 2021 Datasets & Benchmarks; arXiv 2103.14749, opened and verified): algorithmic flagging via confident learning + crowdsourced verification; ResNet-18 beats ResNet-50 on corrected ImageNet if mislabeled prevalence rises just 6%. Cleanlab is the maintained open-source implementation.",
   "implication": "Answers taxonomy Q1's false-agreement question in the only measured form available: several percent of any 'gold' set is wrong, so judge-vs-gold disagreement at that rate is uninformative; the gold set needs its own standing audit channel (confident-learning-style flagging + expert re-adjudication) as a permanent component, and judge certification margins must exceed the gold error floor.",
   "source": {
    "raw": "Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks (NeurIPS 2021 D&B) | https://arxiv.org/abs/2103.14749 | 2021-11-01 | academic",
    "title": "Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks (NeurIPS 2021 D&B)",
    "url": "https://arxiv.org/abs/2103.14749",
    "date": "2021-11-01",
    "type": "academic"
   },
   "venue": {
    "peer_reviewed": true,
    "marker": "NeurIPS"
   },
   "same_source_claims": [],
   "in_corpus_verification": null,
   "external_verification": {
    "verdict": "confirmed",
    "checked": [
     "average of at least 3.3% errors across the 10 datasets",
     "at least 6% label errors in the ImageNet validation set",
     "51% human-validation precision (roughly half of flagged candidates confirmed mislabeled)",
     "10 canonical benchmarks",
     "confident learning algorithms + crowdsourced human validation",
     "ResNet-18 outperforms ResNet-50 on corrected ImageNet if mislabeled prevalence rises by just 6% (rank flip)"
    ],
    "not_visible": [],
    "quote": "estimate an average of at least 3.3% errors across the 10 datasets ... label errors comprise at least 6% of the ImageNet validation set ... ResNet-18 outperforms ResNet-50 ... if the prevalence of originally mislabeled test examples increases by just 6%",
    "notes": "All numbers and the rank-flip finding confirmed from the abstract page. Cleanlab implementation referenced (github.com/cleanlab/label-errors, labelerrors.com).",
    "checked_at": "2026-07-15",
    "retrieval": "abstract"
   },
   "source_retrieval_meta": {
    "retrieval": "abstract",
    "fetched": [
     "https://arxiv.org/abs/2103.14749"
    ],
    "resolved_title": "Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks",
    "resolved_date": "2021-03-26 (v1); 2021-11-07 (v4); NeurIPS 2021 Datasets & Benchmarks",
    "title_match": true
   },
   "used_on": [
    {
     "page": "decisions.html",
     "anchor": "a5",
     "label": "Decisions - permanent audit"
    },
    {
     "page": "decisions.html",
     "anchor": "invariants",
     "label": "Decisions - settled constraints"
    }
   ],
   "decision_use": {
    "grade": "A",
    "basis": "dataset audit",
    "measured_on": "structural",
    "note": ">=3.3% gold-label error floor across canonical test sets. Property of datasets and labeling processes, not of judges.",
    "models_measured": []
   }
  },
  {
   "id": "AQ-11",
   "domain": "annotation-quality",
   "area": null,
   "claim": "Showing annotators LLM suggestions makes them anchor on the suggestions - significantly shifting the label distribution versus unassisted baseline, raising self-reported confidence without making them faster - and evaluating models against LLM-assisted labels significantly inflates reported model performance.",
   "load_bearing": false,
   "evidence": "Schroeder, Roy & Kabbara (Findings of ACL 2025, July 2025; pre-registered, 350 annotators, 7,000 annotations, 4 conditions x 2 models x 2 datasets, opened and verified): annotators 'strongly took the LLM suggestions'; the paper warns homogenization/lessened variation is a direct risk when LLMs are inserted into annotation loops, and that 'human-approved' LLM-assisted labels change downstream conclusions.",
   "implication": "Direct evidence for taxonomy Q8 monoculture and Q9 induced drift: if AutoQA feedback (or any AI verdict) is visible to attempters/reviewers before they judge, distributions converge on judge-pleasing labels and self-evaluation inflates - so human touches must be blind-first, and population-level label-distribution drift is the right monoculture sensor.",
   "source": {
    "raw": "Just Put a Human in the Loop? Investigating LLM-Assisted Annotation for Subjective Tasks (Findings of ACL 2025) | https://aclanthology.org/2025.findings-acl.1323/ | 2025-07-27 | academic",
    "title": "Just Put a Human in the Loop? Investigating LLM-Assisted Annotation for Subjective Tasks (Findings of ACL 2025)",
    "url": "https://aclanthology.org/2025.findings-acl.1323/",
    "date": "2025-07-27",
    "type": "academic"
   },
   "venue": {
    "peer_reviewed": true,
    "marker": "ACL"
   },
   "same_source_claims": [],
   "in_corpus_verification": null,
   "external_verification": {
    "verdict": "confirmed",
    "checked": [
     "annotators anchored on LLM suggestions, significantly changing label distribution vs baseline",
     "did NOT make them faster but improved self-reported confidence",
     "using LLM-assisted labels for evaluation significantly increases reported model performance",
     "pre-registered experiment",
     "350 unique annotators, 7,000 annotations, 4 conditions, 2 models, 2 datasets"
    ],
    "not_visible": [
     "'homogenization/lessened variation is a direct risk' framing (from the evidence narrative) is not present in the abstract; abstract instead says label-distribution shifts can affect conclusions drawn even from human-approved datasets"
    ],
    "quote": "annotators strongly took the LLM suggestions, significantly changing the label distribution compared to the baseline",
    "notes": "Every core assertion and headline number in the claim is confirmed by the abstract. Only the 'homogenization' phrasing (which is in the verifier's evidence rationale, not the claim text) is not visible in the abstract; it may appear in the full paper (not fetched).",
    "checked_at": "2026-07-15",
    "retrieval": "abstract"
   },
   "source_retrieval_meta": {
    "retrieval": "abstract",
    "fetched": [
     "https://aclanthology.org/2025.findings-acl.1323/"
    ],
    "resolved_title": "Just Put a Human in the Loop? Investigating LLM-Assisted Annotation for Subjective Tasks",
    "resolved_date": "July 2025 (Findings of ACL 2025; pp. 25771-25795)",
    "title_match": true
   },
   "used_on": [],
   "decision_use": {
    "grade": "A",
    "basis": "human science",
    "measured_on": "human-work",
    "note": "Anchoring on shown suggestions is human behavior; better suggestions anchor harder. Blind-first design stays mandatory.",
    "models_measured": []
   }
  },
  {
   "id": "AQ-12",
   "domain": "annotation-quality",
   "area": null,
   "claim": "RLVR-style reasoning training degrades an LLM's ability to model human annotator disagreement, while naive chain-of-thought on RLHF models improves it - evaluated across 60 setups on 3 tasks - meaning stronger 'reasoning' judges are measurably worse at knowing when humans legitimately diverge.",
   "load_bearing": false,
   "evidence": "Ni et al., EACL 2026 main (arXiv 2506.19467 v3, Jan 12 2026, opened and verified): systematic evaluation over model sizes, distribution-expression and steering methods using variance-correlation and distributional-alignment metrics; authors explicitly warn about 'the potential risk of replacing human annotators with reasoning LLMs' where disagreement carries signal; author communication adds that small fine-tuned models (ModernBERT-class) with human labels often beat large reasoning models at this.",
   "implication": "For the 'instructions-underdetermine / route-to-human' triage lane, do not assume the strongest reasoning judge is the best ambiguity detector: disagreement-prediction may need a dedicated small fine-tuned model (or CoT-on-RLHF configuration) separate from the verdict judge.",
   "source": {
    "raw": "Can Reasoning Help Large Language Models Capture Human Annotator Disagreement? (EACL 2026) | https://arxiv.org/abs/2506.19467 | 2026-01-12 | academic",
    "title": "Can Reasoning Help Large Language Models Capture Human Annotator Disagreement? (EACL 2026)",
    "url": "https://arxiv.org/abs/2506.19467",
    "date": "2026-01-12",
    "type": "academic"
   },
   "venue": {
    "peer_reviewed": true,
    "marker": "ACL"
   },
   "same_source_claims": [],
   "in_corpus_verification": null,
   "external_verification": {
    "verdict": "confirmed",
    "checked": [
     "RLVR degrades disagreement modeling: 'RLVR-style reasoning degrades performance in disagreement modeling'",
     "naive CoT improves RLHF models: 'naive Chain-of-Thought (CoT) reasoning improves the performance of RLHF LLMs'",
     "'60 experimental setups across 3 tasks' -- matches asserted 60 setups / 3 tasks",
     "warning: 'the potential risk of replacing human annotators with reasoning LLMs'",
     "EACL 2026 Main listed; v3 dated Jan 12 2026 (matches source_date 2026-01-12)"
    ],
    "not_visible": [
     "evidence-field 'author communication' that small fine-tuned ModernBERT-class models with human labels beat large reasoning models -- explicitly external to the paper, not in the retrieved abstract",
     "specific metric names 'variance-correlation' and 'distributional-alignment' not verbatim in retrieved abstract (though 'distribution expression methods, and steering methods' are confirmed)"
    ],
    "quote": "RLVR-style reasoning degrades performance in disagreement modeling ... naive Chain-of-Thought (CoT) reasoning improves the performance of RLHF LLMs ... resulting in 60 experimental setups across 3 tasks",
    "notes": "All core claim assertions and the 60-setups/3-tasks figures are confirmed. The ModernBERT comparison is flagged in the evidence itself as author communication, so it cannot be checked against the source and is not weighed against the claim.",
    "checked_at": "2026-07-15",
    "retrieval": "abstract"
   },
   "source_retrieval_meta": {
    "retrieval": "abstract",
    "fetched": [
     "https://arxiv.org/abs/2506.19467"
    ],
    "resolved_title": "Can Reasoning Help Large Language Models Capture Human Annotator Disagreement?",
    "resolved_date": "2025-06-24 (v1); 2026-01-12 (v3)",
    "title_match": true
   },
   "used_on": [],
   "decision_use": {
    "grade": "B",
    "basis": "current-gen capability",
    "measured_on": "model-outputs",
    "note": "Jan-2026: RLVR training degrades disagreement modeling. Mechanism relevant to judge-model selection at the deployment tier; magnitude tier-bound.",
    "models_measured": []
   }
  },
  {
   "id": "CM-01",
   "domain": "critique-models",
   "area": null,
   "claim": "Human+critic-model teams occupy a strictly better operating point than either alone: in OpenAI's CriticGPT study, Human+CriticGPT teams wrote more comprehensive critiques than unassisted humans while hallucinating and nitpicking less than the model alone, and the comprehensiveness-vs-spurious-claims tradeoff is a tunable inference-time dial (FSBS length penalty), not a fixed property.",
   "load_bearing": true,
   "evidence": "CriticGPT critiques were preferred over human contractor critiques in 63% of cases on naturally occurring LLM errors (>80% Elo win rate on inserted bugs); both ChatGPT and CriticGPT caught substantially more inserted bugs than paid contractors (median ~50 min/review, ~5 yrs Python experience). Humans alone had far fewer hallucinated bugs/nitpicks; Human+CriticGPT teams landed between, 'moving beyond the model-only Pareto frontier.' Force Sampling Beam Search scores candidates by rm_score + LENGTH_MODIFIER x num_highlights, tracing a precision/comprehensiveness Pareto curve selectable at deployment without retraining. Verified directly from the arXiv HTML full text.",
   "implication": "Validates the hybrid one-touch foundation: AI critic drafts findings, human filters hallucinated nitpicks. Make the precision/comprehensiveness operating point an explicit per-project config knob (Q2/Q4/Q7 false-accusation SLO), since the same trained judge can be run strict or comprehensive per project.",
   "source": {
    "raw": "LLM Critics Help Catch LLM Bugs (CriticGPT, OpenAI) | https://arxiv.org/abs/2407.00215 | 2024-06-28 | primary",
    "title": "LLM Critics Help Catch LLM Bugs (CriticGPT, OpenAI)",
    "url": "https://arxiv.org/abs/2407.00215",
    "date": "2024-06-28",
    "type": "primary"
   },
   "venue": {
    "peer_reviewed": false,
    "marker": "preprint"
   },
   "same_source_claims": [
    "CM-02"
   ],
   "in_corpus_verification": {
    "verdict": "confirmed",
    "transcript": "Source verified: arXiv:2407.00215 \"LLM Critics Help Catch LLM Bugs\" (McAleese et al., OpenAI), submitted 2024-06-28 - date matches. Abstract directly confirms: 63% preference for model critiques on naturally occurring LLM errors; models catch more bugs than paid human reviewers; human-machine teams catch similar bug counts to LLM critics while hallucinating less than LLMs alone. Full-text details corroborated via independent secondary coverage: critic-assisted contractors wrote more comprehensive critiques than unassisted contractors while reducing hallucination/nitpick rate relative to the model, described as moving beyond the model-only Pareto frontier; FSBS scores candidates by rm_score + LENGTH_MODIFIER x num_highlights, giving a deployment-time precision/comprehensiveness dial without retraining. Minor reading note: \"strictly better than either alone\" holds in the paper's Pareto sense - teams matched (did not exceed) the model's bug-catch/comprehensiveness while beating its hallucination rate, and exceeded human comprehensiveness; the claim's own wording states exactly this, so no correction needed. Supersession check: no official OpenAI follow-up found; later work (CodeCriticBench 2025; 2026 papers on LLM reviewer overcorrection) extends but does not contradict these findings.",
    "corrected": null
   },
   "external_verification": {
    "verdict": "confirmed",
    "checked": [
     "'model-written critiques are preferred over human critiques in 63% of cases' (naturally occurring LLM errors)",
     "'model critiques are preferred over human critiques more than 80% of the time' on Human Inserted Bugs, scale 'linear in Elo'",
     "'Human+CriticGPT teams move beyond the model-only Pareto frontier'",
     "FSBS selects critiques by 'rm_score + LENGTH_MODIFIER * num_highlights'; 'explored 4 values of LENGTH_MODIFIER' (inference-time dial)",
     "contractors median ~5 years Python experience, ~50 min per critique",
     "human-machine teams 'hallucinating less than LLMs alone'"
    ],
    "not_visible": [],
    "quote": "Human+CriticGPT teams move beyond the model-only Pareto frontier.",
    "notes": "63% preference, >80% Elo on inserted bugs, the FSBS length-penalty dial, and the beyond-the-frontier team result all confirmed from arXiv HTML full text. The critic model is named 'CriticGPT' in the body though the arXiv title is 'LLM Critics Help Catch LLM Bugs'.",
    "checked_at": "2026-07-15",
    "retrieval": "fulltext"
   },
   "source_retrieval_meta": {
    "retrieval": "fulltext",
    "fetched": [
     "https://arxiv.org/abs/2407.00215",
     "https://arxiv.org/html/2407.00215v1"
    ],
    "resolved_title": "LLM Critics Help Catch LLM Bugs",
    "resolved_date": "28 Jun 2024 (v1)",
    "title_match": true
   },
   "used_on": [],
   "decision_use": {
    "grade": "C",
    "basis": "superseded-gen capability + human floor",
    "measured_on": "human-work",
    "note": "CriticGPT (GPT-4-era, 2024). Critic-assisted-beats-either-alone is a floor that should strengthen with critic quality, but the 63% preference number is two generations stale and unreplicated in 24+ months.",
    "models_measured": []
   }
  },
  {
   "id": "CM-02",
   "domain": "critique-models",
   "area": null,
   "claim": "Critic-assisted review of items already rated 'flawless' by a first human found real, rating-changing problems in 24% of flagged cases, versus 6% found by an unassisted second human review of the same data - a 4x lift on the exact 'confident pass' slice.",
   "load_bearing": false,
   "evidence": "CriticGPT (RL-only) was run over ChatGPT training data rated flawless by a first annotator; where the critique flagged a problem, reviewers agreed it substantially lowered the rating in 24% of cases, vs 6% in a no-critique second-rater replication - and this held out-of-distribution (mostly non-code tasks). Prioritizing by critique reward-model score further raised catch rates and cut hallucinations. Verified from arXiv HTML full text.",
   "implication": "Direct evidence for Q1's false-agreement problem and Q6's routing: unassisted human re-review of passed items misses ~3/4 of what critic-assisted review catches, so the standing expert-audit channel over confident passes should itself be critic-assisted, and critic-confidence scores can prioritize the audit queue.",
   "source": {
    "raw": "LLM Critics Help Catch LLM Bugs (CriticGPT, OpenAI) - 'flawless' data experiment | https://arxiv.org/abs/2407.00215 | 2024-06-28 | primary",
    "title": "LLM Critics Help Catch LLM Bugs (CriticGPT, OpenAI) - 'flawless' data experiment",
    "url": "https://arxiv.org/abs/2407.00215",
    "date": "2024-06-28",
    "type": "primary"
   },
   "venue": {
    "peer_reviewed": false,
    "marker": "preprint"
   },
   "same_source_claims": [
    "CM-01"
   ],
   "in_corpus_verification": null,
   "external_verification": {
    "verdict": "confirmed",
    "checked": [
     "With critique: 'In 24% of cases contractors indicated that the critique found a problem that substantially decreased the rating'",
     "Without critique: completions first rated 'flawless' were 'rated similarly poorly by a second only 6% of the time'",
     "Run over ChatGPT training data first rated 'flawless'",
     "Held out-of-distribution ('out-of-distribution for the critic model'), mostly non-code"
    ],
    "not_visible": [],
    "quote": "In 24% of cases contractors indicated that the critique found a problem that substantially decreased the rating",
    "notes": "The 24% (critic-assisted) vs 6% (unassisted second rater) comparison on the 'flawless' slice is confirmed verbatim in Section 3.6 - a 4x lift as the claim states.",
    "checked_at": "2026-07-15",
    "retrieval": "fulltext"
   },
   "source_retrieval_meta": {
    "retrieval": "fulltext",
    "fetched": [
     "https://arxiv.org/abs/2407.00215",
     "https://arxiv.org/html/2407.00215v1"
    ],
    "resolved_title": "LLM Critics Help Catch LLM Bugs",
    "resolved_date": "28 Jun 2024 (v1)",
    "title_match": true
   },
   "used_on": [
    {
     "page": "research.html",
     "anchor": "research-cards",
     "label": "Research - findings by decision weight"
    },
    {
     "page": "architecture.html",
     "anchor": "ontology",
     "label": "System - claim types"
    },
    {
     "page": "pilot.html",
     "anchor": "main",
     "label": "Pilot - gates and objectives"
    },
    {
     "page": "decisions.html",
     "anchor": "a5",
     "label": "Decisions - permanent audit"
    }
   ],
   "decision_use": {
    "grade": "C",
    "basis": "superseded-gen capability + human floor",
    "measured_on": "human-work",
    "note": "The 24%-vs-6% flawless-slice lift is GPT-4-era. Direction (audit confident passes with critic assistance) is the durable design input; the 4x magnitude is not a current estimate.",
    "models_measured": []
   }
  },
  {
   "id": "CM-03",
   "domain": "critique-models",
   "area": null,
   "claim": "In DeepMind's amplified-oversight experiments on fact-verification rating, confidence-based hybridization (AI rates when confident, humans rate the low-confidence slice with evidence-only assistance) reached 91.3% accuracy vs 87.7% AI-alone and 75.1% human-alone - and showing humans the AI's verdict/reasoning/confidence caused measurable over-reliance, while showing only retrieved evidence was the sole format that helped when the AI was right without hurting when it was wrong.",
   "load_bearing": true,
   "evidence": "1,918 expert-labeled fact-verification items; tool-using AI rater with majority-vote-over-50-samples confidence (well calibrated). AI 87.7% vs individual humans 75.1% (majority vote 80.6%). Hybridization at threshold 0.62: 89.3%; on the routed low-confidence slice humans beat the AI 71.3% vs 60.5% (real complementarity region). Evidence-only assistance: 79.3% when AI right (vs 71.3% baseline) with no significant harm when AI wrong; conditions showing Judgments+Reasoning/Confidence produced significant over-reliance decrements (e.g. beta = -0.768, p < .001); AI-debate condition was numerically worst. Verified from arXiv HTML full text.",
   "implication": "Answers Q6 concretely: spend the one human touch on AI-low-confidence items, and give that human the judge's collected evidence spans WITHOUT the judge's verdict or rationale (blind-ish adjudication over curated evidence). Requires the judge to emit calibrated confidence and quoted evidence as first-class outputs (Q3), and argues against 'human verifies AI verdict' as the default interaction.",
   "source": {
    "raw": "Human-AI Complementarity: A Goal for Amplified Oversight (Google DeepMind) | https://arxiv.org/abs/2510.26518 | 2025-10-30 | primary",
    "title": "Human-AI Complementarity: A Goal for Amplified Oversight (Google DeepMind)",
    "url": "https://arxiv.org/abs/2510.26518",
    "date": "2025-10-30",
    "type": "primary"
   },
   "venue": {
    "peer_reviewed": false,
    "marker": "preprint"
   },
   "same_source_claims": [],
   "in_corpus_verification": {
    "verdict": "confirmed",
    "transcript": "Every quantitative element checks out against the arXiv HTML full text (2510.26518). Dataset: 1,918 expert-labeled tuples - confirmed. AI rater 87.7%, individual humans 75.1%, human majority vote 80.6% - confirmed. Confidence via 50 samples/sentence (avg 33.25 passing format check), described as fairly well calibrated - confirmed. Unassisted hybridization at threshold T=0.62: 89.3% (beta=0.413, p=.012) - confirmed. Low-confidence routed slice (280 items): humans 71.3% vs AI 60.5% (p=.006) - confirmed. The headline 91.3% is specifically the ASSISTED hybrid (AI when confident + evidence-assisted humans on the low-confidence slice), vs 89.3% unassisted hybrid - exactly as the claim words it, so 91.3% is correctly attributed. Evidence-only assistance: 79.3% vs 71.3% baseline when AI correct (beta=0.446, p=.009), no significant harm when AI wrong (64.0% vs 61.5%); paper explicitly calls it \"the only form of assistance that achieves the ideal of helping when correct and not hurting when wrong\" - confirmed. Over-reliance when AI wrong: Evidence&Reasoning&Judgments beta=-0.768 (SE=0.174, z=-4.413, p<.001), ER&J&Confidence beta=-0.636 p<.001, Judgments&Confidence beta=-0.356 p=.042 - the quoted beta=-0.768 matches the Evidence&Reasoning&Judgments condition exactly. Debate was numerically worst (64.9%, only format below baseline) - confirmed. Date: v1 submitted 2025-10-30 as claimed. Currency: a v2 was posted 2026-06-25 and the paper was published at ACM FAccT '26 (DOI 10.1145/3805689.3812308); this is publication/revision, not supersession - no later work contradicting the findings found. DeepMind blog notes ongoing follow-up combining hybridization+assistance on an internal rating task, unpublished as of 2026-07-14. Minor caveat only: whoever cites this should prefer the v2/FAccT version in case numbers shifted slightly in revision (v1 numbers verified here).",
    "corrected": null
   },
   "external_verification": {
    "verdict": "confirmed",
    "checked": [
     "AI rater 87.7% vs human raters 75.1%",
     "human majority vote 80.6% (vs 75.1%)",
     "N = 1918 Evaluation Set examples",
     "unassisted hybridization 89.3% (higher than AI-alone 87.7%)",
     "confidence threshold T = 0.62",
     "on low-confidence slice AI 60.5% vs human 71.3%",
     "evidence-assisted hybridization achieves 91.3%",
     "evidence-only helps on items AI gets correct: 79.3% vs 71.3%",
     "over-reliance coefficient -0.768, SE=0.174, z=-4.413, p<.001"
    ],
    "not_visible": [],
    "quote": "rater achieves 87.7% accuracy, performing above human raters who achieved 75.1% accuracy ... confidence threshold of 0.62, results in an accuracy of 91.3%, compared to 89.3% using unassisted ... -0.768, SE = 0.174, z = -4.413, p < .001",
    "notes": "Every headline number in the claim verified verbatim from the PDF full text via pdftotext. AI-alone 87.7%, human-alone 75.1%, majority 80.6%, N=1918, threshold 0.62, unassisted hybrid 89.3%, evidence-assisted 91.3%, low-conf slice 60.5% AI / 71.3% human, evidence-only 79.3% vs 71.3%, beta=-0.768 p<.001 all present. Note: PDF served at /pdf/ is v2 (June 2026); source_date is v1 (Oct 30 2025). All cited numbers match the retrieved v2 text. Authors: Jain, Bridgers, Janzer, Greig, Teh, Mikulik.",
    "checked_at": "2026-07-15",
    "retrieval": "pdf"
   },
   "source_retrieval_meta": {
    "retrieval": "pdf",
    "fetched": [
     "https://arxiv.org/abs/2510.26518",
     "https://arxiv.org/pdf/2510.26518"
    ],
    "resolved_title": "Human-AI Complementarity: A Goal for Amplified Oversight",
    "resolved_date": "v1 2025-10-30; PDF served is v2 (2510.26518v2, 2026-06-25)",
    "title_match": true
   },
   "used_on": [
    {
     "page": "brief.html",
     "anchor": "hybrid",
     "label": "Summary - human role"
    },
    {
     "page": "pilot.html",
     "anchor": "main",
     "label": "Pilot - gates and objectives"
    },
    {
     "page": "decisions.html",
     "anchor": "a4",
     "label": "Decisions - authority boundaries"
    },
    {
     "page": "decisions.html",
     "anchor": "c3",
     "label": "Decisions - default: who sees what"
    }
   ],
   "decision_use": {
    "grade": "B",
    "basis": "current-gen + human science",
    "measured_on": "adjacent-domain",
    "note": "2025 DeepMind amplified-oversight. The format effect (verdict-visible -> over-reliance; evidence-only safe) is human behavior and durable; the 91.3/87.7/75.1 magnitudes are tier- and task-bound (fact verification, not annotation QA).",
    "models_measured": []
   }
  },
  {
   "id": "CM-04",
   "domain": "critique-models",
   "area": null,
   "claim": "A Nature Human Behaviour meta-analysis of 100+ experiments (300+ effect sizes) found human-AI combinations on average performed significantly WORSE than the best of human or AI alone on decision/judgment tasks, with gains only where humans outperformed the AI solo - naive human-checks-AI designs destroy value.",
   "load_bearing": false,
   "evidence": "Vaccaro, Almaatouq & Malone systematic review/meta-analysis: average human-AI combo underperformed the stronger party alone; losses concentrated in decision-making tasks, gains in content-creation tasks; combos helped when humans beat AI alone. DeepMind's complementarity paper explicitly frames its confidence-routing + evidence-only design as the answer to this result. Verified from arXiv abstract.",
   "implication": "The Q6 fork ('is one shallow touch better than zero touch plus deep audits?') is empirically live: an automation-biased shallow verify step is the modal failure in the literature. The touch only pays where the human plausibly beats the judge (routed low-confidence/contested items), which mandates closed-loop routing on observed overturn rates rather than fixed per-item review.",
   "source": {
    "raw": "When combinations of humans and AI are useful: A systematic review and meta-analysis | https://arxiv.org/abs/2405.06087 | 2024-10-28 | contrarian",
    "title": "When combinations of humans and AI are useful: A systematic review and meta-analysis",
    "url": "https://arxiv.org/abs/2405.06087",
    "date": "2024-10-28",
    "type": "contrarian"
   },
   "venue": {
    "peer_reviewed": false,
    "marker": "preprint"
   },
   "same_source_claims": [],
   "in_corpus_verification": null,
   "external_verification": {
    "verdict": "confirmed",
    "checked": [
     "meta-analysis of over 100 recent experimental studies reporting over 300 effect sizes",
     "human-AI combinations performed significantly worse than the best of humans or AI alone",
     "performance losses in tasks that involved making decisions",
     "significantly greater gains in tasks that involved creating content",
     "when humans outperformed AI alone, we found performance gains in the combination",
     "Nat Hum Behav (2024) / Nature Human Behaviour"
    ],
    "not_visible": [
     "exact counts 106 studies / 370 effect sizes (page says 'over 100' and 'over 300')",
     "DeepMind complementarity paper framing this result as its motivation -- an external editorial connection, not on this page"
    ],
    "quote": "human-AI combinations performed significantly worse than the best of humans or AI alone ... when humans outperformed AI alone, we found performance gains in the combination",
    "notes": "Nature Human Behaviour venue, 100+ studies / 300+ effect sizes, decision-task losses, creation-task gains, and the humans-beat-AI condition all confirmed verbatim.",
    "checked_at": "2026-07-15",
    "retrieval": "abstract"
   },
   "source_retrieval_meta": {
    "retrieval": "abstract",
    "fetched": [
     "https://arxiv.org/abs/2405.06087"
    ],
    "resolved_title": "When combinations of humans and AI are useful: A systematic review and meta-analysis",
    "resolved_date": "2024-05-09 (v1); v2 2024-10-29; Nat Hum Behav (2024)",
    "title_match": true
   },
   "used_on": [],
   "decision_use": {
    "grade": "A",
    "basis": "human science (direction)",
    "measured_on": "adjacent-domain",
    "note": "Meta-analytic human-AI complementarity result. As deployment-tier models widen the human-AI gap, the warning against naive verify-the-AI designs strengthens.",
    "models_measured": []
   }
  },
  {
   "id": "CM-05",
   "domain": "critique-models",
   "area": null,
   "claim": "In the largest deployed RCT of AI critiquing human evaluative work (ICLR 2025, feedback on >20,000 peer reviews), 27% of reviewers who received LLM feedback revised their reviews, incorporating >12,000 suggestions, producing reviews +80 words that blinded raters judged more informative, plus higher author-rebuttal engagement - and feedback was only delivered if it passed a suite of automated LLM reliability tests.",
   "load_bearing": true,
   "evidence": "Randomized controlled study at ICLR 2025 (~22.5k feedback-arm vs ~22.4k control reviews): Review Feedback Agent flagged vague comments, content misunderstandings, unprofessional remarks; guardrail LLM tests gated delivery; 26.6-27% update rate, avg +80 words among updaters, blinded informativeness gains, longer author-reviewer discussions. Verified from arXiv abstract page; corroborated by the official ICLR blog.",
   "implication": "Q7's central bet is supported at deployment scale: constructive, item-specific AI feedback measurably changes the behavior of human evaluators, and the shipped pattern held feedback to its own automated verification gate before delivery (fail-closed feedback). Also gives a realistic effect-size prior: ~1 in 4 recipients act on feedback - plan longitudinal metrics around that base rate.",
   "source": {
    "raw": "Can LLM feedback enhance review quality? A randomized study of 20K reviews at ICLR 2025 | https://arxiv.org/abs/2504.09737 | 2025-04-13 | academic",
    "title": "Can LLM feedback enhance review quality? A randomized study of 20K reviews at ICLR 2025",
    "url": "https://arxiv.org/abs/2504.09737",
    "date": "2025-04-13",
    "type": "academic"
   },
   "venue": {
    "peer_reviewed": true,
    "marker": "ICLR"
   },
   "same_source_claims": [],
   "in_corpus_verification": {
    "verdict": "confirmed",
    "transcript": "Source verified directly (arXiv:2504.09737, v1 submitted 2025-04-13 - date matches). Abstract states: Review Feedback Agent deployed at ICLR 2025 as a large randomized controlled study on >20,000 randomly selected reviews; feedback targeted vague comments, content misunderstandings, unprofessional remarks; 27% of reviewers who received feedback updated their reviews; >12,000 suggestions incorporated (12,222 in the paper); +80 words average among updaters; more informative per blinded evaluators; longer author-reviewer discussions; feedback delivered only if it passed a suite of automated LLM reliability tests. Arm sizes corroborated from paper figures: 22,467 selected for feedback vs 22,364 control; 18,946 successfully received feedback, of whom 26.6% updated (the abstract rounds to 27%). Two nuances, neither contradicting the claim: (1) the superlative \"largest deployed RCT of AI critiquing human evaluative work\" is the researcher's framing, not verbatim in the abstract - no larger counterexample found, and it is consistent with the study's scale; (2) not superseded but now formally published: Nature Machine Intelligence, Feb 23, 2026, as \"A large-scale randomized study of large language model feedback in peer review\" (https://www.nature.com/articles/s42256-026-01188-x) - the NMI version is the preferred citation going forward. A related follow-up mixed-methods study on reviewer perceptions exists (arXiv:2602.13817) but does not alter these results.",
    "corrected": null
   },
   "external_verification": {
    "verdict": "confirmed",
    "checked": [
     "feedback to more than 20,000 randomly selected reviews",
     "27% of reviewers who received feedback updated their reviews",
     "over 12,000 feedback suggestions incorporated",
     "average increase of 80 words among updaters",
     "more informative reviews as judged by blinded researchers",
     "longer author-reviewer discussions / more rebuttal engagement",
     "a suite of automated LLM-powered reliability tests acted as guardrails gating delivery"
    ],
    "not_visible": [
     "the specific '26.6%' figure (abstract states 27%; 26.6% appears only in the claim's evidence field)"
    ],
    "quote": "27% of reviewers who received feedback updated their reviews and over 12,000 feedback suggestions from the agent were incorporated by those reviewers ... an average increase of 80 words among those who updated after receiving feedback.",
    "notes": "All headline numbers in the core claim confirmed verbatim from the abstract.",
    "checked_at": "2026-07-15",
    "retrieval": "abstract"
   },
   "source_retrieval_meta": {
    "retrieval": "abstract",
    "fetched": [
     "https://arxiv.org/abs/2504.09737"
    ],
    "resolved_title": "Can LLM feedback enhance review quality? A randomized study of 20K reviews at ICLR 2025",
    "resolved_date": "2025-04-13",
    "title_match": true
   },
   "used_on": [
    {
     "page": "research.html",
     "anchor": "research-cards",
     "label": "Research - findings by decision weight"
    },
    {
     "page": "decisions.html",
     "anchor": "d1",
     "label": "Decisions - enforcement weight"
    },
    {
     "page": "decisions.html",
     "anchor": "c3",
     "label": "Decisions - default: who sees what"
    }
   ],
   "decision_use": {
    "grade": "B",
    "basis": "human response to AI feedback",
    "measured_on": "human-work",
    "note": "2025 RCT on >20k real peer reviews - one of the few human-work measurements in the corpus. Revision behavior is human; feedback quality rises with model tier, so 27% is closer to a floor than a ceiling. Domain transfer (peer review -> paid annotation) untested (E6).",
    "models_measured": []
   }
  },
  {
   "id": "CM-06",
   "domain": "critique-models",
   "area": null,
   "claim": "LLMs are structurally weak at detecting omissions: on AbsenceBench, average F1 drops 56.9 points versus detecting the same content as insertions (best model 71.2% overall, 40.0% on code diffs, vs ~99.5% needle-in-haystack), and inserting explicit placeholders at gap sites recovers ~35.7 points - absence has no attention key to attend to.",
   "load_bearing": true,
   "evidence": "4,302 instances across poetry, numerical sequences, GitHub PRs (~5K token contexts): models must list deliberately deleted elements given original+modified documents. Gemini-2.5-flash (thinking) best at 71.2% avg F1; Mixtral-8x7B scores 99%+ on NIAH but 14.7% here; thinking tokens buy only ~7.9%; '<missing line>' placeholders add up to +81.8%. NeurIPS 2025 Datasets & Benchmarks. Verified from arXiv HTML full text.",
   "implication": "Caps what the judge can verify autonomously (Q3/Q5): 'comprehensiveness'/'complete' praise claims and omission-type attempter failures are the judge's structurally weakest lane. The rubric compiler should convert completeness criteria into explicit enumerable checklists (placeholder effect = give absence a token), and free-form omission judgment belongs in the human or human-assisted lane, not judge-autonomous.",
   "source": {
    "raw": "AbsenceBench: Language Models Can't Tell What's Missing | https://arxiv.org/abs/2506.11440 | 2025-06-13 | academic",
    "title": "AbsenceBench: Language Models Can't Tell What's Missing",
    "url": "https://arxiv.org/abs/2506.11440",
    "date": "2025-06-13",
    "type": "academic"
   },
   "venue": {
    "peer_reviewed": true,
    "marker": "NeurIPS"
   },
   "same_source_claims": [],
   "in_corpus_verification": null,
   "external_verification": {
    "verdict": "partially_confirmed",
    "checked": [
     "'a massive 56.9% drop in F1-score on average' (omission vs insertion)",
     "best overall average F1 71.2 (Gemini-2.5-flash, thinking)",
     "highest GitHub-PR/code-diff score only 40.0% (by Claude-3.7-Sonnet thinking)",
     "nearly 99.5% F1 on poetry under the insertion (NIAH-style) setting",
     "'<missing line>' placeholders boost performance 35.7% on average; +81.8% for Claude-3.7-Sonnet on GitHub PRs",
     "4302 instances total, average context length 5K tokens",
     "Mixtral-8x7B: perfect NIAH score but only 14.7% F1 on AbsenceBench",
     "inference-time compute (thinking) yields modest 7.9% improvement"
    ],
    "not_visible": [
     "NeurIPS 2025 Datasets & Benchmarks acceptance (not in the abstract or HTML full text)"
    ],
    "quote": "AbsenceBench contains 4302 instances in total, with an average context length of 5K tokens. ... observing a massive 56.9% drop in F1-score on average ... This boosts the performance by a dramatic 35.7% on average",
    "notes": "All headline numbers are confirmed verbatim in the HTML full text. Minor attribution nuance: the claim pairs '71.2% overall' and '40.0% on code diffs' as one 'best model', but 71.2 is Gemini-2.5-flash's average while 40.0% is the domain-best set by Claude-3.7-Sonnet (thinking) (Gemini's GitHub-PR score is 30.9). Downgraded to partially_confirmed only because the asserted NeurIPS 2025 venue is not present in retrieved text.",
    "checked_at": "2026-07-15",
    "retrieval": "fulltext"
   },
   "source_retrieval_meta": {
    "retrieval": "fulltext",
    "fetched": [
     "https://arxiv.org/abs/2506.11440",
     "https://arxiv.org/html/2506.11440v1"
    ],
    "resolved_title": "AbsenceBench: Language Models Can't Tell What's Missing",
    "resolved_date": "2025-06-13",
    "title_match": true
   },
   "used_on": [
    {
     "page": "research.html",
     "anchor": "research-cards",
     "label": "Research - findings by decision weight"
    },
    {
     "page": "architecture.html",
     "anchor": "ontology",
     "label": "System - claim types"
    },
    {
     "page": "decisions.html",
     "anchor": "invariants",
     "label": "Decisions - settled constraints"
    }
   ],
   "decision_use": {
    "grade": "C",
    "basis": "superseded-gen capability",
    "measured_on": "model-outputs",
    "note": "AbsenceBench on 2025 flash-tier (Gemini-2.5-flash). Omission-blindness is attention-mechanics and plausibly persists, but deployment-tier recall is unmeasured - E2 measures it before any absence-claim verdict is trusted.",
    "models_measured": [
     "Gemini-2.5-flash"
    ]
   }
  },
  {
   "id": "CM-07",
   "domain": "critique-models",
   "area": null,
   "claim": "Critique quality is itself quantifiable at usable reliability: MetaCritique decomposes critiques into atomic information units and scores precision/recall against references, with GPT-4 AIU-level judgments at 85-89% accuracy and Meta-F1 correlating with human gold at Pearson 0.84-0.89 - and measured baselines show human critiques are high-precision/low-recall (87.6%/48.7%) while LLM critiques are the inverse-ish (71.9%/53.3%).",
   "load_bearing": false,
   "evidence": "ACL Findings 2024 framework: precision = fraction of critique AIUs judged factual; recall = coverage of reference-critique AIUs; each judgment carries a natural-language rationale. Choosing critiques by Meta-F1 yields refinements that win 51%/lose 27% by human eval, vs GPT-4 pairwise picks losing more than winning. ~28% of LLM-critique AIUs were non-factual vs ~12% for humans; LLM critiques carried 2.4x the information volume (8.10 vs 3.31 AIUs). Verified from arXiv HTML full text.",
   "implication": "Q7's reflexive-grounding requirement is buildable today: run AutoQA feedback through AIU-level precision scoring as its own admission gate, and report critique-precision/recall as first-class system metrics. Also sets expectations: the AI side of the pipeline supplies coverage, the human side supplies precision - mirroring CriticGPT.",
   "source": {
    "raw": "The Critique of Critique (MetaCritique) | https://arxiv.org/abs/2401.04518 | 2024-01-09 | academic",
    "title": "The Critique of Critique (MetaCritique)",
    "url": "https://arxiv.org/abs/2401.04518",
    "date": "2024-01-09",
    "type": "academic"
   },
   "venue": {
    "peer_reviewed": true,
    "marker": "ACL"
   },
   "same_source_claims": [],
   "in_corpus_verification": null,
   "external_verification": {
    "verdict": "confirmed",
    "checked": [
     "critiques decomposed into Atomic Information Units (AIUs), precision/recall scored against references with NL rationales",
     "GPT-4 AIU-level accuracy 85-89% (Table 3: 89.12 / 87.96 precision task; 85.47 / 86.82 recall task; 'nearly 90%')",
     "Meta-F1 Pearson with human gold 0.841 (human-written) / 0.886 (LLM-generated)",
     "human critiques high-precision/low-recall 87.61 / 48.72",
     "LLM critiques 71.85 / 53.28",
     "refinement by Meta-F1: 51% Better Critique Wins / 22% Tie / 27% Loses (human eval)",
     "average AIUs 8.10 (LLM/Hypo.l) vs 3.31 (human/Hypo.h) = 2.4x",
     "non-factual AIUs ~28% (LLM) vs ~12% (human), derived from precision 71.85 / 87.61"
    ],
    "not_visible": [
     "the specific comparative that GPT-4 pairwise picks 'lose more than they win' vs MetaCritique (a separate sub-figure not surfaced in retrieval)"
    ],
    "quote": "GPT-4 ... achieves an impressive performance (nearly 90%) ... MetaCritique-GPT4-F1 scores 0.841 ... 0.886 ... 51% Better Critique Wins, 22% Tie, 27% Loses ... Hypo.l = 8.10, Hypo.h = 3.31",
    "notes": "All headline numbers verified verbatim from arXiv HTML full text (Tables 1,3,4,6, Figure 4b). 8.10/3.31 = 2.45 confirms the '2.4x information volume' claim. Non-factual percentages are correctly derived from precision as the claim itself states.",
    "checked_at": "2026-07-15",
    "retrieval": "fulltext"
   },
   "source_retrieval_meta": {
    "retrieval": "fulltext",
    "fetched": [
     "https://arxiv.org/abs/2401.04518",
     "https://arxiv.org/html/2401.04518"
    ],
    "resolved_title": "The Critique of Critique",
    "resolved_date": "2024-01-09 (v2 2024-06-01; Findings of ACL 2024)",
    "title_match": true
   },
   "used_on": [
    {
     "page": "pilot.html",
     "anchor": "main",
     "label": "Pilot - gates and objectives"
    }
   ],
   "decision_use": {
    "grade": "C",
    "basis": "superseded-gen method demo",
    "measured_on": "model-outputs",
    "note": "MetaCritique (GPT-4, 2024). The critique-decomposition METHOD is reusable; its reported reliability numbers are stale.",
    "models_measured": [
     "GPT-4"
    ]
   }
  },
  {
   "id": "CM-08",
   "domain": "critique-models",
   "area": null,
   "claim": "Verdict accuracy and critique validity dissociate badly: in CriticBench-THU data 24.8% of items got the correct verdict with a low-quality critique, and open-loop verdict-agreement metrics compress a 27.1-point real error-identification gap (ProcessBench: o1-mini 88.9 vs Qwen2.5-72B 61.8) into a 1.3-point verdict-F1 gap - so critique quality must be evaluated closed-loop by whether the critique drives a successful correction.",
   "load_bearing": false,
   "evidence": "RealCritic (Jan 2025, arXiv 2501.14492): closed-loop critique-then-correct design over 8 reasoning benchmarks; classical LLMs (incl. GPT-4) lose accuracy under self-critique (-1.8 to -5.1 avg; up to -35.6 domain drops) while o1-mini is the only model with positive self-critique delta; failure modes 'superficial success' (right verdict, wrong analysis) and 'contradictory output' documented. Independently confirmed at scale in 2026: an ICML 2026 paper (RM-NLHF, arXiv 2601.07349) finds outcome-rewarded generative reward models 'guess correct outcomes without sound critiques' and uses similarity-to-human-critique as a process reward to fix it. RealCritic verified via arXiv abstract + HTML (HTML render partially draft-quality; abstract and appendix numbers consistent).",
   "implication": "For Q1 meta-evaluation: never certify the AutoQA on verdict agreement alone - a judge can match pass/fail labels while its rationales are wrong, which poisons the feedback channel and attempter trust. Validation must score the rationale (does the cited evidence entail the finding; does acting on the feedback fix the item), i.e., Q7's explainability contract is also the correct validation instrument.",
   "source": {
    "raw": "RealCritic: Towards Effectiveness-Driven Evaluation of Language Model Critiques | https://arxiv.org/abs/2501.14492 | 2025-01-24 | academic",
    "title": "RealCritic: Towards Effectiveness-Driven Evaluation of Language Model Critiques",
    "url": "https://arxiv.org/abs/2501.14492",
    "date": "2025-01-24",
    "type": "academic"
   },
   "venue": {
    "peer_reviewed": false,
    "marker": "preprint"
   },
   "same_source_claims": [],
   "in_corpus_verification": null,
   "external_verification": {
    "verdict": "partially_confirmed",
    "checked": [
     "closed-loop methodology evaluating quality of corrections generated from critiques",
     "eight challenging reasoning tasks",
     "classical LLMs significantly lag o1-mini across all critique scenarios",
     "classical LLMs can fall below their own baselines under self/iterative critique"
    ],
    "not_visible": [
     "CriticBench-THU 24.8%",
     "27.1-point error-identification gap",
     "ProcessBench o1-mini 88.9 vs Qwen2.5-72B 61.8",
     "1.3-point verdict-F1 gap",
     "-1.8 to -5.1 avg self-critique delta",
     "-35.6 domain drop",
     "'superficial success' failure-mode name",
     "'contradictory output' failure-mode name",
     "RM-NLHF / arXiv 2601.07349 (separate source)"
    ],
    "quote": "our approach employs a closed-loop methodology that evaluates the quality of corrections generated from critiques ... classical LLMs significantly lag behind the advanced reasoning-based model o1-mini across all critique scenarios",
    "notes": "RealCritic's core design (closed-loop critique-then-correct over 8 reasoning tasks; o1-mini only model with positive self-critique; classical LLMs drop below baselines) is confirmed. However essentially every headline number in the claim is drawn from OTHER benchmarks (CriticBench-THU 24.8%, ProcessBench 88.9/61.8) or from RealCritic detailed results not present in the abstract; none are visible in the retrieved text.",
    "checked_at": "2026-07-15",
    "retrieval": "abstract"
   },
   "source_retrieval_meta": {
    "retrieval": "abstract",
    "fetched": [
     "https://arxiv.org/abs/2501.14492"
    ],
    "resolved_title": "RealCritic: Towards Effectiveness-Driven Evaluation of Language Model Critiques",
    "resolved_date": "2025-01-24",
    "title_match": true
   },
   "used_on": [
    {
     "page": "pilot.html",
     "anchor": "main",
     "label": "Pilot - gates and objectives"
    }
   ],
   "decision_use": {
    "grade": "C",
    "basis": "superseded-gen capability",
    "measured_on": "model-outputs",
    "note": "Right-verdict/poor-critique dissociation measured on 2024-25 models; re-check found headline numbers span sibling benchmarks. Keep the failure mode, drop the numbers.",
    "models_measured": [
     "o1-mini",
     "Qwen2.5-72B",
     "GPT-4"
    ]
   }
  },
  {
   "id": "CM-09",
   "domain": "critique-models",
   "area": null,
   "claim": "A June 2026 study of 21 LLM judges (~541k judgments) shows raw percent-agreement overstates chance-corrected agreement by 33.8-41.3 points (85% agreement = kappa ~0.48), judge rankings flip by up to 15 positions across benchmarks, and the most reproducible judges are among the least valid (test-retest 0.99 with position bias 0.19) - leading the authors to prescribe a pre-deployment Minimum Viable Validation Protocol.",
   "load_bearing": false,
   "evidence": "'Reliability without Validity' (UC Berkeley, arXiv 2606.19544, evals run March-April 2026): MT-Bench/JudgeBench/RewardBench, 118 runs at temperature 0; kappa deflation universal across all 21 judges; consistency-bias paradox instantiated by Qwen3-8B and Gemini 2.5 Flash; MVVP = report kappa/Krippendorff alpha as headline, AB+BA position swaps, >=3 replicates, >=2 benchmarks spanning preference- and correctness-style labels, and audit that high stability is not just high bias. Verified from arXiv HTML full text.",
   "implication": "Directly answers Q1's statistics question (ban raw percent-agreement; standardize on chance-corrected stats) and Q5's ship-gate question (perturbation/position-swap + replicate harness is now published prescriptive practice). Run-to-run consistency must never be reported as evidence of validity - a perfectly consistent judge can be laundering a bias.",
   "source": {
    "raw": "Reliability without Validity: A Systematic, Large-Scale Evaluation of LLM-as-a-Judge Models Across Agreement, Consistency, and Bias | https://arxiv.org/abs/2606.19544 | 2026-06-17 | academic",
    "title": "Reliability without Validity: A Systematic, Large-Scale Evaluation of LLM-as-a-Judge Models Across Agreement, Consistency, and Bias",
    "url": "https://arxiv.org/abs/2606.19544",
    "date": "2026-06-17",
    "type": "academic"
   },
   "venue": {
    "peer_reviewed": false,
    "marker": "preprint"
   },
   "same_source_claims": [
    "AJ-03",
    "CT-01"
   ],
   "in_corpus_verification": null,
   "external_verification": {
    "verdict": "confirmed",
    "checked": [
     "exact match overstates chance-corrected agreement by between 33.8 and 41.3 percentage points across the 21 models",
     "a judge reporting '85% agreement' on MT-Bench has kappa approximately 0.48",
     "Llama 3.3 70B which shifts 15 positions, from 5 on MT-Bench to 20 on JudgeBench",
     "Qwen 3 8B (test-retest 0.992, position bias 0.192)",
     "Gemini 2.5 Flash (test-retest 0.988, position bias 0.125)",
     "All evaluations used temperature 0",
     "UC Berkeley School of Information",
     "118 runs",
     "MVVP: report Cohen's kappa or Krippendorff's alpha alongside exact-match, AB+BA position swaps, >=3 replicates, >=2 benchmarks spanning preference- and correctness-style labels, audit that high stability is not high bias"
    ],
    "not_visible": [
     "'March 2026' as an eval start month (page states results hold across 'the April 2026 frontier'; March not explicitly seen)"
    ],
    "quote": "exact match overstates chance-corrected agreement by between 33.8 and 41.3 percentage points across the 21 models",
    "notes": "Every precise figure the claim attributes to the HTML full text is present verbatim. The claim's '15 positions' matches Section 4.3 ('shifts 15 positions'), though the abstract and Figure 2 say 14 positions -- the paper is internally inconsistent on this number. Claim's rounded 0.99/0.19 corresponds to the paper's 0.992/0.192 for Qwen 3 8B.",
    "checked_at": "2026-07-15",
    "retrieval": "fulltext"
   },
   "source_retrieval_meta": {
    "retrieval": "fulltext",
    "fetched": [
     "https://arxiv.org/abs/2606.19544",
     "https://arxiv.org/html/2606.19544v1"
    ],
    "resolved_title": "Reliability without Validity: A Systematic, Large-Scale Evaluation of LLM-as-a-Judge Models Across Agreement, Consistency, and Bias",
    "resolved_date": "2026-06-17",
    "title_match": true
   },
   "used_on": [],
   "decision_use": {
    "grade": "B",
    "basis": "current-gen capability",
    "measured_on": "model-outputs",
    "note": "Same study family as AJ-03; June-2026 window.",
    "models_measured": [
     "Qwen3-8B",
     "Gemini 2.5"
    ]
   }
  },
  {
   "id": "CM-10",
   "domain": "critique-models",
   "area": null,
   "claim": "The closest production analog to the proposed AutoQA already exists: Toloka's deployed 'LLM QA' runs a tool-using agent per quality-metric on every human annotation submission, emits a strict three-way Pass / Fail / Unable-to-verify verdict (the abstain class exists specifically to prevent hallucinated verdicts), coaches annotators with Socratic feedback, and explicitly accepts lower precision on Fail because a false pass costs more than escalating a genuine pass.",
   "load_bearing": false,
   "evidence": "Toloka engineering blog (bylined 2026-03-30): agentic autocheck with web/image/audio/Python/bash tools; 'one metric, one entity, one verdict' scoping (separate agent instance per criterion); internal benchmark of 300+ real submissions from 20+ live projects with senior-QA ground truth (numbers in an image, not extractable); reported constraint: performance 'entirely bounded by task design' - ambiguous guidelines produce walls of Unable-to-verify. Verified by opening the post.",
   "implication": "Validates several foundational choices from production: criterion-scoped judge instances (Q3 decomposition at the per-axis level), a first-class abstain/escalate verdict distinct from pass/fail (Q2 verdict ontology), asymmetric error costs favoring false-flags over false-passes, and rubric quality as the binding constraint (Q4: tighten criteria before writing guidelines).",
   "source": {
    "raw": "LLM QA: Scaling data quality assurance technologically (Toloka) | https://toloka.ai/blog/llm-qa-scaling-data-quality-assurance-technologically/ | 2026-03-30 | practitioner",
    "title": "LLM QA: Scaling data quality assurance technologically (Toloka)",
    "url": "https://toloka.ai/blog/llm-qa-scaling-data-quality-assurance-technologically/",
    "date": "2026-03-30",
    "type": "practitioner"
   },
   "venue": {
    "peer_reviewed": null,
    "marker": "practitioner"
   },
   "same_source_claims": [],
   "in_corpus_verification": null,
   "external_verification": {
    "verdict": "confirmed",
    "checked": [
     "agentic autocheck running on every submission",
     "separate instance launched for every single quality metric",
     "rigid three-way scale: Pass, Fail, or Unable to verify",
     "abstain state exists because otherwise a model will hallucinate a guess",
     "Socratic-style feedback that asks questions rather than pointing to the error",
     "lower precision on Fail is an acceptable tradeoff; false pass costlier than sending a genuine pass to human review",
     "tools: download web pages, view images, process audio/video, run Python, execute bash",
     "'one metric, one entity, one verdict' constraint",
     "internal benchmark of over 300 real submissions from 20-plus live projects",
     "ground truth from Toloka senior QA reviewers",
     "performance entirely bounded by task design; ambiguity yields a wall of Unable to verify"
    ],
    "not_visible": [
     "the actual benchmark performance numbers (claim itself notes these are in an image, not extractable)"
    ],
    "quote": "We enforce a deliberate constraint: one metric, one entity, one verdict.",
    "notes": "All eight sub-claims supported verbatim. Byline Vitaly Moiseev and Mariya Shmatova, dated March 30, 2026. Vendor self-description; benchmark scores are referenced but not displayed as text, consistent with the claim's own caveat.",
    "checked_at": "2026-07-15",
    "retrieval": "fulltext"
   },
   "source_retrieval_meta": {
    "retrieval": "fulltext",
    "fetched": [
     "https://toloka.ai/blog/llm-qa-scaling-data-quality-assurance-technologically/"
    ],
    "resolved_title": "LLM QA: Scaling data quality assurance technologically",
    "resolved_date": "byline 2026-03-30 (page metadata also shows a 2026-07-15 timestamp)",
    "title_match": true
   },
   "used_on": [
    {
     "page": "decisions.html",
     "anchor": "a4",
     "label": "Decisions - authority boundaries"
    }
   ],
   "decision_use": {
    "grade": "A",
    "basis": "production practice",
    "measured_on": "structural",
    "note": "Toloka's deployed pass/fail/unable-to-verify QA with human escalation. Existence proof of the architecture in production - vendor-reported, flagged as such.",
    "models_measured": []
   }
  },
  {
   "id": "CM-11",
   "domain": "critique-models",
   "area": null,
   "claim": "The foundational 2022 result behind the whole lineage: model-written critiques helped human evaluators find flaws in summaries they would otherwise have missed (including planted flaws in deliberately misleading human-written summaries), and models exhibit a discriminator-critique gap - they can often detect that something is wrong better than they can articulate why.",
   "load_bearing": false,
   "evidence": "Saunders et al. (OpenAI, arXiv 2206.05802): topic-based summarization assistance experiments; critiques helped on both model- and human-written summaries; framework comparing generation, discrimination, and critique ability found 'even large models may still have relevant knowledge they cannot or do not articulate as critiques'; larger models critique and self-refine better. Verified from arXiv abstract.",
   "implication": "The discriminator-critique gap argues for a two-stage judge: use cheap discrimination (scores/uncertainty) for routing and prioritization even where articulated critique is unreliable, and hold only the articulated-critique layer to the evidence-quoting bar (Q3/Q7). Judge-assisted human review of HUMAN-written work is validated at the root of this lineage, not just AI-output critique.",
   "source": {
    "raw": "Self-critiquing models for assisting human evaluators (OpenAI) | https://arxiv.org/abs/2206.05802 | 2022-06-12 | primary",
    "title": "Self-critiquing models for assisting human evaluators (OpenAI)",
    "url": "https://arxiv.org/abs/2206.05802",
    "date": "2022-06-12",
    "type": "primary"
   },
   "venue": {
    "peer_reviewed": false,
    "marker": "preprint"
   },
   "same_source_claims": [],
   "in_corpus_verification": null,
   "external_verification": {
    "verdict": "confirmed",
    "checked": [
     "model-written critiques help humans find flaws in summaries they would otherwise have missed",
     "surface intentional flaws in summaries humans wrote to be deliberately misleading",
     "framework comparing generation, discrimination, and critique ability",
     "'even large models may still have relevant knowledge they cannot or do not articulate as critiques' (supports discriminator-critique gap)",
     "larger models write more helpful critiques and self-refine better (topic-based summarization)"
    ],
    "not_visible": [],
    "quote": "even large models may still have relevant knowledge they cannot or do not articulate as critiques",
    "notes": "All elements of the claim and evidence confirmed from the abstract, including the verbatim quote. The 'discriminator-critique gap' is not named by that phrase but is the framework's finding.",
    "checked_at": "2026-07-15",
    "retrieval": "abstract"
   },
   "source_retrieval_meta": {
    "retrieval": "abstract",
    "fetched": [
     "https://arxiv.org/abs/2206.05802"
    ],
    "resolved_title": "Self-critiquing models for assisting human evaluators",
    "resolved_date": "2022-06-12 (v1); v2 2022-06-14",
    "title_match": true
   },
   "used_on": [],
   "decision_use": {
    "grade": "A",
    "basis": "human science (direction)",
    "measured_on": "human-work",
    "note": "2022 foundational result: model critiques help humans find flaws they'd miss. Floor-type: strengthens with critic quality.",
    "models_measured": []
   }
  },
  {
   "id": "CM-12",
   "domain": "critique-models",
   "area": null,
   "claim": "A deployed LLM compliance-checker was gamed by its subjects in a single shot: the NeurIPS 2024 author checklist assistant (234 papers) was rated useful by >70% of authors, but the organizers found the system 'not robust to gaming' by authors and concluded it is a poor substitute for human review.",
   "load_bearing": false,
   "evidence": "Goldberg, Ullah, Guyon, Shah et al. (arXiv 2411.03417; NeurIPS blog Dec 2024): GPT-4-turbo checklist verification offered pre-submission; surveys plus qualitative analysis showed genuine improvements in some submissions alongside demonstrated manipulation of the assistant (authors could satisfy the checker without satisfying the standard); access was restricted to authors partly to avoid biasing review. Verified via arXiv abstract and organizer blog summary.",
   "implication": "Q8's decision-boundary-leak concern has deployment evidence even without repeated feedback cycles: pay/acceptance-motivated subjects will optimize against a visible checker. Defenses (rotating seeded probes, held-out human-graded golden items to detect AutoQA-vs-gold divergence, abstracted rather than checker-revealing feedback) must be designed in from day one, not retrofitted.",
   "source": {
    "raw": "Usefulness of LLMs as an Author Checklist Assistant for Scientific Papers: NeurIPS'24 Experiment | https://arxiv.org/abs/2411.03417 | 2024-11-05 | academic",
    "title": "Usefulness of LLMs as an Author Checklist Assistant for Scientific Papers: NeurIPS'24 Experiment",
    "url": "https://arxiv.org/abs/2411.03417",
    "date": "2024-11-05",
    "type": "academic"
   },
   "venue": {
    "peer_reviewed": true,
    "marker": "NeurIPS"
   },
   "same_source_claims": [],
   "in_corpus_verification": null,
   "external_verification": {
    "verdict": "partially_confirmed",
    "checked": [
     "234 papers voluntarily submitted to the LLM-based Checklist Assistant",
     "over 70% of authors found the assistant useful",
     "gaming vulnerability: 'could be manipulated to enhance scores through fabricated justifications'"
    ],
    "not_visible": [
     "specific model 'GPT-4-turbo' (the abstract does not name the model)",
     "conclusion that it is 'a poor substitute for human review' (abstract instead frames it as 'a promising, but controversial, tool in aiding scientific peer review')",
     "exact phrase 'not robust to gaming' (substance present, phrase not)"
    ],
    "quote": "over 70% of authors found the assistant useful ... could be manipulated to enhance scores through fabricated justifications ... a promising, but controversial, tool in aiding scientific peer review",
    "notes": "Core (234 papers, >70% useful, demonstrated gaming/manipulation) confirmed. But the claim's characterization that organizers 'concluded it is a poor substitute for human review' is not supported by the retrieved abstract, which is more favorable ('promising, but controversial'); and the GPT-4-turbo model is not named on the page. 'Gamed in a single shot' is an editorial gloss not in the text.",
    "checked_at": "2026-07-15",
    "retrieval": "abstract"
   },
   "source_retrieval_meta": {
    "retrieval": "abstract",
    "fetched": [
     "https://arxiv.org/abs/2411.03417"
    ],
    "resolved_title": "Usefulness of LLMs as an Author Checklist Assistant for Scientific Papers: NeurIPS'24 Experiment",
    "resolved_date": "2024-11-05 (v1); v2 2024-11-08",
    "title_match": true
   },
   "used_on": [],
   "decision_use": {
    "grade": "A",
    "basis": "adversarial floor",
    "measured_on": "human-work",
    "note": "Deployed checklist assistant gamed in one shot (NeurIPS 2024). Attackers improve with model tier; the design lesson is permanent.",
    "models_measured": [
     "GPT-4"
    ]
   }
  },
  {
   "id": "GF-01",
   "domain": "grounding-faithfulness",
   "area": null,
   "claim": "Grounded claim-verification is commoditized but ceilinged: the best model on the LLM-AggreFact benchmark (11 datasets of claim-vs-grounding-document verification) is a specialized 7B model, Bespoke-MiniCheck-7B, at 77.4 average balanced accuracy, with 0.4-0.8B specialized checkers (FactCG, MiniCheck-Flan-T5-L) within ~2.5 points of frontier LLMs like Claude-3.5-Sonnet (77.2) and GPT-4o (75.9).",
   "load_bearing": true,
   "evidence": "Fetched the live LLM-AggreFact leaderboard: top-10 shows Bespoke-MiniCheck-7B 77.4, Claude-3.5-Sonnet 77.2, Granite Guardian 3.3 8B 76.5, FactCG-DeBERTa-L (0.4B) 75.6, MiniCheck-Flan-T5-L (0.8B) 75.0, Llama-3.1-405B 74.4. Benchmark measures binary supported/unsupported vs grounding docs. Confirmed independently by Paladin-mini (June 2025 arXiv) citing Bespoke-MiniCheck as leaderboard SOTA.",
   "implication": "A cheap, deterministic, reproducible entailment-check stage (sub-1B to 7B checker) for 'is this attempter statement supported by the cited evidence' is off-the-shelf and costs ~1/100 of frontier calls - but ~22% claim-level error means item verdicts cannot be a naive AND over claim checks; error-tolerant aggregation and confidence-routing to humans are structurally required.",
   "source": {
    "raw": "LLM-AggreFact Leaderboard (MiniCheck project) | https://llm-aggrefact.github.io/ | 2025 (accessed 2026-07-14) | primary",
    "title": "LLM-AggreFact Leaderboard (MiniCheck project)",
    "url": "https://llm-aggrefact.github.io/",
    "date": "2025 (accessed 2026-07-14)",
    "type": "primary"
   },
   "venue": {
    "peer_reviewed": null,
    "marker": "primary"
   },
   "same_source_claims": [],
   "in_corpus_verification": {
    "verdict": "confirmed",
    "transcript": "Fetched https://llm-aggrefact.github.io/ on 2026-07-14; every number matches exactly: Bespoke-MiniCheck-7B 77.4 (rank 1), Claude-3.5 Sonnet 77.2, Granite Guardian 3.3 8B 76.5, gpt-4o-2024-05-13 75.9, FactCG-DeBERTa-L (0.4B) 75.6, MiniCheck-Flan-T5-L (0.8B) 75.0, Llama-3.1-405B 74.4. Benchmark is 11 datasets of grounded factuality (supported/unsupported vs grounding docs), avg balanced accuracy - as claimed. The \"within ~2.5 points\" framing is accurate and even conservative: FactCG (75.6) is only 1.6 below Claude-3.5-Sonnet and above GPT-4o. Supersession search found nothing beating 77.4: ACV (May 2026) reports 76.5 training-free; an ACL 2026 paper ranks second to a post-trained metric; Paladin-mini beats Bespoke-MiniCheck only on its own separate benchmark subsets, not LLM-AggreFact overall. Caveats on the interpretive \"ceilinged\" framing: (1) the leaderboard's frontier entries are 2024-era models (Claude-3.5-Sonnet, gpt-4o-2024-05-13, Llama-3.1-405B) - no Claude 4/GPT-5-class/o3 entries, so \"specialized ~ frontier\" reflects 2024 frontiers; (2) \"Verifying the Verifiers\" (arXiv 2506.13342, June 2025) finds ~16% of benchmark labels ambiguous/incorrect and that few-shot frontier LLMs reach top-tier performance, suggesting the ~77 plateau is partly benchmark label noise rather than a pure task ceiling; (3) \"Verify with Caution\" (arXiv 2501.14883) shows models with similar aggregate BAcc make very different instance-level predictions. None of these contradict the stated facts.",
    "corrected": null
   },
   "external_verification": {
    "verdict": "confirmed",
    "checked": [
     "Bespoke-MiniCheck-7B tops the board at 77.4 average",
     "Claude-3.5 Sonnet 77.2",
     "Granite Guardian 3.3 8B 76.5",
     "gpt-4o-2024-05-13 75.9",
     "FactCG-DeBERTa-L (0.4B) 75.6",
     "MiniCheck-Flan-T5-L (0.8B) 75.0",
     "Llama-3.1-405B-Instruct 74.4",
     "11 datasets; grounded factuality / hallucination (claim supported vs source documents)"
    ],
    "not_visible": [],
    "quote": "Aggregates 11 datasets on grounded factuality (i.e., hallucination) evaluation",
    "notes": "All seven cited leaderboard numbers and the 11-dataset scope match exactly. 0.4-0.8B specialized checkers (75.6, 75.0) are within ~2.4 points of top model (77.4), consistent with the claim's '~2.5 points'.",
    "checked_at": "2026-07-15",
    "retrieval": "landing"
   },
   "source_retrieval_meta": {
    "retrieval": "landing",
    "fetched": [
     "https://llm-aggrefact.github.io/"
    ],
    "resolved_title": "LLM-AggreFact Leaderboard",
    "resolved_date": "no explicit date (live leaderboard; accessed 2026-07-15)",
    "title_match": true
   },
   "used_on": [
    {
     "page": "research.html",
     "anchor": "research-cards",
     "label": "Research - findings by decision weight"
    },
    {
     "page": "architecture.html",
     "anchor": "ontology",
     "label": "System - claim types"
    },
    {
     "page": "decisions.html",
     "anchor": "d4",
     "label": "Decisions - judge sourcing"
    }
   ],
   "decision_use": {
    "grade": "C",
    "basis": "superseded-gen capability",
    "measured_on": "model-outputs",
    "note": "The ~77 bacc AggreFact ceiling and specialists-within-2.5-points are GPT-4o/Llama-3.1-era. Deployment-tier entailment ceilings are unknown; the two-tier screening economics need repricing at the bakeoff.",
    "models_measured": [
     "MiniCheck",
     "Claude-3.5-Sonnet ",
     "GPT-4o",
     "Claude-3.5-Sonnet 77.2",
     "Llama-3.1-405B"
    ]
   }
  },
  {
   "id": "GF-02",
   "domain": "grounding-faithfulness",
   "area": null,
   "claim": "On adversarially-hard hallucination sets, specialized detectors collapse and few-shot anchoring with human-annotated exemplars is the measured fix: on FaithBench, prior detectors hit ~50% accuracy (negligible), HHEM-2.1-Open 66.7% and Bespoke-MiniCheck 71.2% balanced accuracy, zero-shot frontier judges stay below 78%, while FaithJudge - prompting o3-mini-high with human-annotated peer responses to the same source document - reaches 84.0% balanced accuracy / 82.1 F1.",
   "load_bearing": true,
   "evidence": "Fetched full text of arXiv:2505.04847v2 (Vectara, EMNLP 2025 Industry Track, v2 Nov 2025). Table 1: baseline detector numbers; Table 2: FaithJudge results; FaithJudge beats the FACTS Grounding judging prompt head-to-head on all four RAGTruth/FaithBench splits (e.g., FaithBench 70.8 vs 54.3 F1). Documented failure modes: judge underpredicts hallucinations for some generator families; 'benign'/'questionable' ternary labels classified unreliably so only binary is used; specificity drops as more in-context examples are added.",
   "implication": "The strongest 2025-measured grounding-judge architecture is exactly the AutoQA rubric-compilation contract: a per-project pool of human-adjudicated exemplar judgments (annotated spans + labels on comparable items) injected few-shot, not zero-shot judging and not a fixed fine-tuned checker. Also: severity-graded verdicts (benign/questionable) are where both humans and judges lose reliability - keep the machine verdict binary and treat severity as a separate, human-anchored layer.",
   "source": {
    "raw": "Benchmarking LLM Faithfulness in RAG with Evolving Leaderboards (FaithJudge) | https://arxiv.org/abs/2505.04847 | 2025-11 | academic",
    "title": "Benchmarking LLM Faithfulness in RAG with Evolving Leaderboards (FaithJudge)",
    "url": "https://arxiv.org/abs/2505.04847",
    "date": "2025-11",
    "type": "academic"
   },
   "venue": {
    "peer_reviewed": true,
    "marker": "EMNLP"
   },
   "same_source_claims": [],
   "in_corpus_verification": {
    "verdict": "partially_confirmed",
    "transcript": "Verified against arXiv:2505.04847 (v2 dated Nov 6, 2025; EMNLP 2025 Industry Track; Vectara/Waterloo authors - source and date correct). Confirmed: (1) paper states prior detectors, including LLM classifiers, achieved \"near 50% accuracy\" on FaithBench; (2) HHEM-2.1-Open = 66.7% balanced accuracy on FaithBench (claim-wise, Table 1) - though flagged with an asterisk because HHEM was used to adversarially select FaithBench articles; (3) FaithJudge with o3-mini-high = 84.0% balanced accuracy / 82.1 F1-macro (Table 2), best zero-shot judge on FaithBench was o3-mini-high at 68.8%; (4) FaithJudge vs FACTS Grounding head-to-head 70.8 vs 54.3 F1 on FaithBench confirmed (Table 4); (5) all three stated failure modes confirmed (underprediction for Command-R/Mistral/Qwen generators; Benign/Questionable misclassified - only 10/84 Benign labeled correctly; specificity slightly decreases as in-context examples increase). ERRORS: (a) Bespoke-MiniCheck's 71.2 is NOT its FaithBench score - 71.2 is its balanced-accuracy AVERAGE across all four datasets (AggreFact, RAGTruth, TofuEval-MB, FaithBench) in Table 1; its actual FaithBench score is 60.1% claim-wise / 55.7% summary-wise. (b) Minor: the \"below 78% balanced accuracy\" figure for zero-shot judges is the paper's cross-dataset average claim, not a FaithBench-specific figure (on FaithBench zero-shot judges max out at 68.8%, so the claim still holds directionally). Supersession check: searches found no 2026 work surpassing FaithJudge on FaithBench; it remains the reported state of the art as of July 2026.",
    "corrected": "On adversarially-hard hallucination sets, specialized detectors collapse and few-shot anchoring with human-annotated exemplars is the measured fix: on FaithBench, prior detectors hit ~50% accuracy (negligible); fine-tuned detectors stay weak (HHEM-2.1-Open 66.7% and Bespoke-MiniCheck 60.1% claim-wise balanced accuracy on FaithBench; Bespoke-MiniCheck averages 71.2% across the four benchmark datasets); zero-shot frontier judges reach at most 68.8% on FaithBench (and stay below 78% averaged across datasets), while FaithJudge - prompting o3-mini-high with human-annotated peer responses to the same source document - reaches 84.0% balanced accuracy / 82.1 F1-macro."
   },
   "external_verification": {
    "verdict": "confirmed",
    "checked": [
     "prior detectors ~50% on FaithBench: 'current methods ... achieved near 50% accuracy, suggesting negligible ability'",
     "HHEM-2.1-Open 66.7 (FaithBench claim-wise, asterisked value)",
     "Bespoke-MiniCheck 71.2 balanced accuracy (Table 1)",
     "zero-shot frontier judges: 'balanced accuracy below 78% and F1-macro below 72%'",
     "FaithJudge (o3-mini-high) 84.0 balanced accuracy / 82.1 F1-macro (Table 2 and text)",
     "beats FACTS Grounding prompt on all four splits; FaithBench-Summary F1 54.3 (FACTS) vs 70.8 (FaithJudge) (Table 4)",
     "failure modes: underprediction for Command-R/Mistral/Qwen; Benign/Questionable ternary unreliable so binary only; specificity slightly decreases as more examples given",
     "EMNLP Industry Track 2025; v2 Nov 6 2025 (matches source_date 2025-11)"
    ],
    "not_visible": [],
    "quote": "The highest effectiveness is achieved using the o3-mini-high judge, reaching a balanced accuracy of 84% and an F1-macro of 82.1%",
    "notes": "Full-text verification confirms every asserted number, upgrading the prior in_corpus_verdict (partially_confirmed) to confirmed. One nuance: the 66.7 for HHEM-2.1-Open is its FaithBench-specific claim-wise value (asterisked because HHEM helped select FaithBench's adversarial articles); its cross-dataset average is 67.1. Claim asserted 66.7, which matches the FaithBench value.",
    "checked_at": "2026-07-15",
    "retrieval": "fulltext"
   },
   "source_retrieval_meta": {
    "retrieval": "fulltext",
    "fetched": [
     "https://arxiv.org/abs/2505.04847",
     "https://arxiv.org/html/2505.04847v2"
    ],
    "resolved_title": "Benchmarking LLM Faithfulness in RAG with Evolving Leaderboards",
    "resolved_date": "2025-05-07 (v1); 2025-11-06 (v2)",
    "title_match": true
   },
   "used_on": [
    {
     "page": "decisions.html",
     "anchor": "d4",
     "label": "Decisions - judge sourcing"
    }
   ],
   "decision_use": {
    "grade": "B",
    "basis": "current-gen architecture result",
    "measured_on": "model-outputs",
    "note": "FaithJudge (late-2025, o3-mini-era). Exemplar-anchoring-beats-zero-shot is an architecture effect expected to persist; margins are tier-bound.",
    "models_measured": [
     "MiniCheck",
     "o3-mini"
    ]
   }
  },
  {
   "id": "GF-03",
   "domain": "grounding-faithfulness",
   "area": null,
   "claim": "Claim decomposition helps weak verifiers but actively degrades strong ones: with MiniCheck as verifier on WiCE, no-decomposition scores 80.01 balanced accuracy while FActScore-style atomic decomposition drops it to 71.11; with the weaker AlignScore verifier the same decomposition improves results - and gains only reappear as input complexity grows, with best results when sub-claim count does not exceed input complexity.",
   "load_bearing": true,
   "evidence": "Fetched full text of 'Decomposition Dilemmas' (NAACL 2025, arXiv:2411.02400). Four-way decomposition-error taxonomy from manual inspection: (A) omission of context/logical relations, (B) ambiguity (unclear pronouns/references), (C) over-decomposition, (D) alteration of original meaning. FActScore-style atomicity produces the most over-decomposition errors; VeriScore-style tends to omit context. On FELM, decomposition raised MiniCheck F1 48.1->~68 but dropped GPT-4o-mini F1 71.6->54.3.",
   "implication": "Directly answers the granularity design question: there is no universally-best decomposition level. Granularity must be tuned per verifier strength and per item complexity - atomic-claim pipelines with a strong judge are measurably WORSE than judging larger spans. The 'checklist myopia' risk is real and has a named mechanism (context omission + meaning alteration). Budget a granularity-calibration step per project rather than fixing atomic decomposition in the foundation.",
   "source": {
    "raw": "Decomposition Dilemmas: Does Claim Decomposition Boost or Burden Fact-Checking Performance? | https://aclanthology.org/2025.naacl-long.320/ | 2025-05 | academic",
    "title": "Decomposition Dilemmas: Does Claim Decomposition Boost or Burden Fact-Checking Performance?",
    "url": "https://aclanthology.org/2025.naacl-long.320/",
    "date": "2025-05",
    "type": "academic"
   },
   "venue": {
    "peer_reviewed": true,
    "marker": "ACL"
   },
   "same_source_claims": [],
   "in_corpus_verification": {
    "verdict": "confirmed",
    "transcript": "Verified against the full PDF (extracted text at /private/tmp/claude-501/-/640d77df-bd94-4e91-8873-f5bb2df7b27d/scratchpad/paper.txt). (1) Source says exactly this: Table 2 (WiCE) shows MiniCheck baseline BAcc 80.01 vs FActScore decomposition 71.11 (F1 72.32 vs 59.90); all decomposition methods hurt MiniCheck on WiCE. With the weaker AlignScore verifier, FActScore decomposition improves BAcc 54.80->56.87 (and WiCE-style 56.26); note VeriScore decomposition slightly hurt AlignScore too, but the claim as stated (\"the same decomposition,\" i.e., FActScore-style) is accurate. Section 4.2 states verbatim that \"decomposition generally benefits weaker verifiers, while it tends to negatively affect stronger verification systems.\" Section 6.2-6.4 confirms gains reappear as input complexity grows (complexity scale-up experiments; FELMshort scale-down degrades), and Figure 3 discussion states \"for each level, the maximum F1 is observed when the number of the decomposed sub-claim is less than or equal to the complexity level.\" Evidence-summary side facts also check: FELM MiniCheck F1 48.10->67.5-68.1, GPT-4o-mini F1 71.56->54.34 (Table 3); four-way error taxonomy (context omission, ambiguity, over-decomposition, meaning alteration) present in Section 5. (2) Date: ACL Anthology lists April 2025 (NAACL 2025, Albuquerque, pp. 6313-6336); the conference ran Apr 29-May 4, so \"2025-05\" is a trivial one-month imprecision, not a substantive error. (3) Supersession search: later work (e.g., presupposition-free question decomposition, arXiv 2508.16838; \"Alignment Bottleneck in Decomposition-Based Claim Verification,\" Feb 2026) extends the decomposition-tradeoff line but does not refute the verifier-strength finding.",
    "corrected": null
   },
   "external_verification": {
    "verdict": "partially_confirmed",
    "checked": [
     "'Decompose-Then-Verify paradigm' present",
     "authors 'introduce a categorization of decomposition errors and reveal a trade-off between accuracy gains and the noise introduced'"
    ],
    "not_visible": [
     "MiniCheck; WiCE; 80.01 vs 71.11 balanced accuracy",
     "AlignScore; FActScore; VeriScore; FELM",
     "F1 48.1->~68 and 71.6->54.3",
     "the specific 'helps weak verifiers / degrades strong ones' direction",
     "the four-way A/B/C/D decomposition-error taxonomy",
     "'best results when sub-claim count does not exceed input complexity'"
    ],
    "quote": "introduce a categorization of decomposition errors and reveal a trade-off between accuracy gains and the noise introduced",
    "notes": "Fetch constraints allow only the given ACL landing page (no PDF or other same-site page, and the URL contains no arXiv ID). The abstract confirms the general trade-off framing and that decomposition errors are categorized, but none of the specific verifiers, datasets, or numbers appear, and the abstract does not explicitly state the weak-vs-strong-verifier direction. Everything specific in the claim is not visible in the retrieved text.",
    "checked_at": "2026-07-15",
    "retrieval": "landing"
   },
   "source_retrieval_meta": {
    "retrieval": "landing",
    "fetched": [
     "https://aclanthology.org/2025.naacl-long.320/"
    ],
    "resolved_title": "Decomposition Dilemmas: Does Claim Decomposition Boost or Burden Fact-Checking Performance?",
    "resolved_date": "NAACL 2025",
    "title_match": true
   },
   "used_on": [
    {
     "page": "architecture.html",
     "anchor": "ontology",
     "label": "System - claim types"
    }
   ],
   "decision_use": {
    "grade": "C",
    "basis": "superseded-gen capability",
    "measured_on": "model-outputs",
    "note": "MiniCheck/GPT-4o-mini-era decomposition harm. Adaptive granularity stays the design default; the 80->71 magnitude is not a current estimate.",
    "models_measured": [
     "MiniCheck",
     "GPT-4o-mini"
    ]
   }
  },
  {
   "id": "GF-04",
   "domain": "grounding-faithfulness",
   "area": null,
   "claim": "Decomposition and decontextualization are in direct tension - isolating atomic facts strips the context needed to verify them, while adding context back creates multi-fact claims where the verifier may credit or penalize the wrong content; DnDScore resolves this by verifying the original subclaim WITH the added information treated as context rather than as content to be verified.",
   "load_bearing": false,
   "evidence": "Fetched ACL Anthology abstract of DnDScore (Wanner, Van Durme, Dredze; EMNLP 2025 main, pp. 23609-23626). The paper evaluates combinations of decomposition, decontextualization, and verification strategies and finds the strategy choice materially changes factuality scores.",
   "implication": "When the AutoQA checks an attempter's sentence, the verification unit should be 'claim + explicit context annotations' with the verifier told which part is under test - not a free-floating atomic rewrite. This is a concrete spec for the claim-extraction layer's output schema.",
   "source": {
    "raw": "DnDScore: Decontextualization and Decomposition for Factuality Verification in Long-Form Text Generation | https://aclanthology.org/2025.emnlp-main.1205/ | 2025-11 | academic",
    "title": "DnDScore: Decontextualization and Decomposition for Factuality Verification in Long-Form Text Generation",
    "url": "https://aclanthology.org/2025.emnlp-main.1205/",
    "date": "2025-11",
    "type": "academic"
   },
   "venue": {
    "peer_reviewed": true,
    "marker": "ACL"
   },
   "same_source_claims": [],
   "in_corpus_verification": null,
   "external_verification": {
    "verdict": "confirmed",
    "checked": [
     "decomposition and decontextualization have conflicting purposes",
     "decomposition isolates atomic facts while decontextualization inserts relevant information",
     "adding context back creates multi-fact text ('what part of the augmented text should be verified')",
     "DnDScore validates subclaims in the context of contextual information (context, not content)",
     "strategy choice materially changes factuality scores",
     "pages 23609-23626; authors Wanner, Van Durme, Dredze"
    ],
    "not_visible": [],
    "quote": "decomposition isolates atomic facts while decontextualization inserts relevant information ... they present \"DnDScore, a decontextualization aware verification method that validates subclaims in the context of contextual information\" ... \"the choice of strategy matters in the resulting factuality scores.\"",
    "notes": "Core tension and the resolution (verify subclaim WITH added info as context rather than content) confirmed verbatim from the ACL Anthology abstract. Title, authors, pages all match.",
    "checked_at": "2026-07-15",
    "retrieval": "abstract"
   },
   "source_retrieval_meta": {
    "retrieval": "abstract",
    "fetched": [
     "https://aclanthology.org/2025.emnlp-main.1205/"
    ],
    "resolved_title": "DnDScore: Decontextualization and Decomposition for Factuality Verification in Long-Form Text Generation",
    "resolved_date": "EMNLP 2025 (November 2025), pp. 23609-23626",
    "title_match": true
   },
   "used_on": [
    {
     "page": "architecture.html",
     "anchor": "ontology",
     "label": "System - claim types"
    }
   ],
   "decision_use": {
    "grade": "B",
    "basis": "current-gen mechanism",
    "measured_on": "model-outputs",
    "note": "2025 decomposition/decontextualization tension analysis. Design-level tension, partially structural.",
    "models_measured": []
   }
  },
  {
   "id": "GF-05",
   "domain": "grounding-faithfulness",
   "area": null,
   "claim": "Verifiability triage before entailment is established machinery: VeriScore extracts and scores ONLY verifiable claims because FActScore/SAFE 'assume that every claim is verifiable', which breaks on real long-form text containing opinions and unverifiable content; FactBench's VERIFY pipeline (ACL 2025) further labels content units supported/unsupported/UNDECIDABLE against retrieval, with 4,467 human-annotated units released for validation.",
   "load_bearing": false,
   "evidence": "Fetched arXiv:2406.19276 abstract (VeriScore, EMNLP Findings 2024): human evaluation found VeriScore's extracted claims 'more sensible' than competitors across 8 long-form tasks. Fetched launchnlp/FactBench GitHub (ACL 2025 paper 2410.22257): three-way supported/unsupported/undecidable labeling according to retrieval results, benchmarked against FActScore, SAFE, Factcheck-GPT.",
   "implication": "The claim ontology question has a field-standard answer: a front-stage classifier that routes each attempter statement into {verifiable claim, opinion/unverifiable, undecidable-given-evidence} BEFORE any entailment check. Opinion-class statements get consistency-with-cited-evidence checks at most, never truth verdicts; 'undecidable' is a first-class output, not a forced pass/fail.",
   "source": {
    "raw": "VeriScore (EMNLP Findings 2024) + FactBench/VERIFY (ACL 2025) | https://arxiv.org/abs/2406.19276 | 2024-06 (VeriScore, foundational); 2025-07 (FactBench ACL 2025) | academic",
    "title": "VeriScore (EMNLP Findings 2024) + FactBench/VERIFY (ACL 2025)",
    "url": "https://arxiv.org/abs/2406.19276",
    "date": "2024-06 (VeriScore, foundational); 2025-07 (FactBench ACL 2025)",
    "type": "academic"
   },
   "venue": {
    "peer_reviewed": true,
    "marker": "ACL"
   },
   "same_source_claims": [],
   "in_corpus_verification": null,
   "external_verification": {
    "verdict": "partially_confirmed",
    "checked": [
     "prior methods FActScore/SAFE 'assume that every claim is verifiable (i.e., can plausibly be proven true or false)'",
     "VERISCORE handles 'both verifiable and unverifiable content'",
     "human evaluation confirms that VERISCORE's extracted claims are more sensible than those from competing methods",
     "across eight different long-form tasks"
    ],
    "not_visible": [
     "FactBench and its VERIFY pipeline -- the words 'FactBench' and 'VERIFY' do not appear on this page",
     "three-way supported/unsupported/UNDECIDABLE labeling against retrieval",
     "4,467 human-annotated content units",
     "ACL 2025 venue for FactBench",
     "'EMNLP Findings 2024' venue -- the arXiv page shows no conference venue, only arXiv"
    ],
    "quote": "they assume that every claim is verifiable (i.e., can plausibly be proven true or false) ... human evaluation confirms that VERISCORE's extracted claims are more sensible than those from competing methods",
    "notes": "The VeriScore half of this compound claim is fully confirmed. The FactBench/VERIFY half (three-way labeling, 4,467 units, ACL 2025) comes from a separate paper (2410.22257) that the single-URL constraint prohibits fetching, so it is not present in retrieved text. Minor: claim says VeriScore scores 'ONLY verifiable claims' while the abstract frames it as handling 'both verifiable and unverifiable content'. title_match=false because source_title is a composite two-paper label, not the actual paper title, and asserts an EMNLP venue not shown on the page.",
    "checked_at": "2026-07-15",
    "retrieval": "abstract"
   },
   "source_retrieval_meta": {
    "retrieval": "abstract",
    "fetched": [
     "https://arxiv.org/abs/2406.19276"
    ],
    "resolved_title": "VERISCORE: Evaluating the factuality of verifiable claims in long-form text generation",
    "resolved_date": "2024-06-27",
    "title_match": false
   },
   "used_on": [
    {
     "page": "architecture.html",
     "anchor": "ontology",
     "label": "System - claim types"
    },
    {
     "page": "decisions.html",
     "anchor": "c1",
     "label": "Decisions - default: entailment standard"
    }
   ],
   "decision_use": {
    "grade": "A",
    "basis": "method / typology",
    "measured_on": "structural",
    "note": "Verifiability triage before entailment (VeriScore/FactBench). A pipeline design pattern, not a capability claim.",
    "models_measured": []
   }
  },
  {
   "id": "GF-06",
   "domain": "grounding-faithfulness",
   "area": null,
   "claim": "Holistic LLM judging of context-grounded outputs is far weaker than claim-level checking: on ContextualJudgeBench (2,000 pairs, 8 splits over RAG/summarization with conditional criteria like 'faithfulness first, then completeness'), the best of 20 judge models tested (OpenAI o1) barely reaches 55% consistent accuracy.",
   "load_bearing": false,
   "evidence": "Fetched arXiv:2503.15620 (Salesforce; ACL 2025). 11 specialized judge models + 9 general-purpose models evaluated; the benchmark encodes criteria hierarchies and finds contextual assessment poses a significant challenge even to SOTA models. Contrast: the same model class hits 77-84% on single-claim grounding tasks (LLM-AggreFact, FaithJudge).",
   "implication": "Do not architect the AutoQA as one holistic 'judge this writeup against the rubric' call - measured reliability roughly halves versus decomposed per-criterion, per-claim checks. Criteria priority ordering (which humans dispute) must be compiled explicitly into the pipeline structure, because judges fail at applying conditional criteria internally.",
   "source": {
    "raw": "Does Context Matter? ContextualJudgeBench for Evaluating LLM-based Judges in Contextual Settings | https://arxiv.org/abs/2503.15620 | 2025-03 | academic",
    "title": "Does Context Matter? ContextualJudgeBench for Evaluating LLM-based Judges in Contextual Settings",
    "url": "https://arxiv.org/abs/2503.15620",
    "date": "2025-03",
    "type": "academic"
   },
   "venue": {
    "peer_reviewed": true,
    "marker": "ACL"
   },
   "same_source_claims": [],
   "in_corpus_verification": null,
   "external_verification": {
    "verdict": "confirmed",
    "checked": [
     "2,000 challenging response pairs",
     "eight (8) splits",
     "20 judge models (11 judge models + 9 general-purpose models)",
     "OpenAI o1 best-performing, barely reaches 55% consistent accuracy",
     "conditional evaluation criteria (factuality then completeness)",
     "RAG and summarization settings"
    ],
    "not_visible": [
     "Salesforce affiliation (not stated on the abstract page)",
     "ACL 2025 venue (not listed on the abstract page)"
    ],
    "quote": "we propose ContextualJudgeBench, a judge benchmark with 2,000 challenging response pairs across eight splits ... Our comprehensive study across 11 judge models and 9 general purpose models ... OpenAI's o1, the best-performing model, barely reaches 55% consistent accuracy.",
    "notes": "All core numbers confirmed. Minor terminology: the claim's illustrative criterion 'faithfulness first, then completeness' appears in the abstract as 'factuality and then considering completeness' (factuality vs faithfulness). The evidence field's 'Salesforce; ACL 2025' is not visible on the abstract page but is not part of the core claim.",
    "checked_at": "2026-07-15",
    "retrieval": "abstract"
   },
   "source_retrieval_meta": {
    "retrieval": "abstract",
    "fetched": [
     "https://arxiv.org/abs/2503.15620"
    ],
    "resolved_title": "Does Context Matter? ContextualJudgeBench for Evaluating LLM-based Judges in Contextual Settings",
    "resolved_date": "2025-03-19",
    "title_match": true
   },
   "used_on": [],
   "decision_use": {
    "grade": "C",
    "basis": "superseded-gen capability",
    "measured_on": "model-outputs",
    "note": "ContextualJudgeBench ~55% on o1-era judges. Holistic-vs-claim-level gap direction is consistent across studies; magnitude stale.",
    "models_measured": [
     "o1"
    ]
   }
  },
  {
   "id": "GF-07",
   "domain": "grounding-faithfulness",
   "area": null,
   "claim": "Grounding judges have a measured agreement-default asymmetry: in Google's FACTS Leaderboard (Dec 2025), grounding judges score ~85 F1 on the positive (grounded) class but only ~46 F1 on the negative (ungrounded) class, and the grounding metric is gameable by vague, short responses that avoid unsupported claims - countered by a mandatory eligibility gate that scores non-responsive answers as inaccurate.",
   "load_bearing": true,
   "evidence": "Fetched full text of arXiv:2512.10791v1 (The FACTS Leaderboard, Google, Dec 2025). Best judge-prompt combos reach only ~65 macro-F1 (gemini-2.5-flash + v2 prompt: 65.33) against a 320-example human-adjudicated held-out set; positive-class ~85 vs negative-class ~46 F1; eligibility check explicitly added because 'grounding metrics can be hacked via vague, short responses'.",
   "implication": "Two foundational mandates: (1) the sycophancy/agreement-default bias on positive claims is quantified and large - the AutoQA must be evaluated on per-class recall (especially fail-class recall at fixed prevalence), never single-number accuracy; (2) an eligibility/responsiveness gate must precede grounding checks or attempters can pass by writing vacuous hedged annotations - the exact adversarial dynamic the design taxonomy flags.",
   "source": {
    "raw": "The FACTS Leaderboard: A Comprehensive Benchmark for Large Language Model Factuality | https://arxiv.org/abs/2512.10791 | 2025-12 | academic",
    "title": "The FACTS Leaderboard: A Comprehensive Benchmark for Large Language Model Factuality",
    "url": "https://arxiv.org/abs/2512.10791",
    "date": "2025-12",
    "type": "academic"
   },
   "venue": {
    "peer_reviewed": false,
    "marker": "preprint"
   },
   "same_source_claims": [],
   "in_corpus_verification": null,
   "external_verification": {
    "verdict": "confirmed",
    "checked": [
     "top combo gemini-2.5-flash + v2 prompt: Macro-F1 65.33",
     "F1(+) 84.51 vs F1(-) 46.15 (positive/negative class asymmetry, ~85 vs ~46)",
     "held-out human-adjudicated evaluation set N=320 (class ratio 79:19)",
     "eligibility gate: ineligible/non-responsive answers deemed inaccurate",
     "explicit rationale that grounding metrics can be hacked by vague, short responses"
    ],
    "not_visible": [],
    "quote": "gemini-2.5-flash with the v2 prompt, scores Macro-F1 65.33, with F1(+) of 84.51 but F1(-) of only 46.15 ... the final factuality score is adjusted such that ineligible responses are deemed as inaccurate",
    "notes": "All four asserted specifics confirmed in the full text. The claim's mapping of F1(+)/F1(-) to 'grounded'/'ungrounded' class is a reasonable interpretation (majority positive class 79:19); paper labels them F1(+)/F1(-). The ~85/~46 figures are the top combo's, which the claim uses to characterize grounding judges generally.",
    "checked_at": "2026-07-15",
    "retrieval": "fulltext"
   },
   "source_retrieval_meta": {
    "retrieval": "fulltext",
    "fetched": [
     "https://arxiv.org/html/2512.10791v1"
    ],
    "resolved_title": "The FACTS Leaderboard: A Comprehensive Benchmark for Large Language Model Factuality",
    "resolved_date": "2025-12",
    "title_match": true
   },
   "used_on": [
    {
     "page": "index.html",
     "anchor": "evidence",
     "label": "Overview - what the evidence supports"
    },
    {
     "page": "research.html",
     "anchor": "research-cards",
     "label": "Research - findings by decision weight"
    },
    {
     "page": "architecture.html",
     "anchor": "ontology",
     "label": "System - claim types"
    },
    {
     "page": "decisions.html",
     "anchor": "c1",
     "label": "Decisions - default: entailment standard"
    }
   ],
   "decision_use": {
    "grade": "B",
    "basis": "current-gen capability (flash-tier graders)",
    "measured_on": "model-outputs",
    "note": "FACTS (Dec 2025, gemini-2.5-flash graders). The agreement-default asymmetry (85/46) is the single most decision-relevant grounding mechanism - and it was measured on flash-tier graders. Deployment-tier asymmetry is E2's first number.",
    "models_measured": [
     "gemini-2.5-flash"
    ]
   }
  },
  {
   "id": "GF-08",
   "domain": "grounding-faithfulness",
   "area": null,
   "claim": "Cross-family judge ensembles are the standard mitigation for self-preference: FACTS Grounding v1 (Jan 2025) aggregated three frontier judges from different families (Gemini 1.5 Pro, GPT-4o, Claude 3.5 Sonnet) explicitly because 'models are biased towards favorably judging their own outputs', and Grounding v2 (Dec 2025) kept a two-family ensemble (Gemini 2.5 Flash + GPT-5).",
   "load_bearing": false,
   "evidence": "Fetched arXiv:2512.10791v1 which states the v2 judge-ensemble rationale and composition; v1 design documented in arXiv:2501.03200 (Jan 2025) and the DeepMind blog (Dec 17, 2024). FACTS Parametric suite separately validated that a single-judge setup preserved rankings vs a mixed panel (Gemini 2.5 Pro, o3, Grok 4) - i.e., ensemble necessity is task-dependent and testable.",
   "implication": "When attempters critique outputs from a given model family, the AutoQA judge should default to a different family or an ensemble spanning families; but ensemble overhead can be dropped where a single-judge-vs-panel ranking-preservation test passes - a cheap per-project validation, not a dogma.",
   "source": {
    "raw": "FACTS Grounding Leaderboard (v1 paper + v2 in FACTS Leaderboard) | https://arxiv.org/abs/2501.03200 | 2025-01 (v1); 2025-12 (v2) | academic",
    "title": "FACTS Grounding Leaderboard (v1 paper + v2 in FACTS Leaderboard)",
    "url": "https://arxiv.org/abs/2501.03200",
    "date": "2025-01 (v1); 2025-12 (v2)",
    "type": "academic"
   },
   "venue": {
    "peer_reviewed": false,
    "marker": "preprint"
   },
   "same_source_claims": [],
   "in_corpus_verification": null,
   "external_verification": {
    "verdict": "partially_confirmed",
    "checked": [
     "v1 uses three judge models from three providers: Gemini 1.5 Pro (Google), GPT-4o (OpenAI), Claude 3.5 Sonnet (Anthropic)",
     "rationale is self-preference: 'models have been shown to be biased towards favorably judging their own outputs'",
     "aggregate of multiple judge models to mitigate evaluation bias",
     "v1 dated Jan 2025"
    ],
    "not_visible": [
     "FACTS Grounding v2 (Dec 2025) two-family ensemble Gemini 2.5 Flash + GPT-5 (asserted from arXiv:2512.10791, a different paper not fetchable under the constraint)",
     "the DeepMind blog (Dec 17, 2024) v1 design write-up",
     "FACTS Parametric single-judge-preserves-rankings validation (Gemini 2.5 Pro, o3, Grok 4)"
    ],
    "quote": "We use three different judge models in order to reduce the bias of a particular judge model ... as models have been shown to be biased towards favorably judging their own outputs",
    "notes": "The v1 core (cross-provider ensemble as self-preference mitigation) is fully confirmed from the given URL's full text. Paper says 'three different judge models' from three providers rather than literally 'different families' - substantively equivalent. All v2 / Parametric specifics require arXiv:2512.10791 and external pages, which the fetch constraint excludes, so they are marked not_visible rather than checked.",
    "checked_at": "2026-07-15",
    "retrieval": "fulltext"
   },
   "source_retrieval_meta": {
    "retrieval": "fulltext",
    "fetched": [
     "https://arxiv.org/abs/2501.03200",
     "https://arxiv.org/html/2501.03200"
    ],
    "resolved_title": "The FACTS Grounding Leaderboard: Benchmarking LLMs' Ability to Ground Responses to Long-Form Input",
    "resolved_date": "2025-01-06",
    "title_match": true
   },
   "used_on": [],
   "decision_use": {
    "grade": "B",
    "basis": "current-gen practice",
    "measured_on": "model-outputs",
    "note": "FACTS v1->v2 ensemble evolution; includes GPT-5-class by v2. Cross-provider ensembling as self-preference mitigation is practice-level.",
    "models_measured": [
     "Gemini 1.5",
     "GPT-4o",
     "Claude 3.5 Sonnet",
     "Gemini 2.5",
     "GPT-5",
     "o3"
    ]
   }
  },
  {
   "id": "GF-09",
   "domain": "grounding-faithfulness",
   "area": null,
   "claim": "Span-restricted NLI citation checking (the ALCE-style 'does the cited span entail the sentence' standard) is judged a suboptimal proxy by 2025 work: CiteEval (ACL 2025) argues citation quality must be evaluated against the FULL retrieval context, user query, and generated text - not cited sources alone - and its model-based CiteEval-Auto metrics correlate better with human judgments on the multi-domain CiteBench than NLI-based metrics.",
   "load_bearing": false,
   "evidence": "Fetched arXiv:2506.01829 (CiteEval, ACL 2025 main). Explicitly frames binary/ternary NLI vs cited sources as 'a suboptimal proxy for citation evaluation'. Corroborated direction: CiteGuard (ACL 2026, aclanthology.org/2026.acl-long.282) reframes citation evaluation as attribution alignment and notes 'the reliability of LLM-as-a-Judge alone is also in doubt' for citation judging.",
   "implication": "Evidence-closure design: checking attempter claims only against the spans they cite will miss the dominant human failure of citing the wrong/incomplete evidence when better evidence existed in scope. The AutoQA needs two distinct checks - (a) does cited evidence entail the claim (span-closed), and (b) is the citation the right one given the full in-scope corpus (context-closed) - with (b) also catching true-but-wrongly-cited claims.",
   "source": {
    "raw": "CiteEval: Principle-Driven Citation Evaluation for Source Attribution | https://arxiv.org/abs/2506.01829 | 2025-06 | academic",
    "title": "CiteEval: Principle-Driven Citation Evaluation for Source Attribution",
    "url": "https://arxiv.org/abs/2506.01829",
    "date": "2025-06",
    "type": "academic"
   },
   "venue": {
    "peer_reviewed": true,
    "marker": "ACL"
   },
   "same_source_claims": [],
   "in_corpus_verification": null,
   "external_verification": {
    "verdict": "confirmed",
    "checked": [
     "NLI binary/ternary supportiveness judged 'a suboptimal proxy for citation evaluation'",
     "evaluates full retrieval context, user query, and generated text (not cited sources alone)",
     "CiteBench multi-domain human-annotated benchmark",
     "CiteEval-Auto model-based metrics with strong correlation to human judgments"
    ],
    "not_visible": [
     "'ALCE' (not named on page)",
     "explicit head-to-head numeric comparison 'better than NLI-based metrics' (entailed by 'suboptimal proxy' + 'strong correlation' but not stated as a number)",
     "CiteGuard (ACL 2026) corroboration (separate source, out of scope)"
    ],
    "quote": "current frameworks mainly rely on Natural Language Inference (NLI) to assess binary or ternary supportiveness from cited sources, which we argue is a suboptimal proxy for citation evaluation ... not only the cited sources but the full retrieval context, user query, and generated text",
    "notes": "Core assertion and every substantive element (suboptimal NLI proxy, full retrieval context, CiteBench, CiteEval-Auto human-correlation) confirmed verbatim in the abstract.",
    "checked_at": "2026-07-15",
    "retrieval": "abstract"
   },
   "source_retrieval_meta": {
    "retrieval": "abstract",
    "fetched": [
     "https://arxiv.org/abs/2506.01829"
    ],
    "resolved_title": "CiteEval: Principle-Driven Citation Evaluation for Source Attribution",
    "resolved_date": "2025-06-02",
    "title_match": true
   },
   "used_on": [
    {
     "page": "research.html",
     "anchor": "research-cards",
     "label": "Research - findings by decision weight"
    },
    {
     "page": "architecture.html",
     "anchor": "ontology",
     "label": "System - claim types"
    },
    {
     "page": "pilot.html",
     "anchor": "main",
     "label": "Pilot - gates and objectives"
    }
   ],
   "decision_use": {
    "grade": "B",
    "basis": "current-gen method critique",
    "measured_on": "model-outputs",
    "note": "2025: span-restricted citation checking is insufficient. Method-level; motivates the dual-pass citation design.",
    "models_measured": []
   }
  },
  {
   "id": "GF-10",
   "domain": "grounding-faithfulness",
   "area": null,
   "claim": "Fact-verification LLMs are brittle to semantically-minor input perturbations: FactEval (NAACL 2025) tested 17 realistic word- and character-level perturbations plus 4 subpopulations on FEVER across zero-shot/few-shot/CoT setups and found LLMs 'brittle to small input changes' with performance varying across subpopulations.",
   "load_bearing": false,
   "evidence": "Fetched ACL Anthology page for 2025.naacl-long.534 (Mamta & Cocarascu). Abstract is qualitative; per-perturbation numbers are in the full paper. Consistent with FaithJudge's finding that adding more in-context examples shifts sensitivity/specificity, and with FACTS' judge-prompt sensitivity (macro-F1 varies by prompt version).",
   "implication": "A perturbation-robustness harness (paraphrase, reorder, typo-level noise on both claims and evidence) belongs in the per-project ship gate; verdicts that flip under meaning-preserving edits should be auto-routed to human review rather than delivered.",
   "source": {
    "raw": "FactEval: Evaluating the Robustness of Fact Verification Systems in the Era of Large Language Models | https://aclanthology.org/2025.naacl-long.534/ | 2025-04 | academic",
    "title": "FactEval: Evaluating the Robustness of Fact Verification Systems in the Era of Large Language Models",
    "url": "https://aclanthology.org/2025.naacl-long.534/",
    "date": "2025-04",
    "type": "academic"
   },
   "venue": {
    "peer_reviewed": true,
    "marker": "ACL"
   },
   "same_source_claims": [],
   "in_corpus_verification": null,
   "external_verification": {
    "verdict": "confirmed",
    "checked": [
     "17 realistic word-level and character-level perturbations and 4 types of subpopulations",
     "built on FEVER",
     "zero-shot, few-shot, and chain-of-thought prompting",
     "LLMs brittle to small input changes and performance variations across subpopulations",
     "authors Mamta and Oana Cocarascu",
     "NAACL 2025"
    ],
    "not_visible": [
     "per-perturbation quantitative results (abstract is qualitative; claim does not assert specific numbers)"
    ],
    "quote": "LLMs are brittle to small input changes and also exhibit performance variations across different subpopulations.",
    "notes": "Abstract confirms every asserted specific (17 perturbations, 4 subpopulations, FEVER, three prompting setups, brittleness finding). The claim explicitly defers per-perturbation numbers to the full paper, so nothing overclaimed.",
    "checked_at": "2026-07-15",
    "retrieval": "abstract"
   },
   "source_retrieval_meta": {
    "retrieval": "abstract",
    "fetched": [
     "https://aclanthology.org/2025.naacl-long.534/"
    ],
    "resolved_title": "FactEval: Evaluating the Robustness of Fact Verification Systems in the Era of Large Language Models",
    "resolved_date": "NAACL 2025, April 2025, pp. 10647-10660",
    "title_match": true
   },
   "used_on": [
    {
     "page": "pilot.html",
     "anchor": "main",
     "label": "Pilot - gates and objectives"
    }
   ],
   "decision_use": {
    "grade": "B",
    "basis": "current-gen brittleness",
    "measured_on": "model-outputs",
    "note": "FactEval (2025): verifier verdicts flip under meaning-preserving perturbation. Perturbation harness stays a ship gate; deployment-tier flip rates unknown (could be far lower - measure, don't assume).",
    "models_measured": []
   }
  },
  {
   "id": "GF-11",
   "domain": "grounding-faithfulness",
   "area": null,
   "claim": "2026 SOTA joins verdict and explanation in one cheap model: FaithLens (ACL 2026 Findings) is an 8B faithfulness-hallucination detector that jointly outputs a binary prediction AND an explanation, trained via cold-start fine-tuning on filtered synthetic data plus rule-based RL rewarding both prediction correctness and explanation quality, and outperforms GPT-5.2 and o3 across 12 tasks.",
   "load_bearing": false,
   "evidence": "Fetched ACL Anthology entry 2026.findings-acl.689 (Si et al., Findings of ACL 2026, pp. 14068-14099). Abstract claims superiority over GPT-5.2/o3 and 'a distinctive balance of trustworthiness, efficiency, and effectiveness'; per-task numbers require the full PDF.",
   "implication": "Constructive-feedback generation and grounding verdicts do not need separate machinery: an explanation-rewarded detector produces the evidence-citing rationale as a first-class output, at 8B cost. This is the current-generation template for the AutoQA's 'every criticism must carry its evidence' contract, and shows explanation quality can be an explicit training/reward target rather than a post-hoc add-on.",
   "source": {
    "raw": "FaithLens: Detecting and Explaining Faithfulness Hallucination | https://aclanthology.org/2026.findings-acl.689/ | 2026-07 (ACL 2026) | academic",
    "title": "FaithLens: Detecting and Explaining Faithfulness Hallucination",
    "url": "https://aclanthology.org/2026.findings-acl.689/",
    "date": "2026-07 (ACL 2026)",
    "type": "academic"
   },
   "venue": {
    "peer_reviewed": true,
    "marker": "ACL"
   },
   "same_source_claims": [],
   "in_corpus_verification": null,
   "external_verification": {
    "verdict": "confirmed",
    "checked": [
     "8B faithfulness-hallucination detection model",
     "jointly outputs binary predictions plus explanations",
     "fine-tuned on filtered LLM-synthesized data as a cold start",
     "further optimized with rule-based reinforcement learning rewarding prediction correctness and explanation quality",
     "beats GPT-5.2 and o3 across 12 diverse tasks",
     "pages 14068-14099, Findings of ACL 2026, Si et al."
    ],
    "not_visible": [
     "per-task numeric scores (require full PDF; abstract only, consistent with the claim's own caveat)"
    ],
    "quote": "a distinctive balance of trustworthiness, efficiency, and effectiveness",
    "notes": "Every headline assertion in the claim is present in the ACL Anthology abstract; page range and authorship match the evidence.",
    "checked_at": "2026-07-15",
    "retrieval": "landing"
   },
   "source_retrieval_meta": {
    "retrieval": "landing",
    "fetched": [
     "https://aclanthology.org/2026.findings-acl.689/"
    ],
    "resolved_title": "FaithLens: Detecting and Explaining Faithfulness Hallucination",
    "resolved_date": "2026-07 (Findings of ACL 2026, San Diego), pp. 14068-14099",
    "title_match": true
   },
   "used_on": [],
   "decision_use": {
    "grade": "B",
    "basis": "current-gen capability",
    "measured_on": "model-outputs",
    "note": "FaithLens (ACL 2026, GPT-5.2-era comparison) - among the most current capability anchors in the corpus. Still model-output verification, not human-writeup QA.",
    "models_measured": [
     "GPT-5.2",
     "o3"
    ]
   }
  },
  {
   "id": "GF-12",
   "domain": "grounding-faithfulness",
   "area": null,
   "claim": "RAGAS-style prompt-chain faithfulness scoring materially underperforms finetuned/frontier judges on hallucination detection: on the 15K-sample HaluBench, RAGAS Faithfulness scored 66.9% accuracy versus 87.4% for finetuned Lynx-70B and 86.5% for GPT-4o; Patronus has since shipped Lynx 2.0 (8B, long-context, 8 hallucination subtypes including coreference and calculation errors).",
   "load_bearing": false,
   "evidence": "Confirmed via Patronus AI announcement and Lynx paper coverage (July 11, 2024) with the HaluBench table (Lynx-70B 87.4, GPT-4o 86.5, GPT-4-Turbo 85.0, Llama-3-70B 80.1, RAGAS 66.9); Lynx 2.0 details from Patronus docs. Caveat: vendor-reported numbers on the vendor's own benchmark, though HaluBench is public on HuggingFace. RAGAS remains the most widely adopted OSS RAG-eval framework per 2026 practitioner surveys (atlan.com RAG evaluation guide, 2026).",
   "implication": "Do not build the grounding stage on RAGAS-style statement-extraction prompt chains despite their ubiquity in 2026 tooling - the popularity/accuracy gap is ~20 points. If an off-the-shelf component is wanted, current-generation finetuned detectors (Bespoke-MiniCheck, Lynx 2.0, Granite Guardian, FaithLens) dominate at equal or lower cost.",
   "source": {
    "raw": "Patronus AI Lynx / HaluBench results | https://www.patronus.ai/blog/lynx-state-of-the-art-open-source-hallucination-detection-model | 2024-07 (Lynx 1.0, foundational); Lynx 2.0 later update | practitioner",
    "title": "Patronus AI Lynx / HaluBench results",
    "url": "https://www.patronus.ai/blog/lynx-state-of-the-art-open-source-hallucination-detection-model",
    "date": "2024-07 (Lynx 1.0, foundational); Lynx 2.0 later update",
    "type": "practitioner"
   },
   "venue": {
    "peer_reviewed": null,
    "marker": "practitioner"
   },
   "same_source_claims": [],
   "in_corpus_verification": null,
   "external_verification": {
    "verdict": "partially_confirmed",
    "checked": [
     "HaluBench is 15k samples",
     "Lynx is positioned as SOTA open-source hallucination detector, beating frontier judges (relative): Lynx-70B 8.3% more accurate than GPT-4o on PubMedQA; Lynx-8B beat GPT-3.5 by 24.5% on HaluBench"
    ],
    "not_visible": [
     "RAGAS Faithfulness 66.9%",
     "Lynx-70B 87.4%",
     "GPT-4o 86.5%",
     "GPT-4-Turbo 85.0%",
     "Llama-3-70B 80.1%",
     "the HaluBench accuracy comparison table itself (this blog reports only relative % improvements, not the absolute-accuracy table)",
     "Lynx 2.0, 8B long-context model, 8 hallucination subtypes (coreference/calculation)"
    ],
    "quote": "a comprehensive hallucination evaluation benchmark consisting of 15k samples",
    "notes": "The given URL confirms the 15k HaluBench size and the directional finding (finetuned Lynx > frontier judges), but NONE of the specific accuracy numbers cited in the claim (66.9/87.4/86.5/85.0/80.1) appear on this page; they are attributed in the evidence to the Lynx paper/table, not this blog. Lynx 2.0 and its subtypes are not mentioned here at all. Resolved title is the literal page title, which differs from the descriptive source_title.",
    "checked_at": "2026-07-15",
    "retrieval": "landing"
   },
   "source_retrieval_meta": {
    "retrieval": "landing",
    "fetched": [
     "https://www.patronus.ai/blog/lynx-state-of-the-art-open-source-hallucination-detection-model"
    ],
    "resolved_title": "Lynx: State-of-the-Art Open Source Hallucination Detection Model",
    "resolved_date": "not stated on page (July 2024 inferred from linked paper/NeMo note)",
    "title_match": false
   },
   "used_on": [],
   "decision_use": {
    "grade": "C",
    "basis": "superseded-gen capability",
    "measured_on": "model-outputs",
    "note": "Lynx-vs-GPT-4-era comparison (2024). Fine-tuned-specialist economics need current repricing.",
    "models_measured": [
     "Lynx",
     "GPT-4o",
     "GPT-4",
     "Llama-3-70B"
    ]
   }
  },
  {
   "id": "HS-01",
   "domain": "hybrid-statistical",
   "area": null,
   "claim": "Prediction-Powered Inference (PPI/PPI++) combines a small human gold set with a large judge-labeled set to produce unbiased estimates with always-valid confidence intervals - coverage holds regardless of judge quality (a worse judge widens intervals, never invalidates them) - and increased effective human sample size by up to 50% with a GPT-4 judge in 'AutoEval Done Right'.",
   "load_bearing": true,
   "evidence": "arXiv:2403.07008 (Boyeau, Angelopoulos, Yosef, Malik, Jordan, UC Berkeley) abstract confirmed: 'increase the effective human-labeled sample size by up to 50% on experiments with GPT-4'; v3 camera-ready posted 2026-06-01. The unconditional-coverage property is restated and empirically demonstrated in the GLIDE paper (arXiv:2605.31278) and is the foundation of the entire 2026 LLM-judge debiasing literature (PRECISE/AAAI 2026, arXiv:2606.05308 ranking extension).",
   "implication": "The AutoQA meta-evaluation layer should treat human gold labels as bias-correctors for judge-derived population metrics (pass rates, per-attempter/per-criterion error rates), never report raw judge numbers. Because validity survives a bad judge, pooled/hierarchical cross-project validation is statistically safe even where per-project gold sets are tiny - directly resolving taxonomy Q1's gold-set-arithmetic fork toward pooled certification with per-project bias correction.",
   "source": {
    "raw": "AutoEval Done Right: Sample-Efficient Human Evaluation via Prediction-Powered Inference | https://arxiv.org/abs/2403.07008 | 2024-03 (v3 camera-ready 2026-06-01) | academic",
    "title": "AutoEval Done Right: Sample-Efficient Human Evaluation via Prediction-Powered Inference",
    "url": "https://arxiv.org/abs/2403.07008",
    "date": "2024-03 (v3 camera-ready 2026-06-01)",
    "type": "academic"
   },
   "venue": {
    "peer_reviewed": true,
    "marker": "AAAI"
   },
   "same_source_claims": [],
   "in_corpus_verification": {
    "verdict": "partially_confirmed",
    "transcript": "Verified directly against arXiv:2403.07008: authors (Boyeau, Angelopoulos, Yosef, Malik, Jordan) correct; abstract verbatim contains 'increase the effective human-labeled sample size by up to 50% on experiments with GPT-4'; abstract claims methods 'improve sample efficiency while remaining unbiased'; version history confirms v1 2024-03-09 and v3 2026-06-01 with comment 'camera-ready paper version' - dates as claimed. GLIDE (arXiv:2605.31278, Martinon/Merad/Raki, v1 2026-05-29, ICML 2026 workshop) exists and states PPI 'combines both into debiased estimates with valid confidence intervals', consistent with the claim, though the exact phrase 'regardless of judge quality' was not verifiable in its abstract. The 2026 PPI-for-LLM-judge literature is real and active (PRECISE/Amazon Science, arXiv:2606.05308 ranking extension, arXiv:2601.05420, arXiv:2601.20913), so the claim is not superseded - it is being extended. Two corrections: (1) 'always-valid confidence intervals' is a misuse of a term of art - 'always-valid'/anytime-valid refers to sequential inference; PPI/PPI++ intervals are fixed-n ASYMPTOTIC (CLT-based) intervals, so coverage is guaranteed only asymptotically and finite-sample coverage can degrade with very small gold sets. (2) The judge-quality-independence property is a general PPI/PPI++ property (Angelopoulos et al. 2023) rather than something stated in the AutoEval abstract itself; the claim's framing correctly captures its substance (worse judge -> wider intervals via lower correlation, coverage preserved) but slightly overstates its strength and its provenance in this specific source.",
    "corrected": "Prediction-Powered Inference (PPI/PPI++) combines a small human gold set with a large judge-labeled set to produce unbiased estimates with asymptotically valid confidence intervals whose coverage does not depend on judge accuracy (a worse judge widens intervals rather than breaking coverage, subject to standard asymptotic/regularity conditions and a sufficiently large gold set); 'AutoEval Done Right' (arXiv:2403.07008, Boyeau, Angelopoulos, Yosef, Malik & Jordan) reports increasing the effective human-labeled sample size by up to 50% in experiments with GPT-4."
   },
   "external_verification": {
    "verdict": "partially_confirmed",
    "checked": [
     "core method is prediction-powered inference (PPI) and the optimized variant PPI++",
     "combines a small amount of human data with a large amount of synthetic/judge data to increase effective human sample size without compromising validity",
     "estimators remain unbiased; PPI/PPI++ yield valid, calibrated confidence intervals across all labeled set sizes",
     "robust to judge quality: even with a very poor annotator model, PPI++ performs at least as well as the classical approach (worse judge widens, never invalidates)",
     "v3 camera-ready dated 2026-06-01"
    ],
    "not_visible": [
     "the specific figure '50% with a GPT-4 judge': the LLM/judge experiment used gpt-4o-mini and gave only 20-35% ESS improvement; the ~50% ESS gain came from the ImageNet and protein-fitness experiments, not the GPT-judge experiment",
     "GLIDE (arXiv:2605.31278), PRECISE/AAAI 2026, arXiv:2606.05308 -- separate papers, not part of this source"
    ],
    "quote": "even with a very poor annotator model, PPI++ performs at least as well as the classical approach",
    "notes": "TITLE MISMATCH: source_title says 'Sample-Efficient Human Evaluation via Prediction-Powered Inference' but the actual (v3) title is 'Using Synthetic Data for Model Evaluation'. PPI/PPI++, unbiasedness, valid/calibrated CIs, and coverage-robustness-to-judge-quality are all directly confirmed in the full HTML text. The abstract does say 'up to 50% on experiments with GPT-4', but the body attributes ~50% ESS to ImageNet/protein experiments while the actual LLM/GPT-style judge (gpt-4o-mini) experiment yielded only 20-35% -- so the '50% with a GPT-4 judge' specificity is imprecise. Kept at partially_confirmed for that discrepancy plus the title mismatch.",
    "checked_at": "2026-07-15",
    "retrieval": "fulltext"
   },
   "source_retrieval_meta": {
    "retrieval": "fulltext",
    "fetched": [
     "https://arxiv.org/abs/2403.07008",
     "https://arxiv.org/pdf/2403.07008",
     "https://arxiv.org/html/2403.07008v3"
    ],
    "resolved_title": "AutoEval Done Right: Using Synthetic Data for Model Evaluation",
    "resolved_date": "v1 2024-03-09; v3 (camera-ready) 2026-06-01",
    "title_match": false
   },
   "used_on": [
    {
     "page": "index.html",
     "anchor": "decision",
     "label": "Overview - the five authorizations"
    },
    {
     "page": "decisions.html",
     "anchor": "a2",
     "label": "Decisions - first project and gold funding"
    },
    {
     "page": "decisions.html",
     "anchor": "c4",
     "label": "Decisions - default: statistics staging"
    }
   ],
   "decision_use": {
    "grade": "A",
    "basis": "statistical method",
    "measured_on": "structural",
    "note": "PPI/PPI++ unbiasedness and judge-robust intervals are mathematics; they get MORE useful as judges improve. (Empirical demos in the paper are GPT-4-era; the method is the decision input. Title changed on arXiv v3 - noted in re-check.)",
    "models_measured": [
     "GPT-4"
    ]
   }
  },
  {
   "id": "HS-02",
   "domain": "hybrid-statistical",
   "area": null,
   "claim": "Cascaded Selective Evaluation (Trust or Escalate, ICLR 2025 Oral) delivers a provable, user-specified human-agreement guarantee for LLM judges by calibrating a confidence threshold with fixed-sequence testing on a small human calibration set, escalating low-confidence items up a judge cascade and abstaining (to humans) when even the strongest judge is unconfident - achieving guaranteed >80% human agreement at ~80% coverage on a Chatbot Arena subset where GPT-4 alone almost never reaches 80% agreement, with ~88% of judged items handled by much cheaper models.",
   "load_bearing": true,
   "evidence": "arXiv:2407.18370 abstract fetched directly (80% agreement / ~80% coverage claim confirmed); ICLR 2025 Oral status confirmed at iclr.cc/virtual/2025/oral/31838; fixed-sequence-testing calibration mechanism confirmed via OpenReview PDF snippets and secondary reviews (83 citations as of 2026-07). Its 'Simulated Annotators' method (in-context simulation of diverse annotators) is the confidence estimator that makes the guarantee achievable at high coverage.",
   "implication": "This is the direct architectural blueprint for 'which items get the single human touch': the human queue is exactly the abstention set of a confidence-calibrated cascade, and the abstention rate is a tunable dial trading human budget against a guaranteed agreement level. Caveat for taxonomy Q1: the guarantee is agreement-with-human-majority, so this component's validity is capped at human panel reliability - it certifies consistency-of-application, not truth.",
   "source": {
    "raw": "Trust or Escalate: LLM Judges with Provable Guarantees for Human Agreement (ICLR 2025 Oral) | https://arxiv.org/abs/2407.18370 | 2024-07-25 (ICLR 2025-04) | academic",
    "title": "Trust or Escalate: LLM Judges with Provable Guarantees for Human Agreement (ICLR 2025 Oral)",
    "url": "https://arxiv.org/abs/2407.18370",
    "date": "2024-07-25 (ICLR 2025-04)",
    "type": "academic"
   },
   "venue": {
    "peer_reviewed": true,
    "marker": "ICLR"
   },
   "same_source_claims": [],
   "in_corpus_verification": {
    "verdict": "confirmed",
    "transcript": "All load-bearing elements check out against the primary sources. (1) arXiv:2407.18370 abstract (fetched directly) states Cascaded Selective Evaluation provides a provable human-agreement guarantee at a user-specified level, and that on a Chatbot Arena subset \"where GPT-4 almost never achieves 80% human agreement,\" the method \"guarantees over 80% human agreement with almost 80% test coverage\" even using Mistral-7B. (2) Paper body (arxiv.org/html/2407.18370v1) confirms the mechanism: threshold calibration via fixed sequence testing (Bauer, 1991) on a small human calibration set (|D_cal|=500 for ChatArena/TL;DR, 392 for Auto-J), cascade escalation from cheap to strong judges on low confidence, abstention (output empty set) when even the strongest judge is unconfident, and Simulated Annotators as the confidence estimator enabling high coverage. (3) The ~88% figure is exact: \"79.1% of all samples, among which 88.1% are evaluated by substantially cheaper Mistral-7B or GPT-3.5 instead of GPT-4\" (Table 3: GPT-4 handles only 17.5%). (4) ICLR 2025 Oral confirmed at iclr.cc/virtual/2025/oral/31838 (Jung, Brahman, Choi); dates correct (v1 2024-07-25; ICLR 2025 in April 2025). Two trivial nuances, neither rising to a correction: the paper defines abstention as returning empty set / excluding from coverage - deferral \"to humans\" is the intended framing but abstained items are not actually routed to human annotators in the experiments; and the guarantee holds with probability 1-delta over calibration-set sampling (standard for such guarantees, and consistent with \"provable, user-specified\"). Supersession check: not superseded - but a newer alternative exists: SCOPE (arXiv:2602.13110, ICML 2026 poster, Feb 2026) uses conformal calibration with a Bidirectional Preference Entropy signal and reports better calibration/coverage than Simulated Annotators; it complements rather than invalidates the claim about this paper. The \"83 citations\" figure was not independently verified but is plausible and not load-bearing.",
    "corrected": null
   },
   "external_verification": {
    "verdict": "confirmed",
    "checked": [
     "Cascaded Selective Evaluation and Simulated Annotators confirmed (abstract)",
     "provable, user-specified human-agreement guarantee confirmed",
     ">80% agreement at ~80% coverage: 'guarantees over 80% human agreement with almost 80% test coverage'; precise 'covering 79.1% of all samples'",
     "GPT-4 alone almost never reaches 80% agreement on ChatArena subset (abstract)",
     "~88% by cheaper models: '88.1% are evaluated by substantially cheaper Mistral-7B or GPT-3.5 instead of GPT-4'",
     "fixed-sequence testing: 'we adopt fixed-sequence testing instead of Bonferroni correction'",
     "small human calibration set: 'given access to a small calibration set' of human preferences (size 500 / 392 in experiments)"
    ],
    "not_visible": [
     "'ICLR 2025 Oral' venue/status -- arXiv page carries no journal-ref or comments field and lists only v1 (July 2024); the evidence cited iclr.cc, which is outside the allowed fetch scope, so venue cannot be confirmed from the source"
    ],
    "quote": "among which 88.1% are evaluated by substantially cheaper Mistral-7B or GPT-3.5 instead of GPT-4 ... we adopt fixed-sequence testing instead of Bonferroni correction",
    "notes": "Every substantive mechanism and number in the claim is confirmed from the arXiv HTML full text (the PDF returned only compressed binary). The sole unverifiable item is the ICLR 2025 Oral label, which the arXiv record does not carry.",
    "checked_at": "2026-07-15",
    "retrieval": "fulltext"
   },
   "source_retrieval_meta": {
    "retrieval": "fulltext",
    "fetched": [
     "https://arxiv.org/abs/2407.18370",
     "https://arxiv.org/pdf/2407.18370",
     "https://arxiv.org/html/2407.18370v1"
    ],
    "resolved_title": "Trust or Escalate: LLM Judges with Provable Guarantees for Human Agreement",
    "resolved_date": "2024-07-25 (v1, only version listed on arXiv)",
    "title_match": true
   },
   "used_on": [],
   "decision_use": {
    "grade": "A",
    "basis": "statistical method (numbers dated)",
    "measured_on": "model-outputs",
    "note": "Trust-or-Escalate's guarantee construction is method-level and durable. Its reported coverage numbers are GPT-4-era; recalibrate on deployment-tier judges - the framework exists precisely to do that.",
    "models_measured": [
     "GPT-4"
    ]
   }
  },
  {
   "id": "HS-03",
   "domain": "hybrid-statistical",
   "area": null,
   "claim": "The minimum viable human gold set is computable in closed form: with a doubly-robust two-stage design, required human labels converge to a floor of n*(1-rho^2) where n* is the target effective sample size and rho the judge-human correlation - e.g., a target effective n=200 with R^2=0.70 and 2,000 judge ratings needs only 65 human labels - with diminishing returns from adding more judge labels and up to ~13% further savings from stratified allocation when judge reliability varies across evaluation axes.",
   "load_bearing": true,
   "evidence": "arXiv:2605.16354v1 (Jane Paik Kim, Stanford Psychiatry, 2026-05-08) HTML fetched: derives sample-size formulas from the doubly-robust estimator's asymptotic variance; worked examples confirmed (n*=200, R^2=0.7, N=2000 -> 65 human labels; N=400 -> 100). Explicit caveats: pilot overestimation of R^2 breaks precision guarantees (use conservative values), framework assumes human ratings are a reliable gold standard, single-rater scope.",
   "implication": "Answers taxonomy Q1's gold-set question with arithmetic instead of convention: run a small pilot to estimate per-axis judge-human correlation, then size each project's gold set from the formula with a conservative correlation estimate. Because the floor depends on rho, per-axis correlation heterogeneity argues for stratified gold sampling by criterion, and low-rho axes are where gold budget must concentrate.",
   "source": {
    "raw": "Augmenting Human Evaluation with LLM Judges: How Many Human Reviews Do You Need? | https://arxiv.org/html/2605.16354v1 | 2026-05-08 | academic",
    "title": "Augmenting Human Evaluation with LLM Judges: How Many Human Reviews Do You Need?",
    "url": "https://arxiv.org/html/2605.16354v1",
    "date": "2026-05-08",
    "type": "academic"
   },
   "venue": {
    "peer_reviewed": true,
    "marker": "CHI"
   },
   "same_source_claims": [],
   "in_corpus_verification": {
    "verdict": "confirmed",
    "transcript": "All elements verified against the source (https://arxiv.org/html/2605.16354v1, Jane Paik Kim, Stanford Dept. of Psychiatry and Behavioral Sciences, submitted 2026-05-08 - date and attribution correct). (1) Floor formula: paper states required human reviews converge to n*(1-rho^2) as N increases - exact match. (2) Worked example: n*=200, R^2=0.70, N=2000 -> 65 human labels, confirmed verbatim; the N=400/100 figure is stated in the paper as \"with a budget of 100 human ratings, N can be reduced to 400 while achieving target power\" - same numbers, slightly different framing (human budget fixed, N solved) but mathematically equivalent to the claim. (3) Stratification savings: paper reports \"up to 12.9%\" reduction when R^2 gap is large (0.8 vs 0.1) - matches \"~13%\", with the paper adding that moderate gaps yield only ~2%. (4) Diminishing returns from more judge labels: explicitly stated (\"each successive increase in N produces diminishing reductions in n\"). (5) All three caveats in the evidence summary (pilot R^2 overestimation breaking precision guarantees, human-as-gold-standard assumption, single-rater scope) appear in the paper's limitations. Supersession check: two searches found no v2 revision, no citing papers, and no later work replacing or contradicting the result as of 2026-07-14. Minor precision notes only: \"up to ~13%\" should strictly be \"up to 12.9% in the extreme R^2-gap case, ~2% for moderate gaps,\" and the floor is asymptotic (approached as N grows), not attainable at finite N - the claim's own phrasing (\"converge to a floor,\" \"up to\") already reflects both.",
    "corrected": null
   },
   "external_verification": {
    "verdict": "confirmed",
    "checked": [
     "floor formula n*(1-rho^2): 'converges to a floor of n*(1-rho2)'",
     "worked example: n*=200, R2=0.70, N=2000 -> 'only 65 human samples are needed'",
     "N=400 -> 100: 'If the budget can allow for 100 human ratings, then the LLM sample size N can be reduced to 400'",
     "doubly robust two-stage sampling design (missingness known by design)",
     "stratification 'reduces the human budget by up to 12.9%' (claim's ~13%)",
     "caveats: pilot overestimation of R2; human gold standard a 'pragmatic choice rather than a settled fact'; single-rater scope",
     "author Jane Paik Kim, Stanford Dept. of Psychiatry"
    ],
    "not_visible": [],
    "quote": "With an LLM sample size of N=2000, only 65 human samples are needed.",
    "notes": "All specifics verified from the given HTML full text. Claim's '~13%' stratified savings equals the paper's precise 12.9%. Formula, both worked examples (65 and 100 human labels), diminishing returns, and all three caveats match.",
    "checked_at": "2026-07-15",
    "retrieval": "fulltext"
   },
   "source_retrieval_meta": {
    "retrieval": "fulltext",
    "fetched": [
     "https://arxiv.org/html/2605.16354v1"
    ],
    "resolved_title": "Augmenting Human Evaluation with LLM Judges: How Many Human Reviews Do You Need?",
    "resolved_date": "arXiv 2605.16354 (May 2026, per arXiv ID; source_date 2026-05-08 consistent; exact day not surfaced in fetched HTML)",
    "title_match": true
   },
   "used_on": [
    {
     "page": "research.html",
     "anchor": "research-cards",
     "label": "Research - findings by decision weight"
    },
    {
     "page": "decisions.html",
     "anchor": "a2",
     "label": "Decisions - first project and gold funding"
    }
   ],
   "decision_use": {
    "grade": "A",
    "basis": "statistical method",
    "measured_on": "structural",
    "note": "Closed-form gold-set sizing. Mathematics.",
    "models_measured": []
   }
  },
  {
   "id": "HS-04",
   "domain": "hybrid-statistical",
   "area": null,
   "claim": "Using the LLM's own confidence signals to choose WHICH items receive human annotation (Confidence-Driven Inference, building on Active Statistical Inference) cut required human annotations by >25% in all three tested settings while keeping provably valid confidence intervals that remain safe even if the LLM annotations are poor - but a 2026 ICLR paper finds the opposite in the sequential regime, where near-uniform sampling at the budget ceiling beat uncertainty-driven querying.",
   "load_bearing": true,
   "evidence": "arXiv:2408.15204 (Gligoric, Zrnic, Lee, Candes, Jurafsky; NAACL 2025) abstract fetched: '>25% reduction in each of three settings' (politeness, stance, bias) with validity guarantees regardless of LLM quality. Foundation: Active Statistical Inference (Zrnic & Candes, ICML 2024, arXiv:2403.03208): label where the model is uncertain, provably valid CIs, fewer labels than uniform. Counter-evidence verified separately (see disagreements): arXiv:2604.18569 (ICLR 2026).",
   "implication": "Uncertainty-routed allocation of the scarce human touch is the statistically defensible default and comes with a safety net (validity even when the judge is wrong), but the gain over random sampling is regime-dependent and contested - so the foundation should make the routing policy (uncertainty-weighted vs uniform vs stratified) a measured, per-project choice with a built-in A/B against uniform, not a hard-wired assumption.",
   "source": {
    "raw": "Can Unconfident LLM Annotations Be Used for Confident Conclusions? (NAACL 2025) + Active Statistical Inference (ICML 2024) | https://arxiv.org/abs/2408.15204 | 2024-08-27 (v2 2025-02-08; NAACL 2025-04) | academic",
    "title": "Can Unconfident LLM Annotations Be Used for Confident Conclusions? (NAACL 2025) + Active Statistical Inference (ICML 2024)",
    "url": "https://arxiv.org/abs/2408.15204",
    "date": "2024-08-27 (v2 2025-02-08; NAACL 2025-04)",
    "type": "academic"
   },
   "venue": {
    "peer_reviewed": true,
    "marker": "ACL"
   },
   "same_source_claims": [],
   "in_corpus_verification": null,
   "external_verification": {
    "verdict": "partially_confirmed",
    "checked": [
     "Confidence-Driven Inference combines LLM annotations and LLM confidence to pick which items get human annotation",
     "reduces needed human annotations by over 25% in each of three settings (politeness, stance, bias)",
     "safeguards against poor-quality LLM annotations; conclusions valid and no less accurate than human-only",
     "provably valid confidence intervals"
    ],
    "not_visible": [
     "'Active Statistical Inference' foundation (attributed to separate paper arXiv:2403.03208, not on this page)",
     "2026 ICLR counter-finding (arXiv:2604.18569) that near-uniform sampling beats uncertainty-driven querying"
    ],
    "quote": "The paper introduces \"Confidence-Driven Inference: a method that combines LLM annotations and LLM confidence indicators\" ... it reduces \"the needed number of human annotations by over 25% in each\" setting.",
    "notes": "The core claim attributable to this URL (>25% reduction in all three settings, validity regardless of LLM quality) is fully confirmed verbatim. Marked partially_confirmed because two asserted specifics are external to this source by the claim's own admission (Active Statistical Inference foundation = arXiv:2403.03208; the ICLR-2026 counter-finding = arXiv:2604.18569) and neither is visible on the fetched abstract. Authors: Gligoric, Zrnic, Lee, Candes, Jurafsky.",
    "checked_at": "2026-07-15",
    "retrieval": "abstract"
   },
   "source_retrieval_meta": {
    "retrieval": "abstract",
    "fetched": [
     "https://arxiv.org/abs/2408.15204"
    ],
    "resolved_title": "Can Unconfident LLM Annotations Be Used for Confident Conclusions?",
    "resolved_date": "v1 2024-08-27; v2 2025-02-08 (NAACL 2025)",
    "title_match": true
   },
   "used_on": [],
   "decision_use": {
    "grade": "A",
    "basis": "statistical method",
    "measured_on": "structural",
    "note": "Confidence-driven human-annotation allocation. Method.",
    "models_measured": []
   }
  },
  {
   "id": "HS-05",
   "domain": "hybrid-statistical",
   "area": null,
   "claim": "GLIDE (May 2026) is an open-source industrial library unifying PPI++, stratified PPI, predict-then-debias bootstrap, and active inference with a decision tree for method selection, and it quantifies when PPI helps: effective-sample gain is ~1.0x at judge-human correlation rho=0.1, ~2.2x at rho=0.9; CLT-based estimators need roughly >=50 human labels per stratum (below that use bootstrap PTD); in a real agent-safety case (rho=0.59, ~13-point judge bias) 100 human labels became worth 143-157.",
   "load_bearing": false,
   "evidence": "arXiv:2605.31278v2 HTML fetched (Martinon, Merad, Raki; Emerton Data): Monte Carlo with 500 human + 1000 proxy labels shows effective n ~500 at rho=0.1 vs ~1100 at rho=0.9; R-Judge case study with Claude Sonnet as judge; library at github.com/EmertonData/glide with reproducible validation notebooks; limitations: means/proportions only, i.i.d. assumptions, no anytime-valid inference.",
   "implication": "Ready-made tooling exists for the statistical layer; more importantly it publishes the correlation-vs-benefit curve and the >=50-labels-per-stratum threshold the foundation can budget against: if a project's judge-human correlation on an axis is below ~0.3, PPI buys nothing there and gold labels must carry that axis alone.",
   "source": {
    "raw": "Industrializing Prediction-Powered Inference: The GLIDE Library for Reliable GenAI and Agentic Systems Evaluation | https://arxiv.org/html/2605.31278v2 | 2026-05 (v2 2026-06) | practitioner",
    "title": "Industrializing Prediction-Powered Inference: The GLIDE Library for Reliable GenAI and Agentic Systems Evaluation",
    "url": "https://arxiv.org/html/2605.31278v2",
    "date": "2026-05 (v2 2026-06)",
    "type": "practitioner"
   },
   "venue": {
    "peer_reviewed": false,
    "marker": "preprint"
   },
   "same_source_claims": [],
   "in_corpus_verification": null,
   "external_verification": {
    "verdict": "confirmed",
    "checked": [
     "unifies state-of-the-art PPI estimators (PPI++, Stratified PPI, Predict-Then-Debias and its stratified variants, Active Statistical Inference)",
     "an empirically grounded decision tree for method selection",
     "effective sample size grows from ~500 at rho=0.1 to ~1100 at rho=0.9 (a 2.2x effective gain), from 500 human + 1000 proxy labels",
     "at least fifty labeled samples per stratum for the asymptotic intervals to be reliable",
     "R-Judge case study; proxy is claude-sonnet-4-6 run as zero-shot LLM-as-judge",
     "Pearson correlation with expert labels rho approximately 0.59",
     "overshoots the true rate by about 13 percentage points",
     "statistically equivalent to roughly 157, 148, and 143 purely human-labeled trajectories",
     "https://github.com/EmertonData/glide",
     "limitations: mean estimation only, single proxy / i.i.d. assumption, does not provide anytime-valid constructions"
    ],
    "not_visible": [],
    "quote": "the effective sample size grows from approximately 500 at rho=0.1 to approximately 1100 at rho=0.9 ... a 2.2x effective gain at no cost to validity",
    "notes": "Every headline number confirmed verbatim in the HTML full text: the ~1.0x-to-2.2x effective-sample gain, >=50 labels/stratum, and the R-Judge case (rho~0.59, ~13pp bias, 100 labels worth 143-157). Authors Martinon/Merad/Raki, Emerton Data affiliation, and the GitHub repo all present.",
    "checked_at": "2026-07-15",
    "retrieval": "fulltext"
   },
   "source_retrieval_meta": {
    "retrieval": "fulltext",
    "fetched": [
     "https://arxiv.org/html/2605.31278v2"
    ],
    "resolved_title": "Industrializing Prediction-Powered Inference: The GLIDE Library for Reliable GenAI and Agentic Systems Evaluation",
    "resolved_date": "2026-06-04 (v2)",
    "title_match": true
   },
   "used_on": [],
   "decision_use": {
    "grade": "A",
    "basis": "tooling fact",
    "measured_on": "structural",
    "note": "GLIDE library unifying PPI-class estimators (May 2026). Tooling availability, current.",
    "models_measured": []
   }
  },
  {
   "id": "HS-06",
   "domain": "hybrid-statistical",
   "area": null,
   "claim": "For estimating scores from noisy judges, PPI++/efficient-influence-function estimators produce near-identical, shortest confidence intervals - reported 3-15x narrower than Rogan-Gladen-style misclassification-correction estimators depending on the human-labeling ratio, with the advantage largest when the judge is close to random guessing.",
   "load_bearing": false,
   "evidence": "arXiv:2601.05420 (Chen, Lu, Li, Guo, Li; 2026-01-08) abstract fetched: unifies measurement-error correction and PPI classes via semiparametric efficiency theory, derives EIF-based estimators, characterizes when PPI-style strictly dominates; the 3-15x interval-width figure comes from the paper's empirical section as surfaced in search-indexed full text (not independently recomputed). Code: github.com/yiqunchen/debias-llm-as-a-judge.",
   "implication": "Standardize the estimation layer on PPI++/EIF-style residual calibration rather than sensitivity/specificity (confusion-matrix) correction when the target is a rate or mean - but note the certification literature pushes the other way for hypothesis tests (see disagreements).",
   "source": {
    "raw": "Efficient Inference for Noisy LLM-as-a-Judge Evaluation | https://arxiv.org/abs/2601.05420 | 2026-01-08 | academic",
    "title": "Efficient Inference for Noisy LLM-as-a-Judge Evaluation",
    "url": "https://arxiv.org/abs/2601.05420",
    "date": "2026-01-08",
    "type": "academic"
   },
   "venue": {
    "peer_reviewed": false,
    "marker": "preprint"
   },
   "same_source_claims": [],
   "in_corpus_verification": null,
   "external_verification": {
    "verdict": "confirmed",
    "checked": [
     "EIF and PPI++ produce nearly identical and shortest intervals",
     "outperform Rogan-Gladen by a factor of 3-15x",
     "3-15x factor depends on the labeling ratio",
     "advantage most pronounced when judge is closer to random guessing (q0+q1-1 small)",
     "at q0=q1=0.6, RG intervals ~10x wider than EIF/PPI++",
     "code repo github.com/yiqunchen/debias-llm-as-a-judge",
     "authors Chen, Lu, Li, Guo, Li"
    ],
    "not_visible": [],
    "quote": "EIF and PPI++ produce nearly identical and shortest intervals, outperforming Rogan-Gladen by a factor of 3-15x depending on the labeling ratio. The advantage is most pronounced when q0+q1-1 is small (i.e., when the LLM-judge is closer to random guessing).",
    "notes": "The 3-15x figure and the near-random-guessing / labeling-ratio dependence are NOT in the abstract (and the PDF was FlateDecode-compressed and unreadable), but the allowed HTML full text (Figure 4 caption + Section 5.3) confirms them verbatim, including 'PPI++' specifically (the abstract only said 'PPI'). Fully confirmed.",
    "checked_at": "2026-07-15",
    "retrieval": "fulltext"
   },
   "source_retrieval_meta": {
    "retrieval": "fulltext",
    "fetched": [
     "https://arxiv.org/abs/2601.05420",
     "https://arxiv.org/pdf/2601.05420",
     "https://arxiv.org/html/2601.05420v1"
    ],
    "resolved_title": "Efficient Inference for Noisy LLM-as-a-Judge Evaluation",
    "resolved_date": "2026-01-08",
    "title_match": true
   },
   "used_on": [],
   "decision_use": {
    "grade": "A",
    "basis": "statistical theory",
    "measured_on": "structural",
    "note": "EIF/PPI++ efficiency results. Mathematics (interval-width empirics corpus-flagged as not independently recomputed).",
    "models_measured": []
   }
  },
  {
   "id": "HS-07",
   "domain": "hybrid-statistical",
   "area": null,
   "claim": "For certifying that a failure rate is below a threshold (the low-base-rate pass/fail regime), the 'Noisy but Valid' framework (ICLR 2026) uses a small human calibration set to estimate judge TPR/FPR, applies a variance-corrected test to the large judge-labeled stream with finite-sample Type-I error control, and derives exact conditions under which judge-based testing has HIGHER statistical power than direct human evaluation.",
   "load_bearing": false,
   "evidence": "arXiv:2601.20913 (Feng, Shen, Balashankar, Gerner-Beuerle, Rodrigues; submitted 2026-01-28, accepted ICLR 2026 per arXiv page) abstract fetched: variance-corrected critical threshold, finite-sample Type-I control under calibration uncertainty, validation on Jigsaw/Hate Speech/SafeRLHF, and a quantified 'oracle gap' - the power cost of having to estimate judge reliability rather than knowing it.",
   "implication": "For attempter-level or batch-level certification ('this attempter's error rate is below X'), a judge-plus-small-calibration design can be provably MORE powerful than human-only review - reframing the human panel's job as estimating judge error rates, not per-item adjudication. This is direct evidence for the zero-per-item-touch + calibration-audit branch of taxonomy Q6's fork.",
   "source": {
    "raw": "Noisy but Valid: Robust Statistical Evaluation of LLMs with Imperfect Judges | https://arxiv.org/abs/2601.20913 | 2026-01-28 | academic",
    "title": "Noisy but Valid: Robust Statistical Evaluation of LLMs with Imperfect Judges",
    "url": "https://arxiv.org/abs/2601.20913",
    "date": "2026-01-28",
    "type": "academic"
   },
   "venue": {
    "peer_reviewed": true,
    "marker": "ICLR"
   },
   "same_source_claims": [],
   "in_corpus_verification": null,
   "external_verification": {
    "verdict": "confirmed",
    "checked": [
     "small human-labelled calibration set estimates judge True Positive and False Positive Rates",
     "variance-corrected critical threshold applied to large judge-labelled dataset",
     "finite-sample Type-I error control (validity) despite calibration uncertainty",
     "derives exact conditions under which noisy testing yields higher power than direct evaluation",
     "validated on Jigsaw Comment, Hate Speech, SafeRLHF",
     "quantified 'Oracle' gap (cost of estimating judge parameters)",
     "Accepted to ICLR 2026"
    ],
    "not_visible": [],
    "quote": "leveraging a small human-labelled calibration set to estimate the judge's True Positive and False Positive Rates ... the exact conditions under which noisy testing yields higher statistical power than direct evaluation ... Accepted to ICLR2026",
    "notes": "Every element of the claim is present in the abstract, including the low-base-rate certification framing, finite-sample Type-I control under calibration uncertainty, the three datasets, the oracle gap, and ICLR 2026 acceptance. Authors match (Feng, Shen, Balashankar, Gerner-Beuerle, Rodrigues).",
    "checked_at": "2026-07-15",
    "retrieval": "abstract"
   },
   "source_retrieval_meta": {
    "retrieval": "abstract",
    "fetched": [
     "https://arxiv.org/abs/2601.20913"
    ],
    "resolved_title": "Noisy but Valid: Robust Statistical Evaluation of LLMs with Imperfect Judges",
    "resolved_date": "2026-01-28",
    "title_match": true
   },
   "used_on": [
    {
     "page": "pilot.html",
     "anchor": "main",
     "label": "Pilot - gates and objectives"
    }
   ],
   "decision_use": {
    "grade": "A",
    "basis": "statistical method",
    "measured_on": "structural",
    "note": "'Noisy but Valid' certification: judge-assisted testing provably more powerful under conditions. Mathematics.",
    "models_measured": []
   }
  },
  {
   "id": "HS-08",
   "domain": "hybrid-statistical",
   "area": null,
   "claim": "Naive reporting of judge-scored evaluations is systematically biased by judge sensitivity/specificity; a plug-in correction with confidence intervals propagating BOTH test-set and calibration-set uncertainty fixes it, remains unbiased under distribution shift between calibration and test data, includes adaptive allocation of calibration labels, and characterizes regimes where corrected judge-based evaluation is more reliable than human-only evaluation.",
   "load_bearing": false,
   "evidence": "arXiv:2511.21140 (Lee, Zeng, Jeong, Sohn, Kangwook Lee; submitted 2025-11-26, v4 2026-05-31, accepted ICML 2026 per arXiv page) abstract fetched. The calibration-to-test distribution-shift robustness is the distinctive property versus standard PPI's i.i.d. assumption.",
   "implication": "Supports the taxonomy Q1 proposal to ban raw judge-agreement numbers from all reporting: every reported metric should be the corrected estimator with a CI that includes calibration uncertainty. The shift-robustness result matters specifically because attempter adaptation (taxonomy Q8) shifts the live item distribution away from the calibration set over time.",
   "source": {
    "raw": "How to Correctly Report LLM-as-a-Judge Evaluations | https://arxiv.org/abs/2511.21140 | 2025-11-26 (v4 2026-05-31) | academic",
    "title": "How to Correctly Report LLM-as-a-Judge Evaluations",
    "url": "https://arxiv.org/abs/2511.21140",
    "date": "2025-11-26 (v4 2026-05-31)",
    "type": "academic"
   },
   "venue": {
    "peer_reviewed": false,
    "marker": "preprint"
   },
   "same_source_claims": [
    "F26-09"
   ],
   "in_corpus_verification": null,
   "external_verification": {
    "verdict": "confirmed",
    "checked": [
     "imperfect judge sensitivity/specificity biases naive evaluation scores",
     "plug-in framework corrects the bias",
     "confidence intervals propagate uncertainty from BOTH test dataset and human-labeled calibration dataset",
     "remains unbiased under distribution shift between test and calibration datasets",
     "adaptive strategy to allocate calibration samples for tighter intervals",
     "characterizes regimes (by true score, sensitivity, specificity) where judge-based eval beats human-only",
     "accepted ICML 2026; v1 2025-11-26, v4 2026-05-31"
    ],
    "not_visible": [],
    "quote": "imperfect sensitivity and specificity of the LLM judges induce bias in naive evaluation scores ... uncertainty from both the test dataset and a human-labeled calibration dataset ... remains unbiased under distribution shift between the test and calibration datasets",
    "notes": "Every core element of the claim is present in the abstract and confirmed. Distribution-shift robustness (vs standard PPI i.i.d.) is stated explicitly.",
    "checked_at": "2026-07-15",
    "retrieval": "fulltext"
   },
   "source_retrieval_meta": {
    "retrieval": "fulltext",
    "fetched": [
     "https://arxiv.org/abs/2511.21140",
     "https://arxiv.org/html/2511.21140"
    ],
    "resolved_title": "How to Correctly Report LLM-as-a-Judge Evaluations",
    "resolved_date": "2025-11-26 (v4 2026-05-31; ICML 2026)",
    "title_match": true
   },
   "used_on": [
    {
     "page": "decisions.html",
     "anchor": "d6",
     "label": "Decisions - the throughput dial"
    }
   ],
   "decision_use": {
    "grade": "A",
    "basis": "statistical method",
    "measured_on": "structural",
    "note": "Sensitivity/specificity-corrected pass-rate reporting. Mathematics.",
    "models_measured": []
   }
  },
  {
   "id": "HS-09",
   "domain": "hybrid-statistical",
   "area": null,
   "claim": "Conformal prediction now provides per-item uncertainty for LLM-as-judge ratings: EMNLP 2025 work builds guaranteed-coverage score intervals from a single judging run (with an ordinal boundary adjustment and a lower-bias midpoint score), and April 2026 work shows conformal set width is a genuine per-instance reliability signal (r_s=+0.576 with reliability, n=1,918; widths correlate ~0.32-0.38 across different judges) with reliability driven more by the CRITERION than the judge model (relevance avg set size ~3.0 vs fluency/consistency ~4.9 on a 1-5 scale).",
   "load_bearing": false,
   "evidence": "Sheng, Liu, He, Zhao, Kang, 'Analyzing Uncertainty of LLM-as-a-Judge: Interval Evaluations with Conformal Prediction', EMNLP 2025 main (aclanthology.org/2025.emnlp-main.569), abstract confirmed via NeurIPS 2025 workshop listing. Gupta & Kumar, arXiv:2604.15302 (2026-04-16, under review): split conformal on SummEval; also finds 33-67% of documents contain intransitive judge preference cycles despite low aggregate violation rates.",
   "implication": "Conformal set width is a cheap, distribution-free routing statistic for the human-touch queue that transfers across judges, and the criterion-dominates-judge finding is empirical support for taxonomy Q2/Q5's axis-triage: per-criterion reliability ceilings must be measured and some axes routed human-only regardless of judge choice.",
   "source": {
    "raw": "Analyzing Uncertainty of LLM-as-a-Judge (EMNLP 2025) + Diagnosing LLM Judge Reliability (arXiv:2604.15302) | https://arxiv.org/abs/2604.15302 | 2025-11 (EMNLP) / 2026-04-16 | academic",
    "title": "Analyzing Uncertainty of LLM-as-a-Judge (EMNLP 2025) + Diagnosing LLM Judge Reliability (arXiv:2604.15302)",
    "url": "https://arxiv.org/abs/2604.15302",
    "date": "2025-11 (EMNLP) / 2026-04-16",
    "type": "academic"
   },
   "venue": {
    "peer_reviewed": true,
    "marker": "NeurIPS"
   },
   "same_source_claims": [
    "F26-05"
   ],
   "in_corpus_verification": null,
   "external_verification": {
    "verdict": "partially_confirmed",
    "checked": [
     "conformal set width is a per-instance reliability indicator, r_s = +0.576, N = 1,918",
     "cross-judge width agreement r-bar = 0.32-0.38",
     "relevance avg set size ~3.0 vs fluency/consistency ~4.9",
     "criterion matters more than the judge model",
     "split conformal prediction over 1-5 Likert scores on SummEval"
    ],
    "not_visible": [
     "EMNLP 2025 Sheng/Liu/He/Zhao/Kang paper: guaranteed-coverage intervals from a single judging run, ordinal boundary adjustment, lower-bias midpoint score (this is a DIFFERENT paper at aclanthology.org/2025.emnlp-main.569, not fetchable from this arXiv URL)"
    ],
    "quote": "set width serving as a per-instance reliability indicator (r_s = +0.576, N = 1,918, p < 10^-100) ... consistent cross-judge agreement (r-bar = 0.32-0.38)",
    "notes": "The claim bundles TWO papers; only the arXiv:2604.15302 (April 2026, Gupta & Kumar) half is at the fetched URL and it is fully confirmed. The EMNLP 2025 conformal-interval work (Sheng et al.) with the ordinal boundary adjustment and lower-bias midpoint score is a separate source not retrievable here, hence its specifics are not visible.",
    "checked_at": "2026-07-15",
    "retrieval": "abstract"
   },
   "source_retrieval_meta": {
    "retrieval": "abstract",
    "fetched": [
     "https://arxiv.org/abs/2604.15302"
    ],
    "resolved_title": "Diagnosing LLM Judge Reliability: Conformal Prediction Sets and Transitivity Violations",
    "resolved_date": "2026-04-16",
    "title_match": true
   },
   "used_on": [],
   "decision_use": {
    "grade": "B",
    "basis": "current-gen method + measurement",
    "measured_on": "model-outputs",
    "note": "Conformal per-item uncertainty for judge ratings (2025-26). Method durable; measured widths tier-bound.",
    "models_measured": []
   }
  },
  {
   "id": "HS-10",
   "domain": "hybrid-statistical",
   "area": null,
   "claim": "Judge-vs-system drift can be attributed with anytime-valid statistics: a frozen human-labeled anchor set periodically re-scored by the live judge, monitored with betting e-processes, correctly attributed a silent judge version bump in 60/60 runs with zero misattribution (rolling z-test baseline false-alarmed on 75% of drift-free streams), at 0.21-0.64x the cost of strong-judging every item.",
   "load_bearing": false,
   "evidence": "arXiv:2606.15474 (Yitao Li, 2026-06-13) abstract fetched: 'one-way identification (only the judge can move the anchors)'; three-verdict output (none/system/judge); strict-prompt change attributed in 110/120 runs; replicated on TL;DR summarization 240/240. Single-author preprint, not yet peer-reviewed - treat quantitative claims as provisional.",
   "implication": "Taxonomy Q9's three-moving-parts problem has an existing statistical design for one axis: freeze a per-project human-labeled anchor set at onboarding, interleave it continuously into the judge stream, and use e-process alarms to separate 'judge changed' from 'attempter population changed' before triggering rubric revision or de-graduation.",
   "source": {
    "raw": "Who Drifted: the System or the Judge? Anytime-Valid Attribution in LLM Evaluation Pipelines | https://arxiv.org/abs/2606.15474 | 2026-06-13 | academic",
    "title": "Who Drifted: the System or the Judge? Anytime-Valid Attribution in LLM Evaluation Pipelines",
    "url": "https://arxiv.org/abs/2606.15474",
    "date": "2026-06-13",
    "type": "academic"
   },
   "venue": {
    "peer_reviewed": false,
    "marker": "preprint"
   },
   "same_source_claims": [],
   "in_corpus_verification": null,
   "external_verification": {
    "verdict": "confirmed",
    "checked": [
     "fixed human-labeled anchor set re-scored by the current judge at a steady interleave",
     "betting e-process on the judge-versus-human gap",
     "silent version bump detected as judge drift in 60/60 runs with zero judge-to-system misattribution",
     "rolling z-test false-alarms on 75% of drift-free streams",
     "approximately 0.64 of the cost of strong-judging every item, or 0.21 in a cheaper-but-deafer regime",
     "one-way identification (only the judge can move the anchors)",
     "verdict in {none, system, judge}",
     "strict-prompt change attributed on 110 of 120 runs",
     "TL;DR replication attribution perfect (240/240)"
    ],
    "not_visible": [],
    "quote": "a silent version bump is detected as judge drift in 60/60 runs with zero judge-to-system misattribution.",
    "notes": "All 11 asserted specifics verified verbatim in the abstract. Single-author preprint (Yitao Li), submitted 2026-06-13, matching the claim's own not-yet-peer-reviewed caveat.",
    "checked_at": "2026-07-15",
    "retrieval": "abstract"
   },
   "source_retrieval_meta": {
    "retrieval": "abstract",
    "fetched": [
     "https://arxiv.org/abs/2606.15474"
    ],
    "resolved_title": "Who Drifted: the System or the Judge? Anytime-Valid Attribution in LLM Evaluation Pipelines",
    "resolved_date": "arXiv submitted 2026-06-13",
    "title_match": true
   },
   "used_on": [],
   "decision_use": {
    "grade": "A",
    "basis": "statistical method",
    "measured_on": "structural",
    "note": "Anchor-set drift attribution with anytime-valid statistics. Method.",
    "models_measured": []
   }
  },
  {
   "id": "HS-11",
   "domain": "hybrid-statistical",
   "area": null,
   "claim": "In sequential prediction-powered mean estimation, uncertainty-based active querying contributed little: the smallest confidence widths occurred when query probabilities were near-constant at the budget ceiling, and theory shows optimized query probabilities converge to the maximum allowed constant rate when chosen without reference to covariates.",
   "load_bearing": false,
   "evidence": "arXiv:2604.18569 (Sfyraki & Wang, submitted 2026-04-20, accepted ICLR 2026 per arXiv page) abstract fetched: empirical finding that the uncertainty-blend weight near the constant term minimizes CI width, plus non-asymptotic analysis and no-regret query-probability selection corroborating near-uniform optimality in this regime.",
   "implication": "Direct counter-evidence to uncertainty-routed human sampling for population estimation in streaming settings: the foundation must not assume uncertainty routing dominates - measure the routing policy's marginal value per project, and expect uniform-at-budget to be near-optimal when the goal is estimation rather than per-item verdict quality.",
   "source": {
    "raw": "Revisiting Active Sequential Prediction-Powered Mean Estimation | https://arxiv.org/abs/2604.18569 | 2026-04-20 | contrarian",
    "title": "Revisiting Active Sequential Prediction-Powered Mean Estimation",
    "url": "https://arxiv.org/abs/2604.18569",
    "date": "2026-04-20",
    "type": "contrarian"
   },
   "venue": {
    "peer_reviewed": true,
    "marker": "ICLR"
   },
   "same_source_claims": [],
   "in_corpus_verification": null,
   "external_verification": {
    "verdict": "confirmed",
    "checked": [
     "smallest confidence width occurs when the weight on the constant probability is close to one (uncertainty-driven part diminishes)",
     "non-asymptotic analysis yielding a data-dependent CI bound",
     "with a no-regret approach the query probability converges to the maximum-value constraint when set obliviously to covariates",
     "Published as a conference paper at ICLR 2026",
     "submitted 20 Apr 2026 (Sfyraki & Wang)"
    ],
    "not_visible": [],
    "quote": "the smallest confidence width tends to occur when the weight on the constant probability is close to one",
    "notes": "Core claim confirmed. The claim's 'near-constant at the budget ceiling' maps to the paper's 'maximum-value constraint' the oblivious query probability converges to; consistent phrasing.",
    "checked_at": "2026-07-15",
    "retrieval": "abstract"
   },
   "source_retrieval_meta": {
    "retrieval": "abstract",
    "fetched": [
     "https://arxiv.org/abs/2604.18569"
    ],
    "resolved_title": "Revisiting Active Sequential Prediction-Powered Mean Estimation",
    "resolved_date": "2026-04-20",
    "title_match": true
   },
   "used_on": [],
   "decision_use": {
    "grade": "A",
    "basis": "statistical finding",
    "measured_on": "structural",
    "note": "Sequential estimation: uncertainty-based querying bought little vs uniform. Regime-dependent statistics; the design response (A/B routing against uniform) is already encoded.",
    "models_measured": []
   }
  },
  {
   "id": "HS-12",
   "domain": "hybrid-statistical",
   "area": null,
   "claim": "Commercial tooling for confidence-scored LLM judgments with human calibration is mature: Cleanlab's Trustworthy Language Model attaches a real-time trustworthiness score to every LLM response, supports custom evaluation criteria, explicitly documents calibrating trust scores against human quality ratings, and publishes benchmarks (Nov 2025) for auto-flagging incorrect structured outputs for human review.",
   "load_bearing": false,
   "evidence": "Cleanlab TLM documentation (help.cleanlab.ai/tlm/: quickstart, custom-eval calibration-against-human-ratings tutorial, advanced usage with low-score explanations) and 2025-11-18 structured-outputs benchmark blog post, all fetched via search with page content confirmed. Vendor source: benchmark numbers are self-reported.",
   "implication": "The 'score every judgment, calibrate the score against a human-rated set, route low-trust items to humans' pattern is already productized - the foundation can specify this as a component contract (per-verdict trust score + human-calibration hook) rather than inventing it, while keeping vendor-neutrality.",
   "source": {
    "raw": "Cleanlab Trustworthy Language Model (TLM) documentation and benchmark | https://help.cleanlab.ai/tlm/ | 2025-11-18 (benchmark); docs current 2026 | practitioner",
    "title": "Cleanlab Trustworthy Language Model (TLM) documentation and benchmark",
    "url": "https://help.cleanlab.ai/tlm/",
    "date": "2025-11-18 (benchmark); docs current 2026",
    "type": "practitioner"
   },
   "venue": {
    "peer_reviewed": null,
    "marker": "practitioner"
   },
   "same_source_claims": [],
   "in_corpus_verification": null,
   "external_verification": {
    "verdict": "partially_confirmed",
    "checked": [
     "attaches a real-time trustworthiness score to responses from any LLM",
     "supports custom evaluation criteria (a 'Custom Evaluation Criteria' tutorial is listed)"
    ],
    "not_visible": [
     "documentation of calibrating trust scores against human quality ratings (not on this overview page)",
     "the 2025-11-18 structured-outputs benchmark for auto-flagging incorrect outputs for human review (not on this overview page)"
    ],
    "quote": "scores the trustworthiness of responses from any LLM in real-time ... state-of-the-art trustworthiness scores for any LLM application",
    "notes": "The two core capabilities (real-time trust score for any LLM; custom eval criteria) are confirmed on the given overview URL. The human-rating calibration tutorial and the Nov-2025 structured-outputs benchmark cited in the evidence live on sub-pages (custom-eval / benchmark blog) that are outside the given URL and were not fetched; they are not visible on help.cleanlab.ai/tlm/ itself.",
    "checked_at": "2026-07-15",
    "retrieval": "landing"
   },
   "source_retrieval_meta": {
    "retrieval": "landing",
    "fetched": [
     "https://help.cleanlab.ai/tlm/"
    ],
    "resolved_title": "Trustworthy Language Model (TLM) | Cleanlab Documentation",
    "resolved_date": "not dated on overview page",
    "title_match": true
   },
   "used_on": [],
   "decision_use": {
    "grade": "A",
    "basis": "tooling fact",
    "measured_on": "structural",
    "note": "Cleanlab TLM: trust-scored judgments with calibration tooling. Vendor tooling fact, current.",
    "models_measured": []
   }
  },
  {
   "id": "RR-01",
   "domain": "rubrics-recent",
   "area": null,
   "claim": "OpenAI HealthBench meta-evaluates its LLM grader per-criterion against physician majority grades using Macro F1 (met/not-met, class-balanced), and reports the grader as a percentile of individual physicians: GPT-4.1 grader MF1 = 0.709, exceeding the average physician in 5 of 7 themes, while physician-vs-physician MF1 is only 0.569-0.730 with wide individual spread.",
   "load_bearing": true,
   "evidence": "Full paper text (Section 8.1, Tables 5-6) confirmed: 60,896 meta-examples over 34 physician-consensus criteria; 'typical physician' baseline computed by scoring each physician against the others exactly as the model is scored; grader model choice matters (GPT-4.1 0.709 > o4-mini 0.692 > o3 0.681 > GPT-4.1-nano 0.580); prompt AND individual criterion phrasings were tuned 'so that their intent was unmistakable to the grader'. Overall benchmark std across 16 full runs is ~0.002.",
   "implication": "The 2025-26 standard validation design for a rubric judge: per-criterion binary meta-eval against adjudicated expert-consensus labels with a chance-robust class-balanced statistic, reporting the judge as a percentile of individual human raters - not raw accuracy vs. a single reviewer. The human agreement ceiling is empirically low (MF1 ~0.6-0.7 even among physicians on physician-written criteria), so 'exceeds median human' is an attainable and meaningful bar, and consensus-adjudicated labels (not single-reviewer labels) must be the target. Also: rubric criterion text is a tunable judge-alignment artifact, not fixed scripture.",
   "source": {
    "raw": "HealthBench: Evaluating Large Language Models Towards Improved Human Health (OpenAI) | https://arxiv.org/html/2505.08775v1 | 2025-05-13 | primary",
    "title": "HealthBench: Evaluating Large Language Models Towards Improved Human Health (OpenAI)",
    "url": "https://arxiv.org/html/2505.08775v1",
    "date": "2025-05-13",
    "type": "primary"
   },
   "venue": {
    "peer_reviewed": false,
    "marker": "preprint"
   },
   "same_source_claims": [
    "RR-02"
   ],
   "in_corpus_verification": {
    "verdict": "confirmed",
    "transcript": "Fetched https://arxiv.org/html/2505.08775v1 (marked arXiv:2505.08775v1 [cs.CL] 13 May 2025 - date correct). All claim elements verified against Section 8.1 / Tables 5-7: (1) meta-evaluation uses macro F1 on binary met/not-met per consensus criterion, explicitly chosen for class imbalance (random baseline MF1=0.50); (2) GPT-4.1 grader MF1 = 0.709, best of graders tested (o4-mini 0.692, o3 0.681, GPT-4.1 mini 0.661, GPT-4.1 nano 0.580); (3) over 60,896 meta-examples across the 34 physician-consensus criteria (avg 1,791/criterion); (4) 'typical physician' baseline computed by scoring each physician against the others the same way as the model, reported as a percentile - GPT-4.1 exceeds the average physician in 5 of 7 themes (below only expertise-tailored communication 0.610 vs 0.618 and health data tasks 0.683 vs 0.730), upper half in 6/7, above 33rd percentile in all; (5) theme-level weighted-average physician-vs-physician MF1 ranges 0.569 (response depth) to 0.730 (health data tasks), with paper noting physician-physician and model-physician agreement both span ~55-75% and wide individual spread; (6) grading prompt AND individual criterion phrasings were tuned so intent was unmistakable to the grader (paper also cautions GPT-4.1's top rank may be partly explained by its use during prompt tuning); (7) std across 16 full runs ~0.002 (Table 7: 0.0016-0.0029 by model). Supersession check: no arXiv v2 exists; 2026 activity is HealthBench Professional (OpenAI, April 2026) and third-party replications (e.g., open-source graders: Kimi-K2 0.693, Qwen3-235B 0.681 vs GPT-4.1 0.709) - these extend, not contradict, the original meta-eval numbers. Only pedantic nuance: 0.569-0.730 are theme-level weighted averages of physician MF1, not the full range across individual physicians (individual spread is wider) - the claim's own 'wide individual spread' wording already captures this correctly.",
    "corrected": null
   },
   "external_verification": {
    "verdict": "confirmed",
    "checked": [
     "per-criterion Macro F1 (met/not-met, class-balanced) meta-evaluation of the grader",
     "GPT-4.1 grader MF1 = 0.709; exceeds average physician score in five out of seven themes (verbatim)",
     "physician-vs-physician weighted MF1 range 0.569 (Response depth) to 0.730 (Health data tasks)",
     "o4-mini 0.692, o3 0.681, GPT-4.1-nano 0.580 (GPT-4.1-mini 0.661)",
     "over 60,896 meta-examples",
     "typical-physician baseline: each physician scored against the others, excluding their own grades",
     "std ~0.002 across 16 full runs (Table 7: 0.0016-0.0029)"
    ],
    "not_visible": [],
    "quote": "It \"exceeds the average physician score in five out of seven themes\" - verbatim confirmed",
    "notes": "All eight cited numbers match verbatim. NUANCE: the grading ground truth is each INDIVIDUAL physician's grade, not a majority-vote label; majority agreement was used earlier to assign consensus criteria to examples, not as the grading target. The claim's phrase 'physician majority grades' is therefore a minor mischaracterization, but does not affect any headline number or the core method.",
    "checked_at": "2026-07-15",
    "retrieval": "fulltext"
   },
   "source_retrieval_meta": {
    "retrieval": "fulltext",
    "fetched": [
     "https://arxiv.org/html/2505.08775v1"
    ],
    "resolved_title": "HealthBench: Evaluating Large Language Models Towards Improved Human Health",
    "resolved_date": "2025-05-13 (arXiv 2505.08775v1)",
    "title_match": true
   },
   "used_on": [],
   "decision_use": {
    "grade": "B",
    "basis": "current-gen practice + capability",
    "measured_on": "model-outputs",
    "note": "HealthBench grader meta-evaluation (2025, GPT-4.1/o3-era graders). The per-criterion meta-evaluation PATTERN is the durable input; MF1 0.709 is tier-bound.",
    "models_measured": [
     "GPT-4.1",
     "o4-mini",
     "o3",
     "GPT-4.1-nano"
    ]
   }
  },
  {
   "id": "RR-02",
   "domain": "rubrics-recent",
   "area": null,
   "claim": "HealthBench's production rubric format is per-item (conversation-specific) criteria written by 262 physicians (48,562 criteria), each an independently-judged binary met/not-met check with a nonzero weight from -10 to +10 (negative criteria encode pitfalls), aggregated as a weighted sum normalized by max attainable positive points.",
   "load_bearing": false,
   "evidence": "Confirmed in fetched full text: 'Each rubric criterion has an associated nonzero point value between -10 and 10, with negative points used for criteria that are undesirable'; grader judges each criterion independently; item score can go negative; the vast majority of criteria were written specifically for one example.",
   "implication": "Reference design for the AutoQA verdict aggregation question: per-item compiled criteria (not one generic project rubric), explicit negative-weighted pitfall criteria as first-class citizens, independent binary judging per criterion, and a transparent weighted-sum aggregation - i.e., claim-level findings with an arithmetic, auditable roll-up rather than a holistic judge verdict.",
   "source": {
    "raw": "HealthBench (rubric structure, Section 3) | https://arxiv.org/html/2505.08775v1 | 2025-05-13 | primary",
    "title": "HealthBench (rubric structure, Section 3)",
    "url": "https://arxiv.org/html/2505.08775v1",
    "date": "2025-05-13",
    "type": "primary"
   },
   "venue": {
    "peer_reviewed": false,
    "marker": "preprint"
   },
   "same_source_claims": [
    "RR-01"
   ],
   "in_corpus_verification": null,
   "external_verification": {
    "verdict": "confirmed",
    "checked": [
     "rubrics created by 262 physicians",
     "48,562 unique criteria across all conversations; conversation-specific",
     "each criterion nonzero point value between -10 and 10; negative for undesirable/pitfalls",
     "grader judges each criterion independently, binary met/not-met",
     "aggregated as weighted sum divided by maximum possible score; per-example score can go negative"
    ],
    "not_visible": [],
    "quote": "an associated nonzero point value between -10 and 10, with negative points used for criteria that are undesirable",
    "notes": "Full text confirms all elements: 262 physicians, 48,562 criteria, -10..10 point range, independent binary judging, weighted-sum normalization by max attainable, and negative-capable item scores.",
    "checked_at": "2026-07-15",
    "retrieval": "fulltext"
   },
   "source_retrieval_meta": {
    "retrieval": "fulltext",
    "fetched": [
     "https://arxiv.org/html/2505.08775v1"
    ],
    "resolved_title": "HealthBench: Evaluating Large Language Models Towards Improved Human Health",
    "resolved_date": "2025-05-13 (arXiv 2505.08775v1)",
    "title_match": true
   },
   "used_on": [
    {
     "page": "decisions.html",
     "anchor": "a2",
     "label": "Decisions - first project and gold funding"
    }
   ],
   "decision_use": {
    "grade": "A",
    "basis": "production practice",
    "measured_on": "structural",
    "note": "48,562 physician-written per-item criteria: what expert-grounded rubric authoring costs at production scale. Practice fact.",
    "models_measured": []
   }
  },
  {
   "id": "RR-04",
   "domain": "rubrics-recent",
   "area": null,
   "claim": "RubricBench (Mar 2026) measures a stable ~26-27 point preference-accuracy gap between self-generated and human-annotated rubrics across all judge backbones (e.g., DeepSeek-v3.2 57.8% vs 84.9%; Gemini-3-Flash 58.0% vs 85.3%), and scaling test-time compute (more sampled rubrics, refinement) does not close it - 'rubric formation', not judge reasoning, is the bottleneck.",
   "load_bearing": true,
   "evidence": "Full HTML (arXiv 2603.01562) confirmed: 1,147 adversarial pairwise comparisons where surface cues mislead (rejected responses longer/better-formatted/more confident); expert rubrics are 2-10 atomic binary checks derived only from the instruction; LLM-generated rubrics over-produce unnecessary (17.9% vs 10.1%) and overly rigid (12.8% vs 7.7%) rules; checklists of 13+ items add noise ('attention displacement'); even with human rubrics accuracy plateaus ~85% because judges treat must-have constraints as soft signals; safety domain worst for self-generated rubrics (~25-30% vs >90% human).",
   "implication": "Directly bounds the per-project customization contract: instruction-only automated rubric compilation (no expert exemplars/references) leaves a large fidelity gap that more compute cannot buy back - an expert-grounding step (reference answers, adjudicated exemplars, or human rubric sign-off) is structurally required, not optional. Also: cap compiled checklist length (~<13 active criteria per judge call), and give hard constraints enforcement semantics (gating) rather than listing them as soft criteria.",
   "source": {
    "raw": "RubricBench: Aligning Model-Generated Rubrics with Human Standards | https://arxiv.org/html/2603.01562v1 | 2026-03-02 | academic",
    "title": "RubricBench: Aligning Model-Generated Rubrics with Human Standards",
    "url": "https://arxiv.org/html/2603.01562v1",
    "date": "2026-03-02",
    "type": "academic"
   },
   "venue": {
    "peer_reviewed": false,
    "marker": "preprint"
   },
   "same_source_claims": [],
   "in_corpus_verification": {
    "verdict": "partially_confirmed",
    "transcript": "Source check (arXiv 2603.01562v1, fetched): the RubricBench paper exists, is dated 2026-03-02 (v2 2026-03-03), and confirms nearly every specific: 1,147 adversarial pairwise comparisons; DeepSeek-v3.2 57.8% self-generated vs 84.9% human (vanilla 38.8%); Gemini-3-Flash 58.0% vs 85.3% (vanilla 56.4%); test-time scaling fails (Rub@4->Rub@32 flat or declining, refinement depth declining, while scaling HUMAN rubrics helps 75.4->85.3); unnecessary-rule 17.9% vs 10.1% and rigid-rule 12.8% vs 7.7%; 13+-item checklists / 'Attention Displacement' with >70% hallucination rates; ~85% human-rubric plateau; safety ~25-30% self-generated vs >90% human. Two corrections: (1) the gap across ALL backbones is ~22-28 points, not a tight '~26-27' - 26-27 fits only the two cited models; (2) the 'rubric formation is the bottleneck' framing is the paper's interpretation and, more importantly, has been operationally SUPERSEDED in part: 'Support Vector Rubrics' (arXiv 2606.08077, June 2026, PKU/USTC) closes the gap on RubricBench to 0.3 points (82.8 vs 83.1 human-oracle, vs 59.0 self-generated with same GPT-OSS-120B judge) via max-margin contrastive rubric-bank learning over preference data - so the gap is stable under naive self-generation and test-time compute, but not under trained discriminative rubric construction. Safety remains the residual weak domain even for SVR (83.8 vs 92.5). Claim is accurate as a description of the March paper; the 'does not close' generalization no longer holds as of June 2026.",
    "corrected": "RubricBench (arXiv 2603.01562, Mar 2026) measures a stable ~22-28 point preference-accuracy gap between self-generated and human-annotated rubrics across judge backbones (DeepSeek-v3.2 57.8% vs 84.9%; Gemini-3-Flash 58.0% vs 85.3%), and scaling test-time compute (more sampled rubrics, refinement depth) does not close it - the paper attributes the bottleneck to rubric formation, not judge reasoning. However, follow-up work (Support Vector Rubrics, arXiv 2606.08077, Jun 2026) closes the gap to ~0.3 points on RubricBench via max-margin rubric-bank learning from preference data, showing the gap is specific to naive self-generation rather than fundamental."
   },
   "external_verification": {
    "verdict": "confirmed",
    "checked": [
     "gap 'remains stable at ~26%' / 'a severe 27% accuracy gap' (Table 3 Delta range +22.1 to +28.3)",
     "DeepSeek-v3.2: Self-Gen 57.8 vs Human 84.9 (Delta +27.1)",
     "Gemini-3-Flash: Self-Gen 58.0 vs Human 85.3 (Delta +27.3)",
     "1,147 adversarial pairwise comparisons",
     "unnecessary rules N=1 17.9% vs 10.1%; overly rigid R=5 12.8% vs 7.7%",
     "checklists ~13.2/15.4/14.6 items; 'Attention Displacement'",
     "'even with human rubrics, accuracy plateaus around 85%'",
     "safety self-generated ~25-30% vs human >90%",
     "test-time compute does NOT close gap (GPT-4o-mini 48.0%->46.8%; refinement 46.7->46.4->45.7)"
    ],
    "not_visible": [],
    "quote": "even with human rubrics, accuracy plateaus around 85% rather than approaching 100%",
    "notes": "Every specific number in the claim matches the full-text tables. The claim's '~26-27 point... across all judge backbones' aligns with the paper's own ~26%/27% summary, though per-backbone Delta ranges +22.1 to +28.3. This is stronger than the batch's in_corpus 'partially_confirmed' - full text confirms all specifics, so verdict is confirmed.",
    "checked_at": "2026-07-15",
    "retrieval": "fulltext"
   },
   "source_retrieval_meta": {
    "retrieval": "fulltext",
    "fetched": [
     "https://arxiv.org/html/2603.01562v1"
    ],
    "resolved_title": "RubricBench: Aligning Model-Generated Rubrics with Human Standards",
    "resolved_date": "arXiv 2603.01562 (March 2026, per arXiv ID; source_date 2026-03-02 consistent)",
    "title_match": true
   },
   "used_on": [],
   "decision_use": {
    "grade": "B",
    "basis": "current-gen capability",
    "measured_on": "model-outputs",
    "note": "RubricBench (Mar 2026, DeepSeek-v3.2/Gemini-3-era): 26-27pp self-vs-expert rubric gap that test-time compute doesn't close. Current-window; deployment-tier gap unknown.",
    "models_measured": [
     "DeepSeek-v3.2",
     "Gemini-3-Flash"
    ]
   }
  },
  {
   "id": "RR-05",
   "domain": "rubrics-recent",
   "area": null,
   "claim": "EvalGen ('Who Validates the Validators', UIST 2024) established criteria drift: evaluation criteria cannot be fully specified a priori because grading outputs itself changes the criteria - 'users need criteria to grade outputs, but grading outputs helps users define criteria' - and some criteria are dependent on the specific outputs observed.",
   "load_bearing": true,
   "evidence": "Fetched abstract/paper confirmed the catch-22 formulation, the finding that criteria appear dependent on observed outputs rather than definable in advance, and the conclusion that judge-alignment is inherently iterative and subjective; EvalGen operationalizes this by generating candidate assertions/judge prompts and selecting those best aligned with accumulated human grades.",
   "implication": "The rubric compilation pipeline must be a loop, not a compile step: initial rubric -> human grades a sample -> criteria are revised -> re-validate. Foundational consequence for the AutoQA: budget a per-project calibration phase where criteria are expected to change, version every rubric, and treat early judge-human disagreement partly as rubric-specification signal (rewrite) rather than judge noise (vote harder) - the two-builds fork in design question 5 has an evidence-backed default.",
   "source": {
    "raw": "Who Validates the Validators? Aligning LLM-Assisted Evaluation of LLM Outputs with Human Preferences | https://arxiv.org/abs/2404.12272 | 2024-10-13 | academic",
    "title": "Who Validates the Validators? Aligning LLM-Assisted Evaluation of LLM Outputs with Human Preferences",
    "url": "https://arxiv.org/abs/2404.12272",
    "date": "2024-10-13",
    "type": "academic"
   },
   "venue": {
    "peer_reviewed": false,
    "marker": "preprint"
   },
   "same_source_claims": [
    "CT-04"
   ],
   "in_corpus_verification": null,
   "external_verification": {
    "verdict": "partially_confirmed",
    "checked": [
     "criteria drift established as a named phenomenon",
     "'users need criteria to grade outputs, but grading outputs helps users define criteria'",
     "some criteria are dependent on the specific LLM outputs observed",
     "alignment is subjective and iterative; EvalGen generates candidate assertions/judge prompts"
    ],
    "not_visible": [
     "'UIST 2024' venue attribution (arXiv page lists no conference/journal venue)"
    ],
    "quote": "\"users need criteria to grade outputs, but grading outputs helps users define criteria\" ... \"some criteria appears dependent on the specific LLM outputs observed\" ... the study \"underscores the subjectivity and iterative process of alignment\".",
    "notes": "Criteria-drift core confirmed verbatim. Only the venue label 'UIST 2024' is not visible on the arXiv abstract page (Comments: 16 pages, 4 figures, 2 tables; no venue listed).",
    "checked_at": "2026-07-15",
    "retrieval": "abstract"
   },
   "source_retrieval_meta": {
    "retrieval": "abstract",
    "fetched": [
     "https://arxiv.org/abs/2404.12272"
    ],
    "resolved_title": "Who Validates the Validators? Aligning LLM-Assisted Evaluation of LLM Outputs with Human Preferences",
    "resolved_date": "v1 2024-04-18 (arXiv page shows no venue; source_date 2024-10-13 not confirmed on page)",
    "title_match": true
   },
   "used_on": [
    {
     "page": "pilot.html",
     "anchor": "main",
     "label": "Pilot - gates and objectives"
    },
    {
     "page": "decisions.html",
     "anchor": "d2",
     "label": "Decisions - rubric authority"
    }
   ],
   "decision_use": {
    "grade": "A",
    "basis": "human science",
    "measured_on": "human-work",
    "note": "EvalGen criteria drift: human graders' criteria change as they grade. Human behavior; durable.",
    "models_measured": []
   }
  },
  {
   "id": "RR-06",
   "domain": "rubrics-recent",
   "area": null,
   "claim": "RIFT (Snorkel AI, Apr 2026) provides the first rubric failure-mode taxonomy - 8 modes in 3 categories (Reliability: Subjective, Non-Atomic, Ungrounded; Content Validity: Misaligned/Rigid, Missing Criteria; Consequential Validity: Hackable, Low Signal, Redundant) - and shows failure-mode density correlates with judge-human misalignment (r=0.162, p=0.0021); synthetic rubrics skew Subjective (86.7% vs 52.6% for human-crafted) while human rubrics skew Misaligned/Rigid (63.2% vs 20.0%); automated LLM detection reaches F1 0.925 for Subjective but ~0.000 for Hackable.",
   "load_bearing": false,
   "evidence": "Full HTML (arXiv 2604.01375v2) confirmed all numbers: 85 rubrics / 255 expert annotations across 5 sources (human: AdvancedIF, ResearchRubrics; synthetic: WildChecklists, OpenRubrics, AutoRubrics); taxonomy built via grounded theory, mean Cohen's kappa 0.64; diagnostics include LLM-as-judge classifiers (GPT-5.2, Gemini 3 Pro), inter-rater reliability across labeler models, alignment to a strong reference, and reward-variance probes.",
   "implication": "Rubric QA/linting is now a named, tooled 2026 practice: run compiled rubrics through a RIFT-style lint before go-live. Two structural lessons: (1) human and LLM rubric authors fail in complementary ways, motivating co-authoring rather than either alone; (2) hackability/gameability is the one failure mode automated diagnostics cannot detect - the anti-gaming review of the AutoQA's compiled rubrics must be a human red-team task, permanently.",
   "source": {
    "raw": "RIFT: A RubrIc Failure Mode Taxonomy and Automated Diagnostics (Snorkel AI) | https://arxiv.org/html/2604.01375v2 | 2026-04-20 | academic",
    "title": "RIFT: A RubrIc Failure Mode Taxonomy and Automated Diagnostics (Snorkel AI)",
    "url": "https://arxiv.org/html/2604.01375v2",
    "date": "2026-04-20",
    "type": "academic"
   },
   "venue": {
    "peer_reviewed": false,
    "marker": "preprint"
   },
   "same_source_claims": [],
   "in_corpus_verification": null,
   "external_verification": {
    "verdict": "confirmed",
    "checked": [
     "8 modes in 3 categories: Subjective, Non-Atomic, Ungrounded (Reliability); Misaligned/Rigid, Missing Criteria (Content Validity); Hackable, Low Signal, Redundant Criteria (Consequential)",
     "Pearson's r = 0.162, p = 0.0021",
     "Table 2: Subjective Human 52.6% vs Synthetic 86.7%",
     "Table 2: Misaligned/Rigid Human 63.2% vs Synthetic 20.0%",
     "Table 3: Subjective F1 0.925; Hackable F1 0.000",
     "85 rubrics with 255 expert annotations drawn from five data sources",
     "AdvancedIF, ResearchRubrics, WildChecklists, OpenRubrics, and AutoRubrics",
     "developed using grounded theory",
     "0.64 average Cohen's kappa",
     "classifiers GPT-5.2 and Gemini 3 Pro"
    ],
    "not_visible": [],
    "quote": "Pearson's r = 0.162, p = 0.0021 ... we analyze 85 rubrics with 255 expert annotations drawn from five data sources",
    "notes": "All 8 modes/3 categories, both correlation stats, both Table 2 prevalence pairs, both Table 3 F1 extremes, the 85/255 counts, five sources, grounded theory, kappa 0.64, and both classifier models confirmed. Minor internal inconsistency in the paper: abstract says 'up to 0.925 F1' while a contributions bullet says 'up to 0.86 F1'; the claim uses the 0.925 (Table 3) figure.",
    "checked_at": "2026-07-15",
    "retrieval": "fulltext"
   },
   "source_retrieval_meta": {
    "retrieval": "fulltext",
    "fetched": [
     "https://arxiv.org/html/2604.01375v2"
    ],
    "resolved_title": "RIFT: A RubrIc Failure Mode Taxonomy and Automated Diagnostics",
    "resolved_date": "2026-04-20 (v2)",
    "title_match": true
   },
   "used_on": [],
   "decision_use": {
    "grade": "B",
    "basis": "current-gen capability",
    "measured_on": "model-outputs",
    "note": "RIFT (Apr 2026, GPT-5.2/Gemini-3-era linting): Subjective detectable (F1 .925), Hackable not (~0). Current-window; the human-red-team conclusion for anti-gaming review stands until a later tier changes it.",
    "models_measured": [
     "GPT-5.2",
     "Gemini 3"
    ]
   }
  },
  {
   "id": "RR-07",
   "domain": "rubrics-recent",
   "area": null,
   "claim": "Support Vector Rubrics (Jun 2026) identifies the mechanism behind the LLM-vs-human rubric gap - 'self-generated rubrics describe good responses, whereas effective criteria must discriminate between close candidates' - and closes the RubricBench gap from 24.1 to 0.3 points by mining contrastive rubrics from preference pairs (support-pair selection plus adversarial probing of hard negatives).",
   "load_bearing": false,
   "evidence": "Fetched abstract (arXiv 2606.08077, Sun et al.) confirmed the objective-mismatch framing, the max-margin method (rubric bank mined from preference data, prompt-conditioned selector, iterative refinement on hard negatives), transfer of the rubric bank across judges without retraining, and competitiveness with dedicated reward models on RewardBench 1/2 and RM-Bench.",
   "implication": "The highest-value compilation artifacts are contrastive: borderline pass/fail adjudication pairs, not descriptions of quality. For the AutoQA's exemplar question (5 vs 50 vs 500), this says invest in hard-negative/near-miss pairs per criterion and derive discriminative criteria from them - descriptive 'what good looks like' statements alone are the documented failure pattern.",
   "source": {
    "raw": "Support Vector Rubrics: Closing the Gap Between Self-Generated and Human Rubrics | https://arxiv.org/abs/2606.08077 | 2026-06-06 | academic",
    "title": "Support Vector Rubrics: Closing the Gap Between Self-Generated and Human Rubrics",
    "url": "https://arxiv.org/abs/2606.08077",
    "date": "2026-06-06",
    "type": "academic"
   },
   "venue": {
    "peer_reviewed": false,
    "marker": "preprint"
   },
   "same_source_claims": [],
   "in_corpus_verification": null,
   "external_verification": {
    "verdict": "confirmed",
    "checked": [
     "objective mismatch: self-generated rubrics describe good responses, effective criteria must discriminate between close candidates",
     "narrows gap to human reference rubrics from 24.1 to 0.3 points",
     "mines contrastive features from preference pairs into a rubric bank",
     "support-pair selection and adversarial probing of hard negatives",
     "max-margin boundary learning over preference data",
     "learned bank transfers across judges without retraining",
     "competitive with dedicated reward models on RewardBench 1&2 and RM-Bench"
    ],
    "not_visible": [],
    "quote": "self-generated rubrics describe good responses, whereas effective criteria must discriminate between close candidates ... SVR narrows the gap to human reference rubrics from 24.1 to 0.3 points ... the learned bank transfers across judges without retraining.",
    "notes": "Every element of the claim confirmed verbatim from the abstract, including the exact 24.1 -> 0.3 figure. Authors Sun et al. match.",
    "checked_at": "2026-07-15",
    "retrieval": "abstract"
   },
   "source_retrieval_meta": {
    "retrieval": "abstract",
    "fetched": [
     "https://arxiv.org/abs/2606.08077"
    ],
    "resolved_title": "Support Vector Rubrics: Closing the Gap Between Self-Generated and Human Rubrics",
    "resolved_date": "2026-06-06",
    "title_match": true
   },
   "used_on": [],
   "decision_use": {
    "grade": "B",
    "basis": "current-gen result",
    "measured_on": "model-outputs",
    "note": "Support Vector Rubrics (Jun 2026): contrastive mining closes the rubric gap to ~0.3pp given preference pairs. Current-window; depends on adjudication data the pilot creates.",
    "models_measured": []
   }
  },
  {
   "id": "RR-08",
   "domain": "rubrics-recent",
   "area": null,
   "claim": "PaperBench (OpenAI, Apr 2025) demonstrates the extreme end of decomposition: 8,316 individually-graded binary leaf criteria across 20 papers, sibling-relative manual weights, rubrics co-developed with each paper's original authors over multiple weeks per rubric; an o3-mini judge then grades leaves at F1 = 0.83 vs expert human judgments for ~$66/paper versus ~12 hours of expert grading time.",
   "load_bearing": false,
   "evidence": "Full paper (arXiv 2504.01848) confirmed: each leaf is pass/fail; parent scores are weighted averages propagating to a root score; weights reflect importance relative to siblings, not difficulty; JudgeEval is a separate benchmark for judges (o1 scored 0.84 but ~$830/paper - o3-mini chosen as cost-effective); human expert grading estimated at 'tens of hours' per paper.",
   "implication": "Atomic binary decomposition of genuinely expert judgments is feasible and auto-gradable at usable-but-imperfect agreement (F1 0.83), with ~100x cost reduction over expert grading - but rubric authoring cost is weeks of joint expert time per complex artifact. For AutoQA: fine decomposition pays at grading time and the judge gets its own meta-benchmark (JudgeEval pattern), but the compilation budget, not the judging budget, is the binding constraint per project.",
   "source": {
    "raw": "PaperBench: Evaluating AI's Ability to Replicate AI Research (OpenAI) | https://arxiv.org/html/2504.01848v2 | 2025-04-02 | primary",
    "title": "PaperBench: Evaluating AI's Ability to Replicate AI Research (OpenAI)",
    "url": "https://arxiv.org/html/2504.01848v2",
    "date": "2025-04-02",
    "type": "primary"
   },
   "venue": {
    "peer_reviewed": false,
    "marker": "preprint"
   },
   "same_source_claims": [],
   "in_corpus_verification": null,
   "external_verification": {
    "verdict": "confirmed",
    "checked": [
     "8,316 leaf nodes / individually gradable tasks across 20 papers",
     "each leaf scored binary (1 pass / 0 fail)",
     "node weight = importance relative to siblings, not implementation difficulty",
     "rubrics co-developed with the paper's author(s); took multiple weeks per paper",
     "o3-mini SimpleJudge: F1 0.83 at $66 USD/paper (most cost-effective)",
     "human expert grading on the order of tens of hours per paper (cost modeled at 12 hrs x $100/hr)",
     "JudgeEval is a separate benchmark for evaluating automated judges",
     "o1 judge: 0.84 F1 at $830 USD/paper"
    ],
    "not_visible": [],
    "quote": "Across the 20 papers in PaperBench there are 8,316 leaf nodes. ... o3-mini with the SimpleJudge scaffolding is the most cost-effective, with an F1 score of 0.83 at $66 USD per paper",
    "notes": "All decomposition and cost figures confirmed verbatim in the full text, including sibling-relative weighting, weeks-long author co-development, the o3-mini 0.83/$66 vs o1 0.84/$830 tradeoff, JudgeEval as a separate judge benchmark, and both the 'tens of hours' and '12 hours at $100/hr' human-grading figures.",
    "checked_at": "2026-07-15",
    "retrieval": "fulltext"
   },
   "source_retrieval_meta": {
    "retrieval": "fulltext",
    "fetched": [
     "https://arxiv.org/html/2504.01848v2"
    ],
    "resolved_title": "PaperBench: Evaluating AI's Ability to Replicate AI Research",
    "resolved_date": "2025-04-02",
    "title_match": true
   },
   "used_on": [],
   "decision_use": {
    "grade": "B",
    "basis": "current-gen practice + capability",
    "measured_on": "model-outputs",
    "note": "PaperBench (2025, o1/o3-era judge F1 0.83 at ~1/100 cost). Decomposition-at-scale pattern durable; judge F1 and cost tier-bound.",
    "models_measured": [
     "o3-mini",
     "o1"
    ]
   }
  },
  {
   "id": "RR-09",
   "domain": "rubrics-recent",
   "area": null,
   "claim": "OpenAI's official grader documentation prescribes that every model grader be validated on its own eval built from expert-graded answers (grader score ordering must reproduce expert ranking), and names the canonical reward-hacking detection signature: the graded model scores well on the model-grader eval while doing poorly on expert human evaluations.",
   "load_bearing": false,
   "evidence": "Fetched live doc (developers.openai.com/api/docs/guides/graders, accessed 2026-07-14): grader taxonomy = string_check, text_similarity, score_model (LLM judge with numeric range), python (sandboxed), multigrader (formula-combined sub-graders, RFT only); guidance: iterate grader prompts, use few-shot examples of good/fair/poor answers, prefer smooth over binary scores, add edge cases to the grader eval over time, 'guard against reward hacking'. Page carries a deprecation note for graders within the evals/fine-tuning workflows they support.",
   "implication": "Lab-published operational doctrine matching the AutoQA's needs: (1) mixed grader stacks - deterministic checks composed with LLM judges via explicit formulas - are the production pattern for the objective/subjective split; (2) a standing expert-audit channel is the designed-in hack detector (grader-vs-expert divergence), i.e., the independent expert-audit sampling channel is a permanent component in the reference architecture, not a bootstrap phase.",
   "source": {
    "raw": "OpenAI API documentation: Graders | https://developers.openai.com/api/docs/guides/graders | 2026-07-14 | primary",
    "title": "OpenAI API documentation: Graders",
    "url": "https://developers.openai.com/api/docs/guides/graders",
    "date": "2026-07-14",
    "type": "primary"
   },
   "venue": {
    "peer_reviewed": null,
    "marker": "primary"
   },
   "same_source_claims": [],
   "in_corpus_verification": null,
   "external_verification": {
    "verdict": "confirmed",
    "checked": [
     "grader taxonomy: string check, text similarity, score model (LLM judge, numeric), python (code execution), multigrader (combine/nest, reinforcement fine-tuning only)",
     "validate a model grader on its own eval built from expert-graded answers",
     "grader scores must reproduce the expert ranking (answer_1 > answer_2 > answer_3)",
     "reward/grader hacking signature: scores high on model grader evals but poorly on expert human evaluations",
     "guidance: smooth score not pass/fail; few-shot great/fair/poor examples; add edge cases over time; guard against reward hacking",
     "deprecation note: OpenAI deprecating graders within evals/fine-tuning workflows"
    ],
    "not_visible": [],
    "quote": "A model that's hacked the grader will score highly on model grader evals but score poorly on expert human evaluations ... Produce a smooth score, not a pass/fail stamp ... In reinforcement fine-tuning, you can nest and combine graders by using multigraders",
    "notes": "All prescribed elements of the claim are present verbatim in the live doc, including the named reward-hacking signature and the deprecation note.",
    "checked_at": "2026-07-15",
    "retrieval": "landing"
   },
   "source_retrieval_meta": {
    "retrieval": "landing",
    "fetched": [
     "https://developers.openai.com/api/docs/guides/graders"
    ],
    "resolved_title": "Graders (OpenAI API guide)",
    "resolved_date": "living doc; no publication date shown, accessed 2026-07-14",
    "title_match": true
   },
   "used_on": [],
   "decision_use": {
    "grade": "A",
    "basis": "production practice",
    "measured_on": "structural",
    "note": "OpenAI grader doctrine: every model grader validated on its own expert-graded eval. Practice prescription, current.",
    "models_measured": []
   }
  },
  {
   "id": "RR-10",
   "domain": "rubrics-recent",
   "area": null,
   "claim": "A cluster of 2026 papers (Rubric-ARM Feb 2026; EvoRubrics, ARBOR, AMARIS Jun 2026) converges on the position that static rubrics are an exploitable reward specification under optimization pressure and must be adapted/co-evolved during training, with Rubric-ARM jointly optimizing a rubric generator and judge via alternating RL from preference feedback.",
   "load_bearing": false,
   "evidence": "Rubric-ARM abstract opened and confirmed (prior approaches 'rely on static rubrics or disjoint training pipelines, which limits their adaptability'; alternating optimization with variance-reduction analysis; SOTA on judge benchmarks and downstream policy alignment). EvoRubrics/ARBOR/AMARIS confirmed at title/abstract level via search only - trend triangulated across 5+ 2026 arXiv titles, not deep-read.",
   "implication": "Attempters paid per accepted item exert the same optimization pressure as an RL policy: the 2026 consensus is that a frozen compiled rubric will be Goodharted, so rubric revision must be a designed, continuous loop (with regression gates), not an exceptional maintenance event. This directly supports treating criteria drift response and anti-gaming rubric rotation as foundational architecture.",
   "source": {
    "raw": "Rubric-ARM: Alternating Reinforcement Learning for Rubric-Based Reward Modeling (plus 2026 co-evolution cluster) | https://arxiv.org/abs/2602.01511 | 2026-02-02 | academic",
    "title": "Rubric-ARM: Alternating Reinforcement Learning for Rubric-Based Reward Modeling (plus 2026 co-evolution cluster)",
    "url": "https://arxiv.org/abs/2602.01511",
    "date": "2026-02-02",
    "type": "academic"
   },
   "venue": {
    "peer_reviewed": false,
    "marker": "preprint"
   },
   "same_source_claims": [],
   "in_corpus_verification": null,
   "external_verification": {
    "verdict": "partially_confirmed",
    "checked": [
     "Rubric-ARM jointly optimizes a rubric generator and a judge via reinforcement learning from preference feedback",
     "prior methods 'rely on static rubrics or disjoint training pipelines'",
     "alternating optimization strategy to mitigate non-stationarity",
     "reduces gradient variance; SOTA on benchmarks and improved policy alignment"
    ],
    "not_visible": [
     "'adaptability' wording (not found on page)",
     "'co-evolve' framing",
     "'exploitable reward specification under optimization pressure' framing",
     "EvoRubrics, ARBOR, AMARIS (separate 2026 papers, not at this URL; claim states these are search-only)"
    ],
    "quote": "jointly optimizes a rubric generator and a judge using reinforcement learning from preference feedback ... Unlike existing methods that rely on static rubrics or disjoint training pipelines",
    "notes": "Only Rubric-ARM is at this URL and its mechanism is confirmed verbatim. The 'cluster of 2026 papers converging' (EvoRubrics/ARBOR/AMARIS) is not verifiable here and the claim itself flags it as search-only. Actual title lacks the 'Rubric-ARM:' prefix and the '(plus 2026 co-evolution cluster)' suffix of source_title, but the substantive title matches. v1 2026-02-02, v2 2026-02-11.",
    "checked_at": "2026-07-15",
    "retrieval": "abstract"
   },
   "source_retrieval_meta": {
    "retrieval": "abstract",
    "fetched": [
     "https://arxiv.org/abs/2602.01511"
    ],
    "resolved_title": "Alternating Reinforcement Learning for Rubric-Based Reward Modeling in Non-Verifiable LLM Post-Training",
    "resolved_date": "2026-02-02",
    "title_match": true
   },
   "used_on": [],
   "decision_use": {
    "grade": "B",
    "basis": "current-gen literature position",
    "measured_on": "model-outputs",
    "note": "2026 RL-rubrics cluster: static rubrics are exploitable specifications under optimization. Conceptual; strengthens with optimizer capability.",
    "models_measured": []
   }
  },
  {
   "id": "RR-11",
   "domain": "rubrics-recent",
   "area": null,
   "claim": "Autorubric (Feb 2026, Rao & Callison-Burch) consolidates scattered rubric-evaluation techniques into one open-source framework with opinionated defaults - binary/ordinal/nominal criterion types, judge ensembles, few-shot calibration, bias mitigations, and psychometric reliability metrics - reporting 87% binary accuracy with moderate-to-substantial kappa on CHARM-100, and showing per-criterion explanations double as improvement signals (peer-review agent 0.47 -> 0.85, beating a 0.82 expert-curated baseline).",
   "load_bearing": false,
   "evidence": "Fetched abstract (arXiv 2603.00077, v2 2026-04-03, 52 pages) confirmed the unification claim ('scattered across papers with inconsistent terminology and partial implementations'), the three validation benchmarks (RiceChem 80% with 5-shot calibration; ResearcherBench 931 criteria cross-judge agreement; CHARM-100), and RL-reward use (+0.039 AdvancedIF, Wilcoxon p=0.032, positive IFEval transfer).",
   "implication": "By early 2026 rubric-based judging has a consolidated reference implementation whose defaults (ensembles, few-shot calibration on human-graded examples, psychometric reliability reporting per criterion type) can be adopted wholesale rather than reinvented; and the same per-criterion explanation trace that justifies a verdict is demonstrated to work as constructive improvement feedback - verdict machinery and feedback generation can share one artifact.",
   "source": {
    "raw": "Autorubric: Unifying Rubric-based LLM Evaluation | https://arxiv.org/abs/2603.00077 | 2026-02-13 | academic",
    "title": "Autorubric: Unifying Rubric-based LLM Evaluation",
    "url": "https://arxiv.org/abs/2603.00077",
    "date": "2026-02-13",
    "type": "academic"
   },
   "venue": {
    "peer_reviewed": false,
    "marker": "preprint"
   },
   "same_source_claims": [
    "F26-12"
   ],
   "in_corpus_verification": null,
   "external_verification": {
    "verdict": "confirmed",
    "checked": [
     "unifies techniques scattered across papers with inconsistent terminology and partial implementations",
     "binary, ordinal, and nominal criteria; single-judge and ensemble; few-shot calibration; bias mitigations; psychometric reliability metrics",
     "CHARM-100: 87% binary accuracy, moderate-to-substantial kappa",
     "peer review agent 0.47 to 0.85, above the 0.82 expert-curated baseline",
     "RiceChem 80% accuracy with 5-shot calibration",
     "ResearcherBench 931 criteria, cross-judge agreement",
     "RL reward AdvancedIF +0.039, Wilcoxon p=0.032, positive transfer to IFEval",
     "authors Delip Rao, Chris Callison-Burch"
    ],
    "not_visible": [],
    "quote": "raise a peer review agent's score from 0.47 to 0.85 (above the 0.82 expert-curated baseline)",
    "notes": "All nine asserted specifics verified verbatim in the abstract. v1 2026-02-13 matches source_date; v2 2026-04-03, 52 pages, both confirmed.",
    "checked_at": "2026-07-15",
    "retrieval": "abstract"
   },
   "source_retrieval_meta": {
    "retrieval": "abstract",
    "fetched": [
     "https://arxiv.org/abs/2603.00077"
    ],
    "resolved_title": "Autorubric: Unifying Rubric-based LLM Evaluation",
    "resolved_date": "arXiv v1 2026-02-13; v2 2026-04-03; 52 pages",
    "title_match": true
   },
   "used_on": [],
   "decision_use": {
    "grade": "A",
    "basis": "tooling fact",
    "measured_on": "structural",
    "note": "Autorubric open-source framework. Tooling availability.",
    "models_measured": []
   }
  },
  {
   "id": "RR-12",
   "domain": "rubrics-recent",
   "area": null,
   "claim": "LangChain's Align Evals (LangSmith, Jul 2025) productized judge validation as a standard workflow: humans grade a representative golden set per criterion, an 'alignment score' measures judge-vs-human match, unaligned cases are surfaced for prompt iteration, and each judge-prompt version is compared against a saved baseline alignment score.",
   "load_bearing": false,
   "evidence": "Fetched launch blog confirmed the four-step workflow (select criteria; select representative good+bad examples; human-grade expected scores; iterate evaluator prompt against alignment score), the stated problem ('Our evaluation scores don't match what we'd expect a human on our team to say'), inspiration from Eugene Yan's AlignEval, and roadmap items (alignment tracking over time, automatic prompt optimization).",
   "implication": "The industry-default minimum-viable gold set practice: small per-criterion human-graded golden sets with versioned alignment baselines, refreshed as unaligned cases surface - evidence that per-criterion (not per-item-holistic) alignment scoring with regression baselines is the practical validation loop the AutoQA should ship with from day one.",
   "source": {
    "raw": "Introducing Align Evals: Streamlining LLM Application Evaluation (LangChain) | https://www.langchain.com/blog/introducing-align-evals | 2025-07-29 | practitioner",
    "title": "Introducing Align Evals: Streamlining LLM Application Evaluation (LangChain)",
    "url": "https://www.langchain.com/blog/introducing-align-evals",
    "date": "2025-07-29",
    "type": "practitioner"
   },
   "venue": {
    "peer_reviewed": null,
    "marker": "practitioner"
   },
   "same_source_claims": [],
   "in_corpus_verification": null,
   "external_verification": {
    "verdict": "confirmed",
    "checked": [
     "Align Evals is a LangSmith feature (launched Jul 29 2025)",
     "four-step workflow: select criteria; select representative good+bad examples; human-grade expected scores (golden set); iterate evaluator prompt against alignment score",
     "'alignment score' measures evaluator vs human match; unaligned cases surfaced by sorting",
     "saved baseline alignment score to compare new prompt versions",
     "inspired by Eugene Yan's AlignEval",
     "roadmap: analytics for tracking over time; automatic prompt optimization"
    ],
    "not_visible": [],
    "quote": "Our evaluation scores don't match what we'd expect a human on our team to say",
    "notes": "Every element of the claim and evidence confirmed from the launch blog.",
    "checked_at": "2026-07-15",
    "retrieval": "landing"
   },
   "source_retrieval_meta": {
    "retrieval": "landing",
    "fetched": [
     "https://www.langchain.com/blog/introducing-align-evals"
    ],
    "resolved_title": "Introducing Align Evals: Streamlining LLM Application Evaluation",
    "resolved_date": "2025-07-29",
    "title_match": true
   },
   "used_on": [],
   "decision_use": {
    "grade": "A",
    "basis": "tooling fact",
    "measured_on": "structural",
    "note": "LangSmith Align Evals: judge-calibration workflow productized. Tooling fact.",
    "models_measured": []
   }
  },
  {
   "id": "IP-02",
   "domain": "industry-practice",
   "area": null,
   "claim": "Handshake AI's BVB meta-eval (3,204 practicing-banker pass/fail labels, 89.5% inter-annotator agreement on a dual-coded adjudicated subset) shows the best automated verifier they could build (Gandalf, a reactive agent-judge running in the same environment as the work) reaches only F1 0.633-0.664 against expert labels on artifact-heavy tasks, while hitting F1 0.951 on a simpler stateful-tools benchmark, and that verifier architecture (evidence view + evidence path) matters more than judge model choice.",
   "load_bearing": true,
   "evidence": "Fetched the full Handshake research post (2026-05-27). BVB is explicitly a meta-eval of verifiers, not agents; text-only 'autorubric' judges fail because evidence lives in spreadsheet formulas/tool state they cannot see; snapshot serialization loses formula references; Gandalf (open-sourced, OpenHands SDK) opens files and follows references at inference time. Cheapest Gandalf config (~$42) beat the best baseline (~$422) by ~3 F1 at one-tenth cost; within one model family the architecture gap was 9.5 F1. They prescribe measuring every verifier against adjudicated expert labels before trusting it to score work or produce rewards.",
   "implication": "Direct evidence for three taxonomy questions: (Q1) meta-evaluation must use adjudicated expert labels, and a ~89.5% human agreement floor is achievable with dual-coding + adjudication; (Q5) the judge-vs-adjudicated-human ceiling is wildly domain-dependent (F1 0.63 vs 0.95), so per-axis judge-autonomous/assisted/human-only lanes are mandatory, not optional; (Q3) evidence closure is the binding constraint - the judge must see the same evidence surface the attempter's claims are about, or it verifies against a flattened proxy.",
   "source": {
    "raw": "Your verifier is probably the bottleneck. We built one that isn't. (Gandalf the Grader / BVB) | https://joinhandshake.com/research/ai/gandalf-the-grader/ | 2026-05-27 | primary",
    "title": "Your verifier is probably the bottleneck. We built one that isn't. (Gandalf the Grader / BVB)",
    "url": "https://joinhandshake.com/research/ai/gandalf-the-grader/",
    "date": "2026-05-27",
    "type": "primary"
   },
   "venue": {
    "peer_reviewed": null,
    "marker": "primary"
   },
   "same_source_claims": [
    "F26-02"
   ],
   "in_corpus_verification": {
    "verdict": "confirmed",
    "transcript": "Fetched https://joinhandshake.com/research/ai/gandalf-the-grader/ directly. Every element of the claim checks out against the source: (1) dated May 27, 2026; (2) BVB (BankerVerifierBench) is explicitly a meta-eval of verifiers built on 21 BankerToolBench tasks with 3,204 pass/fail judgments from practicing bankers; 89.5% inter-annotator agreement on a dual-coded subset, with disagreements adjudicated before inclusion; (3) Gandalf is a reactive agent-as-judge running in the same rollout environment (shared filesystem, interpreter, MCP tools), open-sourced on the OpenHands SDK; (4) F1 0.633-0.664 on BVB (artifact-heavy) vs F1 0.951 on OpenClaw, their internal stateful-tools productivity benchmark (200-criterion sample, +7.5 over next-best); (5) architecture (evidence view: text/snapshot/live env; evidence path: fixed vs inference-time) beats model choice - within GPT-5.4 family, Gandalf/Nano beats Archipelago/GPT-5.4 by 9.5 F1; (6) cheapest Gandalf config (~$42, GPT-5.4 Nano) beats best baseline (Archipelago/Gemini 3 Pro, F1 0.604, ~$422) by ~3 F1 at ~one-tenth cost; (7) the post prescribes meta-eval of verifiers against adjudicated expert labels before trusting them for scoring/rewards. Supersession check (two searches, July 2026): no newer contradicting work found; the full Gandalf paper and BVB dataset release are still pending, so numbers could later be refined, but as of today the post is the authoritative source and the claim matches it. Minor nuance only: the 0.951 benchmark is named OpenClaw and is internal (not public), which the claim's phrasing (\"simpler stateful-tools benchmark\") accurately reflects.",
    "corrected": null
   },
   "external_verification": {
    "verdict": "confirmed",
    "checked": [
     "BVB is a meta-eval of verifiers, not an agent benchmark",
     "3,204 expert-graded (practicing-banker) pass/fail judgments",
     "inter-annotator agreement 89.5%; disagreements adjudicated; a subset dual-coded",
     "Gandalf F1 range 0.633-0.664 on artifact-heavy banking tasks",
     "Gandalf F1 0.951 on OpenClaw (stateful tools), 7.5 points ahead of next-best",
     "cheapest Gandalf (GPT-5.4 Nano ~$42) beats best baseline (Archipelago/Gemini 3 Pro, F1 0.604, ~$422) by ~3 F1 at ~one-tenth cost",
     "within-family architecture gap 9.5 F1 (Gandalf/GPT-5.4 Nano vs Archipelago/GPT-5.4)",
     "Gandalf is a reactive agent-as-judge running inside the rollout environment; open-sourced; OpenHands SDK harness",
     "verifier architecture (view + path) matters more than backing model"
    ],
    "not_visible": [],
    "quote": "F1 range, from 0.633 to 0.664, sits above the highest-F1 non-Gandalf run",
    "notes": "Every cited figure is present in the post. Minor framing: 89.5% is the overall inter-annotator agreement (with a separate dual-coded subset for rubric-application consistency), which is consistent with the claim's wording.",
    "checked_at": "2026-07-15",
    "retrieval": "landing"
   },
   "source_retrieval_meta": {
    "retrieval": "landing",
    "fetched": [
     "https://joinhandshake.com/research/ai/gandalf-the-grader/"
    ],
    "resolved_title": "Your verifier is probably the bottleneck. We built one that isn't. (Gandalf the Grader / BankerVerifierBench)",
    "resolved_date": "not surfaced verbatim in fetch (claimed 2026-05-27); content references 2026-era models (GPT-5.4, Gemini 3 Pro)",
    "title_match": true
   },
   "used_on": [
    {
     "page": "research.html",
     "anchor": "research-cards",
     "label": "Research - findings by decision weight"
    }
   ],
   "decision_use": {
    "grade": "B",
    "basis": "current-gen vendor eval",
    "measured_on": "human-work",
    "note": "BVB (May 2026): 3,204 practicing-banker labels - a rare human-work meta-eval. Vendor research; benchmark release pending per corpus.",
    "models_measured": []
   }
  },
  {
   "id": "IP-03",
   "domain": "industry-practice",
   "area": null,
   "claim": "The 2026 eval-tooling vendors (LangSmith, Braintrust, Galileo) have converged on one judge-calibration pattern - small human-labeled reference sets (LangSmith: start at ~20 balanced examples; Braintrust: 50-100 per scoring dimension) with an 'alignment score' defined as RAW PERCENT AGREEMENT with the human expert, iterate the judge prompt against it, and route low-confidence or judge-disagreement items to humans - and none of them ships chance-corrected statistics.",
   "load_bearing": true,
   "evidence": "Fetched LangSmith's official 'Improve LLM-as-judge evaluators using human feedback' doc: annotation queue -> reference dataset (>=20 examples, balanced 0/1 labels) -> Evaluator Playground -> alignment score = 'percentage of examples where the evaluator's judgment matches that of the human expert' -> cluster misalignments into failure modes and encode as prompt instructions. Fetched Braintrust's 2026-04-03 article: 3-tier stack (deterministic checks -> LLM judge -> human), 50-100 labels per dimension before trusting a judge, multi-judge disagreement routed to humans, worst failure mode named as 'silent overconfidence' (judge always returns a score even without adequate context), and explicit warning that human review is not automatically gold standard (fatigue, vague rubrics can be worse than a decent judge). Galileo's CLHF (2025-02-10 blog) auto-adapts metrics from expert corrections.",
   "implication": "Our foundation can adopt the proven mechanics (annotation-queue-to-reference-set loop, failure-mode clustering into rubric amendments, disagreement routing) but must NOT copy the metric: the entire tooling industry standardized on raw percent agreement over tiny unbalanced sets - exactly what taxonomy Q1 proposes banning. Differentiation and validity headroom lie in chance-corrected stats, per-class recall at fixed prevalence, and minimum reference-set designs; also budget expectations: tens of labels per axis is the industry's plateau assumption, not hundreds.",
   "source": {
    "raw": "LangSmith: Improve LLM-as-judge evaluators using human feedback (+ Braintrust: LLM-as-a-judge vs human-in-the-loop, 2026-04-03) | https://docs.langchain.com/langsmith/improve-judge-evaluator-feedback | retrieved 2026-07-14 (Braintrust article 2026-04-03) | primary",
    "title": "LangSmith: Improve LLM-as-judge evaluators using human feedback (+ Braintrust: LLM-as-a-judge vs human-in-the-loop, 2026-04-03)",
    "url": "https://docs.langchain.com/langsmith/improve-judge-evaluator-feedback",
    "date": "retrieved 2026-07-14 (Braintrust article 2026-04-03)",
    "type": "primary"
   },
   "venue": {
    "peer_reviewed": null,
    "marker": "primary"
   },
   "same_source_claims": [],
   "in_corpus_verification": {
    "verdict": "partially_confirmed",
    "transcript": "VERIFIED: (1) LangSmith doc (https://docs.langchain.com/langsmith/improve-judge-evaluator-feedback) says exactly what was reported - annotation queue -> reference dataset (~20 examples minimum, balanced 0/1) -> Evaluator Playground -> alignment score quoted verbatim as \"the percentage of examples where the evaluator's judgment matches that of the human expert\" -> cluster misalignments into failure modes as prompt instructions. No kappa/alpha anywhere; doc is undated. (2) Braintrust article (https://www.braintrust.dev/articles/llm-as-a-judge-vs-human-in-the-loop-evals, confirmed 2026-04-03) matches on all six checked points: 3-tier stack, 50-100 labels per scoring dimension, multi-judge disagreement routed to humans, \"silent overconfidence\" named as the most dangerous failure mode, human-review-not-gold-standard warning (reviewer fatigue, vague rubrics), and agreement tracked as raw agreement rate with no chance-corrected statistic. (3) Galileo CLHF exists and auto-adapts metrics from expert corrections (feedback -> LLM-generated few-shot examples appended to the metric prompt, capped at most recent 15, claimed 20-30% accuracy gain). REFUTED/OVERSTATED PARTS: (a) The \"convergence on raw percent agreement\" framing fails for Galileo. Galileo's own guidance - \"How to Calibrate Your LLM Judge With Human Annotations\" (https://galileo.ai/blog/calibrate-llm-judge-human-annotations, May 15/Jul 6 2026, i.e. NEWER than the cited 2025-02-10 CLHF post) - explicitly recommends AGAINST raw percent agreement (imbalanced-label example: 90% raw agreement, kappa ~ -0.05) and prescribes Cohen's kappa, Fleiss' kappa, and Krippendorff's alpha with a recalibrate-when-kappa<0.60 trigger. Galileo also has a dedicated Cohen's Kappa explainer (https://galileo.ai/blog/cohens-kappa-metric, 2025-03-12). (b) Galileo's CLHF is also not the same pattern as LangSmith/Braintrust - it has no reference set + alignment score loop; it is few-shot auto-improvement, so the three-vendor \"one pattern\" claim overreaches. (c) The narrower claim \"none of them SHIPS chance-corrected statistics\" as an in-product feature survives verification: Galileo's kappa content points users to scikit-learn/statsmodels/R for computation, and its product features (Autotune, Annotations, Signals, Experiments, CLHF) are not described as computing kappa; LangSmith and Braintrust expose only percent-agreement metrics. But the claim as worded implies the vendors don't even acknowledge chance correction, which is false for Galileo. SUPERSESSION CHECK: newer relevant work exists - arXiv 2606.00093 (Rao & Callison-Burch, 2026-05-25, \"Agreement Metrics for LLM-as-Judge Evaluation: What to Report and Why\") argues for reporting kappa alongside accuracy, and LangChain published an Align Evals calibration resource (2026-03-10) still framed around human-correction agreement.",
    "corrected": "LangSmith and Braintrust have converged on a judge-calibration pattern of small human-labeled reference sets (LangSmith: ~20 balanced examples; Braintrust: 50-100 per scoring dimension) with alignment measured as raw percent agreement with human experts, prompt iteration against misalignments, and routing of low-confidence/judge-disagreement items to humans; neither ships a chance-corrected statistic. Galileo differs: its CLHF (2025) turns expert corrections into few-shot prompt examples (max 15) rather than a reference-set alignment score, and its newer 2026 calibration guidance explicitly rejects raw percent agreement in favor of Cohen's kappa / Fleiss' kappa / Krippendorff's alpha (recalibrate below kappa 0.60) - though it directs users to external tools (scikit-learn, R) rather than computing these in-product. So the \"no chance-corrected stats shipped in-product\" observation holds across all three, but there is no three-vendor convergence on raw percent agreement as the alignment definition."
   },
   "external_verification": {
    "verdict": "partially_confirmed",
    "checked": [
     "annotation queue -> reference dataset workflow confirmed",
     "'We recommend starting with at least 20 examples' and examples 'balanced in both 0 and 1 labels'",
     "alignment score = 'the percentage of examples where the evaluator's judgment matches that of the human expert' (raw percent agreement, no chance correction)",
     "Evaluator Playground confirmed ('Start Alignment')",
     "cluster misaligned examples 'into common failure modes' then encode as prompt instructions"
    ],
    "not_visible": [
     "Braintrust specifics (50-100 labels per scoring dimension, 3-tier deterministic/judge/human stack, 'silent overconfidence' failure mode, human-review-not-automatically-gold-standard, 2026-04-03 date) -- on a separate Braintrust URL not in scope for this source",
     "Galileo CLHF auto-adaptation from expert corrections (2025-02-10 blog) -- separate URL, not fetchable under the source-only constraint",
     "cross-vendor generalization 'none of them ships chance-corrected statistics' -- only LangSmith's raw-percent-agreement design is verifiable from the single given URL"
    ],
    "quote": "The alignment score is the percentage of examples where the evaluator's judgment matches that of the human expert. ... We recommend starting with at least 20 examples",
    "notes": "The LangSmith half of the claim (>=20 balanced examples, alignment score = raw percent agreement, iterate prompt, cluster failure modes, route disagreements) is fully confirmed. The Braintrust and Galileo assertions live on other pages the hard constraint forbids fetching, so the multi-vendor convergence claim is only partially verifiable.",
    "checked_at": "2026-07-15",
    "retrieval": "fulltext"
   },
   "source_retrieval_meta": {
    "retrieval": "fulltext",
    "fetched": [
     "https://docs.langchain.com/langsmith/improve-judge-evaluator-feedback"
    ],
    "resolved_title": "Improve LLM-as-judge evaluators using human feedback (LangSmith docs)",
    "resolved_date": "no publication date visible on the doc page",
    "title_match": true
   },
   "used_on": [],
   "decision_use": {
    "grade": "A",
    "basis": "market practice",
    "measured_on": "structural",
    "note": "2026 eval-tooling convergence on human-calibrated judges. Market fact.",
    "models_measured": []
   }
  },
  {
   "id": "IP-05",
   "domain": "industry-practice",
   "area": null,
   "claim": "Handshake acquired Cleanlab (announced 2026-01-28) specifically for algorithms that 'flag incorrect data without a second human reviewer,' after competing bids from other data-labeling companies - signaling that algorithmic second-pass QA replacing the human review layer is now an acquisition-grade core capability for human-data vendors.",
   "load_bearing": false,
   "evidence": "TechCrunch (2026-01-28), Cleanlab CEO letter, and Handshake's own announcement triangulate: 9 Cleanlab staff including all three MIT-PhD co-founders join Handshake research; rationale quotes name confident learning, data-centric AI, and LLM evaluation as the acquired capabilities; Handshake serves 8 top labs at ~$300M annualized run-rate (end 2025). Cleanlab's core product audited human-labeler output quality algorithmically.",
   "implication": "The market has already voted for our thesis: augment/replace the human QA layer over the mandatory human attempter layer. Confident-learning-style label-noise detection (statistical, model-based, no second human) is a complementary channel to LLM-judge QA and should be considered in the foundation's filtering stage.",
   "source": {
    "raw": "TechCrunch: AI data labeler Handshake buys Cleanlab, an acquisition target of multiple others | https://techcrunch.com/2026/01/28/ai-data-labeler-handshake-buys-cleanlab-an-acquisition-target-of-multiple-others/ | 2026-01-28 | practitioner",
    "title": "TechCrunch: AI data labeler Handshake buys Cleanlab, an acquisition target of multiple others",
    "url": "https://techcrunch.com/2026/01/28/ai-data-labeler-handshake-buys-cleanlab-an-acquisition-target-of-multiple-others/",
    "date": "2026-01-28",
    "type": "practitioner"
   },
   "venue": {
    "peer_reviewed": null,
    "marker": "practitioner"
   },
   "same_source_claims": [],
   "in_corpus_verification": null,
   "external_verification": {
    "verdict": "partially_confirmed",
    "checked": [
     "Handshake acquired Cleanlab (announced 2026-01-28)",
     "Cleanlab algorithms 'flag incorrect data without a second human reviewer'",
     "competing acquisition interest from other AI data-labeling companies",
     "nine key Cleanlab employees join Handshake research",
     "three co-founders (Northcutt, Mueller, Athalye) earned PhDs from MIT",
     "Handshake provided data for eight top AI labs, including OpenAI",
     "$300 million annualized revenue run-rate (end of 2025)"
    ],
    "not_visible": [
     "the phrases 'confident learning', 'data-centric AI', and 'LLM evaluation' (not used in the TechCrunch article; claim attributes them to Cleanlab CEO letter / Handshake announcement)"
    ],
    "quote": "\"Cleanlab's researchers are experts in developing algorithms that flag incorrect data without a second human reviewer.\" ... \"The company has provided data for eight top AI labs, including OpenAI.\"",
    "notes": "Core assertion and every headline number/fact drawn from TechCrunch confirmed verbatim, including the key 'flag incorrect data without a second human reviewer' quote, competing bids, 9 staff, 3 MIT co-founders, 8 labs, $300M ARR. Only the capability-name phrases in the claim's evidence ('confident learning', 'data-centric AI', 'LLM evaluation') are absent from the article; the claim itself attributes those to other documents.",
    "checked_at": "2026-07-15",
    "retrieval": "landing"
   },
   "source_retrieval_meta": {
    "retrieval": "landing",
    "fetched": [
     "https://techcrunch.com/2026/01/28/ai-data-labeler-handshake-buys-cleanlab-an-acquisition-target-of-multiple-others/"
    ],
    "resolved_title": "AI data labeler Handshake buys Cleanlab, an acquisition target of multiple others",
    "resolved_date": "2026-01-28 (11:00 AM PST), by Marina Temkin",
    "title_match": true
   },
   "used_on": [],
   "decision_use": {
    "grade": "A",
    "basis": "market fact",
    "measured_on": "structural",
    "note": "Handshake acquired Cleanlab for no-second-human flagging (Jan 2026). Market signal, current.",
    "models_measured": []
   }
  },
  {
   "id": "IP-06",
   "domain": "industry-practice",
   "area": null,
   "claim": "Surge AI (~$1.2B revenue 2024, ~50k contractors, bootstrapped) runs QA as real-time dashboards over gold-standard accuracy, inter-annotator agreement, and per-worker trust ratings, with flagged labels automatically reassigned to other annotators; its published AdvancedIF verifier work with Meta reports F1 0.728 versus human judgments and +13% RL gains from rubric-based rewards.",
   "load_bearing": false,
   "evidence": "Fetched Sacra's Surge AI company report (updated 2026-04-21; figures are Sacra estimates). QA description: gold-standard accuracy + IAA + per-worker trust ratings on live dashboards; automated reassignment of low-quality labels; premium per-minute pay (~$0.30-0.40/min) as a quality incentive; vetting = domain tests + background checks + ongoing performance evaluation. Hemingway-bench used 5,000+ blind pairwise expert comparisons. Risks noted: ~12-customer concentration, CA misclassification suit alleging unpaid training time and tight task timers.",
   "implication": "The premium-quality market leader's QA is continuous gold-task injection + IAA + per-worker trust telemetry with automatic reassignment - i.e., population-level statistical QC, not per-item semantic review. Their own verifier tops out at F1 0.728 vs humans on instruction-following, a second independent datapoint (with Handshake's 0.63-0.66) that ~0.65-0.75 F1 is the current judge ceiling on expert-graded open-ended work (taxonomy Q5).",
   "source": {
    "raw": "Sacra: Surge AI revenue, funding & news | https://sacra.com/c/surge-ai | 2026-04-21 | practitioner",
    "title": "Sacra: Surge AI revenue, funding & news",
    "url": "https://sacra.com/c/surge-ai",
    "date": "2026-04-21",
    "type": "practitioner"
   },
   "venue": {
    "peer_reviewed": null,
    "marker": "practitioner"
   },
   "same_source_claims": [],
   "in_corpus_verification": null,
   "external_verification": {
    "verdict": "confirmed",
    "checked": [
     "Sacra estimates that Surge AI hit $1.2B in annualized revenue in 2024",
     "approximately 50,000 expert contractors and 130 full-time employees",
     "bootstrapped company",
     "Real-time dashboards monitor annotation quality using ... gold-standard accuracy, inter-annotator agreement ... and per-worker trust ratings",
     "Labels identified as low quality are automatically reassigned to other annotators",
     "AdvancedIF ... in partnership with Meta Superintelligence Labs",
     "its verifier achieved 0.728 F1 versus human judgments",
     "using human-written rubrics as RL reward signals yields 13% performance gains",
     "premium rates of 30-40 cents per working minute",
     "rigorous vetting, including domain-specific tests, background checks, and ongoing performance evaluations",
     "Hemingway-bench ... built from 5,000+ blind pairwise comparisons by expert human judges",
     "reliance on 12 customers for over $1 billion in revenue",
     "class action lawsuit filed in California ... alleges Surge misclassified data annotators"
    ],
    "not_visible": [
     "source_date 2026-04-21 -- no last-updated date is shown on the page",
     "the lawsuit's specific allegations of 'unpaid training time and tight task timers' (not surfaced in retrieved text)"
    ],
    "quote": "Real-time dashboards monitor annotation quality using metrics such as gold-standard accuracy, inter-annotator agreement ... per-worker trust ratings ... Labels identified as low quality are automatically reassigned to other annotators",
    "notes": "All revenue/contractor/QA/AdvancedIF figures confirmed. Minor: page states '13% performance gains' (no literal '+' sign). Figures are explicitly Sacra estimates. Page shows no date, so the 2026-04-21 source_date could not be corroborated.",
    "checked_at": "2026-07-15",
    "retrieval": "landing"
   },
   "source_retrieval_meta": {
    "retrieval": "landing",
    "fetched": [
     "https://sacra.com/c/surge-ai"
    ],
    "resolved_title": "Surge AI revenue, funding & news | Sacra",
    "resolved_date": "not shown on page",
    "title_match": true
   },
   "used_on": [],
   "decision_use": {
    "grade": "A",
    "basis": "market practice",
    "measured_on": "structural",
    "note": "Surge QA-as-dashboards practice. Vendor-reported.",
    "models_measured": []
   }
  },
  {
   "id": "IP-08",
   "domain": "industry-practice",
   "area": null,
   "claim": "Braintrust's published cost arithmetic makes the human-budget fork concrete: an LLM judge over 10,000 outputs costs ~$5-15 while a domain expert reviewing 500 of them costs ~$800-1,800 (~100x per output), which is why its prescribed architecture is deterministic checks on everything, LLM judges on everything, and humans only on flagged/low-confidence/disagreement slices plus random spot-checks of confident passes.",
   "load_bearing": false,
   "evidence": "Fetched Braintrust article (2026-04-03) and human-review docs. The hybrid loop is explicit: trace production -> continuous deterministic + judge scoring -> flag low-confidence/scorer-disagreement -> human review -> reviewed findings become new scorers/rubric refinements -> CI regression gates. They name the biggest practitioner mistake as treating judge scores as ground truth without human validation, and note untrained reviewers with vague rubrics can produce labels worse than a decent judge.",
   "implication": "Supports taxonomy Q6's routing answer as consensus practice: spend the single human touch on judge-flagged and judge-disagreement items, keep a randomized audit slice of confident passes, and convert every human review into a compiled-config improvement (new scorer/exemplar) rather than a one-off verdict.",
   "source": {
    "raw": "Braintrust: LLM-as-a-judge vs human-in-the-loop evals: When to use each | https://www.braintrust.dev/articles/llm-as-a-judge-vs-human-in-the-loop-evals | 2026-04-03 | practitioner",
    "title": "Braintrust: LLM-as-a-judge vs human-in-the-loop evals: When to use each",
    "url": "https://www.braintrust.dev/articles/llm-as-a-judge-vs-human-in-the-loop-evals",
    "date": "2026-04-03",
    "type": "practitioner"
   },
   "venue": {
    "peer_reviewed": true,
    "marker": "Findings"
   },
   "same_source_claims": [],
   "in_corpus_verification": null,
   "external_verification": {
    "verdict": "confirmed",
    "checked": [
     "LLM judge on 10,000 outputs (Claude Sonnet) might cost $5-15",
     "domain expert reviewing 500 outputs at $50-75/hr, 2-3 min each, runs $800-1,800",
     "roughly 100x more expensive per output for human review",
     "three-tier architecture: deterministic checks on every output, LLM judges for continuous coverage, humans only on flagged/low-confidence/disagreement cases + random spot-checks",
     "biggest mistake: 'Treating judge scores as ground truth without validating them against human labels'",
     "untrained reviewers / vague rubrics can yield labels 'worse than a decent LLM judge'"
    ],
    "not_visible": [],
    "quote": "Running an LLM judge on 10,000 outputs using a model like Claude Sonnet might cost $5-15 ... review 500 of those same outputs at $50-75/hour, spending 2-3 minutes per output, runs $800-1,800. That is roughly 100x more expensive per output",
    "notes": "Cost arithmetic, tiered architecture (deterministic + judge on everything, humans on flagged/disagreement slices plus random spot-checks), the named biggest-mistake, and the weak-human-review caveat all confirmed. Published date 3 April 2026 matches.",
    "checked_at": "2026-07-15",
    "retrieval": "landing"
   },
   "source_retrieval_meta": {
    "retrieval": "landing",
    "fetched": [
     "https://www.braintrust.dev/articles/llm-as-a-judge-vs-human-in-the-loop-evals"
    ],
    "resolved_title": "LLM-as-a-judge vs human-in-the-loop evals: When to use each",
    "resolved_date": "2026-04-03",
    "title_match": true
   },
   "used_on": [
    {
     "page": "decisions.html",
     "anchor": "d4",
     "label": "Decisions - judge sourcing"
    }
   ],
   "decision_use": {
    "grade": "B",
    "basis": "market pricing",
    "measured_on": "structural",
    "note": "Braintrust cost arithmetic (judge $5-15/10k vs expert $800-1,800/500). Pricing is current but moves with model economics - reprice at bakeoff; the routing logic it supports is durable.",
    "models_measured": []
   }
  },
  {
   "id": "IP-09",
   "domain": "industry-practice",
   "area": null,
   "claim": "Fine-tuned specialist judges have commoditized binary grounding checks: Patronus Lynx (fine-tuned Llama-3-70B) beats GPT-4o on HaluBench faithfulness detection and lifted one customer's hallucination detection from 0.375 to 0.69, while Galileo's Luna-2 (arXiv 2026-02-20) runs hundreds of per-metric LoRA adapters on one small-model backbone at >80x lower cost and >20x lower latency than LLM-as-judge at claimed parity accuracy, in production on 100M+ sessions.",
   "load_bearing": false,
   "evidence": "Patronus launch post + SiliconANGLE (2024-07-11) give HaluBench numbers (Lynx 70B +8.3% vs GPT-4o on PubMedQA inaccuracies; 8B +24.5% vs GPT-3.5); Patronus evaluators page gives the Algomo 0.375->0.69 customer number and Gamma's 1,000+ manual eval hours saved; Percival extends this to tracing where in a multi-step agent run the failure occurred. Fetched Luna-2 arXiv abstract (2602.18583): single-token SLM evaluation, >80x cost / >20x latency reduction, 100M+ sessions and 100B tokens/month in production, $30M+ claimed annual eval savings. Vendor-reported numbers; treat magnitudes, not decimals.",
   "implication": "For taxonomy Q2/Q3: the 'compile to binary checks' bucket (groundedness, faithfulness, toxicity, tool-selection) can run on cheap specialized models per axis at near-zero marginal cost, reserving frontier-model judges for entailment/calibration/judgment-laden axes - a two-tier judge economy is standard 2026 practice, not an exotic design.",
   "source": {
    "raw": "Luna-2: Scalable Single-Token Evaluation with Small Language Models (arXiv 2602.18583) + Patronus Lynx launch materials | https://arxiv.org/abs/2602.18583 | 2026-02-20 (Lynx 2024-07-11) | primary",
    "title": "Luna-2: Scalable Single-Token Evaluation with Small Language Models (arXiv 2602.18583) + Patronus Lynx launch materials",
    "url": "https://arxiv.org/abs/2602.18583",
    "date": "2026-02-20 (Lynx 2024-07-11)",
    "type": "primary"
   },
   "venue": {
    "peer_reviewed": false,
    "marker": "preprint"
   },
   "same_source_claims": [],
   "in_corpus_verification": null,
   "external_verification": {
    "verdict": "partially_confirmed",
    "checked": [
     "Luna-2: single-token SLM evaluation vs multi-token LLM-as-judge",
     "each metric = a lightweight LoRA/PEFT head on a shared SLM backbone (hundreds concurrently)",
     "cost reduced by over 80x, latency by over 20x",
     "matches accuracy of state-of-the-art LLM-based evaluators (parity)",
     "production scale: 100M+ AI sessions, 100B+ tokens/month",
     "eval cost savings of over $30M annually"
    ],
    "not_visible": [
     "all Patronus Lynx claims: Lynx-70B beats GPT-4o on HaluBench (+8.3% PubMedQA), Lynx-8B +24.5% vs GPT-3.5",
     "Algomo customer 0.375 -> 0.69 hallucination detection",
     "Gamma 1,000+ manual eval hours saved; Percival tracing",
     "SiliconANGLE / Patronus launch-post sourcing (separate URLs, not fetched)"
    ],
    "quote": "matches the accuracy of state-of-the-art LLM-based evaluators ... by over 80x ... by over 20x ... protecting 100M+ AI sessions ... eval cost savings of over $30M annually",
    "notes": "The Luna-2 half of the claim is fully confirmed from the given arXiv abstract. The Patronus Lynx half is sourced to Patronus/SiliconANGLE pages that are not the given URL and thus not verifiable under the fetch constraint; those specifics are marked not_visible, not checked.",
    "checked_at": "2026-07-15",
    "retrieval": "abstract"
   },
   "source_retrieval_meta": {
    "retrieval": "abstract",
    "fetched": [
     "https://arxiv.org/abs/2602.18583"
    ],
    "resolved_title": "Luna-2: Scalable Single-Token Evaluation with Small Language Models",
    "resolved_date": "2026-02-20",
    "title_match": true
   },
   "used_on": [],
   "decision_use": {
    "grade": "C",
    "basis": "superseded-gen capability",
    "measured_on": "model-outputs",
    "note": "Lynx vs GPT-4o-era comparisons. Specialist-vs-frontier margins need current repricing.",
    "models_measured": [
     "Lynx",
     "Llama-3-70B",
     "GPT-4o",
     "GPT-3.5"
    ]
   }
  },
  {
   "id": "IP-10",
   "domain": "industry-practice",
   "area": null,
   "claim": "DeepEval's recommended progression - start with holistic G-Eval, then move to DAG (a deterministic decision tree whose nodes are narrow LLM-judge calls) 'for more control' - is the open-source codification of decomposed judging: verdicts assembled from atomic, individually-checkable steps rather than one holistic score.",
   "load_bearing": false,
   "evidence": "DeepEval docs (metrics-introduction, metrics-dag; DAG introduced 2025-02, docs current 2026-07) describe DAGMetric as 'deterministic decision trees' for objective/mixed criteria producing deterministic scores, explicitly positioned above G-Eval when control matters, including a conversational variant for multi-turn work. Confident-AI's 2025-02-09 engineering post explains the motivation: holistic LLM scores are too noisy for objective criteria.",
   "implication": "Taxonomy Q3's decomposition-granularity question has a live open-source reference implementation: per-criterion decision trees with narrow judge nodes and deterministic aggregation. Adopt the pattern (and its admission that decomposition is for objective/mixed criteria - judgment-laden axes stay holistic/G-Eval-style) rather than inventing the machinery.",
   "source": {
    "raw": "DeepEval docs: DAG (Deep Acyclic Graph) metric + Confident AI engineering post | https://deepeval.com/docs/metrics-dag | 2025-02-09, docs verified 2026-07-10 | practitioner",
    "title": "DeepEval docs: DAG (Deep Acyclic Graph) metric + Confident AI engineering post",
    "url": "https://deepeval.com/docs/metrics-dag",
    "date": "2025-02-09, docs verified 2026-07-10",
    "type": "practitioner"
   },
   "venue": {
    "peer_reviewed": null,
    "marker": "practitioner"
   },
   "same_source_claims": [],
   "in_corpus_verification": null,
   "external_verification": {
    "verdict": "partially_confirmed",
    "checked": [
     "DAG lets you 'easily build deterministic decision trees for evaluation'",
     "positioned above G-Eval: 'the DAGMetric will give you much greater control'",
     "'more deterministic control' over GEval; produces deterministic scores",
     "nodes are LLM-judge calls (BinaryJudgementNode, NonBinaryJudgementNode; 'reduce LLM-judge variance')",
     "breakdown into more atomic units for evaluation"
    ],
    "not_visible": [
     "'conversational' variant for multi-turn work (word not found on page)",
     "Confident-AI 2025-02-09 engineering post ('holistic LLM scores too noisy') - separate URL, not fetched",
     "exact phrase 'for more control' (page says 'much greater control' / 'more deterministic control')",
     "docs page publication/last-updated date"
    ],
    "quote": "You can still use GEval in the DAGMetric, but the DAGMetric will give you much greater control ... easily build deterministic decision trees for evaluation",
    "notes": "Core G-Eval -> DAG progression and decomposed/deterministic-node judging confirmed on the docs page. The asserted 'conversational variant' and the Confident-AI blog motivation (source_date 2025-02-09) are not verifiable from this docs URL.",
    "checked_at": "2026-07-15",
    "retrieval": "fulltext"
   },
   "source_retrieval_meta": {
    "retrieval": "fulltext",
    "fetched": [
     "https://deepeval.com/docs/metrics-dag"
    ],
    "resolved_title": "DAG (Deep Acyclic Graph)",
    "resolved_date": "not shown on docs page",
    "title_match": true
   },
   "used_on": [],
   "decision_use": {
    "grade": "A",
    "basis": "tooling practice",
    "measured_on": "structural",
    "note": "DeepEval G-Eval->DAG progression. Tooling doctrine.",
    "models_measured": []
   }
  },
  {
   "id": "IP-11",
   "domain": "industry-practice",
   "area": null,
   "claim": "The human reviewer layer at scale vendors is itself unreliable for structural/incentive reasons the AutoQA must not inherit: Outlier's attempter->reviewer->senior-reviewer pyramid runs on timed, capped per-task pay, opaque promotion/removal, uneven task allocation, and internal competition, per 2025 contributor reports across Indeed/Glassdoor/Trustpilot.",
   "load_bearing": false,
   "evidence": "Aggregated 2025 worker reports: pay only within a capped time window (unpaid overage), removal from projects without stated reason, alleged 300-reviewer queues where ~10 got unlimited tasks, queue managers running webinars/'war rooms' as the calibration mechanism, quality undermined by contributors and reviewers 'pitted against each other'. Community-sourced and unverifiable individually, but consistent across platforms and consistent with the Inc. documents' picture of overwhelmed review capacity.",
   "implication": "Reviewer variance (our root causes 1-3) is partly manufactured by incentive design: time-capped pay punishes careful review, and opaque consequences destroy calibration motivation. The AutoQA foundation should specify the human-touch interaction's incentive contract (untimed or accuracy-paid adjudication, visible appeal outcomes) as part of the design, not just the routing logic.",
   "source": {
    "raw": "Indeed/Glassdoor/Trustpilot Outlier reviewer reports (aggregated) | https://www.indeed.com/cmp/Outlier-Ai/reviews?fjobtitle=Reviewer&fcountry=US&ftext=outlier | 2025 (various) | community",
    "title": "Indeed/Glassdoor/Trustpilot Outlier reviewer reports (aggregated)",
    "url": "https://www.indeed.com/cmp/Outlier-Ai/reviews?fjobtitle=Reviewer&fcountry=US&ftext=outlier",
    "date": "2025 (various)",
    "type": "community"
   },
   "venue": {
    "peer_reviewed": null,
    "marker": "community"
   },
   "same_source_claims": [],
   "in_corpus_verification": null,
   "external_verification": {
    "verdict": "partially_confirmed",
    "checked": [
     "removal from projects without stated reason (multiple reviews 2024-2026)",
     "uneven / unpredictable task allocation (feast-or-famine workload)",
     "opaque eligibility/onboarding decisions felt arbitrary with no transparency",
     "no communication from higher-ups; contact only via Slack/Discord",
     "pay deterioration after Outlier merger; unpaid training",
     "overall 2.4/5 from 774 reviews"
    ],
    "not_visible": [
     "explicit attempter -> reviewer -> senior-reviewer pyramid",
     "timed / capped per-task pay with unpaid overage window",
     "specific '300-reviewer queues where ~10 got unlimited tasks'",
     "queue managers running webinars / 'war rooms' as calibration",
     "contributors and reviewers 'pitted against each other' / internal competition"
    ],
    "quote": "you can also be cut from a project with no warning or reason, placed on another completely different project",
    "notes": "Only the Indeed page could be fetched; source aggregates Indeed/Glassdoor/Trustpilot but constraints forbid the other two sites. The general thrust (opaque removal, unstable allocation, opacity, pay complaints) is supported, but the specific structural mechanisms in the claim (pyramid, capped timed pay, 300/10 queues, war rooms, pitted-against-each-other) are not present on this page. title_match false: source_title is a synthetic three-platform aggregate label, not this single Indeed page's title; reviews span 2024-2026, not solely 2025.",
    "checked_at": "2026-07-15",
    "retrieval": "fulltext"
   },
   "source_retrieval_meta": {
    "retrieval": "fulltext",
    "fetched": [
     "https://www.indeed.com/cmp/Outlier-Ai/reviews?fjobtitle=Reviewer&fcountry=US&ftext=outlier"
    ],
    "resolved_title": "Outlier AI Employee Reviews (Indeed, Reviewer filter) - 2.4/5 from 774 reviews",
    "resolved_date": "reviews span 2024-2026",
    "title_match": false
   },
   "used_on": [],
   "decision_use": {
    "grade": "A",
    "basis": "structural/incentive analysis",
    "measured_on": "human-work",
    "note": "Scale-vendor reviewer layer unreliability for incentive reasons. Structural; the AutoQA must not inherit it.",
    "models_measured": []
   }
  },
  {
   "id": "CT-01",
   "domain": "contrarian",
   "area": null,
   "claim": "Raw percent-agreement systematically overstates LLM-judge ability: in the largest judge meta-evaluation to date (21 judges, 9 providers, ~541,000 judgments, including April-2026 frontier models), Cohen's kappa runs 33-41 percentage points below exact-match agreement on MT-Bench, judge rankings shift by up to 14 positions across benchmarks, and two production-deployed judges combine test-retest reliability >0.95 with severe position bias >0.10 (a 'consistency-bias paradox').",
   "load_bearing": true,
   "evidence": "Norman, Rivera & Hughes ran 118 runs across MT-Bench, JudgeBench, RewardBench under three protocols (agreement, consistency, bias audit); kappa deflation was universal across the cohort; they distill a 'Minimum Viable Validation Protocol'. Verified from the arXiv abstract page. Corroborated by the NeurIPS 2025 position paper 'Neither Valid nor Reliable? Investigating the Use of LLMs as Judges' (arXiv 2508.18076, Chehbouni et al.), which argues via social-science measurement theory that LLJ adoption has outpaced scrutiny of validity/reliability and that human-agreement proxying is an unvalidated assumption.",
   "implication": "Directly answers Q1's metric-standardization question: ban raw percent-agreement and single 'accuracy vs humans' numbers from all AutoQA meta-evaluation; require chance-corrected statistics per benchmark. And for Q5: run-to-run consistency must never be reported as evidence of validity - a judge can be perfectly repeatable and severely biased simultaneously, so a separate perturbation/bias audit is a mandatory ship gate independent of the consistency gate.",
   "source": {
    "raw": "Reliability without Validity: A Systematic, Large-Scale Evaluation of LLM-as-a-Judge Models Across Agreement, Consistency, and Bias | https://arxiv.org/abs/2606.19544 | 2026-06-17 | academic",
    "title": "Reliability without Validity: A Systematic, Large-Scale Evaluation of LLM-as-a-Judge Models Across Agreement, Consistency, and Bias",
    "url": "https://arxiv.org/abs/2606.19544",
    "date": "2026-06-17",
    "type": "academic"
   },
   "venue": {
    "peer_reviewed": true,
    "marker": "NeurIPS"
   },
   "same_source_claims": [
    "AJ-03",
    "CM-09"
   ],
   "in_corpus_verification": {
    "verdict": "confirmed",
    "transcript": "Verified directly against the arXiv abstract page (https://arxiv.org/abs/2606.19544, v1 submitted 2026-06-17, only version). Every quantitative element matches: 21 judge models from 9 providers, 118 runs, ~541,000 judgments across MT-Bench/JudgeBench/RewardBench under three protocols (agreement, consistency, bias audit); exact-match vs Cohen's kappa gap of 33-41 pp on MT-Bench; judge rankings shifting up to 14 positions across benchmarks; two production-deployed judges with test-retest reliability >0.95 and position bias >0.10, framed as a 'consistency-bias paradox'; findings 'consistent across the full cohort, including the April 2026 frontier'; a 'Minimum Viable Validation Protocol' is distilled. Authors are Justin D. Norman, Michael U. Rivera, D. Alex Hughes as claimed. The corroborating paper (arXiv 2508.18076, Chehbouni et al.) is confirmed as a NeurIPS 2025 poster (https://neurips.cc/virtual/2025/poster/121914) making exactly the measurement-theory validity/reliability argument described. Supersession check (Exa, freshness=month): no later paper superseding or contradicting it found as of 2026-07-14; only secondary coverage (e.g., DEV Community 2026-07-01) restating its findings, and one adjacent study (arXiv 2606.13685) on run-to-run instability that complements rather than contradicts. Two caveats, neither rising to a correction: (a) 'largest judge meta-evaluation to date' is the authors' own self-characterization ('largest systematic evaluation of the paradigm so far'), not independently established; (b) verification is abstract-level - I did not audit the paper body. Note: two WebSearch calls were blocked by stochastic model safeguards; Exa search substituted successfully.",
    "corrected": null
   },
   "external_verification": {
    "verdict": "confirmed",
    "checked": [
     "21 judges from nine providers",
     "approximately 541,000 individual judgments",
     "kappa deflation ... 33--41 pp on MT-Bench",
     "judge rankings shift by up to 14 positions across benchmarks",
     "test--retest reliability (>0.95) with severe position bias (>0.10) in two production-deployed judges",
     "consistency--bias paradox",
     "118 runs across MT-Bench, JudgeBench, RewardBench",
     "three protocols (agreement, consistency, bias)",
     "Minimum Viable Validation Protocol",
     "April 2026 frontier"
    ],
    "not_visible": [
     "corroborating paper arXiv 2508.18076 (Chehbouni et al., 'Neither Valid nor Reliable?') -- a separate source outside the fetched URL; not verifiable under fetch constraints"
    ],
    "quote": "we evaluate 21 judges from nine providers across MT-Bench, JudgeBench, and RewardBench ... 118 runs and approximately 541,000 individual judgments",
    "notes": "Core assertion and all headline numbers confirmed. The claim's secondary corroborating citation (2508.18076) is a different paper and was not fetched per the single-URL constraint.",
    "checked_at": "2026-07-15",
    "retrieval": "fulltext"
   },
   "source_retrieval_meta": {
    "retrieval": "fulltext",
    "fetched": [
     "https://arxiv.org/abs/2606.19544",
     "https://arxiv.org/html/2606.19544v1"
    ],
    "resolved_title": "Reliability without Validity: A Systematic, Large-Scale Evaluation of LLM-as-a-Judge Models Across Agreement, Consistency, and Bias",
    "resolved_date": "2026-06-17",
    "title_match": true
   },
   "used_on": [],
   "decision_use": {
    "grade": "B",
    "basis": "current-gen capability",
    "measured_on": "model-outputs",
    "note": "Contrarian read of the June-2026 meta-evaluation (same source as AJ-03; not independent).",
    "models_measured": []
   }
  },
  {
   "id": "CT-02",
   "domain": "contrarian",
   "area": null,
   "claim": "Across 106 experiments and 370 effect sizes, human-AI combinations performed significantly worse than the best of human or AI alone (Hedges g = -0.23, 95% CI -0.39 to -0.07), with losses concentrated in decision-making tasks and specifically when the AI outperforms the human alone - and a 2025 AIES study found human reviewers followed severely race-biased AI hiring recommendations ~90% of the time.",
   "load_bearing": true,
   "evidence": "Vaccaro, Almaatouq & Malone (Nature Human Behaviour, preregistered meta-analysis; direction, decision-task losses, and AI-better-than-human losses verified from the arXiv abstract; g value corroborated across MIT Sloan, Nature's own summary, and ResearchGate). UW's 'No Thoughts Just AI' (AIES 2025, DOI 10.1609/aies.v8i3.36749, 528 participants, 16 job types) found participants matched biased AI picks even at moderate bias and ~90% at severe bias, vs equal selection rates with no/neutral AI. The 2026 practitioner literature (e.g., tianpan.co HITL rubber-stamp essay, 2026-04-15) reports the same mechanism in production review pipelines.",
   "implication": "Q6's fork is real and the evidence leans against the default: 'one shallow human touch verifying the AI verdict' is precisely the decision-task, AI-better-than-human configuration where the meta-analysis finds negative synergy and rubber-stamping. The foundation should treat zero-per-item-touch plus randomized BLIND deep audits (human judges never see the AI verdict before committing their own) as the null design to beat, and any verify-the-AI-verdict interaction must be measured against its own overturn rate to prove the human is adding information.",
   "source": {
    "raw": "When combinations of humans and AI are useful: A systematic review and meta-analysis | https://www.nature.com/articles/s41562-024-02024-1 | 2024-10-28 | academic",
    "title": "When combinations of humans and AI are useful: A systematic review and meta-analysis",
    "url": "https://www.nature.com/articles/s41562-024-02024-1",
    "date": "2024-10-28",
    "type": "academic"
   },
   "venue": {
    "peer_reviewed": true,
    "marker": "Nature"
   },
   "same_source_claims": [],
   "in_corpus_verification": {
    "verdict": "confirmed",
    "transcript": "All load-bearing facts verified against primary sources. (1) Meta-analysis: full text of Vaccaro, Almaatouq & Malone (Nature Human Behaviour, published 2024-10-28; verified via arXiv:2405.06087v2 accepted version) states verbatim \"370 unique effect sizes from 106 different experiments\" and \"Hedges' g = -0.23, 95% confidence interval -0.39 to -0.07\" for human-AI combos vs best of human or AI alone; decision-task losses and losses-when-AI-outperforms-human both confirmed. Date correct. (2) AIES 2025: \"No Thoughts Just AI\" (Wilson, Sim, Gueorguieva & Caliskan, UW), AIES Proceedings 8(3):2692-2704, article 36749 (DOI 10.1609/aies.v8i3.36749 matches), N=528, 16 occupations; equal selection with no/neutral AI, and participants favored AI-preferred candidates \"up to 90% of the time\" with severely biased AI - claim's \"~90%\" is a fair paraphrase (study says \"up to 90%\", with a simulated LLM, not a deployed system). (3) Supersession: only later work found is a 2026 Open MIND/NHB-collaboration reproduction of the meta-analysis by Brodeur et al. (osf.io/u6gea) - a reproduction effort, not a refutation; no retraction, correction, or contradicting update located. The tianpan.co practitioner-essay corroboration was not independently checked but is non-load-bearing.",
    "corrected": null
   },
   "external_verification": {
    "verdict": "partially_confirmed",
    "checked": [
     "106 experimental studies reporting 370 effect sizes",
     "human-AI combinations performed significantly worse than the best of humans or AI alone, Hedges' g = -0.23; 95% CI -0.39 to -0.07",
     "performance losses concentrated in decision-making tasks",
     "losses specifically when AI outperformed humans alone ('when AI outperformed humans alone, we found losses')"
    ],
    "not_visible": [
     "the AIES 2025 'No Thoughts Just AI' finding that human reviewers followed race-biased AI hiring recommendations ~90% of the time (a separate source, not part of this Nature article)"
    ],
    "quote": "on average, human-AI combinations performed significantly worse than the best of humans or AI alone (Hedges' g = -0.23; 95% confidence interval, -0.39 to -0.07) ... when AI outperformed humans alone, we found losses",
    "notes": "Every fact attributed to this Nature meta-analysis is confirmed verbatim from the abstract (retrieved via curl; WebFetch was blocked by a 303 login redirect). Marked partially_confirmed only because the claim bundles a second, distinct finding (the ~90% AIES 2025 hiring-bias result) that is not from this URL and cannot be verified here.",
    "checked_at": "2026-07-15",
    "retrieval": "landing"
   },
   "source_retrieval_meta": {
    "retrieval": "landing",
    "fetched": [
     "https://www.nature.com/articles/s41562-024-02024-1 (via curl; WebFetch hit a 303 auth redirect)"
    ],
    "resolved_title": "When combinations of humans and AI are useful: A systematic review and meta-analysis",
    "resolved_date": "2024-10-28",
    "title_match": true
   },
   "used_on": [
    {
     "page": "pilot.html",
     "anchor": "main",
     "label": "Pilot - gates and objectives"
    },
    {
     "page": "decisions.html",
     "anchor": "a4",
     "label": "Decisions - authority boundaries"
    }
   ],
   "decision_use": {
    "grade": "A",
    "basis": "human science (direction)",
    "measured_on": "adjacent-domain",
    "note": "g=-0.23 across 370 effect sizes, mostly older-gen AI and adjacent tasks (incl. hiring deference). The mechanism - humans defer when AI outperforms them - is human behavior and strengthens as the capability gap widens. Magnitudes are not current estimates.",
    "models_measured": []
   }
  },
  {
   "id": "CT-03",
   "domain": "contrarian",
   "area": null,
   "claim": "There is a proven theoretical ceiling on judge-based validation: when the judge is no more accurate than the model/content being evaluated, no debiasing method using gold labels can cut the required amount of ground-truth data by more than a factor of two, and empirical savings are smaller than the 2x bound.",
   "load_bearing": true,
   "evidence": "Dorner, Nastl & Hardt (ICLR 2025) prove the bound for methods that combine cheap judge scores with a small gold-label set to correct judge biases such as self-preference; verified from the arXiv abstract page including the central quote 'when the judge is no more accurate than the evaluated model, no debiasing method can decrease the required amount of ground truth' by more than half.",
   "implication": "Q1's gold-set arithmetic cannot be escaped by judge cleverness: on exactly the axes where the expertise-gap root cause bites (attempter competence near or above judge competence), certifying the AutoQA still costs at least half the gold labels a judge-free validation would - so hierarchical/pooled validation across projects and a permanent (not bootstrap) expert-audit channel are structural requirements, and any vendor claim that the judge 'validates itself' at scale should be treated as mathematically impossible in the regime that matters.",
   "source": {
    "raw": "Limits to scalable evaluation at the frontier: LLM as Judge won't beat twice the data | https://arxiv.org/abs/2410.13341 | 2024-10-17 | academic",
    "title": "Limits to scalable evaluation at the frontier: LLM as Judge won't beat twice the data",
    "url": "https://arxiv.org/abs/2410.13341",
    "date": "2024-10-17",
    "type": "academic"
   },
   "venue": {
    "peer_reviewed": true,
    "marker": "ICLR"
   },
   "same_source_claims": [],
   "in_corpus_verification": {
    "verdict": "confirmed",
    "transcript": "Source verified: Dorner, Nastl & Hardt, \"Limits to scalable evaluation at the frontier: LLM as Judge won't beat twice the data,\" arXiv:2410.13341, v1 submitted 2024-10-17 (date matches), accepted as an ICLR 2025 Oral (venue matches, slightly stronger than claimed). The abstract states verbatim that \"when the judge is no more accurate than the evaluated model, no debiasing method can decrease the required amount of ground truth labels by more than half,\" and that empirical sample-size savings are \"even more modest\" than the 2x bound - matching all three parts of the claim (theoretical 2x ceiling, the accuracy condition, and smaller empirical savings). Search for 2025-2026 follow-up work found no refutation or supersession; the authors' later papers (e.g., ROC-n-reroll, ICLR 2026) cover different topics. Minor scoping caveat only: the bound is proven within the paper's statistical framework for debiasing methods combining judge scores with gold labels, and it does not apply when the judge IS more accurate than the evaluated model - the claim already states this condition correctly.",
    "corrected": null
   },
   "external_verification": {
    "verdict": "confirmed",
    "checked": [
     "proven bound on label savings from judge-based debiasing",
     "when the judge is no more accurate than the evaluated model, no debiasing method can decrease required ground truth by more than half (factor of two)",
     "'won't beat twice the data' appears verbatim in the title",
     "self-preferencing named as a distorting bias",
     "empirical/practical savings even more modest than the theoretical limit",
     "authors Dorner, Nastl, Hardt; ICLR 2025"
    ],
    "not_visible": [],
    "quote": "no debiasing method can decrease the required amount of ground truth labels by more than half",
    "notes": "Core assertion (2x ceiling), the no-more-accurate-than-evaluated-model condition, self-preferencing, and 'empirical savings smaller than the bound' all confirmed from the abstract. Authors and ICLR 2025 venue match.",
    "checked_at": "2026-07-15",
    "retrieval": "abstract"
   },
   "source_retrieval_meta": {
    "retrieval": "abstract",
    "fetched": [
     "https://arxiv.org/abs/2410.13341"
    ],
    "resolved_title": "Limits to scalable evaluation at the frontier: LLM as Judge won't beat twice the data",
    "resolved_date": "v1 2024-10-17; v3 2026-01-06 (Comments: ICLR 2025)",
    "title_match": true
   },
   "used_on": [
    {
     "page": "index.html",
     "anchor": "decision",
     "label": "Overview - the five authorizations"
    },
    {
     "page": "research.html",
     "anchor": "research-cards",
     "label": "Research - findings by decision weight"
    },
    {
     "page": "decisions.html",
     "anchor": "a2",
     "label": "Decisions - first project and gold funding"
    },
    {
     "page": "decisions.html",
     "anchor": "a5",
     "label": "Decisions - permanent audit"
    }
   ],
   "decision_use": {
    "grade": "A",
    "basis": "theorem",
    "measured_on": "structural",
    "note": "The 2x ceiling is mathematics. Note the flip side: if deployment-tier judges materially EXCEED attempter accuracy on a lane, the constraint relaxes there - a bakeoff-measurable condition, not an assumption.",
    "models_measured": []
   }
  },
  {
   "id": "CT-04",
   "domain": "contrarian",
   "area": null,
   "claim": "Human evaluation criteria are output-dependent and unstable: even when graders define criteria before grading, the act of grading changes their criteria and they retroactively revise earlier grades ('criteria drift'), implying evaluation criteria for LLM-output quality cannot be fully determined prior to observing outputs.",
   "load_bearing": true,
   "evidence": "Shankar, Zamfirescu-Pereira, Hartmann, Parameswaran & Arawjo (UIST 2024, EvalGen system study); abstract verified via arXiv crawl: 'we identify a phenomenon we dub criteria drift: users need criteria to grade outputs, but grading outputs helps users define criteria'; the authors argue there is reason to believe criteria never fully settle because they adapt to the observed output distribution.",
   "implication": "Falsifies Q4's implicit one-shot compilation contract: 'instructions compile once into artifacts, fresh judge hits target agreement with no conversation' is unstable because the project owner's own criteria will drift once they see real attempter submissions and judge verdicts. Rubric compilation must be designed as a versioned re-compilation loop with an explicit retroactive re-scoring policy, and owner sign-off must happen on graded real items, not on abstract criteria.",
   "source": {
    "raw": "Who Validates the Validators? Aligning LLM-Assisted Evaluation of LLM Outputs with Human Preferences | https://arxiv.org/abs/2404.12272 | 2024-04-18 | academic",
    "title": "Who Validates the Validators? Aligning LLM-Assisted Evaluation of LLM Outputs with Human Preferences",
    "url": "https://arxiv.org/abs/2404.12272",
    "date": "2024-04-18",
    "type": "academic"
   },
   "venue": {
    "peer_reviewed": true,
    "marker": "UIST"
   },
   "same_source_claims": [
    "RR-05"
   ],
   "in_corpus_verification": null,
   "external_verification": {
    "verdict": "partially_confirmed",
    "checked": [
     "criteria drift: grading outputs changes graders' criteria",
     "'users need criteria to grade outputs, but grading outputs helps users define criteria'",
     "criteria appear output-dependent, not definable in advance",
     "authors Shankar, Zamfirescu-Pereira, Hartmann, Parameswaran, Arawjo"
    ],
    "not_visible": [
     "'UIST 2024' venue attribution (not on arXiv page)",
     "graders 'retroactively revise earlier grades' (specific mechanism not stated in the fetched abstract)"
    ],
    "quote": "\"we identify a phenomenon we dub criteria drift: users need criteria to grade outputs, but grading outputs helps users define criteria\"; \"some criteria appears dependent on the specific LLM outputs observed\".",
    "notes": "Core (output-dependent, unstable criteria; criteria drift) confirmed verbatim. Two specifics not visible in the retrieved abstract: the 'UIST 2024' venue and the explicit 'retroactively revise earlier grades' behavior (may appear in the paper body, not the abstract).",
    "checked_at": "2026-07-15",
    "retrieval": "abstract"
   },
   "source_retrieval_meta": {
    "retrieval": "abstract",
    "fetched": [
     "https://arxiv.org/abs/2404.12272"
    ],
    "resolved_title": "Who Validates the Validators? Aligning LLM-Assisted Evaluation of LLM Outputs with Human Preferences",
    "resolved_date": "v1 2024-04-18 (arXiv page shows no venue; source_date 2024-10-13 not confirmed on page)",
    "title_match": true
   },
   "used_on": [
    {
     "page": "pilot.html",
     "anchor": "main",
     "label": "Pilot - gates and objectives"
    },
    {
     "page": "decisions.html",
     "anchor": "d2",
     "label": "Decisions - rubric authority"
    }
   ],
   "decision_use": {
    "grade": "A",
    "basis": "human science",
    "measured_on": "human-work",
    "note": "Grader criteria instability. Human behavior.",
    "models_measured": []
   }
  },
  {
   "id": "CT-05",
   "domain": "contrarian",
   "area": null,
   "claim": "LLM judges favor models whose mistakes resemble their own (a generalization of self-preference, measured by chance-adjusted mistake-overlap CAPA), and model errors across the industry are becoming MORE correlated as capabilities improve - undermining the assumption that AI oversight of AI-assisted work catches failures.",
   "load_bearing": false,
   "evidence": "Goel et al., ICML 2025; abstract verified via arXiv crawl: 'LLM-as-a-judge scores favor models similar to the judge' and 'model mistakes are becoming more similar with increasing capabilities, pointing to risks from correlated failures'. Related: 'Preference Leakage' (arXiv 2502.01534) documents contamination when generator and judge are related.",
   "implication": "Q5: judge model family must differ from the generator being critiqued AND from whatever assistant attempters plausibly used - per-project judge routing plus similarity reporting. Q1: the false-agreement rate (judge passes an LLM-assisted attempter because both share priors) is predicted to RISE over time, which converts the independent human expert-audit channel from a bootstrap phase into a permanent structural component.",
   "source": {
    "raw": "Great Models Think Alike and this Undermines AI Oversight | https://arxiv.org/abs/2502.04313 | 2025-02-06 | academic",
    "title": "Great Models Think Alike and this Undermines AI Oversight",
    "url": "https://arxiv.org/abs/2502.04313",
    "date": "2025-02-06",
    "type": "academic"
   },
   "venue": {
    "peer_reviewed": false,
    "marker": "preprint"
   },
   "same_source_claims": [],
   "in_corpus_verification": null,
   "external_verification": {
    "verdict": "confirmed",
    "checked": [
     "judge self-preference generalization: 'LLM-as-a-judge scores favor models similar to the judge, generalizing recent self-preference results'",
     "CAPA = 'Chance Adjusted Probabilistic Agreement (CAPA): a metric for LM similarity based on overlap in model mistakes' (matches 'chance-adjusted mistake-overlap')",
     "correlated failures: 'model mistakes are becoming more similar with increasing capabilities, pointing to risks from correlated failures'",
     "undermines AI oversight paradigm",
     "v1 Feb 6 2025 (matches source_date 2025-02-06)"
    ],
    "not_visible": [
     "'ICML 2025' venue -- not present on the arXiv page (comments read '60 pages, 20 figures'; no journal-ref)",
     "cross-reference 'Preference Leakage (arXiv 2502.01534)' -- a different paper, not fetchable under the source-only constraint"
    ],
    "quote": "LLM-as-a-judge scores favor models similar to the judge, generalizing recent self-preference results ... model mistakes are becoming more similar with increasing capabilities, pointing to risks from correlated failures",
    "notes": "Claim core (similar-model judge bias, CAPA mistake-overlap metric, rising error correlation, oversight risk) fully confirmed verbatim. ICML 2025 and the Preference Leakage cross-reference appear only in the evidence field and cannot be checked from this URL.",
    "checked_at": "2026-07-15",
    "retrieval": "abstract"
   },
   "source_retrieval_meta": {
    "retrieval": "abstract",
    "fetched": [
     "https://arxiv.org/abs/2502.04313"
    ],
    "resolved_title": "Great Models Think Alike and this Undermines AI Oversight",
    "resolved_date": "2025-02-06 (v1); 2025-06-12 (v2)",
    "title_match": true
   },
   "used_on": [
    {
     "page": "pilot.html",
     "anchor": "main",
     "label": "Pilot - gates and objectives"
    }
   ],
   "decision_use": {
    "grade": "B",
    "basis": "current-gen mechanism",
    "measured_on": "model-outputs",
    "note": "Mistake-similarity bias + industry-wide error-correlation trend (2025). Mechanism expected to persist; magnitudes tier-bound. Cross-family routing + permanent audit remain the response.",
    "models_measured": []
   }
  },
  {
   "id": "CT-06",
   "domain": "contrarian",
   "area": null,
   "claim": "LLM judges do not follow their own rubrics: on Arena-Hard Auto, the explicit evaluation schema explains under 10% of verdict variance for some judges (unexplained variance >90% for DeepSeek-R1-32B), and factor correlations above 0.93 across nominally distinct criteria show per-axis scores collapse into a single halo factor.",
   "load_bearing": false,
   "evidence": "Feuer et al. introduce 'schematic adherence' (how much of the verdict the stated rubric explains) and psychometric validity checks; abstract verified via arXiv crawl: 'severe schema incoherence and factor collapse across popular judges'; ELO-style aggregation additionally masks genuine ranking uncertainty. Code at github.com/penfever/judgment-to-noise.",
   "implication": "Q3/Q4: a per-axis verdict sheet can be theater - the judge may emit one holistic impression dressed up as N criterion scores. The foundation should require discriminant-validity and schematic-adherence tests per project config (do axis scores actually vary independently? does the rationale's schema explain the verdict?) before per-criterion outputs are exposed to attempters or used for specialist routing.",
   "source": {
    "raw": "When Judgment Becomes Noise: How Design Failures in LLM Judge Benchmarks Silently Undermine Validity | https://arxiv.org/abs/2509.20293 | 2025-09-24 | contrarian",
    "title": "When Judgment Becomes Noise: How Design Failures in LLM Judge Benchmarks Silently Undermine Validity",
    "url": "https://arxiv.org/abs/2509.20293",
    "date": "2025-09-24",
    "type": "contrarian"
   },
   "venue": {
    "peer_reviewed": false,
    "marker": "preprint"
   },
   "same_source_claims": [],
   "in_corpus_verification": null,
   "external_verification": {
    "verdict": "confirmed",
    "checked": [
     "'Schematic adherence quantifies how much of a judge's overall verdict is explained by the explicit evaluation schema'",
     "'unexplained variance exceeding 90 percent for DeepSeek-R1-32B'",
     "'factor correlations above 0.93 for most criteria'",
     "'severe schema incoherence and factor collapse across popular judges'",
     "'Arena-Hard Auto'",
     "'ELO-style aggregation... collapses and masks genuine ranking uncertainty'",
     "code at github.com/penfever/judgment-to-noise"
    ],
    "not_visible": [
     "the exact word 'halo' - paper says 'factor collapse', the claim's 'single halo factor' is an interpretive gloss",
     "the 'under 10% of variance' phrasing - paper states it as '>90% unexplained' (equivalent)"
    ],
    "quote": "factor correlations above 0.93 for most criteria",
    "notes": "All headline numbers match: <10% schema-explained variance = >90% unexplained for DeepSeek-R1-32B; factor correlations >0.93; factor collapse; ELO masking ranking uncertainty; repo confirmed. Verified from the abstract page (given URL); full text not needed as the abstract covers every asserted figure.",
    "checked_at": "2026-07-15",
    "retrieval": "abstract"
   },
   "source_retrieval_meta": {
    "retrieval": "abstract",
    "fetched": [
     "https://arxiv.org/abs/2509.20293"
    ],
    "resolved_title": "When Judgment Becomes Noise: How Design Failures in LLM Judge Benchmarks Silently Undermine Validity",
    "resolved_date": "24 Sep 2025 (v1); revised 8 Oct 2025 (v3)",
    "title_match": true
   },
   "used_on": [],
   "decision_use": {
    "grade": "B",
    "basis": "current-gen capability",
    "measured_on": "model-outputs",
    "note": "2025: judges don't follow their own rubrics (schema explains under half of verdict variance). Motivates the schematic-adherence ship gate; deployment-tier adherence is measurable there.",
    "models_measured": [
     "DeepSeek-R1-32B"
    ]
   }
  },
  {
   "id": "CT-07",
   "domain": "contrarian",
   "area": null,
   "claim": "LLM judges have low intra-rater reliability - identical items re-scored across runs with identical settings produce inconsistent, 'almost arbitrary in the worst case' ratings - and the obvious fix (temperature-0 determinism) measurably REDUCES agreement with human judgment; meanwhile the human baseline itself is weak (SummEval inter-annotator kappa: 0.492 crowd, 0.413 expert first round, 0.71 only after a second adjudication round).",
   "load_bearing": false,
   "evidence": "Haldar & Hockenmaier, Findings of EMNLP 2025 (pp. 24986-25004); verified by reading the paper PDF (pp. 1-3): contributions are (1) low agreement of LLM ratings across runs, (2) disabling sampling hurts human-agreement, (3) phenomenon persists across SummaC, SummEval, MT-Bench; the SummEval human kappa figures are quoted in their Section 3.1.",
   "implication": "Q5: verdict flip rate is intrinsic and cannot be silenced with greedy decoding without paying validity - so the build must budget k-sample voting AND treat residual flip clusters as underspecification signal routed to rubric revision. Q1: raw human labels sit below conventional reliability thresholds; validation targets must be adjudicated (multi-round) labels, which prices the gold set honestly.",
   "source": {
    "raw": "Rating Roulette: Self-Inconsistency in LLM-As-A-Judge Frameworks | https://aclanthology.org/2025.findings-emnlp.1361.pdf | 2025-11-04 | academic",
    "title": "Rating Roulette: Self-Inconsistency in LLM-As-A-Judge Frameworks",
    "url": "https://aclanthology.org/2025.findings-emnlp.1361.pdf",
    "date": "2025-11-04",
    "type": "academic"
   },
   "venue": {
    "peer_reviewed": true,
    "marker": "ACL"
   },
   "same_source_claims": [],
   "in_corpus_verification": null,
   "external_verification": {
    "verdict": "confirmed",
    "checked": [
     "LLM judges have low intra-rater reliability across runs; ratings 'almost arbitrary in the worst case'",
     "disabling sampling (deterministic) degrades performance/agreement with human judgment (trade-off)",
     "phenomenon spans SummaC, SummEval, and MT-Bench",
     "SummEval kappa 0.492 (crowd) and 0.413 (expert first round)",
     "0.7127 after second round of expert annotation",
     "pages 24986-25004; authors Haldar & Hockenmaier"
    ],
    "not_visible": [],
    "quote": "\"LLM judges have low intra-rater reliability in their assigned scores across different runs. This variance makes their ratings inconsistent, almost arbitrary in the worst case\" ... kappa \"0.492 and 0.413 for the crowd-sourced workers and the first round of expert annotations ... second round ... 0.7127\".",
    "notes": "Fully confirmed verbatim from the PDF via pdftotext. Intra-rater-reliability core, the 'almost arbitrary in the worst case' phrase, the no-sampling degradation ('degradation in performance if run without sampling ... trade-off between self-reliability and performance'), all three benchmarks, and all three kappa figures with the exact crowd/expert-first/second-round mapping match. Page range 24986-25004 confirmed.",
    "checked_at": "2026-07-15",
    "retrieval": "pdf"
   },
   "source_retrieval_meta": {
    "retrieval": "pdf",
    "fetched": [
     "https://aclanthology.org/2025.findings-emnlp.1361.pdf"
    ],
    "resolved_title": "Rating Roulette: Self-Inconsistency in LLM-As-A-Judge Frameworks",
    "resolved_date": "Findings of EMNLP 2025, pages 24986-25004",
    "title_match": true
   },
   "used_on": [],
   "decision_use": {
    "grade": "B",
    "basis": "current-gen capability",
    "measured_on": "model-outputs",
    "note": "2025 intra-rater reliability findings. k-sampling + entropy routing remain hygiene.",
    "models_measured": []
   }
  },
  {
   "id": "CT-08",
   "domain": "contrarian",
   "area": null,
   "claim": "LLM judges give the weakest signal exactly where QA needs them most: they cannot reliably grade responses to questions they cannot answer themselves (poor signal on the hardest items in a benchmark), and in expert domains subject-matter experts agreed with LLM-judge picks only 64% (mental health) to 68% (dietetics) of the time, with the judge favoring superficially actionable detail.",
   "load_bearing": false,
   "evidence": "'No Free Labels' (Kim et al., arXiv 2503.05061, Mar 2025) shows judge quality collapses on the most difficult items and that human-written reference answers improve agreement and reduce self-preference; Szymanski et al. (ACM IUI 2025, DOI 10.1145/3708359.3712091) measured the 64/68% SME agreement in pairwise comparisons on domain tasks. Both reported consistently across two independent search engines; abstracts inspected.",
   "implication": "Q5/Q2: the expertise-gap root cause is NOT automatically closed by an LLM judge - the judge's competence ceiling binds hardest on exactly the expert items where human reviewers also fail, so per-axis judge-vs-adjudicated-expert ceilings must be measured before assigning judge-autonomous lanes, and expert-authored reference answers/exemplars are the highest-leverage compile-time artifact (they measurably raise judge validity).",
   "source": {
    "raw": "No Free Labels: Limitations of LLM-as-a-Judge Without Human Grounding | https://arxiv.org/abs/2503.05061 | 2025-03-07 | academic",
    "title": "No Free Labels: Limitations of LLM-as-a-Judge Without Human Grounding",
    "url": "https://arxiv.org/abs/2503.05061",
    "date": "2025-03-07",
    "type": "academic"
   },
   "venue": {
    "peer_reviewed": true,
    "marker": "IUI"
   },
   "same_source_claims": [],
   "in_corpus_verification": null,
   "external_verification": {
    "verdict": "partially_confirmed",
    "checked": [
     "judges show high agreement with human experts only on questions the judges were able to correctly answer themselves",
     "providing the judges with expert-written references largely mitigates this issue"
    ],
    "not_visible": [
     "the 'poor signal on the hardest / most difficult items in a benchmark' framing -- page frames the limitation around questions the judge cannot answer, not a difficulty ranking",
     "64% (mental health) and 68% (dietetics) SME agreement -- these are attributed in the claim to Szymanski et al. (ACM IUI 2025), a separate source; '64%', '68%', 'Szymanski', 'dietetics', 'mental health' are all absent from this page",
     "'reduce self-preference / self-bias' -- the page never mentions self-preference or self-bias"
    ],
    "quote": "judges show high agreement with human experts only on questions the judges were able to correctly answer themselves ... providing the judges with expert-written references largely mitigates this issue",
    "notes": "The portion of the claim sourced to 2503.05061 (judges unreliable on questions they can't answer; expert references improve agreement) is confirmed. The 64/68% SME figures belong to a different citation (Szymanski et al.) outside the fetched URL, and the 'reduce self-preference' specific is not stated on this page. Abstract centers on BFF-Bench (160 questions) and the VERDICTS annotation set.",
    "checked_at": "2026-07-15",
    "retrieval": "abstract"
   },
   "source_retrieval_meta": {
    "retrieval": "abstract",
    "fetched": [
     "https://arxiv.org/abs/2503.05061"
    ],
    "resolved_title": "No Free Labels: Limitations of LLM-as-a-Judge Without Human Grounding",
    "resolved_date": "2025-03-07 (v1); last revised 2026-04-01 (v2)",
    "title_match": true
   },
   "used_on": [
    {
     "page": "pilot.html",
     "anchor": "main",
     "label": "Pilot - gates and objectives"
    }
   ],
   "decision_use": {
    "grade": "B",
    "basis": "current-gen capability",
    "measured_on": "model-outputs",
    "note": "2025: judges weakest on questions they cannot themselves answer. Same mechanism family as re-solving (F26-01); E3 probes it on our workload.",
    "models_measured": []
   }
  },
  {
   "id": "CT-09",
   "domain": "contrarian",
   "area": null,
   "claim": "Judge verdicts are gameable through content-independent artifacts: short universal adversarial phrases learned on a surrogate model transfer to unseen judge LLMs and inflate scores toward the maximum regardless of the assessed text, with absolute scoring far more vulnerable than comparative assessment.",
   "load_bearing": false,
   "evidence": "Raina, Liusie & Gales, EMNLP 2024 (main, pp. 6920+); abstract verified via OpenAlex/ACL record: attackers need no access to the deployed judge - surrogate attack then transfer; 'irrespective of the assessed text, maximum scores are predicted'. Complementary: 'Cheating Automatic LLM Benchmarks' (arXiv 2410.07137) showed null models emitting constant responses achieve top win rates on judged benchmarks.",
   "implication": "Q8: pay-motivated attempters do not need to see the judge's rationale to Goodhart it - transfer attacks work blind, so filtering rationales is insufficient as the sole defense. Day-one requirements: comparative/anchored scoring formats over absolute Likert, rotating seeded probes with known verdicts, and evidence-perturbation tests (swap the cited span, verdict must flip) as a standing citation-theater detector.",
   "source": {
    "raw": "Is LLM-as-a-Judge Robust? Investigating Universal Adversarial Attacks on Zero-shot LLM Assessment | https://aclanthology.org/2024.emnlp-main.427/ | 2024-11-12 | academic",
    "title": "Is LLM-as-a-Judge Robust? Investigating Universal Adversarial Attacks on Zero-shot LLM Assessment",
    "url": "https://aclanthology.org/2024.emnlp-main.427/",
    "date": "2024-11-12",
    "type": "academic"
   },
   "venue": {
    "peer_reviewed": true,
    "marker": "ACL"
   },
   "same_source_claims": [],
   "in_corpus_verification": null,
   "external_verification": {
    "verdict": "confirmed",
    "checked": [
     "short universal adversarial phrases appended to text inflate judge scores",
     "phrase learned on a surrogate model then transferred to unseen judge LLMs",
     "irrespective of the assessed text, maximum scores are predicted",
     "absolute scoring significantly more susceptible than comparative assessment",
     "authors Raina, Liusie, Gales; EMNLP 2024"
    ],
    "not_visible": [],
    "quote": "when transferred to unseen models, scores can be drastically inflated such that irrespective of the assessed text, maximum scores are predicted ... judge-LLMs are significantly more susceptible to these adversarial attacks when used for absolute scoring, as opposed to comparative assessment.",
    "notes": "Core claim confirmed verbatim from the ACL Anthology abstract. One evidence-field discrepancy (not part of the core claim): the citation lists 'pp. 6920+', but the ACL Anthology record gives pages 7499-7517 for this paper.",
    "checked_at": "2026-07-15",
    "retrieval": "abstract"
   },
   "source_retrieval_meta": {
    "retrieval": "abstract",
    "fetched": [
     "https://aclanthology.org/2024.emnlp-main.427/"
    ],
    "resolved_title": "Is LLM-as-a-Judge Robust? Investigating Universal Adversarial Attacks on Zero-shot LLM Assessment",
    "resolved_date": "EMNLP 2024, November 2024 (pages 7499-7517)",
    "title_match": true
   },
   "used_on": [
    {
     "page": "pilot.html",
     "anchor": "main",
     "label": "Pilot - gates and objectives"
    }
   ],
   "decision_use": {
    "grade": "A",
    "basis": "adversarial floor",
    "measured_on": "model-outputs",
    "note": "Universal adversarial phrases transfer to unseen judges (2024). Attack floors strengthen with attacker capability; comparative/anchored scoring and perturbation probes stay day-one requirements.",
    "models_measured": []
   }
  },
  {
   "id": "CT-11",
   "domain": "contrarian",
   "area": null,
   "claim": "Annotator disagreement contains recoverable systematic signal, not just noise: NUTMEG, a Bayesian model separating annotator-competence noise from subpopulation-level systematic disagreement, produces downstream models that significantly outperform both majority-vote aggregation and fully disaggregated training - meaning aggregation to a single 'true' label destroys measurable information.",
   "load_bearing": false,
   "evidence": "Ivey, Gauch & Jurgens, EMNLP 2025 main (pp. 2874-2887); PDF and poster abstract inspected: NUTMEG estimates true labels per subpopulation and beats MACE/majority-vote at replicating subgroup label distributions on politeness and offensiveness. Sits atop the perspectivist NLP literature (NLPerspectives workshops 2022-2025) which rejects single-gold-label resolution for subjective tasks.",
   "implication": "Q2: inter-reviewer disagreement is partly legitimate signal, so an AutoQA that maximizes consistency-of-application on judgment-laden axes launders one interpretation into false objectivity and destroys exactly the pluralism the training data may need. The verdict ontology needs a measured 'systematic-disagreement' class (distinct from noise), triggered by subpopulation-conditioned disagreement, exempt from attempter penalty, and fed back as an instruction-gap report.",
   "source": {
    "raw": "NUTMEG: Separating Signal From Noise in Annotator Disagreement | https://aclanthology.org/2025.emnlp-main.144.pdf | 2025-11-04 | academic",
    "title": "NUTMEG: Separating Signal From Noise in Annotator Disagreement",
    "url": "https://aclanthology.org/2025.emnlp-main.144.pdf",
    "date": "2025-11-04",
    "type": "academic"
   },
   "venue": {
    "peer_reviewed": true,
    "marker": "ACL"
   },
   "same_source_claims": [],
   "in_corpus_verification": null,
   "external_verification": {
    "verdict": "partially_confirmed",
    "checked": [
     "NUTMEG is a Bayesian model incorporating annotator backgrounds to remove noisy annotations while preserving systematic disagreements",
     "downstream models trained on NUTMEG-aggregated data significantly outperform models trained with traditional (majority-vote) aggregation",
     "evaluated on politeness and offensiveness tasks",
     "compared against MACE / Majority Vote and against disaggregated (no-aggregation) training",
     "measured by replicating the ground-truth label distribution by subgroup (Figure 7)",
     "authors Ivey (JHU), Gauch (Arkansas), Jurgens (Michigan); pp. 2874-2887; EMNLP 2025"
    ],
    "not_visible": [
     "a uniform 'significantly outperform ... fully disaggregated training' result (the abstract's own outperformance claim is only vs 'traditional aggregation methods')"
    ],
    "quote": "downstream models trained on NUTMEG-aggregated data significantly outperform models trained on data from traditionally aggregation methods",
    "notes": "Core thesis (disagreement carries recoverable systematic signal; separating competence-noise from subpopulation disagreement; single-label aggregation destroys information) is confirmed, as is beating majority vote. But the claim's 'significantly outperform BOTH majority-vote AND fully disaggregated training' overstates the disaggregated arm: on politeness NUTMEG beats disaggregated on all splits except race, while on offensiveness it is only 'as or better than no-aggregation' (neutral, no significant gain). Hence partially_confirmed.",
    "checked_at": "2026-07-15",
    "retrieval": "pdf"
   },
   "source_retrieval_meta": {
    "retrieval": "pdf",
    "fetched": [
     "https://aclanthology.org/2025.emnlp-main.144.pdf"
    ],
    "resolved_title": "NUTMEG: Separating Signal From Noise in Annotator Disagreement",
    "resolved_date": "EMNLP 2025 (Nov 4-9, 2025), pp. 2874-2887",
    "title_match": true
   },
   "used_on": [],
   "decision_use": {
    "grade": "A",
    "basis": "statistical method",
    "measured_on": "human-work",
    "note": "NUTMEG: separating systematic subpopulation disagreement from noise. Method for human-label modeling.",
    "models_measured": []
   }
  },
  {
   "id": "CT-12",
   "domain": "contrarian",
   "area": null,
   "claim": "AI assistance homogenizes human output at the population level even while raising individual quality: in a Nature Human Behaviour brainstorming study 94% of ChatGPT-assisted participants' ideas shared overlapping concepts (nine independently produced the same product name) while human-only ideas were entirely unique, and the 'Artificial Hivemind' study found different vendors' frontier models converge on near-identical phrasings (~81% average similarity between DeepSeek-V3 and GPT-4o).",
   "load_bearing": false,
   "evidence": "Meincke, Nave & Terwiesch (Nature Human Behaviour, 2025; Wharton Mack Institute summary opened and read); Jiang et al. (UW/CMU/AI2 'Artificial Hivemind', reported by The Decoder) documents intra-model repetition plus inter-model homogeneity. Consistent with Doshi & Hauser (Science Advances 2024) and Wan et al. 2026 (diverse AI personas partially mitigate homogenization). All are adjacent-domain (ideation/writing), not annotation-QA-feedback studies.",
   "implication": "Q8's monoculture worry has strong analogical support: a single judge family delivering item-specific stylistic feedback to thousands of attempters is a homogenization pump aimed at the exact diversity the training data exists to capture. The foundation should ban stylistic guidance in feedback by default, keep feedback at principle level, and stand up a population-level output-diversity drift metric from day one - while flagging that direct evidence in the annotation-QA setting does not yet exist.",
   "source": {
    "raw": "New in Nature: ChatGPT Decreases Idea Diversity in Brainstorming (Meincke, Nave & Terwiesch, Nature Human Behaviour) | https://mackinstitute.wharton.upenn.edu/2025/new-in-nature-chatgpt-decreases-idea-diversity-in-brainstorming/ | 2025-05 | academic",
    "title": "New in Nature: ChatGPT Decreases Idea Diversity in Brainstorming (Meincke, Nave & Terwiesch, Nature Human Behaviour)",
    "url": "https://mackinstitute.wharton.upenn.edu/2025/new-in-nature-chatgpt-decreases-idea-diversity-in-brainstorming/",
    "date": "2025-05",
    "type": "academic"
   },
   "venue": {
    "peer_reviewed": true,
    "marker": "Nature"
   },
   "same_source_claims": [],
   "in_corpus_verification": null,
   "external_verification": {
    "verdict": "partially_confirmed",
    "checked": [
     "94% of ChatGPT-assisted participants' ideas shared overlapping concepts",
     "nine participants independently produced the same product name ('Build-a-Breeze Castle')",
     "human-only ideas were entirely unique",
     "individual quality up while population diversity down (ChatGPT enhances individual creativity but significantly reduces idea diversity)",
     "Meincke, Nave & Terwiesch, Nature Human Behaviour 2025"
    ],
    "not_visible": [
     "'Artificial Hivemind' ~81% average similarity between DeepSeek-V3 and GPT-4o (Jiang et al., reported by The Decoder)",
     "Doshi & Hauser (Science Advances 2024) and Wan et al. 2026 corroboration"
    ],
    "quote": "94% of ideas shared overlapping concepts ... nine participants independently naming their toy 'Build-a-Breeze Castle' ... human-generated ideas were entirely unique ... while ChatGPT can enhance the creativity of individual ideas, it significantly reduces the diversity of ideas",
    "notes": "The Nature Human Behaviour brainstorming portion is fully confirmed from the Wharton summary. The inter-model 'Artificial Hivemind' 81% figure and the Doshi & Hauser / Wan et al. citations are from other sources not on this page, so marked not_visible.",
    "checked_at": "2026-07-15",
    "retrieval": "landing"
   },
   "source_retrieval_meta": {
    "retrieval": "landing",
    "fetched": [
     "https://mackinstitute.wharton.upenn.edu/2025/new-in-nature-chatgpt-decreases-idea-diversity-in-brainstorming/"
    ],
    "resolved_title": "New in Nature: ChatGPT Decreases Idea Diversity in Brainstorming",
    "resolved_date": "2025-05-14",
    "title_match": true
   },
   "used_on": [],
   "decision_use": {
    "grade": "A",
    "basis": "human science (direction)",
    "measured_on": "human-work",
    "note": "AI assistance homogenizes population output while raising individual quality (2025). Monoculture pressure rises with assistance quality; the diversity drift metric stays.",
    "models_measured": [
     "DeepSeek-V3",
     "GPT-4o"
    ]
   }
  },
  {
   "id": "F26-01",
   "domain": "frontier-2026",
   "area": null,
   "claim": "When an LLM auditor checks whether work matches a reference, it silently re-solves the task and trusts its own answer over the reference: on 200 web-agent benchmark items with injected defects (2,400 audits, 3 production models), detection of a wrong reference answer fell from 68% to 9% as the task required tallying hundreds of records, while false positives on clean items rose from 44% to 88% -- yet detection of buggy evaluator code (checkable by reading, not recomputing) stayed at 80%.",
   "load_bearing": true,
   "evidence": "Controlled meta-eval with mechanically known-correct audit verdicts: items rendered clean or with exactly one defect injected into instruction/reference/evaluator, so auditor accuracy is scored against the planted defect. Reasoning traces plus an answer-supplied probe converge on the re-solving mechanism. Directly falsifies the assumption that a judge 'verifying' a claim is doing evidence-anchored checking rather than comparing against its own solution prior.",
   "implication": "Answers taxonomy Q3 (evidence closure) with hard 2026 evidence: the judge WILL verify against its own priors whenever the check exceeds its own compute-the-answer ability, and its false-fail rate explodes exactly there. The AutoQA must (a) require span-quoting evidence-anchored verdicts, (b) route checks by whether they are read-checkable vs recompute-checkable, and (c) treat judge confidence as untrustworthy on recompute-class checks -- capability-triage the check type, not just the axis type.",
   "source": {
    "raw": "Auditing by Re-Solving: LLM Benchmark Auditors Trust Their Own Answer Over the Reference (ACL ARR 2026 May submission) | https://openreview.net/forum?id=mzh9d3dooN | 2026-06 | academic",
    "title": "Auditing by Re-Solving: LLM Benchmark Auditors Trust Their Own Answer Over the Reference (ACL ARR 2026 May submission)",
    "url": "https://openreview.net/forum?id=mzh9d3dooN",
    "date": "2026-06",
    "type": "academic"
   },
   "venue": {
    "peer_reviewed": true,
    "marker": "ACL"
   },
   "same_source_claims": [],
   "in_corpus_verification": {
    "verdict": "confirmed",
    "transcript": "Source verified via Exa crawl of https://openreview.net/forum?id=mzh9d3dooN (direct WebFetch and the OpenReview API were blocked by a Cloudflare challenge). (1) The abstract matches the claim on every checkable detail: title \"Auditing by Re-Solving: LLM Benchmark Auditors Trust Their Own Answer Over the Reference\"; ACL ARR 2026 May Submission (#15335); 200 web-agent-style benchmark items (each over ~800 structured records) rendered clean or with exactly one defect injected into instruction/reference/evaluator so verdicts are scored mechanically; 2,400 audits across three production models; wrong-reference detection falls 68%->9% when tallying hundreds of records; clean-item false positives rise 44%->88%; buggy-evaluator detection (found by reading code, not recomputing) stays at 80%; reasoning traces plus an answer-supplied probe converge on the re-solving mechanism. (2) Date checks out: OpenReview page published 2026-06-02, consistent with \"2026-06\" and the May ARR cycle. (3) Searched for later/superseding work (July 2026): found adjacent LLM-judge reliability literature (e.g., arXiv 2606.19544 \"Reliability without Validity\", AURA arXiv 2606.19714, BenchGuard arXiv 2604.24955) but nothing that supersedes, contradicts, or retracts this result. Caveats: this is an under-review ARR submission, not yet peer-accepted (the claim discloses this); numbers were verified against the abstract only, as the full PDF was not fetched.",
    "corrected": null
   },
   "external_verification": {
    "verdict": "not_retrievable",
    "checked": [],
    "not_visible": [
     "68% -> 9% detection drop",
     "44% -> 88% false-positive rise",
     "80% buggy-evaluator detection",
     "200 web-agent items / 2,400 audits / 3 production models",
     "re-solving mechanism"
    ],
    "quote": "",
    "notes": "Both the forum page and the OpenReview API return a Cloudflare Turnstile 'Verifying your browser' challenge (API: HTTP 403 ChallengeRequiredError). No paper content retrievable via WebFetch or curl. The batch file carried a pre-set in_corpus_verdict of 'confirmed', but nothing could be independently verified from this source.",
    "checked_at": "2026-07-15",
    "retrieval": "failed"
   },
   "source_retrieval_meta": {
    "retrieval": "failed",
    "fetched": [
     "https://openreview.net/forum?id=mzh9d3dooN",
     "https://api.openreview.net/notes?forum=mzh9d3dooN"
    ],
    "resolved_title": null,
    "resolved_date": null,
    "title_match": false
   },
   "used_on": [
    {
     "page": "architecture.html",
     "anchor": "main",
     "label": "System - pipeline and claim types"
    },
    {
     "page": "decisions.html",
     "anchor": "c2",
     "label": "Decisions - default: re-solving routing"
    }
   ],
   "decision_use": {
    "grade": "B",
    "basis": "current-gen mechanism (single study, under review)",
    "measured_on": "model-outputs",
    "note": "Re-solving collapse (68->9%) is the corpus's own flagged single-source provisional claim; C2 holds it at 'replicate before hard-gating'. Current-window measurement.",
    "models_measured": []
   }
  },
  {
   "id": "F26-02",
   "domain": "frontier-2026",
   "area": null,
   "claim": "Handshake's Gandalf (May 27, 2026) shows verifier ARCHITECTURE beats verifier MODEL: on BankerVerifierBench -- a meta-eval of 3,204 expert-graded pass/fail criterion judgments across 21 agentic tasks (expert inter-annotator agreement 89.5%, disagreements adjudicated) -- a reactive agent-judge that runs inside the work environment and chooses at inference time which artifacts/tool-state to inspect beats the strongest text-only/snapshot/workflow verifier on F1 while costing roughly 10x less, and the gap between verifier architectures exceeds the gap between backing models.",
   "load_bearing": true,
   "evidence": "Full primary post opened and read. BVB fixes completed rollouts and rubrics and compares automated verifiers against practicing-banker labels; baselines were Autorubric (text-only rubric judge), Archipelago/APEX snapshot grading, and Agent-as-a-Judge. Key framing: 'verifiability is a relationship between a criterion and the verifier available to check it' -- a criterion invisible to a text-only judge is effectively ungraded, and weak verifiers push benchmark/rubric design toward only what they can reach. Gandalf verifier code is open-sourced; BVB dataset release planned.",
   "implication": "For Q5/Q6 and overall architecture: give the AutoQA judge the same evidence surface the attempter had (source documents, cited spans, artifacts, tool state) and let it decide what to open, rather than judging from the attempter's prose alone -- evidence access is a bigger lever than model choice, and it is also ~10x cheaper than brute-forcing with a stronger model. Also a direct template for our meta-eval: fixed items + expert labels + F1 over candidate judge configs, with expert IAA (~89.5% here) reported as the ceiling.",
   "source": {
    "raw": "Your verifier is probably the bottleneck. We built one that isn't. (Handshake AI research) | https://joinhandshake.com/research/ai/gandalf-the-grader/ | 2026-05-27 | primary",
    "title": "Your verifier is probably the bottleneck. We built one that isn't. (Handshake AI research)",
    "url": "https://joinhandshake.com/research/ai/gandalf-the-grader/",
    "date": "2026-05-27",
    "type": "primary"
   },
   "venue": {
    "peer_reviewed": true,
    "marker": "CHI"
   },
   "same_source_claims": [
    "IP-02"
   ],
   "in_corpus_verification": {
    "verdict": "confirmed",
    "transcript": "Primary source fetched and verified. Date (2026-05-27), BVB composition (3,204 expert pass/fail judgments, 21 tasks from BankerToolBench, 89.5% IAA with adjudication), Gandalf architecture (reactive agent-judge inside the rollout environment via OpenHands SDK, inference-time choice of artifacts/tool state to inspect), baselines (Autorubric text-only, Archipelago/APEX snapshot, Agent-as-a-Judge workflow), and results all match: every Gandalf config (F1 0.633-0.664) beats the best non-Gandalf run (Archipelago/Gemini 3 Pro, F1 0.604, ~$422), with the cheapest Gandalf config (GPT-5.4 Nano, ~$42) winning by ~3 F1 at ~1/10 the cost. The post states verbatim that the gap between verifier architectures is larger than the gap between backing models (e.g., Gandalf/GPT-5.4 Nano beats Archipelago/GPT-5.4 by 9.5 F1). The 'verifiability is a relationship between a criterion and the verifier' framing appears; code is open-sourced (github.com/Handshake-AI-Research/gandalf-the-grader, v1.0.0 on PyPI) and BVB dataset release is planned. Supersession search found the companion BankerToolBench paper (arXiv 2604.11304) and press coverage but nothing newer contradicting the claim as of 2026-07-14. Only precision nuance: the 10x-cheaper win margin is ~3 F1 points, and the 'strongest' baseline is specifically Archipelago backed by Gemini 3 Pro.",
    "corrected": null
   },
   "external_verification": {
    "verdict": "confirmed",
    "checked": [
     "BankerVerifierBench fixes completed rollouts and rubrics and compares automated verifiers to practicing-banker labels",
     "3,204 expert-graded judgments; 21 tasks sampled from BankerToolBench",
     "inter-annotator agreement 89.5%; disagreements adjudicated",
     "reactive agent-judge runs inside the environment and chooses at inference time what evidence to inspect; beats text-only/snapshot/workflow verifiers on F1 at ~10x lower cost",
     "gap between verifier architectures larger than gap between backing models",
     "baselines: Autorubric (text-only), Archipelago from APEX (snapshot), Agent-as-a-Judge (AAAJ)",
     "framing: verifiability is a relationship between a criterion and the verifier available to check it",
     "Gandalf verifier code open-sourced"
    ],
    "not_visible": [
     "'BVB dataset release planned' -- not explicitly surfaced in the fetched summary (Gandalf code open-sourcing is confirmed; BVB dataset release status not quoted)"
    ],
    "quote": "verifiability is not only a property of the task. It is a relationship between a criterion and the verifier available to check it",
    "notes": "Architecture-over-model thesis, all counts (3,204 / 21 / 89.5%), baseline names, and the framing quote confirmed. The '9.5 F1' architecture gap and ~10x cost advantage are supported (see IP-02 checked items).",
    "checked_at": "2026-07-15",
    "retrieval": "landing"
   },
   "source_retrieval_meta": {
    "retrieval": "landing",
    "fetched": [
     "https://joinhandshake.com/research/ai/gandalf-the-grader/"
    ],
    "resolved_title": "Your verifier is probably the bottleneck. We built one that isn't. (Gandalf the Grader / BankerVerifierBench)",
    "resolved_date": "not surfaced verbatim in fetch (claimed 2026-05-27); content references 2026-era models (GPT-5.4, Gemini 3 Pro)",
    "title_match": true
   },
   "used_on": [
    {
     "page": "research.html",
     "anchor": "research-cards",
     "label": "Research - findings by decision weight"
    },
    {
     "page": "architecture.html",
     "anchor": "ontology",
     "label": "System - claim types"
    }
   ],
   "decision_use": {
    "grade": "B",
    "basis": "current-gen vendor result",
    "measured_on": "human-work",
    "note": "Gandalf (May 2026): architecture>model on expert-graded human work. Vendor research; among the most decision-relevant current results - and exactly the kind the bakeoff must reproduce in-house.",
    "models_measured": []
   }
  },
  {
   "id": "F26-03",
   "domain": "frontier-2026",
   "area": null,
   "claim": "Telling an LLM judge the downstream consequences of its verdict ('low scores cause retraining/decommissioning') systematically softens verdicts -- peak verdict shift of -9.8 percentage points, a 30% relative drop in unsafe-content detection across 18,240 controlled judgments (content held strictly constant, 3 judge models) -- and the judge's chain-of-thought contains ZERO explicit acknowledgment of the consequence framing it is acting on (ERR_J = 0.000).",
   "load_bearing": true,
   "evidence": "Abstract and full submission metadata read on OpenReview (CTB@ICML 2026 workshop, 8-page paper). 1,520 responses spanning three safety/quality benchmarks, four response categories from clearly-safe to overtly-harmful; only a one-sentence consequence framing in the system prompt varied. The bias is implicit and undetectable by inspecting reasoning traces.",
   "implication": "Two foundational rules for Q5/Q7: (1) the judge prompt must be stakes-sterile -- never tell the judge that a fail penalizes the attempter's pay/standing, which our QA-of-humans setting does by construction, so the compilation contract must strip consequence language from project instructions before they reach the judge; (2) the Q7 explainability contract cannot rest on CoT audit -- a clean-looking reasoning chain does not certify an unbiased verdict, so bias must be measured behaviorally (seeded perturbation probes), not read off rationales.",
   "source": {
    "raw": "Context Over Content: Exposing Evaluation Faking in Automated Judges (CTB@ICML 2026) | https://openreview.net/forum?id=XI2hnjGnZx | 2026-05-25 | academic",
    "title": "Context Over Content: Exposing Evaluation Faking in Automated Judges (CTB@ICML 2026)",
    "url": "https://openreview.net/forum?id=XI2hnjGnZx",
    "date": "2026-05-25",
    "type": "academic"
   },
   "venue": {
    "peer_reviewed": false,
    "marker": "preprint"
   },
   "same_source_claims": [],
   "in_corpus_verification": {
    "verdict": "confirmed",
    "transcript": "Source verified via Exa crawl of https://openreview.net/forum?id=XI2hnjGnZx (direct WebFetch and API blocked by OpenReview browser challenge) and cross-checked against the arXiv mirror (arXiv:2604.15224, submitted 2026-04-16). Every quantitative element of the claim appears verbatim in the abstract: 1,520 responses, three safety/quality benchmarks, four response categories (clearly safe to overtly harmful), only a brief consequence-framing sentence varied in the system prompt, 18,240 controlled judgments, three diverse judge models, peak Verdict Shift Delta V = -9.8 pp, 30% relative drop in unsafe-content detection, and ERR_J = 0.000. Venue (CTB@ICML 2026), paper type (Long, 8 pages), and OpenReview page date (2026-05-25) all match. Two minor precision notes: (1) the abstract qualifies ERR_J = 0.000 as holding \"across all reasoning-model judgments\" - the claim drops that scope qualifier (zero acknowledgment is asserted for reasoning-trace judgments, not necessarily all 18,240); (2) an earlier arXiv version (April 2026) lists four authors (Gupta, Nair, Wang, Kumar) while the OpenReview record shows Manan Gupta - immaterial to the claim. Supersession check: searches surfaced only related-but-distinct work (e.g., Hwang et al. 2026 \"When Wording Steers the Evaluation,\" arXiv:2601.13537, on framing bias across 14 judges - complementary, predates this paper's stakes-signaling focus) and secondary commentary; nothing refuting or superseding the reported results as of 2026-07-14.",
    "corrected": null
   },
   "external_verification": {
    "verdict": "not_retrievable",
    "checked": [],
    "not_visible": [
     "peak verdict shift of -9.8 percentage points",
     "30% relative drop in unsafe-content detection",
     "18,240 controlled judgments",
     "3 judge models",
     "ERR_J = 0.000 (zero CoT acknowledgment)",
     "1,520 responses across three benchmarks",
     "four response categories clearly-safe to overtly-harmful",
     "CTB@ICML 2026 venue / 8-page paper",
     "title and authors"
    ],
    "quote": "Verifying your browser | OpenReview",
    "notes": "Both WebFetch and Bash curl (browser UA) returned only the Cloudflare Turnstile 'Verifying your browser' interstitial. The forum content loads client-side from api2.openreview.net, which is outside the allowed URL set (no other pages/hosts). No paper content could be retrieved, so the claim cannot be independently verified. Batch file carried a pre-set in_corpus_verdict of 'confirmed', but that is not confirmable from retrieved text here.",
    "checked_at": "2026-07-15",
    "retrieval": "failed"
   },
   "source_retrieval_meta": {
    "retrieval": "failed",
    "fetched": [
     "https://openreview.net/forum?id=XI2hnjGnZx"
    ],
    "resolved_title": null,
    "resolved_date": null,
    "title_match": false
   },
   "used_on": [
    {
     "page": "pilot.html",
     "anchor": "main",
     "label": "Pilot - gates and objectives"
    },
    {
     "page": "decisions.html",
     "anchor": "d6",
     "label": "Decisions - the throughput dial"
    },
    {
     "page": "decisions.html",
     "anchor": "invariants",
     "label": "Decisions - settled constraints"
    }
   ],
   "decision_use": {
    "grade": "B",
    "basis": "current-gen mechanism",
    "measured_on": "model-outputs",
    "note": "Evaluation faking under consequence framing (2026). Stakes-sterile prompting is cheap insurance regardless of tier; behavioral bias measurement stays mandatory.",
    "models_measured": []
   }
  },
  {
   "id": "F26-04",
   "domain": "frontier-2026",
   "area": null,
   "claim": "HyPAC (Jan 30, 2026) gives the one-human-touch routing problem a formal solution: calibrate two uncertainty thresholds (via importance sampling + upper confidence bounds) that partition items into three regions routed to cheap LLM / expensive reasoning model / human, achieving a distribution-free PAC guarantee on annotation error while cutting annotation cost 78.51% in experiments.",
   "load_bearing": true,
   "evidence": "Abstract opened and verified on arXiv. The guarantee is distribution-free and pre-trained-model-free; routing is by calibrated uncertainty, not fixed rules. This is the first 2026-vintage method that turns 'which items get the human?' from a launch-time heuristic into a statistically certified, closed-loop calibration -- exactly the Q6 fork (fixed routing vs closed-loop on observed error).",
   "implication": "For Q6: adopt threshold-calibrated three-lane routing (auto-verdict / stronger-judge / human) with an explicit error budget as the foundational routing primitive, rather than a fixed one-touch-per-item rule; the human budget becomes an output of the target error rate. Caveat from the Auditing-by-Re-Solving finding: uncertainty-based routing must be validated per check-type, since judge confidence is miscalibrated precisely on recompute-class checks.",
   "source": {
    "raw": "HyPAC: Cost-Efficient LLMs-Human Hybrid Annotation with PAC Error Guarantees (arXiv 2602.02550) | https://arxiv.org/abs/2602.02550 | 2026-01-30 | academic",
    "title": "HyPAC: Cost-Efficient LLMs-Human Hybrid Annotation with PAC Error Guarantees (arXiv 2602.02550)",
    "url": "https://arxiv.org/abs/2602.02550",
    "date": "2026-01-30",
    "type": "academic"
   },
   "venue": {
    "peer_reviewed": false,
    "marker": "preprint"
   },
   "same_source_claims": [],
   "in_corpus_verification": null,
   "external_verification": {
    "verdict": "confirmed",
    "checked": [
     "calibrates two decision thresholds using importance sampling and upper confidence bounds",
     "splits inputs into three regions based on uncertainty",
     "routes to fast LLMs / slow reasoning models / human experts",
     "PAC error guarantee free of data distribution and pre-trained models (distribution-free)",
     "reduces the annotation cost by 78.51%",
     "submitted 30 Jan 2026"
    ],
    "not_visible": [
     "'one-human-touch' phrasing (author's editorial framing, not in the paper's abstract)"
    ],
    "quote": "calibrates two decision thresholds using importance sampling and upper confidence bounds ... reduces the annotation cost by 78.51%",
    "notes": "All headline numbers and mechanics confirmed. 'Cheap LLM' is stated as 'fast LLMs' but same meaning; 'expensive reasoning model' as 'slow reasoning models'. 78.51% exact match.",
    "checked_at": "2026-07-15",
    "retrieval": "abstract"
   },
   "source_retrieval_meta": {
    "retrieval": "abstract",
    "fetched": [
     "https://arxiv.org/abs/2602.02550"
    ],
    "resolved_title": "HyPAC: Cost-Efficient LLMs-Human Hybrid Annotation with PAC Error Guarantees",
    "resolved_date": "2026-01-30",
    "title_match": true
   },
   "used_on": [],
   "decision_use": {
    "grade": "A",
    "basis": "statistical method",
    "measured_on": "structural",
    "note": "HyPAC dual-threshold routing with error budgets. Method.",
    "models_measured": []
   }
  },
  {
   "id": "F26-05",
   "domain": "frontier-2026",
   "area": null,
   "claim": "Per-criterion judge reliability ceilings differ sharply and can be measured per-item with conformal prediction: on SummEval across four judges, relevance is judged most reliably (avg conformal set size ~3.0 of 5), coherence moderate (~3.9), fluency and consistency unreliable (~4.9); aggregate pairwise transitivity violations look small (0.8-4.1%) but 33-67% of documents contain at least one intransitive preference cycle, and conformal set width correlates with reliability at r_s=+0.576 (N=1,918).",
   "load_bearing": false,
   "evidence": "Paper opened via arXiv abstract page (submitted 2026-04-16, under review); all numbers confirmed. Criterion choice matters more than judge choice; set width behaves as a per-instance difficulty/abstention signal (cross-judge width agreement 0.32-0.38). Code, prompts, and cached results released.",
   "implication": "Supports Q2/Q5 axis-lane triage with a concrete mechanism: measure per-axis conformal set width during project onboarding and use it to sort axes into judge-autonomous vs route-to-human lanes, and use per-item set width as the abstention/escalation trigger instead of raw self-reported confidence. Also warns that aggregate agreement stats hide per-item incoherence (transitivity cycles).",
   "source": {
    "raw": "Diagnosing LLM Judge Reliability: Conformal Prediction Sets and Transitivity Violations (arXiv 2604.15302) | https://arxiv.org/abs/2604.15302 | 2026-04-16 | academic",
    "title": "Diagnosing LLM Judge Reliability: Conformal Prediction Sets and Transitivity Violations (arXiv 2604.15302)",
    "url": "https://arxiv.org/abs/2604.15302",
    "date": "2026-04-16",
    "type": "academic"
   },
   "venue": {
    "peer_reviewed": false,
    "marker": "preprint"
   },
   "same_source_claims": [
    "HS-09"
   ],
   "in_corpus_verification": null,
   "external_verification": {
    "verdict": "confirmed",
    "checked": [
     "relevance most reliable, avg conformal set size ~3.0 of 5",
     "coherence moderate ~3.9",
     "fluency and consistency unreliable ~4.9",
     "aggregate transitivity violations 0.8-4.1%",
     "33-67% of documents contain at least one directed 3-cycle (intransitive cycle)",
     "conformal set width correlates with reliability r_s = +0.576, N = 1,918",
     "cross-judge width agreement 0.32-0.38",
     "four judges, four criteria on SummEval; criterion > judge"
    ],
    "not_visible": [
     "verbatim word 'intransitive' (paper expresses it as 'directed 3-cycle')"
    ],
    "quote": "relevance judged most reliably (avg. set size approx 3.0) ... fluency and consistency remain unreliable (avg. set size approx 4.9) ... low aggregate violation rates (rho-bar = 0.8-4.1%) ... 33-67% of documents exhibiting at least one directed 3-cycle",
    "notes": "All headline numbers confirmed verbatim in the retrieved abstract. 'Intransitive preference cycle' corresponds to the paper's 'directed 3-cycle'.",
    "checked_at": "2026-07-15",
    "retrieval": "abstract"
   },
   "source_retrieval_meta": {
    "retrieval": "abstract",
    "fetched": [
     "https://arxiv.org/abs/2604.15302"
    ],
    "resolved_title": "Diagnosing LLM Judge Reliability: Conformal Prediction Sets and Transitivity Violations",
    "resolved_date": "2026-04-16",
    "title_match": true
   },
   "used_on": [],
   "decision_use": {
    "grade": "B",
    "basis": "current-gen method + measurement",
    "measured_on": "model-outputs",
    "note": "Per-criterion conformal reliability ceilings (2026). Method durable; measured ceilings tier-bound - recompute per project.",
    "models_measured": []
   }
  },
  {
   "id": "F26-06",
   "domain": "frontier-2026",
   "area": null,
   "claim": "Rubric-level judging is far from solved as of March 2026: on RubricEval, the first rubric-level (per-criterion) meta-evaluation benchmark for instruction following (3,486 quality-controlled instances), GPT-4o scores only 55.97% on the Hard subset; rubric-level evaluation outperforms checklist-level, explicit reasoning improves accuracy, and combining both reduces inter-judge variance.",
   "load_bearing": false,
   "evidence": "Abstract opened and verified on arXiv (2603.25133, Mar 26, 2026). Prior meta-evals scored judges at the response level only; RubricEval scores the fine-grained per-criterion judgments our AutoQA verdicts would be built from, and includes a rubric taxonomy of common judge failure modes. Part of a Q1 2026 wave of criterion-granularity meta-evals (IF-RewardBench arXiv 2603.04738, RubricBench arXiv 2603.01562, MCJudgeBench, AJ-Bench).",
   "implication": "For Q1/Q3/Q4: meta-evaluation infrastructure at exactly our granularity (criterion-level verdict vs human label) now exists and shows even strong judges hover near chance on hard criteria -- so per-criterion judge-vs-adjudicated agreement must be measured per project before go-live, and 'compile rubric then trust judge' is not defensible without it. Also 2026 evidence for the decomposition question: structured rubric-level framing beats flat checklists, and reasoning + rubric structure jointly cut inter-judge variance.",
   "source": {
    "raw": "RubricEval: A Rubric-Level Meta-Evaluation Benchmark for LLM Judges in Instruction Following (arXiv 2603.25133) | https://arxiv.org/abs/2603.25133 | 2026-03-26 | academic",
    "title": "RubricEval: A Rubric-Level Meta-Evaluation Benchmark for LLM Judges in Instruction Following (arXiv 2603.25133)",
    "url": "https://arxiv.org/abs/2603.25133",
    "date": "2026-03-26",
    "type": "academic"
   },
   "venue": {
    "peer_reviewed": false,
    "marker": "preprint"
   },
   "same_source_claims": [
    "AJ-04"
   ],
   "in_corpus_verification": null,
   "external_verification": {
    "verdict": "confirmed",
    "checked": [
     "RubricEval title exact match",
     "3,486 quality-controlled instances",
     "GPT-4o achieves only 55.97% on the Hard subset",
     "rubric-level evaluation outperforms checklist-level",
     "explicit reasoning improves accuracy",
     "combining both reduces inter-judge variance",
     "first rubric-level (per-criterion) meta-evaluation benchmark for instruction following; includes a rubric taxonomy of judge failure modes"
    ],
    "not_visible": [
     "sibling-paper corroboration (IF-RewardBench 2603.04738, RubricBench 2603.01562, MCJudgeBench, AJ-Bench) - external references, not on this page"
    ],
    "quote": "a substantial set of 3,486 quality-controlled instances ... GPT-4o achieves only 55.97% on Hard subset ... rubric-level evaluation outperforms checklist-level, explicit reasoning improves accuracy, and both together reduce inter-judge variance",
    "notes": "All headline numbers and qualitative findings confirmed from the abstract. Submission date 26 Mar 2026 matches source_date.",
    "checked_at": "2026-07-15",
    "retrieval": "abstract"
   },
   "source_retrieval_meta": {
    "retrieval": "abstract",
    "fetched": [
     "https://arxiv.org/abs/2603.25133"
    ],
    "resolved_title": "RubricEval: A Rubric-Level Meta-Evaluation Benchmark for LLM Judges in Instruction Following",
    "resolved_date": "2026-03-26",
    "title_match": true
   },
   "used_on": [
    {
     "page": "index.html",
     "anchor": "evidence",
     "label": "Overview - what the evidence supports"
    },
    {
     "page": "research.html",
     "anchor": "research-cards",
     "label": "Research - findings by decision weight"
    }
   ],
   "decision_use": {
    "grade": "B",
    "basis": "current-gen capability",
    "measured_on": "model-outputs",
    "note": "RubricEval 2026 restatement (same underlying source family as AJ-04).",
    "models_measured": [
     "GPT-4o"
    ]
   }
  },
  {
   "id": "F26-07",
   "domain": "frontier-2026",
   "area": null,
   "claim": "LLM-as-a-Verifier (July 6, 2026; Mirhoseini/Finn/Pavone/Stoica groups) reframes verification as a scaling axis: computing the expectation over scoring-token logits yields continuous scores that improve monotonically along three dimensions -- finer score granularity, repeated evaluation, and criteria decomposition -- reaching reported SOTA on verification benchmarks (Terminal-Bench V2 86.5%, SWE-Bench Verified 78.2%, RoboRewardBench 87.4%, MedAgentBench 73.3%) without any judge training.",
   "load_bearing": false,
   "evidence": "arXiv abstract fetched and verified (2607.05391, v2 July 7, 2026). Finer granularity separates good from bad solutions better; repeated evaluation and decomposition boost accuracy via variance and complexity reduction; includes a cost-efficient candidate-ranking algorithm and progress-estimation use.",
   "implication": "For Q3/Q5: prefer continuous logit-expectation scores over discrete verdict tokens as the judge's raw output (thresholds applied downstream per project), and treat k-sample repeated evaluation plus criteria decomposition as the default variance-reduction stack -- the run-to-run flip-rate question in Q5 partially dissolves if the primitive is a continuous expectation rather than a sampled token.",
   "source": {
    "raw": "LLM-as-a-Verifier: A General-Purpose Verification Framework (arXiv 2607.05391) | https://arxiv.org/abs/2607.05391 | 2026-07-06 | academic",
    "title": "LLM-as-a-Verifier: A General-Purpose Verification Framework (arXiv 2607.05391)",
    "url": "https://arxiv.org/abs/2607.05391",
    "date": "2026-07-06",
    "type": "academic"
   },
   "venue": {
    "peer_reviewed": false,
    "marker": "preprint"
   },
   "same_source_claims": [],
   "in_corpus_verification": null,
   "external_verification": {
    "verdict": "confirmed",
    "checked": [
     "reframes verification as a new scaling axis",
     "computes the expectation over the distribution of scoring-token logits to generate continuous scores",
     "three dimensions: (1) score granularity, (2) repeated evaluation, (3) criteria decomposition",
     "SOTA: Terminal-Bench V2 86.5%, SWE-Bench Verified 78.2%, RoboRewardBench 87.4%, MedAgentBench 73.3%",
     "no additional judge training required",
     "cost-efficient candidate-ranking (best-of-N) algorithm and task-progress estimation",
     "authors include Finn, Pavone, Stoica, Mirhoseini"
    ],
    "not_visible": [
     "the word 'monotonically': the abstract says repeated evaluation and decomposition 'consistently lead to additional gains' -- 'monotonically improving' is the claim's paraphrase, not the abstract's wording"
    ],
    "quote": "Terminal-Bench V2 (86.5%), SWE-Bench Verified (78.2%), RoboRewardBench (87.4%), and MedAgentBench (73.3%)",
    "notes": "All four SOTA numbers, the logit-expectation mechanism, the three scaling dimensions, the no-training property, and the named authors match. Only the 'monotonically' descriptor is a slight paraphrase of 'consistently lead to additional gains'.",
    "checked_at": "2026-07-15",
    "retrieval": "abstract"
   },
   "source_retrieval_meta": {
    "retrieval": "abstract",
    "fetched": [
     "https://arxiv.org/abs/2607.05391"
    ],
    "resolved_title": "LLM-as-a-Verifier: A General-Purpose Verification Framework",
    "resolved_date": "v1 2026-07-06; v2 2026-07-07",
    "title_match": true
   },
   "used_on": [],
   "decision_use": {
    "grade": "B",
    "basis": "current-gen result",
    "measured_on": "model-outputs",
    "note": "LLM-as-a-Verifier (Jul 2026): verification as a scaling axis; decomposition monotonically improves it. The most current capability signal in the corpus; still model-output verification.",
    "models_measured": []
   }
  },
  {
   "id": "F26-08",
   "domain": "frontier-2026",
   "area": null,
   "claim": "OpenAI's CoVal (Jan 14, 2026) operationalizes 'disagreement is signal, not noise' at lab scale: crowd-written, prompt-specific rubrics (~1,000 participants, ~1,000 prompts) make WHY evaluators disagree auditable; rubric-derived scores reach only 0.58-0.61 pairwise concordance with crowd preference, are reliable only when score gaps are large, and OpenAI explicitly warns that treating rubric scores as optimization targets incentivizes checklist-style gaming.",
   "load_bearing": false,
   "evidence": "Full blog post opened and read. CoVal-full preserves conflicting criteria as a record of legitimate disagreement; CoVal-core distills 4 mutually compatible criteria per prompt. OpenAI states lower alignment scores can indicate 'no single stable target to predict' rather than model failure, and that for split preferences any single aggregated score encodes a particular compromise. Dataset released in both forms.",
   "implication": "Directly supports the Q2 third-verdict-class design: the frontier lab position in 2026 is that contested items should surface WHICH criteria drive disagreement (conditional/pluralist reporting) rather than force a scalar verdict, and that 'instructions underdetermine this case' is a real, detectable state. Also a warning for Q8: rubric-visible scoring invites checklist gaming, so per-item rubric details should not be fully leaked to attempters.",
   "source": {
    "raw": "CoVal: Learning values-aware rubrics from the crowd (OpenAI Alignment blog) | https://alignment.openai.com/coval | 2026-01-14 | primary",
    "title": "CoVal: Learning values-aware rubrics from the crowd (OpenAI Alignment blog)",
    "url": "https://alignment.openai.com/coval",
    "date": "2026-01-14",
    "type": "primary"
   },
   "venue": {
    "peer_reviewed": null,
    "marker": "primary"
   },
   "same_source_claims": [],
   "in_corpus_verification": null,
   "external_verification": {
    "verdict": "partially_confirmed",
    "checked": [
     "'crowd-written, prompt-specific rubrics' and 'Our results are based on ~1,000 participants and ~1,000 prompts'",
     "reliable only at large gaps: 'small gaps correspond to near-ties, while large gaps more consistently pick out the human-preferred completion'; 'For small gaps, preference prediction is only slightly better than chance'",
     "gaming warning: 'Treating the score as an optimization target can also incentivize gaming (e.g., verbosity or checklist-style responses'",
     "CoVal-full preserves conflicting criteria; CoVal-core 'retains 4 highly rated, mutually compatible criteria per prompt'",
     "'no single stable target to predict'",
     "dataset released in two complementary forms (Hugging Face openai/coval)"
    ],
    "not_visible": [],
    "quote": "CoVal-full achieves 0.61 concordance with the crowd and CoVal-core achieves 0.58 ... [out-of-sample] CoVal-full achieves .75 and CoVal-core achieves .76",
    "notes": "Discrepancy on the headline number: the claim states rubric scores 'reach only 0.58-0.61 pairwise concordance.' Those values are the IN-SAMPLE appendix benchmark; the paper's OUT-OF-SAMPLE validation concordance is higher (CoVal-full .75 / CoVal-core .76). The 0.58-0.61 figures are genuinely in the paper and the 'reliable only at large gaps' framing is accurate, but citing the in-sample number as the concordance understates the reported result. All other specifics confirmed verbatim.",
    "checked_at": "2026-07-15",
    "retrieval": "fulltext"
   },
   "source_retrieval_meta": {
    "retrieval": "fulltext",
    "fetched": [
     "https://alignment.openai.com/coval"
    ],
    "resolved_title": "CoVal: Learning values-aware rubrics from the crowd",
    "resolved_date": "2026-01-14",
    "title_match": true
   },
   "used_on": [
    {
     "page": "architecture.html",
     "anchor": "ontology",
     "label": "System - claim types"
    },
    {
     "page": "decisions.html",
     "anchor": "a1",
     "label": "Decisions - the standard"
    }
   ],
   "decision_use": {
    "grade": "B",
    "basis": "current-gen practice",
    "measured_on": "human-work",
    "note": "CoVal (Jan 2026): rubric-vs-collective-eval concordance on crowd judgments. Re-check flagged in-sample vs out-of-sample number nuance - see ledger entry.",
    "models_measured": []
   }
  },
  {
   "id": "F26-09",
   "domain": "frontier-2026",
   "area": null,
   "claim": "A plug-in statistical framework now exists for reporting judge-based pass rates correctly: because an imperfect judge (sensitivity q1, specificity q0) biases the naive pass-rate estimate (upward at low true rates, downward at high), the Rogan-Gladen-style corrected estimator plus a confidence interval propagating BOTH test-set and calibration-set uncertainty should replace raw judge-agreement numbers, with an adaptive algorithm allocating calibration labels between true-pass and true-fail classes (roughly m0 ~ (1/p-1)*sqrt(kappa)*m1) -- e.g., ~200 labels per class for CI width <0.1 in the paper's worked regime.",
   "load_bearing": false,
   "evidence": "Full paper text crawled and read, including derivations, Monte Carlo validation (naive CI coverage near zero; corrected CI holds 95% across all true accuracies), and released Python implementation (UW-Madison-Lee-Lab/LLM-judge-reporting). Posted Nov 26, 2025 -- just before the charter window -- but it is the current statistical-reporting layer being circulated in 2026 (alphaXiv feature May 2026) and no 2026 successor supersedes it.",
   "implication": "Answers the Q1 metrics-standardization question concretely: ban raw judge pass rates; report bias-corrected estimates with dual-source CIs as the standard artifact; and size/refresh the per-project gold set using the adaptive allocation rule -- which also says gold sets must deliberately oversample the rare class (true fails) rather than sample uniformly, directly addressing the low-failure-base-rate problem.",
   "source": {
    "raw": "How to Correctly Report LLM-as-a-Judge Evaluations (arXiv 2511.21140) | https://arxiv.org/abs/2511.21140 | 2025-11-26 | academic",
    "title": "How to Correctly Report LLM-as-a-Judge Evaluations (arXiv 2511.21140)",
    "url": "https://arxiv.org/abs/2511.21140",
    "date": "2025-11-26",
    "type": "academic"
   },
   "venue": {
    "peer_reviewed": false,
    "marker": "preprint"
   },
   "same_source_claims": [
    "HS-08"
   ],
   "in_corpus_verification": null,
   "external_verification": {
    "verdict": "partially_confirmed",
    "checked": [
     "imperfect judge (sensitivity q1, specificity q0) biases naive pass rate",
     "direction: positive bias at low true rate, negative bias at high (overestimates below theta=0.75, underestimates above)",
     "Rogan-Gladen (1978) corrected estimator: theta_hat = (p_hat + q0_hat - 1)/(q0_hat + q1_hat - 1)",
     "CI propagating both test-set and calibration-set uncertainty",
     "adaptive allocation formula m0 ~ (1/p - 1)*sqrt(kappa)*m1 with kappa=(1-q0)/(1-q1)",
     "Monte Carlo: naive CI near-zero coverage; corrected CI holds ~95% across true accuracies",
     "released Python implementation github.com/UW-Madison-Lee-Lab/LLM-judge-reporting"
    ],
    "not_visible": [
     "alphaXiv feature May 2026 and 'no 2026 successor supersedes it' (external, unverifiable from this source)"
    ],
    "quote": "achieving an interval shorter than 0.1 requires m approximately 200 calibration examples ... m_0 approximately (1/p - 1)*sqrt(kappa)*m_1 ... p_hat has positive bias at low values of theta and negative bias at high values of theta",
    "notes": "DISCREPANCY on one number: the claim says '~200 labels PER CLASS for CI width <0.1', but the paper's worked regime (p_hat=0.3, q0_hat=0.7, q1_hat=0.9) requires m approximately 200 calibration examples in TOTAL under symmetric allocation (~100 per class), not 200 per class. Every other specific (bias direction, q1/q0, Rogan-Gladen estimator, the m0 allocation formula, Monte Carlo coverage, GitHub repo) is confirmed verbatim from the full text.",
    "checked_at": "2026-07-15",
    "retrieval": "fulltext"
   },
   "source_retrieval_meta": {
    "retrieval": "fulltext",
    "fetched": [
     "https://arxiv.org/abs/2511.21140",
     "https://arxiv.org/html/2511.21140"
    ],
    "resolved_title": "How to Correctly Report LLM-as-a-Judge Evaluations",
    "resolved_date": "2025-11-26 (v4 2026-05-31; ICML 2026)",
    "title_match": true
   },
   "used_on": [
    {
     "page": "decisions.html",
     "anchor": "d6",
     "label": "Decisions - the throughput dial"
    }
   ],
   "decision_use": {
    "grade": "A",
    "basis": "statistical method",
    "measured_on": "structural",
    "note": "Corrected pass-rate certification under imperfect judges. Mathematics (re-check corrected a per-class-vs-total sizing detail - see ledger entry).",
    "models_measured": []
   }
  },
  {
   "id": "F26-10",
   "domain": "frontier-2026",
   "area": null,
   "claim": "FindTheFlaws (AAAI 2026, published March 2026) provides five datasets of long-form expert-annotated flawed vs correct solutions across medicine, math, and science, built specifically to test whether models can detect flawed reasoning and to support scalable-oversight experiments where the evaluator is weaker than the solution author.",
   "load_bearing": false,
   "evidence": "AAAI-40 proceedings record confirmed (DOI 10.1609/aaai.v40i44.41123, Recchia et al.). This is the 2026-venue instantiation of the seeded-error-set methodology: known-flaw items with expert annotations of where and why the reasoning fails, enabling measurement of flaw-detection recall rather than agreement-with-consensus.",
   "implication": "For Q1/Q5: the false-agreement-rate question (judge and attempter both wrong) is only measurable with seeded known-flaw items, and FindTheFlaws is both a ready-made calibration resource and the template for building per-project seeded flaw sets -- measure the judge's flaw-detection recall on planted errors, not just its agreement with noisy human reviewers.",
   "source": {
    "raw": "FindTheFlaws: Annotated Errors for Detecting Flawed Reasoning and Scalable Oversight Research (AAAI 2026) | https://doi.org/10.1609/aaai.v40i44.41123 | 2026-03-14 | academic",
    "title": "FindTheFlaws: Annotated Errors for Detecting Flawed Reasoning and Scalable Oversight Research (AAAI 2026)",
    "url": "https://doi.org/10.1609/aaai.v40i44.41123",
    "date": "2026-03-14",
    "type": "academic"
   },
   "venue": {
    "peer_reviewed": true,
    "marker": "AAAI"
   },
   "same_source_claims": [],
   "in_corpus_verification": null,
   "external_verification": {
    "verdict": "confirmed",
    "checked": [
     "'we present FindTheFlaws, a group of five diverse datasets spanning medicine, mathematics, science, coding, and the Lojban language'",
     "'(1) long-form expert-verified correct solutions and (2) long-form flawed solutions with annotations highlighting specific errors'",
     "scalable oversight with weaker evaluator: 'models performing more poorly on particular datasets can serve as judges/verifiers for more capable models'",
     "author Gabriel Recchia (Recchia, G., Mangat, C.S., Li, I., Krishnakumar, G.)",
     "published 2026-03-14, AAAI 40(44)"
    ],
    "not_visible": [],
    "quote": "we present FindTheFlaws, a group of five diverse datasets spanning medicine, mathematics, science, coding, and the Lojban language",
    "notes": "DOI resolves (302) to the AAAI OJS landing page (allowed redirect variant a). Claim says 'five datasets... across medicine, math, and science' - accurate, though the full set also includes coding and Lojban. The weaker-evaluator scalable-oversight design and expert-annotated flawed/correct solutions are confirmed.",
    "checked_at": "2026-07-15",
    "retrieval": "landing"
   },
   "source_retrieval_meta": {
    "retrieval": "landing",
    "fetched": [
     "https://doi.org/10.1609/aaai.v40i44.41123",
     "https://ojs.aaai.org/index.php/AAAI/article/view/41123"
    ],
    "resolved_title": "FindTheFlaws: Annotated Errors for Detecting Flawed Reasoning and Scalable Oversight Research",
    "resolved_date": "2026-03-14; Proceedings of the AAAI Conference on Artificial Intelligence, 40(44)",
    "title_match": true
   },
   "used_on": [],
   "decision_use": {
    "grade": "A",
    "basis": "dataset availability",
    "measured_on": "structural",
    "note": "FindTheFlaws: expert-annotated flawed-solution datasets exist for seeded-recall measurement (E2's raw material).",
    "models_measured": []
   }
  },
  {
   "id": "F26-11",
   "domain": "frontier-2026",
   "area": null,
   "claim": "The 2026 production norm for trusting an LLM judge is a closed calibration loop on sampled human corrections: humans review a slice of judge verdicts, corrections become ground truth for measuring alignment and are recycled as few-shot exemplars into the judge prompt (LangChain, March 10, 2026); practitioner guides converge on LLM-judge-first with a ~5-10% sampled human verification rate and periodic kappa re-checks as the default large-volume labeling pattern.",
   "load_bearing": false,
   "evidence": "LangChain resource (dated March 10, 2026) describes the correction-collection -> alignment-measurement -> few-shot-rebuild loop as learned from production evaluator deployments; FutureAGI's 2026 guide independently states the '5-10% human verification' default and monthly kappa re-sampling for judge drift. Practitioner-grade evidence, not controlled studies.",
   "implication": "For Q6: industry has already voted against per-item human touch and for sampled-correction calibration -- the human interaction that compounds is the one that writes back into judge config (corrections as exemplars), supporting the 'upstream compounding work' branch of the Q6 fork. Our design should treat the single human touch as config-improving adjudication, not per-item verification.",
   "source": {
    "raw": "How to Calibrate LLM-as-Judge with Human Corrections (LangChain) | https://www.langchain.com/resources/llm-as-a-judge | 2026-03-10 | practitioner",
    "title": "How to Calibrate LLM-as-Judge with Human Corrections (LangChain)",
    "url": "https://www.langchain.com/resources/llm-as-a-judge",
    "date": "2026-03-10",
    "type": "practitioner"
   },
   "venue": {
    "peer_reviewed": null,
    "marker": "practitioner"
   },
   "same_source_claims": [],
   "in_corpus_verification": null,
   "external_verification": {
    "verdict": "partially_confirmed",
    "checked": [
     "closed calibration loop: humans review a sample of judge verdicts and correct disagreements",
     "corrections become ground truth for measuring/improving evaluator alignment",
     "corrections recycled as few-shot examples into the judge prompt",
     "LLM-judge-first pattern (judges produce initial scores; humans correct disagreements)",
     "dated March 10, 2026"
    ],
    "not_visible": [
     "'5-10%' sampled human verification rate (claim attributes this to FutureAGI's 2026 guide, not LangChain)",
     "periodic/monthly kappa re-checks (claim attributes this to FutureAGI; LangChain page mentions no kappa)"
    ],
    "quote": "\"human reviewers examine a sample of those judgments and correct any they disagree with\" ... \"These corrections become the ground truth for measuring and improving evaluator alignment.\" ... \"From those corrections, you build few-shot examples that calibrate the judge.\"",
    "notes": "The LangChain-attributable core (correction-collection -> alignment-measurement -> few-shot-rebuild loop, dated 2026-03-10) confirmed verbatim. The '5-10% human verification' default and 'monthly kappa re-sampling' are attributed by the claim to FutureAGI's guide (a different source) and are not present on the LangChain page.",
    "checked_at": "2026-07-15",
    "retrieval": "landing"
   },
   "source_retrieval_meta": {
    "retrieval": "landing",
    "fetched": [
     "https://www.langchain.com/resources/llm-as-a-judge"
    ],
    "resolved_title": "How to Calibrate LLM-as-Judge with Human Corrections",
    "resolved_date": "2026-03-10",
    "title_match": true
   },
   "used_on": [],
   "decision_use": {
    "grade": "A",
    "basis": "production practice",
    "measured_on": "structural",
    "note": "2026 norm: closed calibration loop on sampled human corrections. Practice pattern.",
    "models_measured": []
   }
  },
  {
   "id": "F26-12",
   "domain": "frontier-2026",
   "area": null,
   "claim": "Autorubric (arXiv 2603.00077, March 2026) is an open-source unified rubric-evaluation framework that packages the 2026 judge-reliability defaults -- per-criterion atomic evaluation, binary/ordinal/nominal criteria with weights, multi-judge ensembles with configurable vote aggregation, few-shot calibration with verdict-balanced sampling, and built-in mitigations for position bias (option shuffling), verbosity bias (length penalties), and criterion conflation.",
   "load_bearing": false,
   "evidence": "arXiv HTML abstract read; feature list confirmed, validated on three benchmarks including a contributed 100-sample ground-truth dataset (CHARM-100). Notably, Handshake's BVB evaluation used Autorubric as the representative text-only rubric-judge family -- and found it architecture-limited (cannot inspect artifacts), placing a measured ceiling on this whole verifier class for evidence-bearing tasks.",
   "implication": "For Q5: the perturbation-robustness/debiasing harness the taxonomy asks about is now a packaged, open-source commodity -- position shuffling, length penalties, per-criterion atomic checks, and verdict-balanced few-shot calibration should be table-stakes defaults in our judge config schema, not research work; but per the Gandalf finding, this text-only class is insufficient alone whenever the attempter's claims reference artifacts the judge cannot open.",
   "source": {
    "raw": "Autorubric: A Unified Framework for Rubric-Based LLM Evaluation (arXiv 2603.00077) | https://arxiv.org/html/2603.00077v1 | 2026-03-02 | academic",
    "title": "Autorubric: A Unified Framework for Rubric-Based LLM Evaluation (arXiv 2603.00077)",
    "url": "https://arxiv.org/html/2603.00077v1",
    "date": "2026-03-02",
    "type": "academic"
   },
   "venue": {
    "peer_reviewed": true,
    "marker": "CHI"
   },
   "same_source_claims": [
    "RR-11"
   ],
   "in_corpus_verification": null,
   "external_verification": {
    "verdict": "confirmed",
    "checked": [
     "Autorubric, an open-source Python library",
     "per-criterion atomic evaluation with natural language explanations",
     "binary, ordinal, and nominal criteria with configurable weights",
     "multi-judge ensemble evaluation with majority, weighted, unanimous, and any-vote aggregation",
     "few-shot calibration with verdict-balanced sampling",
     "mitigations for position bias (option shuffling)",
     "verbosity bias (length penalties)",
     "reduced criterion conflation (independent scoring prevents halo effects)",
     "three benchmarks spanning educational assessment, deep research evaluation, and chatbot quality",
     "CHARM-100, a 100-sample chatbot evaluation dataset with per-sample ground truth labels"
    ],
    "not_visible": [
     "Handshake's BVB evaluation using Autorubric as the representative text-only rubric-judge family -- 'Handshake' and 'BVB' do not appear on this page (external editorial claim)",
     "submission date -- no date present in the fetched HTML"
    ],
    "quote": "Autorubric supports binary, ordinal, and nominal criteria with configurable weights ... single-judge and multi-judge ensemble evaluation with majority, weighted, unanimous, and any-vote aggregation ... CHARM-100, a 100-sample chatbot evaluation dataset",
    "notes": "Every feature the claim attributes to Autorubric, plus the three-benchmark validation and the contributed CHARM-100 dataset, confirmed verbatim. The Handshake/BVB 'measured ceiling' connection is an external editorial addition not present in this paper.",
    "checked_at": "2026-07-15",
    "retrieval": "fulltext"
   },
   "source_retrieval_meta": {
    "retrieval": "fulltext",
    "fetched": [
     "https://arxiv.org/html/2603.00077v1"
    ],
    "resolved_title": "Autorubric: A Unified Framework for Rubric-Based LLM Evaluation",
    "resolved_date": "not shown in fetched HTML",
    "title_match": true
   },
   "used_on": [],
   "decision_use": {
    "grade": "A",
    "basis": "tooling fact",
    "measured_on": "structural",
    "note": "Autorubric framework packaging 2026 judge-reliability techniques.",
    "models_measured": []
   }
  },
  {
   "id": "CG-01",
   "domain": "critic-and-gapfill",
   "area": "Psychometrics and rater-training science (educational/language assessment, many-facet Rasch measurement, generalizability theory). The sweep's 9 domains are all 2020s LLM/annotation literature, but human-rater variance has 40+ years of quantitative treatment: rater severity/leniency modeling, rater drift over scoring sessions, frame-of-reference vs rater-error training effect sizes, G-theory decomposition of variance into rater/item/criterion facets. This directly answers taxonomy Q1's variance-attribution question ('would perfect rubrics leave variance intact?') and Q9's reviewer-skill-decay question with measured effect sizes, instead of leaving them as unanchored pilots.",
   "claim": "Adjusting for rater severity/centrality with a Many-Facet Rasch Model (MFRM) materially changed which AI systems ranked best in the OpenAI RLHF summarization dataset: trained raters' agreement was only QWK .31-.50, and MFRM-adjusted scores flipped raw-mean rankings so two human-feedback policies rose above human-written reference summaries.",
   "load_bearing": false,
   "evidence": "Casabianca & Beiting-Parrish (LAK 2026 LLM Psychometrics workshop) fit MFRM to 6,312 ratings from 15 trained raters on 639 summaries across 19 policies. Full text opened and confirmed: raters R02/R08/R06 systematically lenient, R04/R10/R12 severe, R10 aberrantly central; raw-score rankings 'misrepresent the relative standing of models.' Paper also specifies the estimability requirement: 3-5 ratings per output with rater-item overlap/linkage, and proposes percentile-based flagging (top/bottom 2.5%) for rater remediation.",
   "implication": "AutoQA should treat each human reviewer's severity, centrality, and criterion-specific thresholds as modeled parameters, not noise: use adjusted (fair) scores for pass/fail decisions and use per-rater parameter estimates as the constructive-feedback channel. This forces a data-collection constraint into the foundation: some deliberate rater-item overlap must be designed in, because a single-review-per-item regime with disjoint rater pools is statistically unidentifiable.",
   "source": {
    "raw": "Correcting Human Labels for Rater Effects in AI Evaluation: An Item Response Theory Approach (arXiv 2602.22585) | https://arxiv.org/html/2602.22585v1 | 2026-02 | academic",
    "title": "Correcting Human Labels for Rater Effects in AI Evaluation: An Item Response Theory Approach (arXiv 2602.22585)",
    "url": "https://arxiv.org/html/2602.22585v1",
    "date": "2026-02",
    "type": "academic"
   },
   "venue": {
    "peer_reviewed": true,
    "marker": "LAK"
   },
   "same_source_claims": [],
   "in_corpus_verification": null,
   "external_verification": {
    "verdict": "confirmed",
    "checked": [
     "Many-Facet Rasch Model (MFRM) fit to OpenAI RLHF summarization dataset",
     "6,312 ratings, 639 unique articles/summaries, 19 policies, 15 raters",
     "trained raters' agreement QWK ranged from .31 to .50",
     "MFRM adjustment flips raw-mean rankings: two human-feedback policies (sup4_6b_ppo_rm4_6b, sup4_ppo_rm4) outperform human-generated reference summaries",
     "R02/R08/R06 lenient (negative estimates); R04/R10/R12 severe; R10 aberrantly central",
     "3-5 ratings per output with rater-item overlap/linkage for estimability",
     "flag raters below 2.5th / above 97.5th percentile"
    ],
    "not_visible": [],
    "quote": "The dataset contained 6,312 ratings based on 639 unique articles, summarized by 19 policies which were evaluated by 15 raters ... QWK ranged from .31 to .50 ... two human feedback models, sup4_6b_ppo_rm4_6b and sup4_ppo_rm4, outperform all other models and the human-generated reference summaries.",
    "notes": "Full HTML retrieved; all numbers, rater profiles, ranking flip, estimability requirement, and percentile flagging confirmed verbatim.",
    "checked_at": "2026-07-15",
    "retrieval": "fulltext"
   },
   "source_retrieval_meta": {
    "retrieval": "fulltext",
    "fetched": [
     "https://arxiv.org/html/2602.22585v1"
    ],
    "resolved_title": "Correcting Human Labels for Rater Effects in AI Evaluation: An Item Response Theory Approach",
    "resolved_date": "2026-02 (arXiv 2602; LAK 2026 LLM Psychometrics workshop, April 2026)",
    "title_match": true
   },
   "used_on": [],
   "decision_use": {
    "grade": "A",
    "basis": "psychometric method",
    "measured_on": "human-work",
    "note": "MFRM adjustment changes rankings built on human ratings; estimability needs 3-5 ratings/output linkage. Method + human-rater behavior; applies unchanged.",
    "models_measured": []
   }
  },
  {
   "id": "CG-02",
   "domain": "critic-and-gapfill",
   "area": "Psychometrics and rater-training science (educational/language assessment, many-facet Rasch measurement, generalizability theory). The sweep's 9 domains are all 2020s LLM/annotation literature, but human-rater variance has 40+ years of quantitative treatment: rater severity/leniency modeling, rater drift over scoring sessions, frame-of-reference vs rater-error training effect sizes, G-theory decomposition of variance into rater/item/criterion facets. This directly answers taxonomy Q1's variance-attribution question ('would perfect rubrics leave variance intact?') and Q9's reviewer-skill-decay question with measured effect sizes, instead of leaving them as unanchored pilots.",
   "claim": "Rater training on real scoring behavior improves within-rater consistency but does NOT equalize severity: after formal training of 16 raters in an operational ESL writing exam, significant between-rater severity differences remained even though consistency improved for most raters.",
   "load_bearing": false,
   "evidence": "Weigle 1998 (Language Testing, 363 citations; abstract opened and confirmed verbatim): 'rater training is more successful in helping raters give more predictable scores (i.e., intra-rater reliability) than in getting them to give identical scores (i.e., inter-rater reliability).' This is the canonical, repeatedly replicated finding of the language-testing rater literature (Lumley & McNamara 1995 same conclusion via MFRM).",
   "implication": "Directly answers taxonomy Q1: perfect rubrics plus training will still leave stable between-reviewer severity variance intact. Budget rubric/training work for what it actually buys (self-consistency, fewer misfitting raters) and route severity alignment to statistical adjustment in the AutoQA layer instead of more calibration meetings.",
   "source": {
    "raw": "Using FACETS to model rater training effects (Weigle, Language Testing 15(2)) | https://doi.org/10.1177/026553229801500205 | 1998-04 | academic",
    "title": "Using FACETS to model rater training effects (Weigle, Language Testing 15(2))",
    "url": "https://doi.org/10.1177/026553229801500205",
    "date": "1998-04",
    "type": "academic"
   },
   "venue": {
    "peer_reviewed": true,
    "marker": "Language Testing"
   },
   "same_source_claims": [],
   "in_corpus_verification": null,
   "external_verification": {
    "verdict": "not_retrievable",
    "checked": [],
    "not_visible": [
     "the asserted verbatim quote about intra- vs inter-rater reliability",
     "the 16-raters / operational ESL writing exam detail",
     "the finding that between-rater severity differences remained after training"
    ],
    "quote": "",
    "notes": "DOI redirect resolved (server 302: doi.org -> journals.sagepub.com/doi/10.1177/026553229801500205), confirming the DOI is valid and maps to the expected SAGE article path. However the SAGE landing page is behind a Cloudflare managed challenge: WebFetch returned only the SAGE homepage shell (no article text) and curl with browser headers returned HTTP 403 with a 'Just a moment...' JS challenge. No abstract or article content could be retrieved, so none of the claim's specifics can be verified from retrieved text. Per protocol, both WebFetch and curl were attempted.",
    "checked_at": "2026-07-15",
    "retrieval": "failed"
   },
   "source_retrieval_meta": {
    "retrieval": "failed",
    "fetched": [
     "https://doi.org/10.1177/026553229801500205",
     "https://journals.sagepub.com/doi/10.1177/026553229801500205"
    ],
    "resolved_title": "unavailable (Cloudflare challenge; no article content served)",
    "resolved_date": "unavailable",
    "title_match": false
   },
   "used_on": [
    {
     "page": "index.html",
     "anchor": "problem",
     "label": "Overview - the problem"
    },
    {
     "page": "index.html",
     "anchor": "decision",
     "label": "Overview - the five authorizations"
    },
    {
     "page": "research.html",
     "anchor": "research-cards",
     "label": "Research - findings by decision weight"
    },
    {
     "page": "decisions.html",
     "anchor": "a3",
     "label": "Decisions - baseline first"
    }
   ],
   "decision_use": {
    "grade": "A",
    "basis": "human science",
    "measured_on": "human-work",
    "note": "Weigle 1998: training buys self-consistency, not severity equality. Canonical, replicated; humans unchanged.",
    "models_measured": []
   }
  },
  {
   "id": "CG-03",
   "domain": "critic-and-gapfill",
   "area": "Psychometrics and rater-training science (educational/language assessment, many-facet Rasch measurement, generalizability theory). The sweep's 9 domains are all 2020s LLM/annotation literature, but human-rater variance has 40+ years of quantitative treatment: rater severity/leniency modeling, rater drift over scoring sessions, frame-of-reference vs rater-error training effect sizes, G-theory decomposition of variance into rater/item/criterion facets. This directly answers taxonomy Q1's variance-attribution question ('would perfect rubrics leave variance intact?') and Q9's reviewer-skill-decay question with measured effect sizes, instead of leaving them as unanchored pilots.",
   "claim": "Frame-of-reference (FOR) training, the best-validated rater-training method, has a meta-analytic effect on rating accuracy of about Cohen's d = 0.50 (moderate) - down from the d = 0.83 estimated by the earlier, smaller Woehr & Huffcutt 1994 meta-analysis - i.e., training helps but removes nowhere near all rater error.",
   "load_bearing": false,
   "evidence": "Roch, Woehr, Mishra & Kieszczynska (J. Occup. Organ. Psychol., 2012; abstract opened) is the updated meta-analysis with 4x the studies of Woehr & Huffcutt 1994; it finds FOR training effective, with Borman's differential accuracy and behavioural accuracy most improved. The d=0.50 overall figure confirmed via two independent citing sources (Tsai et al. 2019 SMU full-text PDF: 'medium-to-large effect on improving rating accuracy (d = 0.50, Roch et al., 2012; d = 0.83, Woehr & Huffcutt 1994)'; Loignon et al. 2017). Note: classic 'rater error training' (teaching raters to avoid halo/leniency) reduces those errors but can reduce accuracy - FOR (anchoring raters to a shared performance theory with practice+feedback) is the variant that works.",
   "implication": "Sets the ceiling on root-cause-1 fixes: even the best-in-class training intervention delivers ~half a standard deviation of accuracy gain. If the AutoQA program funds training, fund FOR-style training (shared exemplars, practice ratings, feedback against reference judgments) - which the AutoQA itself can generate as a byproduct - and plan for the residual variance to be caught by per-item verification.",
   "source": {
    "raw": "Rater training revisited: An updated meta-analytic review of frame-of-reference training (Roch et al.) | https://doi.org/10.1111/j.2044-8325.2011.02045.x | 2012-06 | academic",
    "title": "Rater training revisited: An updated meta-analytic review of frame-of-reference training (Roch et al.)",
    "url": "https://doi.org/10.1111/j.2044-8325.2011.02045.x",
    "date": "2012-06",
    "type": "academic"
   },
   "venue": {
    "peer_reviewed": true,
    "marker": "J. Occup"
   },
   "same_source_claims": [],
   "in_corpus_verification": null,
   "external_verification": {
    "verdict": "not_retrievable",
    "checked": [],
    "not_visible": [
     "overall FOR-training effect size Cohen's d approximately 0.50 on rating accuracy",
     "comparison to Woehr & Huffcutt 1994 d approximately 0.83",
     "~4x the studies of the earlier meta-analysis",
     "differential accuracy and behavioural accuracy most improved",
     "resolved title / authors / journal / year"
    ],
    "quote": "",
    "notes": "The DOI issues a server 302 to bpspsychub.onlinelibrary.wiley.com (correct Wiley/BPS host for J. Occup. Organ. Psychol.), confirming the DOI resolves to the intended article. But WebFetch returned HTTP 402 Payment Required and curl hit a Cloudflare 'Just a moment... Enable JavaScript' challenge, so no title, abstract, or numbers could be retrieved. None of the claim's figures (d=0.50, d=0.83, 4x studies) are verifiable from allowed sources.",
    "checked_at": "2026-07-15",
    "retrieval": "failed"
   },
   "source_retrieval_meta": {
    "retrieval": "failed",
    "fetched": [
     "https://doi.org/10.1111/j.2044-8325.2011.02045.x",
     "https://bpspsychub.onlinelibrary.wiley.com/doi/10.1111/j.2044-8325.2011.02045.x"
    ],
    "resolved_title": "not retrievable (DOI resolves to Wiley/BPS article page; content blocked)",
    "resolved_date": "not retrievable",
    "title_match": false
   },
   "used_on": [],
   "decision_use": {
    "grade": "A",
    "basis": "human science",
    "measured_on": "human-work",
    "note": "FOR training d~0.50 ceiling on the best rater-training intervention. Durable.",
    "models_measured": []
   }
  },
  {
   "id": "CG-04",
   "domain": "critic-and-gapfill",
   "area": "Psychometrics and rater-training science (educational/language assessment, many-facet Rasch measurement, generalizability theory). The sweep's 9 domains are all 2020s LLM/annotation literature, but human-rater variance has 40+ years of quantitative treatment: rater severity/leniency modeling, rater drift over scoring sessions, frame-of-reference vs rater-error training effect sizes, G-theory decomposition of variance into rater/item/criterion facets. This directly answers taxonomy Q1's variance-attribution question ('would perfect rubrics leave variance intact?') and Q9's reviewer-skill-decay question with measured effect sizes, instead of leaving them as unanchored pilots.",
   "claim": "G-theory decompositions of rated writing assessments show the rater MAIN effect (global severity) is a minor variance source while high-order interactions dominate: in a 120-student EFL study, student-by-task-by-method (31.8%), student-by-rater-by-task-by-method (26.5%), and student-by-rater-by-method (17.6%) interactions were the largest components, and 99.2% of error variance came from the student-by-rater interaction.",
   "load_bearing": false,
   "evidence": "Khodi 2021 (Language Testing in Asia, open access; PDF opened and percentages confirmed from abstract/results). Consistent with the broader G-theory writing literature (e.g., Gao, Brennan & Guo's GMAT AWA report, 2015): rater main-effect variance is small once training exists; person-by-rater and residual interaction variance is what limits reliability. Interpretation: most rater-related variance is idiosyncratic per-item disagreement (which rater on which submission on which criterion), not a stable harshness offset.",
   "implication": "This is the strongest psychometric argument for the AutoQA's core design: global severity calibration and better rubrics attack only the small main-effect components. The dominant variance lives at the rater-times-item-times-criterion level, which can only be attacked by evidence-grounded evaluation of each specific claim on each specific item - exactly the 'is this statement grounded in the cited evidence' positive-verification layer. It also means severity-adjustment (finding 1) is necessary but not sufficient.",
   "source": {
    "raw": "The affectability of writing assessment scores: a G-theory analysis of rater, task, and scoring method contribution (Khodi) | https://doi.org/10.1186/s40468-021-00134-5 | 2021-11 | academic",
    "title": "The affectability of writing assessment scores: a G-theory analysis of rater, task, and scoring method contribution (Khodi)",
    "url": "https://doi.org/10.1186/s40468-021-00134-5",
    "date": "2021-11",
    "type": "academic"
   },
   "venue": {
    "peer_reviewed": true,
    "marker": "Language Testing"
   },
   "same_source_claims": [],
   "in_corpus_verification": null,
   "external_verification": {
    "verdict": "confirmed",
    "checked": [
     "student-by-task-by-method (STM:B) = 31.8% of total variance",
     "student-by-rater-by-task-by-method (SRTM:B) = 26.5% of total variance",
     "student-by-rater-by-method (SRM:B) = 17.6% of total variance",
     "these three interactions are the major/largest variance sources",
     "student-by-rater (SR:B) = 99.2% of error variance (rater-by-background = 0.8%)",
     "120 students ('One hundred and twenty students, 90 females and 30 males')",
     "EFL learners; G-theory / generalizability theory"
    ],
    "not_visible": [
     "the literal labeled term 'rater main effect' (rater-related variance appears as SR/RB interactions; RB ~8% of error variance)"
    ],
    "quote": "the student by task by method of scoring (nested in background of education) interaction (STM:B) with 31.8% contribution to the total variance ... (SR:B) and rater by background of education with 99.2% and 0.8% contribution to the error variance",
    "notes": "All four headline percentages confirmed with their exact interaction mappings, plus the 120-student sample. The claim's interpretation that the rater MAIN effect is minor is supported: dominant components are all high-order student x rater x task x method interactions, and 99.2% of error variance is the student-by-rater interaction; severity/bias shows up as the SR/RB interaction (~8% of error). Reached via server-redirect chain (doi -> springeropen BMC host) fetched by curl; WebFetch on the springer.com leg bounced to an idp auth wall. Published 2021-10-01 (source_date 2021-11).",
    "checked_at": "2026-07-15",
    "retrieval": "fulltext"
   },
   "source_retrieval_meta": {
    "retrieval": "fulltext",
    "fetched": [
     "https://doi.org/10.1186/s40468-021-00134-5",
     "https://languagetestingasia.springeropen.com/articles/10.1186/s40468-021-00134-5"
    ],
    "resolved_title": "The affectability of writing assessment scores: a G-theory analysis of rater, task, and scoring method contribution",
    "resolved_date": "2021-10-01",
    "title_match": true
   },
   "used_on": [
    {
     "page": "index.html",
     "anchor": "problem",
     "label": "Overview - the problem"
    },
    {
     "page": "research.html",
     "anchor": "research-cards",
     "label": "Research - findings by decision weight"
    },
    {
     "page": "decisions.html",
     "anchor": "a3",
     "label": "Decisions - baseline first"
    }
   ],
   "decision_use": {
    "grade": "A",
    "basis": "human science",
    "measured_on": "human-work",
    "note": "G-theory: interactions dominate rater main effects. Durable; the strongest independent argument for per-claim verification.",
    "models_measured": []
   }
  },
  {
   "id": "CG-05",
   "domain": "critic-and-gapfill",
   "area": "Psychometrics and rater-training science (educational/language assessment, many-facet Rasch measurement, generalizability theory). The sweep's 9 domains are all 2020s LLM/annotation literature, but human-rater variance has 40+ years of quantitative treatment: rater severity/leniency modeling, rater drift over scoring sessions, frame-of-reference vs rater-error training effect sizes, G-theory decomposition of variance into rater/item/criterion facets. This directly answers taxonomy Q1's variance-attribution question ('would perfect rubrics leave variance intact?') and Q9's reviewer-skill-decay question with measured effect sizes, instead of leaving them as unanchored pilots.",
   "claim": "Rater severity is nonstationary within a single scoring session: in an OSCE, later time-slots received systematically higher ratings (regression coefficient 0.88, 95% CI 0.38-1.38, p=.001), with the drift 2.4x larger on difficult stations (1.24 vs 0.52), and excluding warm-up stations did not remove it; the psychometric field has named machinery (DRIFT - differential rater functioning over time, Wolfe et al. 2001; Myford & Wolfe monitoring frameworks) for detecting it.",
   "load_bearing": false,
   "evidence": "PubMed abstract of McLaughlin et al., Medical Education 2009, opened and confirmed (coefficients above). Wolfe & Moulder 2001 (J. Applied Measurement) established Rasch-based procedures for detecting drift types (primacy/recency, centrality drift, practice/fatigue). Drift within and across scoring days is a repeated finding in operational scoring programs (Lunz & Stahl 1990 across grading periods).",
   "implication": "Answers taxonomy Q9 with measured evidence: reviewer calibration decays on the timescale of a single session, not just months, and decays faster on hard items. AutoQA should compute reviewer-effect estimates in rolling windows (per session/day) and alert on drift, rather than certifying a reviewer once at onboarding; item-order randomization and hard-item interleaving are cheap design mitigations.",
   "source": {
    "raw": "The effect of differential rater function over time (DRIFT) on objective structured clinical examination ratings (Medical Education 43(10)) | https://pubmed.ncbi.nlm.nih.gov/19769648 | 2009-10 | academic",
    "title": "The effect of differential rater function over time (DRIFT) on objective structured clinical examination ratings (Medical Education 43(10))",
    "url": "https://pubmed.ncbi.nlm.nih.gov/19769648",
    "date": "2009-10",
    "type": "academic"
   },
   "venue": {
    "peer_reviewed": null,
    "marker": "academic"
   },
   "same_source_claims": [],
   "in_corpus_verification": null,
   "external_verification": {
    "verdict": "confirmed",
    "checked": [
     "DRIFT in OSCE ratings; later time-slots rated higher",
     "regression coefficient 0.88, 95% CI 0.38-1.38, P=0.001",
     "difficult stations coefficient 1.24 vs 0.52 for less difficult (ratio ~2.4x)",
     "removing the first two (warm-up) stations did not correct DRIFT",
     "Med Educ 2009 Oct;43(10):989-92",
     "authors McLaughlin, Ainslie, Coderre, Wright, Violato"
    ],
    "not_visible": [
     "secondary citations in the claim's evidence (Wolfe et al. 2001; Myford & Wolfe; Lunz & Stahl 1990) - external to this source, not on this page"
    ],
    "quote": "Removing the first two stations from our analyses did not correct DRIFT.",
    "notes": "All primary-source specifics confirmed. Minor caveat: the 0.52 difficult-vs-easy comparison coefficient (less difficult stations) was not statistically significant (P=0.09); the '2.4x' is the claimant's derived ratio of the two reported coefficients (1.24/0.52). Paper title uses 'differential rater function over time'.",
    "checked_at": "2026-07-15",
    "retrieval": "abstract"
   },
   "source_retrieval_meta": {
    "retrieval": "abstract",
    "fetched": [
     "https://pubmed.ncbi.nlm.nih.gov/19769648"
    ],
    "resolved_title": "The effect of differential rater function over time (DRIFT) on objective structured clinical examination ratings",
    "resolved_date": "Med Educ. 2009 Oct;43(10):989-92",
    "title_match": true
   },
   "used_on": [],
   "decision_use": {
    "grade": "A",
    "basis": "human science",
    "measured_on": "human-work",
    "note": "Within-session rater severity drift. Durable; rolling-window reviewer effects stay.",
    "models_measured": []
   }
  },
  {
   "id": "CG-06",
   "domain": "critic-and-gapfill",
   "area": "Psychometrics and rater-training science (educational/language assessment, many-facet Rasch measurement, generalizability theory). The sweep's 9 domains are all 2020s LLM/annotation literature, but human-rater variance has 40+ years of quantitative treatment: rater severity/leniency modeling, rater drift over scoring sessions, frame-of-reference vs rater-error training effect sizes, G-theory decomposition of variance into rater/item/criterion facets. This directly answers taxonomy Q1's variance-attribution question ('would perfect rubrics leave variance intact?') and Q9's reviewer-skill-decay question with measured effect sizes, instead of leaving them as unanchored pilots.",
   "claim": "The MFRM toolkit is now being applied symmetrically to LLM judges: a 2025 study fit MFRM to 10 LLMs plus human expert raters scoring the same writing tasks and found LLMs exhibit measurable severity/centrality rater effects that differ by model, with GPT-4o, Gemini 1.5 Pro, and Claude 3.5 Sonnet showing the highest accuracy and fewest rater effects.",
   "load_bearing": false,
   "evidence": "Jiao, Song & Lee, arXiv 2505.18486 (May 2025), abstract opened and confirmed: QWK vs human scores, Cronbach alpha across prompts, and MFRM rater-effect estimates computed identically for human and LLM raters. Parallel 2025 work (Wang et al., Computers and Education: AI, Sept 2025) builds a full psychometric reliability/validity framework for LLM raters in large-scale writing assessment. This is the literature bridge: the same model audits both layers.",
   "implication": "The AutoQA's own LLM-judge component should be enrolled as just another 'rater' facet in the same measurement model as the humans it reviews. This gives a project-agnostic, self-auditing structure: one MFRM per project instruction-set yields severity/centrality/fit for every judge, human or machine, and flags when the LLM layer itself drifts after a prompt or model change. - DISAGREEMENT: Size of the training effect: Woehr & Huffcutt 1994 estimated FOR training accuracy gain at d=0.83; Roch et al. 2012, with 4x the studies, revised it down to d~0.50. Separately, Noh & Matore 2022 (Frontiers in Psychology, 164 teacher-raters, MFRM) found prior training experience made NO difference to rating quality on a speaking assessment while rating and teaching experience did - generic/one-shot training may buy nothing; the meta-analytic d~0.50 applies to structured FOR protocols specifically. - DISAGREEMENT: How pathological rater pools are: Casabianca & Beiting-Parrish 2026 cite evidence (Nieto & Casabianca 2019) that large professional testing-organization rater pools show minimal rater effects, versus their own demonstration that in a small trained pool (OpenAI's 15 raters) a few aberrant raters materially distorted system rankings. The severity of the problem depends on pool size, professionalization, and tenure - crowdsourced/DaaS annotation pools sit at the bad end of this spectrum. - DISAGREEMENT: Where the fixable variance lives: MFRM studies foreground the rater severity main effect (a stable, correctable offset), while G-theory decompositions (Khodi 2021; GMAT AWA tradition) show the main effect is small and idiosyncratic rater-by-item interaction dominates. These are complementary lenses, but they imply different remedies (statistical adjustment vs per-item verification), and the literature does not agree on the split's typical proportions across contexts. - GAP: No meta-analysis isolates how much variance rubric operationalization ALONE removes (rubric vs no-rubric variance-component comparison). Individual studies exist (rating-scale vs rubric vs holistic; criteria-order effects on halo, Kim 2020) but no pooled effect size - so 'root cause 1 vs 2-3' cannot be settled from literature alone, though the Weigle/Khodi pattern strongly suggests substantial rater-intrinsic residual. - GAP: Long-horizon reviewer skill decay (months-scale, the Q9 timescale) is not measured in this literature: DRIFT studies cover within-session and across-day/grading-period drift. No published effect sizes for calibration decay over weeks-to-months in annotation-industry settings were found through 2026-07. - GAP: Minimal-linkage requirements under annotation-economics constraints are underspecified: MFRM needs rater-item overlap, but no published guidance found on the cheapest overlap design (e.g., what fraction of items double-reviewed, seeded common items) sufficient for stable severity estimates in pools with high rater churn - this likely needs a small in-house simulation, which is cheap since the models are standard. - GAP: Semantic Scholar returned HTTP 429 during the sweep; coverage relied on OpenAlex/Crossref/Exa/Tavily. A dedicated pass over Language Testing and Assessing Writing 2025-2026 issues could surface additional recent MFRM-for-annotation work.",
   "source": {
    "raw": "Comparing Human and AI Rater Effects Using the Many-Facet Rasch Model (Jiao, Song, Lee) | https://arxiv.org/abs/2505.18486 | 2025-05-24 | academic",
    "title": "Comparing Human and AI Rater Effects Using the Many-Facet Rasch Model (Jiao, Song, Lee)",
    "url": "https://arxiv.org/abs/2505.18486",
    "date": "2025-05-24",
    "type": "academic"
   },
   "venue": {
    "peer_reviewed": false,
    "marker": "preprint"
   },
   "same_source_claims": [],
   "in_corpus_verification": null,
   "external_verification": {
    "verdict": "partially_confirmed",
    "checked": [
     "fit Many-Facet Rasch Model to ten LLMs plus human expert raters on the same writing tasks",
     "two types of writing tasks (holistic and analytic scores)",
     "GPT-4o (ChatGPT 4o), Gemini 1.5 Pro, and Claude 3.5 Sonnet recommended with high accuracy and fewest rater effects",
     "QWK vs human scores and Cronbach alpha across prompts used",
     "authors Jiao, Song & Lee; submitted May 24 2025"
    ],
    "not_visible": [
     "'severity/centrality' as the specific rater-effect types (abstract says 'rater effects' generally; severity/centrality not named)",
     "'rater effects that differ by model' (differ-by-model direction not stated verbatim in abstract)",
     "Parallel Wang et al. (Computers and Education: AI, Sept 2025) framework - a separate source not fetched"
    ],
    "quote": "high scoring accuracy, better rater reliability, and less rater effects",
    "notes": "Core (10 LLMs + human raters via MFRM; GPT-4o/Gemini 1.5 Pro/Claude 3.5 Sonnet as top performers) confirmed. The specific severity/centrality effect vocabulary is not visible in the retrieved abstract.",
    "checked_at": "2026-07-15",
    "retrieval": "abstract"
   },
   "source_retrieval_meta": {
    "retrieval": "abstract",
    "fetched": [
     "https://arxiv.org/abs/2505.18486"
    ],
    "resolved_title": "Comparing Human and AI Rater Effects Using the Many-Facet Rasch Model",
    "resolved_date": "2025-05-24 (v1); v2 2025-05-28",
    "title_match": true
   },
   "used_on": [
    {
     "page": "decisions.html",
     "anchor": "c4",
     "label": "Decisions - default: statistics staging"
    }
   ],
   "decision_use": {
    "grade": "B",
    "basis": "current-gen method application",
    "measured_on": "model-outputs",
    "note": "MFRM applied to 2025 LLM judges. The symmetric-measurement idea is method-level; the specific judge rankings are tier-bound.",
    "models_measured": [
     "GPT-4o",
     "Gemini 1.5",
     "Claude 3.5 Sonnet showing the highest accuracy and fewest rater effects. Jiao"
    ]
   }
  },
  {
   "id": "CG-07",
   "domain": "critic-and-gapfill",
   "area": "Legal/regulatory constraints on algorithmic management of the attempter workforce, plus data-quality standards. The AutoQA issues automated pass/fail verdicts affecting paid workers' income and proposes process telemetry (timing, edit traces) as a provenance contract - this is squarely regulated territory: GDPR Article 22 (automated decisions with significant effects require human review - which may legally mandate the one-human-touch the design treats as optional), the EU Platform Work Directive (transposition deadline December 2026, explicitly governs algorithmic management, worker access to decision logic, and human oversight of automated account/pay decisions), and ISO/IEC 5259 (data quality for analytics and ML, published 2024-2025) which large buyers may contractually require. Zero regulatory/standards sources appear in any domain.",
   "claim": "Platform Work Directive Article 10(5) requires that any decision to restrict, suspend, or terminate the contractual relationship or account of a person performing platform work - or any other decision of equivalent detriment - be taken by a human being, with no consent or contractual-necessity exception (stricter than GDPR Art 22(2)).",
   "load_bearing": false,
   "evidence": "Verified against the directive text and multiple law-firm/academic analyses (Wolters Kluwer, Taylor Wessing, ETUI). Article 10 covers decisions 'taken or supported' by automated systems, deliberately closing the GDPR gap for hybrid human+algorithm pipelines. Unlike GDPR Art 22, the PWD permits no consent-based or contract-based exception for these decisions.",
   "implication": "For EU-based attempters, 'zero-touch plus randomized audits' is illegal for verdicts that cost workers access to work, pay, or their account once national transposition applies (from 2 Dec 2026). The one-human-touch is not a design option for adverse consequences - it is a legal floor. AutoQA can auto-pass but must route auto-fail-with-consequences through a human decision-maker. This is stricter than GDPR: consent cannot buy the exception back.",
   "source": {
    "raw": "Directive (EU) 2024/2831, Official Journal 11 Nov 2024 (EUR-Lex), corroborated by Wolters Kluwer Global Workplace Law & Policy analysis | https://eur-lex.europa.eu/legal-content/EN/TXT/?uri=CELEX:32024L2831 | 2024-11-11 | primary",
    "title": "Directive (EU) 2024/2831, Official Journal 11 Nov 2024 (EUR-Lex), corroborated by Wolters Kluwer Global Workplace Law & Policy analysis",
    "url": "https://eur-lex.europa.eu/legal-content/EN/TXT/?uri=CELEX:32024L2831",
    "date": "2024-11-11",
    "type": "primary"
   },
   "venue": {
    "peer_reviewed": null,
    "marker": "regulation/standard"
   },
   "same_source_claims": [
    "CG-08",
    "CG-10"
   ],
   "in_corpus_verification": null,
   "external_verification": {
    "verdict": "confirmed",
    "checked": [
     "Article 10(5) requires that decisions to restrict/suspend/terminate the contractual relationship or account, or any decision of equivalent detriment, be taken by a human being",
     "no exception listed in the article text"
    ],
    "not_visible": [
     "the comparison 'stricter than GDPR Art 22(2) / no consent or contractual-necessity exception' (interpretive gloss from law-firm analyses, not stated in the directive text itself)"
    ],
    "quote": "Any decision to restrict, suspend or terminate the contractual relationship or the account of a person performing platform work or any other decision of equivalent detriment shall be taken by a human being.",
    "notes": "Article 10(5) confirmed verbatim. The GDPR Art 22(2) contrast is an accurate legal characterization but is not in the directive text (attributed in evidence to Wolters Kluwer / Taylor Wessing / ETUI, not fetchable here).",
    "checked_at": "2026-07-15",
    "retrieval": "fulltext"
   },
   "source_retrieval_meta": {
    "retrieval": "fulltext",
    "fetched": [
     "https://eur-lex.europa.eu/legal-content/EN/TXT/?uri=CELEX:32024L2831"
    ],
    "resolved_title": "Directive (EU) 2024/2831 of the European Parliament and of the Council of 23 October 2024 on improving working conditions in platform work",
    "resolved_date": "2024-10-23 (adopted); OJ 2024-11-11",
    "title_match": true
   },
   "used_on": [
    {
     "page": "index.html",
     "anchor": "proposal",
     "label": "Overview - the proposal"
    },
    {
     "page": "index.html",
     "anchor": "decision",
     "label": "Overview - the five authorizations"
    },
    {
     "page": "research.html",
     "anchor": "research-cards",
     "label": "Research - findings by decision weight"
    },
    {
     "page": "architecture.html",
     "anchor": "ontology",
     "label": "System - claim types"
    },
    {
     "page": "pilot.html",
     "anchor": "main",
     "label": "Pilot - gates and objectives"
    },
    {
     "page": "decisions.html",
     "anchor": "a4",
     "label": "Decisions - authority boundaries"
    },
    {
     "page": "decisions.html",
     "anchor": "d1",
     "label": "Decisions - enforcement weight"
    },
    {
     "page": "decisions.html",
     "anchor": "d5",
     "label": "Decisions - where it runs and which law binds"
    },
    {
     "page": "decisions.html",
     "anchor": "c5",
     "label": "Decisions - default: the pay-denial question"
    }
   ],
   "decision_use": {
    "grade": "A",
    "basis": "law",
    "measured_on": "structural",
    "note": "Platform Work Directive Art 10(5). Binding for EU workers from transposition (2026-12-02); capability-independent by definition.",
    "models_measured": []
   }
  },
  {
   "id": "CG-08",
   "domain": "critic-and-gapfill",
   "area": "Legal/regulatory constraints on algorithmic management of the attempter workforce, plus data-quality standards. The AutoQA issues automated pass/fail verdicts affecting paid workers' income and proposes process telemetry (timing, edit traces) as a provenance contract - this is squarely regulated territory: GDPR Article 22 (automated decisions with significant effects require human review - which may legally mandate the one-human-touch the design treats as optional), the EU Platform Work Directive (transposition deadline December 2026, explicitly governs algorithmic management, worker access to decision logic, and human oversight of automated account/pay decisions), and ISO/IEC 5259 (data quality for analytics and ML, published 2024-2025) which large buyers may contractually require. Zero regulatory/standards sources appear in any domain.",
   "claim": "Platform Work Directive Article 11 grants persons performing platform work the right to (a) a plain-language explanation of ANY decision taken or supported by an automated decision-making system, (b) a designated competent human contact person to discuss the decision, and (c) human review with a substantiated written reply within two weeks of request; decisions that infringed rights must be rectified within two weeks or compensated.",
   "load_bearing": false,
   "evidence": "Verified verbatim from directive text via extraction of the EUR-Lex page: explanation 'in a transparent and intelligible manner, using clear and plain language'; contact persons must have 'the competence, training and authority necessary'; substantiated reply 'in any event within two weeks of receipt of the request'; rectification/compensation duty plus obligation to modify or discontinue the offending automated system. Carve-out: for workers who are 'business users' under the P2B Regulation (EU) 2019/1150, that regulation's human-review provisions prevail - but P2B itself mandates statements of reasons and an internal complaint-handling system for restriction/suspension/termination.",
   "implication": "Contest/appeal channels are legally mandatory, not optional, with a hard two-week SLA. The AutoQA's constructive-feedback output doubles as the legally required explanation - so verdict rationales must be written in plain language grounded in the project criteria (not model logits or opaque scores), and the system must log enough evidence per verdict to support a substantiated human reply on appeal. The 'modify or discontinue the system' duty means systematic AutoQA errors surfaced via appeals trigger a legal remediation obligation.",
   "source": {
    "raw": "Directive (EU) 2024/2831, Article 11 (EUR-Lex full text) | https://eur-lex.europa.eu/legal-content/EN/TXT/?uri=CELEX:32024L2831 | 2024-11-11 | primary",
    "title": "Directive (EU) 2024/2831, Article 11 (EUR-Lex full text)",
    "url": "https://eur-lex.europa.eu/legal-content/EN/TXT/?uri=CELEX:32024L2831",
    "date": "2024-11-11",
    "type": "primary"
   },
   "venue": {
    "peer_reviewed": null,
    "marker": "regulation/standard"
   },
   "same_source_claims": [
    "CG-07",
    "CG-10"
   ],
   "in_corpus_verification": null,
   "external_verification": {
    "verdict": "confirmed",
    "checked": [
     "Article 11 right to explanation of any decision taken or supported by an ADM system, in transparent/intelligible manner using clear and plain language",
     "designated competent human contact person with 'the competence, training and authority necessary'",
     "human review with a substantiated written reply within two weeks of receipt of the request",
     "decisions that infringe rights rectified without delay and within two weeks of adoption, or adequate compensation",
     "P2B carve-out: Article 11(5) does not apply to persons who are also business users under Regulation (EU) 2019/1150"
    ],
    "not_visible": [
     "the added detail that 'P2B itself mandates statements of reasons and an internal complaint-handling system' (that is content of Regulation 2019/1150, not this directive)"
    ],
    "quote": "The explanation shall be provided in a transparent and intelligible manner, using clear and plain language ... a sufficiently precise and adequately substantiated reply in the form of a written document ... within two weeks of receipt of the request",
    "notes": "All three limbs (a/b/c) plus the rectify-or-compensate-within-two-weeks duty and the business-user carve-out (Art 11(5), Reg 2019/1150) confirmed verbatim from the directive text.",
    "checked_at": "2026-07-15",
    "retrieval": "fulltext"
   },
   "source_retrieval_meta": {
    "retrieval": "fulltext",
    "fetched": [
     "https://eur-lex.europa.eu/legal-content/EN/TXT/?uri=CELEX:32024L2831"
    ],
    "resolved_title": "Directive (EU) 2024/2831 of the European Parliament and of the Council of 23 October 2024 on improving working conditions in platform work",
    "resolved_date": "2024-10-23 (adopted); OJ 2024-11-11",
    "title_match": true
   },
   "used_on": [
    {
     "page": "research.html",
     "anchor": "research-cards",
     "label": "Research - findings by decision weight"
    },
    {
     "page": "decisions.html",
     "anchor": "d5",
     "label": "Decisions - where it runs and which law binds"
    }
   ],
   "decision_use": {
    "grade": "A",
    "basis": "law",
    "measured_on": "structural",
    "note": "PWD Art 11 explanation/review rights.",
    "models_measured": []
   }
  },
  {
   "id": "CG-09",
   "domain": "critic-and-gapfill",
   "area": "Legal/regulatory constraints on algorithmic management of the attempter workforce, plus data-quality standards. The AutoQA issues automated pass/fail verdicts affecting paid workers' income and proposes process telemetry (timing, edit traces) as a provenance contract - this is squarely regulated territory: GDPR Article 22 (automated decisions with significant effects require human review - which may legally mandate the one-human-touch the design treats as optional), the EU Platform Work Directive (transposition deadline December 2026, explicitly governs algorithmic management, worker access to decision logic, and human oversight of automated account/pay decisions), and ISO/IEC 5259 (data quality for analytics and ML, published 2024-2025) which large buyers may contractually require. Zero regulatory/standards sources appear in any domain.",
   "claim": "The directive's scope explicitly covers online-performed work - the 'digital labour platform' definition applies 'irrespective of whether that work is performed online or in a certain location' - recital 19 names tagging/crowdwork, and Chapter III algorithmic-management rules (Arts 7-11) apply to genuinely self-employed persons performing platform work, from the start of recruitment, regardless of where the platform is established.",
   "load_bearing": false,
   "evidence": "Definition verified verbatim from EUR-Lex (Art 2: service provided at a distance by electronic means, at the recipient's request, involving organisation of work for payment, using automated monitoring/decision systems). Article 7(2) verified: 'shall apply to all persons performing platform work from the start of the recruitment or selection procedure.' Chapter III's application to the genuinely self-employed rests on the Art 16(2) TFEU data-protection legal basis (ETUI, Countouris & De Stefano 2025). Lund AI Policy Lab (Mar 2025) analysis specifically concludes AI labelers are covered. Extraterritoriality confirmed by Ius Laboris and Omnivoo (2026): applies to platforms established outside the EU when the work is performed in the EU.",
   "implication": "Classifying attempters as independent contractors does NOT exempt the AutoQA from the algorithmic-management chapter - the human-decision, explanation, appeal, and data-limitation rules attach to self-employed annotators too, and to a US-based platform with EU attempters. The foundational philosophy must treat these as baseline constraints for any EU-touching deployment, effective 2 Dec 2026 (about 5 months after this design ships).",
   "source": {
    "raw": "Directive (EU) 2024/2831 Art 2 & 7 (EUR-Lex) + AI Policy Lab, 'Potential impact of the EU Platform Work Directive on AI labelers' | https://aipolicylab.se/2025/03/25/potential-impact-of-the-eu-platform-work-directive-on-ai-labelers/ | 2025-03-25 | academic",
    "title": "Directive (EU) 2024/2831 Art 2 & 7 (EUR-Lex) + AI Policy Lab, 'Potential impact of the EU Platform Work Directive on AI labelers'",
    "url": "https://aipolicylab.se/2025/03/25/potential-impact-of-the-eu-platform-work-directive-on-ai-labelers/",
    "date": "2025-03-25",
    "type": "academic"
   },
   "venue": {
    "peer_reviewed": null,
    "marker": "academic"
   },
   "same_source_claims": [],
   "in_corpus_verification": null,
   "external_verification": {
    "verdict": "partially_confirmed",
    "checked": [
     "the post concludes AI labelers/annotators are covered by the directive when work is conducted through a digital platform in the EU",
     "recital/intro 19 mentions tagging as a form of crowd work that can be conducted remotely (page labels it 'Article 19 of Introduction')"
    ],
    "not_visible": [
     "verbatim Art 2 'digital labour platform' definition language ('irrespective of whether that work is performed online or in a certain location') -- not on this page (attributed to EUR-Lex, not fetched)",
     "Article 7(2) quote 'shall apply to all persons performing platform work from the start of the recruitment or selection procedure' -- does not appear on this page",
     "Chapter III / Articles 7-11 applying to genuinely self-employed persons from the start of recruitment -- not stated on this page (post discusses Arts 10 and 12 only)",
     "Article 16(2) TFEU data-protection legal basis -- not mentioned",
     "extraterritoriality (applies to platforms established outside the EU) -- attributed to Ius Laboris/Omnivoo, not on this page"
    ],
    "quote": "The EU Directive recognizes AI labeling as a form of platform work if it is conducted through a digital platform within the EU",
    "notes": "Only the AI Policy Lab page is fetchable per constraints. It supports the core conclusion (AI labelers are covered) and the tagging-as-crowdwork-in-recital-19 point. The many verbatim legal specifics in the compound claim (Art 2 definition, Art 7(2), Chapter III scope to self-employed, Art 16(2) TFEU, extraterritoriality) are attributed to EUR-Lex/ETUI/Ius Laboris/Omnivoo and are NOT present on this page. The page also conflates recitals with articles ('Article 19 of Introduction', 'Article 47/44'), weakening any verbatim-quote reliance. Partially_confirmed on the retrieved page alone.",
    "checked_at": "2026-07-15",
    "retrieval": "landing"
   },
   "source_retrieval_meta": {
    "retrieval": "landing",
    "fetched": [
     "https://aipolicylab.se/2025/03/25/potential-impact-of-the-eu-platform-work-directive-on-ai-labelers/"
    ],
    "resolved_title": "Potential impact of the EU Platform Work Directive on AI labelers",
    "resolved_date": "2025-03-25 (byline Mariia Lesina, Lund University)",
    "title_match": true
   },
   "used_on": [
    {
     "page": "decisions.html",
     "anchor": "d5",
     "label": "Decisions - where it runs and which law binds"
    }
   ],
   "decision_use": {
    "grade": "A",
    "basis": "law",
    "measured_on": "structural",
    "note": "PWD scope covers online-performed platform work.",
    "models_measured": []
   }
  },
  {
   "id": "CG-10",
   "domain": "critic-and-gapfill",
   "area": "Legal/regulatory constraints on algorithmic management of the attempter workforce, plus data-quality standards. The AutoQA issues automated pass/fail verdicts affecting paid workers' income and proposes process telemetry (timing, edit traces) as a provenance contract - this is squarely regulated territory: GDPR Article 22 (automated decisions with significant effects require human review - which may legally mandate the one-human-touch the design treats as optional), the EU Platform Work Directive (transposition deadline December 2026, explicitly governs algorithmic management, worker access to decision logic, and human oversight of automated account/pay decisions), and ISO/IEC 5259 (data quality for analytics and ML, published 2024-2025) which large buyers may contractually require. Zero regulatory/standards sources appear in any domain.",
   "claim": "Platform Work Directive Article 7 prohibits automated systems from processing personal data on workers' emotional or psychological state, private conversations (including worker-to-worker exchanges), any data collected while the person is not offering or performing platform work, data predicting exercise of fundamental rights (e.g., organizing), or inferring protected characteristics; Article 8 additionally mandates a data protection impact assessment.",
   "load_bearing": false,
   "evidence": "Verified verbatim from the directive text via EUR-Lex extraction (Art 7(1)(a)-(c) and recital 40). Art 7(3) extends these limits to ANY automated system 'taking or supporting decisions that affect persons performing platform work in any manner' - not just formally designated monitoring systems.",
   "implication": "The proposed process-telemetry provenance contract (timing, edit traces) is lawful only if scoped to on-task activity: no collection while the attempter is off-task, no sentiment/frustration inference from edit patterns or communications, no monitoring of attempter-to-attempter channels. Telemetry feeding the AutoQA requires a DPIA. Design the telemetry schema as an explicit allowlist (task-session-bounded behavioral signals) rather than a general activity log.",
   "source": {
    "raw": "Directive (EU) 2024/2831, Article 7 (EUR-Lex full text) | https://eur-lex.europa.eu/legal-content/EN/TXT/?uri=CELEX:32024L2831 | 2024-11-11 | primary",
    "title": "Directive (EU) 2024/2831, Article 7 (EUR-Lex full text)",
    "url": "https://eur-lex.europa.eu/legal-content/EN/TXT/?uri=CELEX:32024L2831",
    "date": "2024-11-11",
    "type": "primary"
   },
   "venue": {
    "peer_reviewed": null,
    "marker": "regulation/standard"
   },
   "same_source_claims": [
    "CG-07",
    "CG-08"
   ],
   "in_corpus_verification": null,
   "external_verification": {
    "verdict": "confirmed",
    "checked": [
     "Article 7(1)(a) prohibits processing personal data on emotional or psychological state",
     "7(1)(b) private conversations including exchanges with other persons performing platform work",
     "7(1)(c) collecting data while the person is not offering or performing platform work",
     "7(1)(d) predicting the exercise of fundamental rights including freedom of association",
     "7(1)(e) inferring protected characteristics (racial/ethnic origin, religious beliefs, disability, health, trade union membership, sex life/orientation)",
     "Article 7(3) extends limits to automated systems 'taking or supporting decisions that affect persons performing platform work in any manner'",
     "Article 8 mandates a data-protection impact assessment (Art 35(1) GDPR high-risk processing)"
    ],
    "not_visible": [],
    "quote": "this Article shall also apply where digital labour platforms use automated systems taking or supporting decisions that affect persons performing platform work in any manner ... Article 8 Data-protection impact assessment",
    "notes": "Article 7(1)(a)-(e), Article 7(3), and Article 8 all confirmed verbatim. Every element of the claim (including the 'any manner' extension and the DPIA mandate) is present.",
    "checked_at": "2026-07-15",
    "retrieval": "fulltext"
   },
   "source_retrieval_meta": {
    "retrieval": "fulltext",
    "fetched": [
     "https://eur-lex.europa.eu/legal-content/EN/TXT/?uri=CELEX:32024L2831"
    ],
    "resolved_title": "Directive (EU) 2024/2831 of the European Parliament and of the Council of 23 October 2024 on improving working conditions in platform work",
    "resolved_date": "2024-10-23 (adopted); OJ 2024-11-11",
    "title_match": true
   },
   "used_on": [
    {
     "page": "decisions.html",
     "anchor": "d3",
     "label": "Decisions - AI-assistance policy"
    }
   ],
   "decision_use": {
    "grade": "A",
    "basis": "law",
    "measured_on": "structural",
    "note": "PWD Art 7 telemetry prohibitions + DPIA.",
    "models_measured": []
   }
  },
  {
   "id": "CG-11",
   "domain": "critic-and-gapfill",
   "area": "Legal/regulatory constraints on algorithmic management of the attempter workforce, plus data-quality standards. The AutoQA issues automated pass/fail verdicts affecting paid workers' income and proposes process telemetry (timing, edit traces) as a provenance contract - this is squarely regulated territory: GDPR Article 22 (automated decisions with significant effects require human review - which may legally mandate the one-human-touch the design treats as optional), the EU Platform Work Directive (transposition deadline December 2026, explicitly governs algorithmic management, worker access to decision logic, and human oversight of automated account/pay decisions), and ISO/IEC 5259 (data quality for analytics and ML, published 2024-2025) which large buyers may contractually require. Zero regulatory/standards sources appear in any domain.",
   "claim": "Under GDPR Article 22 as interpreted in CJEU SCHUFA (C-634/21, 7 Dec 2023), an automated score that plays a 'determining role' in a downstream decision is itself an Article 22 decision, and human involvement only removes a decision from Art 22 scope if it is meaningful - a reviewer with real authority, data access, and competence to override; rubber-stamping does not count. Enforcement is active: Italy's DPA sanctioned automated rider deactivations, and Hamburg's DPA fined a company ~EUR 490,000 in September 2025 for automated rejections without adequate explanation of the logic.",
   "load_bearing": false,
   "evidence": "SCHUFA three-condition test and 'determining role' doctrine confirmed across TLT, Bird & Bird, IAPP, and Oxford Industrial Law Journal ('Scores as Decisions?', 2024, applying it to the labour context and Uber litigation). EDPB-endorsed WP29 guidance holds that consent is generally an invalid basis in work contexts due to power imbalance, and that human involvement must be substantive. Hamburg fine and Italian rider cases reported in 2025 enforcement roundups; regulator attention to automated rejections continuing into mid-2026.",
   "implication": "Two consequences beyond the PWD: (1) even where AutoQA runs as a vendor scoring layer feeding a client's pass/fail call, the SCHUFA logic makes the score itself an Art 22 decision if the client defers to it - outsourcing the final click doesn't escape regulation; (2) the one-human-touch must be designed as genuine review (authority to overturn, access to the evidence, calibrated competence), or it legally counts as zero-touch. Worker consent cannot serve as the lawful basis for telemetry or automated verdicts.",
   "source": {
    "raw": "CJEU Case C-634/21 SCHUFA analysis + 'Scores as Decisions? Article 22 GDPR ... in the Labour Context', Industrial Law Journal | https://academic.oup.com/ilj/article/53/4/840/7745471 | 2024-09 | academic",
    "title": "CJEU Case C-634/21 SCHUFA analysis + 'Scores as Decisions? Article 22 GDPR ... in the Labour Context', Industrial Law Journal",
    "url": "https://academic.oup.com/ilj/article/53/4/840/7745471",
    "date": "2024-09",
    "type": "academic"
   },
   "venue": {
    "peer_reviewed": null,
    "marker": "academic"
   },
   "same_source_claims": [],
   "in_corpus_verification": null,
   "external_verification": {
    "verdict": "partially_confirmed",
    "checked": [
     "CJEU SCHUFA C-634/21, 7 Dec 2023 cited (footnote 1)",
     "core doctrine present: a probability value is a decision where a third party 'draws strongly' on it to 'establish, implement or terminate a contractual relationship' (paras 48, 73)",
     "general requirement of 'meaningful human involvement' discussed (citing Article 29 Working Party)",
     "Art 22(1) 'lays down a prohibition in principle' (para 52); 'lacuna in legal protection' if scores were merely a 'preparatory act'"
    ],
    "not_visible": [
     "exact phrase 'determining role' -- the source uses 'draws strongly', not 'determining role' (wording discrepancy)",
     "the three-part meaningful-human-involvement test (reviewer with real authority, data access, competence to override) -- not stated in this article as a formulation",
     "'rubber-stamping' term -- not used; closest material is the Amazon/Hannover Administrative Court ruling",
     "Italy DPA sanction of automated rider deactivations -- not in this article",
     "Hamburg DPA ~EUR 490,000 fine, September 2025, for automated rejections without adequate explanation -- not in this 2024 article (and chronologically outside its scope)"
    ],
    "quote": "the CJEU held a probability value is a decision where a third party 'draws strongly' on it to 'establish, implement or terminate a contractual relationship'",
    "notes": "The underlying legal doctrine (a strongly-relied-upon automated score can itself be an Art 22 decision; human involvement must be meaningful) is supported. But the claim's signature phrase 'determining role' is not in this source (it says 'draws strongly'), the three-part reviewer test and 'rubber-stamping' framing are absent, and BOTH enforcement facts (Italian rider deactivations; Hamburg EUR 490k Sept 2025 fine) do not appear in this 2024 academic article -- the evidence field itself attributes them to separate 2025 enforcement roundups.",
    "checked_at": "2026-07-15",
    "retrieval": "fulltext"
   },
   "source_retrieval_meta": {
    "retrieval": "fulltext",
    "fetched": [
     "https://academic.oup.com/ilj/article/53/4/840/7745471"
    ],
    "resolved_title": "Scores as Decisions? Article 22 GDPR and the Judgment of the CJEU in SCHUFA Holding (Scoring) in the Labour Context (Asymina Aza)",
    "resolved_date": "Industrial Law Journal 53(4):840-858; published online 2024-08-29 (Dec 2024 issue)",
    "title_match": true
   },
   "used_on": [
    {
     "page": "index.html",
     "anchor": "decision",
     "label": "Overview - the five authorizations"
    },
    {
     "page": "architecture.html",
     "anchor": "ontology",
     "label": "System - claim types"
    },
    {
     "page": "pilot.html",
     "anchor": "main",
     "label": "Pilot - gates and objectives"
    },
    {
     "page": "decisions.html",
     "anchor": "a4",
     "label": "Decisions - authority boundaries"
    },
    {
     "page": "decisions.html",
     "anchor": "c5",
     "label": "Decisions - default: the pay-denial question"
    }
   ],
   "decision_use": {
    "grade": "A",
    "basis": "case law",
    "measured_on": "structural",
    "note": "SCHUFA doctrine; active enforcement. (Re-check: doctrine confirmed; the phrase 'determining role' traces to the CJEU judgment rather than the fetched commentary - see ledger entry.)",
    "models_measured": []
   }
  },
  {
   "id": "CG-12",
   "domain": "critic-and-gapfill",
   "area": "Legal/regulatory constraints on algorithmic management of the attempter workforce, plus data-quality standards. The AutoQA issues automated pass/fail verdicts affecting paid workers' income and proposes process telemetry (timing, edit traces) as a provenance contract - this is squarely regulated territory: GDPR Article 22 (automated decisions with significant effects require human review - which may legally mandate the one-human-touch the design treats as optional), the EU Platform Work Directive (transposition deadline December 2026, explicitly governs algorithmic management, worker access to decision logic, and human oversight of automated account/pay decisions), and ISO/IEC 5259 (data quality for analytics and ML, published 2024-2025) which large buyers may contractually require. Zero regulatory/standards sources appear in any domain.",
   "claim": "ISO/IEC 5259 'Data quality for analytics and ML' is now a complete five-part published series - Parts 1-4 published 2024 (overview/terminology; data quality measures; data quality management requirements & guidelines; process framework) and Part 5 (data quality governance framework) published 2025 - sold by ISO as an 'AI data quality management bundle'.",
   "load_bearing": false,
   "evidence": "Confirmed from ISO.org catalog pages: ISO/IEC 5259-1:2024 (81088), 5259-2:2024 (81860, 'specifies a data quality model, data quality measures and guidance on reporting data quality'), 5259-3:2024 (81092, management requirements), 5259-4:2024 (81093, process framework covering data labeling among ML data processes), 5259-5:2025 (84150, governance). Full text paywalled; no direct evidence found of large AI-data buyers contractually mandating 5259 yet, but 5259-3/-4 are certifiable-style requirements documents a buyer can reference in contract.",
   "implication": "The AutoQA philosophy document should map its quality dimensions and verdict evidence onto 5259-2's measurement/reporting vocabulary and 5259-4's process framework, so a per-project customization can be presented as 5259-conformant when a buyer asks - cheap to do now, expensive to retrofit. Treat 5259 as the neutral shared vocabulary for 'what a quality measure is' across projects. - DISAGREEMENT: Scope breadth of the algorithmic-management rules: CXC Global (Apr 2026) claims the directive's algorithmic-management rules 'apply to all workers, not only those engaged through digital labour platforms,' but the directive text limits Chapter III to 'persons performing platform work'; ETUC's transposition manual (Mar 2026) confirms extension to all workers is only an ADVOCACY position for national transposition, not the directive's requirement. Some member states may gold-plate; the EU floor covers platform work only. - DISAGREEMENT: Whether online annotation workers will effectively get the employment presumption: AI Policy Lab (Mar 2025) says AI labelers benefit when platforms control workflows/evaluation, but commentators (Verfassungsblog, ILO working papers) caution the control criteria are anchored in traditional notions that fit poorly with subtle microtask monitoring, and Fairwork warns weak transpositions (e.g., Italy) may leave annotators 'self-employed'. Note this dispute concerns the employment presumption only - Chapter III algorithmic-management rights attach regardless of status. - DISAGREEMENT: Strictness of national implementations: law-firm trackers (Ius Laboris Apr 2026, employsome May 2026) expect strict transpositions in France/Germany/Italy/Spain/Netherlands but narrow ones elsewhere (Hungary had taken no steps as of the CMS review; Ireland has no draft legislation), creating a 2026-2027 patchwork - sources disagree on whether to design to the strictest member state or per-country. - GAP: No enforcement action or national transposition text specifically addressing AI data-annotation QA verdicts was found - the directive's application to a QA layer rejecting individual work items (vs. account suspension) is untested; whether a single item pass/fail (payment denial for one task) is a 'decision of equivalent detriment' under Art 10(5) versus merely an Art 11 reviewable decision is unresolved in the sources. - GAP: Whether a BPO/vendor arrangement (annotators employed by an outsourcing firm, AutoQA run by the AI lab or its vendor) falls within the 'digital labour platform' definition is unexamined in available sources; recital 20 excludes platforms that merely connect providers to clients, and the intermediary/subcontractor liability question is flagged by ETUC as a transposition battleground. - GAP: ISO/IEC 5259 full text is paywalled - could not verify its specific annotation-quality measures (e.g., inter-annotator agreement treatment) or confirm any AI-data buyer contractually requiring it; the 'buyers may require it' premise remains plausible but unevidenced. - GAP: No dedicated EDPB guideline on automated decision-making in employment exists as of July 2026 - the operative guidance is still the 2018 WP29 guidelines (EDPB-endorsed); the EDPB Work Programme 2026-2027 (adopted 11 Feb 2026) was not verified in detail for a pending ADM-in-employment item. - GAP: Non-EU jurisdictions were out of scope of this pass: no findings on US state algorithmic-management bills, UK's DUAA ADM regime (expected effective 2026), or Brazil/India annotation-workforce rules - relevant if the attempter pool is global.",
   "source": {
    "raw": "ISO/IEC 5259 series catalog pages (ISO.org) / AI data quality management bundle | https://www.iso.org/publication/PUB200525.html | 2025 | primary",
    "title": "ISO/IEC 5259 series catalog pages (ISO.org) / AI data quality management bundle",
    "url": "https://www.iso.org/publication/PUB200525.html",
    "date": "2025",
    "type": "primary"
   },
   "venue": {
    "peer_reviewed": null,
    "marker": "regulation/standard"
   },
   "same_source_claims": [],
   "in_corpus_verification": null,
   "external_verification": {
    "verdict": "confirmed",
    "checked": [
     "'This bundle includes five essential components'",
     "ISO/IEC 5259-1:2024 - 'Part 1: Overview, terminology, and examples'",
     "5259-2:2024 - 'data quality model, measures, and guidance on reporting data quality'",
     "5259-3:2024 - 'management requirements and guidelines'",
     "5259-4:2024 - 'process framework'",
     "5259-5:2025 - 'governance framework to enable organizations to direct and oversee data quality'",
     "page title 'AI data quality management bundle'"
    ],
    "not_visible": [
     "internal ISO catalog IDs 81088/81860/81092/81093/84150 (from evidence, not shown on this page)",
     "Part 4 'covering data labeling' specifically (page shows 'process framework... guidance on data quality processes' but does not surface 'data labeling')"
    ],
    "quote": "This bundle includes five essential components: ISO/IEC 5259-1:2024 - Artificial intelligence - Data quality for analytics and machine learning (ML) - Part 1: Overview, terminology, and examples",
    "notes": "Retrieved via curl with browser user-agent (WebFetch returned 403). Five-part series with Parts 1-4 (2024) and Part 5 (2025), each part's subject, and the 'AI data quality management bundle' packaging all confirmed. Full standard text is paywalled; the claim's own contract-mandate observation is explicitly hedged.",
    "checked_at": "2026-07-15",
    "retrieval": "landing"
   },
   "source_retrieval_meta": {
    "retrieval": "landing",
    "fetched": [
     "https://www.iso.org/publication/PUB200525.html"
    ],
    "resolved_title": "ISO - AI data quality management bundle",
    "resolved_date": "2025 (bundle; component parts dated 2024 and 2025)",
    "title_match": true
   },
   "used_on": [
    {
     "page": "decisions.html",
     "anchor": "d5",
     "label": "Decisions - where it runs and which law binds"
    },
    {
     "page": "decisions.html",
     "anchor": "d6",
     "label": "Decisions - the throughput dial"
    }
   ],
   "decision_use": {
    "grade": "A",
    "basis": "standard",
    "measured_on": "structural",
    "note": "ISO/IEC 5259 published series. Standards fact.",
    "models_measured": []
   }
  },
  {
   "id": "CG-13",
   "domain": "critic-and-gapfill",
   "area": "Content-moderation QA programs - the largest deployed analog of humans-reviewing-human-judgment at scale. Meta/TikTok/Google have run decade-old programs auditing moderator decisions: golden-set seeding into live queues, audit sampling rates, overturn/appeal-rate monitoring, QA-of-the-QA layers, measured reviewer base-rate and fatigue effects, and published transparency-report methodology plus academic studies (e.g., on moderator agreement and audit design). The industry-practice domain covered only AI-training-data vendors; this adjacent industry has already answered several taxonomy Q6/Q9 questions operationally (seeded known-verdict items, easy-case dilution, de-graduation triggers).",
   "claim": "Facebook's production QA at Cognizant ran a three-layer blind audit: ~50-60 of each moderator's ~1,500 weekly decisions (~3-4%) randomly re-reviewed by a dedicated QA worker (paid $1/hr more), with full-time Facebook employees auditing a subset of QA decisions; 'accuracy' was computed purely as agreement with the auditor against a 95% target (other sites reported 98%), actual scores ran high-80s to 92, and misses triggered a remediation program that often ended in termination.",
   "load_bearing": false,
   "evidence": "Verified in The Verge's original text (fetched): 'From Miguel's 1,500 or so weekly decisions, Facebook will randomly select 50 or 60 to audit... Full-time Facebook employees then audit a subset of QA decisions.' Moderator quote: 'Accuracy is only judged by agreement. If me and the auditor both allow the obvious sale of heroin, Cognizant was correct... This number is fake.' Also confirmed: moderators lobbied QAs off-book to reverse decisions despite prohibition, and pre-firing 'coaching' often served as pretext for managing workers out.",
   "implication": "Directly answers taxonomy Q6/Q9 with a decade-scale postmortem: (a) ~3-4% random audit + a QA-of-QA layer is the deployed baseline; (b) agreement-with-auditor as the accuracy metric is blind to correlated errors and gameable via dispute lobbying; (c) punitive per-item scoring makes the QA relationship adversarial. AutoQA should score attempter statements against grounded evidence (not reviewer agreement), log auditor-attempter deliberations, and decouple constructive feedback from de-graduation triggers.",
   "source": {
    "raw": "The Trauma Floor: The secret lives of Facebook moderators in America (The Verge, Casey Newton) | https://www.theverge.com/2019/2/25/18229714/cognizant-facebook-content-moderator-interviews-trauma-working-conditions-arizona | 2019-02-25 | primary",
    "title": "The Trauma Floor: The secret lives of Facebook moderators in America (The Verge, Casey Newton)",
    "url": "https://www.theverge.com/2019/2/25/18229714/cognizant-facebook-content-moderator-interviews-trauma-working-conditions-arizona",
    "date": "2019-02-25",
    "type": "primary"
   },
   "venue": {
    "peer_reviewed": null,
    "marker": "primary"
   },
   "same_source_claims": [],
   "in_corpus_verification": null,
   "external_verification": {
    "verdict": "partially_confirmed",
    "checked": [
     "three-layer audit: moderator -> QA re-review -> full-time Facebook employees audit a subset of QA decisions",
     "~50-60 of ~1,500 weekly decisions randomly audited (~3-4%)",
     "QA worker makes $1 per hour more",
     "accuracy = agreement with auditor; target 95 percent",
     "actual scores in high 80s / low 90s, ~92 at press time",
     "misses lead to coaching/remedial program, often a pretext for managing workers out (termination)",
     "moderators prohibited from lobbying QAs to reverse decisions but do so regularly"
    ],
    "not_visible": [
     "'other sites reported 98%' (string '98' appears 0 times in the article)",
     "the descriptor 'blind' (word 'blind' appears 0 times; audit structure described but not called blind)"
    ],
    "quote": "\"From Miguel's 1,500 or so weekly decisions, Facebook will randomly select 50 or 60 to audit ... a quality assurance worker, known internally as a QA, who also makes $1 per hour more ... Full-time Facebook employees then audit a subset of QA decisions.\" ... target \"95 percent ... floats in the high 80s or low 90s ... around 92\".",
    "notes": "Every material fact confirmed verbatim, including the '~1,500 weekly / 50-60 audited', '$1 per hour more' QA, Facebook-audits-a-subset-of-QA third layer, 95% target, high-80s-to-92 actuals, the 'Accuracy is only judged by agreement / obvious sale of heroin / This number is fake' quotes, off-book QA lobbying, and coaching-as-pretext-for-termination. Marked partially_confirmed only because the peripheral '98% at other sites' comparison is absent (no '98' anywhere) and the audit is not literally called 'blind'.",
    "checked_at": "2026-07-15",
    "retrieval": "fulltext"
   },
   "source_retrieval_meta": {
    "retrieval": "fulltext",
    "fetched": [
     "https://www.theverge.com/2019/2/25/18229714/cognizant-facebook-content-moderator-interviews-trauma-working-conditions-arizona"
    ],
    "resolved_title": "The secret lives of Facebook moderators in America (on-page headline: 'The Trauma Floor')",
    "resolved_date": "2019-02-25 (matches URL path; byline Casey Newton)",
    "title_match": true
   },
   "used_on": [
    {
     "page": "decisions.html",
     "anchor": "a5",
     "label": "Decisions - permanent audit"
    }
   ],
   "decision_use": {
    "grade": "A",
    "basis": "production practice",
    "measured_on": "adjacent-domain",
    "note": "Facebook/Cognizant audit architecture. Decade-scale deployed practice for humans-reviewing-human-judgment.",
    "models_measured": []
   }
  },
  {
   "id": "CG-14",
   "domain": "critic-and-gapfill",
   "area": "Content-moderation QA programs - the largest deployed analog of humans-reviewing-human-judgment at scale. Meta/TikTok/Google have run decade-old programs auditing moderator decisions: golden-set seeding into live queues, audit sampling rates, overturn/appeal-rate monitoring, QA-of-the-QA layers, measured reviewer base-rate and fatigue effects, and published transparency-report methodology plus academic studies (e.g., on moderator agreement and audit design). The industry-practice domain covered only AI-training-data vendors; this adjacent industry has already answered several taxonomy Q6/Q9 questions operationally (seeded known-verdict items, easy-case dilution, de-graduation triggers).",
   "claim": "The Trust & Safety Professional Association (industry body staffed by Meta/Google/TikTok alumni) codifies exactly two complementary QA channels - forward audit sampling (weighted re-review samples by peers or a dedicated quality team) and 'reverse quality sampling' (pre-reviewed golden/known-verdict items injected through the regular review process) - with the binding constraint that seeded items 'must look exactly like regular reviews to be effective', plus appeals/overturns as a cheap but population-biased third signal, and a four-way error taxonomy (false positive, false negative, wrong-selection, technical error).",
   "load_bearing": false,
   "evidence": "Fetched both TSPA curriculum pages. QA page: golden seeding gives 'full control' to probe chosen edge cases but fails when user history is part of the review or when indistinguishability breaks; appeals surface false positives 'at a low cost' but 'the population of appeals is often different from the population of all decisions.' Metrics page: overturn rate 'can be a difficult metric to interpret' alone (low overturn = correct decisions OR users gave up); 'consistency' (multi-review mismatch rate) is tracked as a distinct metric from error rate; prevalence measurement via random sampling is called a 'premium metric' due to cost.",
   "implication": "Imports the settled industry answer to blind gold injection: it is standard practice, its failure mode is detectability (seeded items must be indistinguishable - hard when context/history is part of the task), and it is complementary to (not a substitute for) forward sampling. The wrong-selection error class (right verdict, wrong cited criterion) maps directly to the AutoQA requirement to check that an attempter's statement is aligned with the specific criterion they invoke, not just that the verdict is right.",
   "source": {
    "raw": "TSPA: Content Moderation Quality Assurance + Metrics for Content Moderation | https://www.tspa.org/curriculum/ts-fundamentals/content-moderation-and-operations/content-moderation-quality-assurance/ | 2022-09-16 | practitioner",
    "title": "TSPA: Content Moderation Quality Assurance + Metrics for Content Moderation",
    "url": "https://www.tspa.org/curriculum/ts-fundamentals/content-moderation-and-operations/content-moderation-quality-assurance/",
    "date": "2022-09-16",
    "type": "practitioner"
   },
   "venue": {
    "peer_reviewed": null,
    "marker": "practitioner"
   },
   "same_source_claims": [],
   "in_corpus_verification": null,
   "external_verification": {
    "verdict": "partially_confirmed",
    "checked": [
     "reviews by a dedicated quality team (and peer reviews); Samples are often weighted",
     "Reverse quality sampling involves taking a set of pre-reviewed examples ... sending them through the regular review process",
     "Examples must look exactly like regular reviews to be effective",
     "appeals surface false positives in particular at a low cost",
     "the population of appeals is often different from the population of all decisions",
     "four-way error taxonomy: False positives, False negatives, Wrong selection, Technical errors"
    ],
    "not_visible": [
     "overturn rate being 'a difficult metric to interpret' -- absent from this QA page (on the separate 'Metrics for Content Moderation' page, which the fetch constraint prohibits)",
     "'consistency' (multi-review mismatch rate) as a distinct tracked metric -- absent from this QA page",
     "prevalence measurement called a 'premium metric' -- absent from this QA page"
    ],
    "quote": "Reverse quality sampling involves taking a set of pre-reviewed examples ... Examples must look exactly like regular reviews to be effective",
    "notes": "The two QA channels, the seeded-item realism requirement, the appeals cost/population-bias points, and the four-way error taxonomy are all confirmed on the fetched QA page. Discrepancy: the claim quotes golden seeding as giving 'full control', but the page actually says 'complete control' (same meaning, different word). The three overturn/consistency/premium-metric items are attributed to a separate 'Metrics for Content Moderation' page (a different URL) and are not present on this page; the single-URL constraint bars fetching it.",
    "checked_at": "2026-07-15",
    "retrieval": "landing"
   },
   "source_retrieval_meta": {
    "retrieval": "landing",
    "fetched": [
     "https://www.tspa.org/curriculum/ts-fundamentals/content-moderation-and-operations/content-moderation-quality-assurance/"
    ],
    "resolved_title": "Content Moderation Quality Assurance (TSPA T&S Curriculum)",
    "resolved_date": "not shown on page",
    "title_match": true
   },
   "used_on": [
    {
     "page": "decisions.html",
     "anchor": "a5",
     "label": "Decisions - permanent audit"
    }
   ],
   "decision_use": {
    "grade": "A",
    "basis": "industry codification",
    "measured_on": "adjacent-domain",
    "note": "TSPA QA curriculum: seeding indistinguishability, appeals bias, wrong-selection errors.",
    "models_measured": []
   }
  },
  {
   "id": "CG-15",
   "domain": "critic-and-gapfill",
   "area": "Content-moderation QA programs - the largest deployed analog of humans-reviewing-human-judgment at scale. Meta/TikTok/Google have run decade-old programs auditing moderator decisions: golden-set seeding into live queues, audit sampling rates, overturn/appeal-rate monitoring, QA-of-the-QA layers, measured reviewer base-rate and fatigue effects, and published transparency-report methodology plus academic studies (e.g., on moderator agreement and audit design). The industry-practice domain covered only AI-training-data vendors; this adjacent industry has already answered several taxonomy Q6/Q9 questions operationally (seeded known-verdict items, easy-case dilution, de-graduation triggers).",
   "claim": "Pinterest's deployed Decision Quality Evaluation Framework (Feb 2026) scores both human moderators and LLM agents against an SME-adjudicated, versioned, immutable golden set selected by inverse-propensity sampling (XGBoost on PinCLIP embeddings, prioritizing low-propensity/novel items), using a two-axis diagnostic - Cohen's kappa for reliability plus correctness vs the golden set - where 'high reliability + low correctness' is read as systematic policy misunderstanding; measured results show 3x-human majority vote adds only +3.6pp accuracy over a single non-expert human, and current LLMs perform on par with a single non-expert human (GPT-5 the only config with positive accuracy delta, +0.9pp).",
   "load_bearing": false,
   "evidence": "Fetched full HTML of arXiv 2602.15809. Additional verified mechanics: policy changes handled by dual-labeling the existing golden set under old and new guidelines and visualizing the 'policy delta' as label flips; QA-of-the-QA is two continuous monitors - content-drift (evaluate on newest golden items each release) and system-stability (re-run fixed prompt on fixed golden version to catch pipeline non-determinism); switching prevalence measurement from 3x-human to LLM gave 'over 30x cost savings and 10x turnaround reduction.' Gemini 2.5 configs showed +22pp recall but +48-57pp false-positive rate.",
   "implication": "The closest 2026 production analog to the AutoQA foundation: (a) reliability and correctness must be separated - kappa alone cannot distinguish shared misunderstanding from noise; (b) golden sets should be deliberately non-representative (oversample rare/hard cases) with coverage measured in embedding space; (c) policy evolution requires versioned relabeling, not patching; (d) an LLM QA layer needs both a drift monitor and a determinism monitor; (e) expect the LLM judge to be ~single-non-expert-human quality with an FPR skew unless calibrated.",
   "source": {
    "raw": "Decision Quality Evaluation Framework at Pinterest (arXiv 2602.15809) | https://arxiv.org/abs/2602.15809 | 2026-02-17 | primary",
    "title": "Decision Quality Evaluation Framework at Pinterest (arXiv 2602.15809)",
    "url": "https://arxiv.org/abs/2602.15809",
    "date": "2026-02-17",
    "type": "primary"
   },
   "venue": {
    "peer_reviewed": true,
    "marker": "CHI"
   },
   "same_source_claims": [],
   "in_corpus_verification": null,
   "external_verification": {
    "verdict": "confirmed",
    "checked": [
     "3x-human majority vote adds +3.6pp accuracy over single non-expert human",
     "LLMs demonstrate quality on par with a single non-expert human (gap to SME-level remains)",
     "GPT-5 (balanced) +0.9 accuracy delta - only LLM config in the positive",
     "inverse propensity sampling; XGBoost trained on PinCLIP embeddings to predict propensity",
     "Cohen's Kappa for reliability + correctness vs GDS; high reliability + low correctness = systematic policy misunderstanding",
     "shift from 3x-human to LLM gave over 30x cost savings and 10x turnaround reduction",
     "Gemini/flash +22.6 recall (22.5% gain) but +57.2 FPR / +47.7% FPR",
     "GDS published as immutable and versioned dataset; SMEs relabel under new policy to produce 'policy delta'",
     "content-drift monitor + system-stability monitor"
    ],
    "not_visible": [],
    "quote": "This shift resulted in over 30x in cost savings and a 10x reduction in labeling turnaround time ... the LLMs demonstrate quality on par with a single non-expert human, but a gap still remains to reach SME-level quality ... High reliability paired with low correctness indicates that labelers are all making the same mistake consistently.",
    "notes": "Abstract alone did not carry the numbers; the allowed HTML full text confirms all nine specifics verbatim (Table 1 + prose). Minor nuance: '+0.9pp' GPT-5 is the only *LLM* config with a positive accuracy delta; the 3x-human-majority row is also positive (+3.6) but is a human baseline, so 'only config' is slightly imprecise. Gemini FPR stated as +47.7% in prose and +57.2 in Table 1, consistent with the claimed +48-57pp range.",
    "checked_at": "2026-07-15",
    "retrieval": "fulltext"
   },
   "source_retrieval_meta": {
    "retrieval": "fulltext",
    "fetched": [
     "https://arxiv.org/abs/2602.15809",
     "https://arxiv.org/html/2602.15809v1"
    ],
    "resolved_title": "Decision Quality Evaluation Framework at Pinterest",
    "resolved_date": "2026-02-17",
    "title_match": true
   },
   "used_on": [],
   "decision_use": {
    "grade": "B",
    "basis": "current-gen practice",
    "measured_on": "human-work",
    "note": "Pinterest decision-quality framework (Feb 2026, GPT-5-class agents in scope): reliability-vs-correctness separation, dual-labeling on policy change. Most current deployed analog.",
    "models_measured": [
     "GPT-5",
     "Gemini 2.5"
    ]
   }
  },
  {
   "id": "CG-16",
   "domain": "critic-and-gapfill",
   "area": "Content-moderation QA programs - the largest deployed analog of humans-reviewing-human-judgment at scale. Meta/TikTok/Google have run decade-old programs auditing moderator decisions: golden-set seeding into live queues, audit sampling rates, overturn/appeal-rate monitoring, QA-of-the-QA layers, measured reviewer base-rate and fatigue effects, and published transparency-report methodology plus academic studies (e.g., on moderator agreement and audit design). The industry-practice domain covered only AI-training-data vendors; this adjacent industry has already answered several taxonomy Q6/Q9 questions operationally (seeded known-verdict items, easy-case dilution, de-graduation triggers).",
   "claim": "Auditing raters by rewarding agreement with the eventual consensus outcome ('consensus-based auditing', as X's Community Notes has done since Sept 2022 by tying participation eligibility to agreement with the final aggregate) measurably induces strategic conformity - minority contributors' evaluations drift toward the majority and their participation share falls precisely on controversial items - and a two-stage alternative that weights contributors by the stability of their past residuals (predictability relative to a latent-factor model) rather than majority-agreement improves out-of-sample predictive performance.",
   "load_bearing": false,
   "evidence": "Fetched arXiv 2603.18053 (Alimohammadi, Huang, Borgs, Chayes, Mar 2026) abstract and framing via Exa crawl: empirical evidence from Community Notes data plus a behavioral model where contributors trade off private beliefs against anticipated penalties for disagreement; the proposed method gives influence to consistently-informative contributors 'even when they disagree with the prevailing consensus.' Senior authors (Borgs h-54, Chayes h-60). Preprint, not yet peer-reviewed.",
   "implication": "Directly constrains de-graduation trigger design: if attempter or QA standing is scored by agreement-with-final-verdict, the AutoQA will train its best dissenters to conform exactly where independent judgment matters most (the subjective-interpretation root cause #1). Score contributors on the stability/informativeness of their residuals against a grounded reference, and never penalize evidence-backed disagreement per se.",
   "source": {
    "raw": "Auditing the Auditors: Does Community-based Moderation Get It Right? (arXiv 2603.18053) | https://arxiv.org/html/2603.18053 | 2026-03-17 | academic",
    "title": "Auditing the Auditors: Does Community-based Moderation Get It Right? (arXiv 2603.18053)",
    "url": "https://arxiv.org/html/2603.18053",
    "date": "2026-03-17",
    "type": "academic"
   },
   "venue": {
    "peer_reviewed": false,
    "marker": "preprint"
   },
   "same_source_claims": [],
   "in_corpus_verification": null,
   "external_verification": {
    "verdict": "confirmed",
    "checked": [
     "'consensus-based auditing' = rewarding agreement with the final aggregate outcome",
     "in September 2022 a platform (Community Notes) adopted consensus-based auditing tying eligibility for participation to agreement",
     "strategic conformity: minority contributors' evaluations drift toward the majority",
     "their participation share falls on controversial topics",
     "proposed two-stage algorithm weights contributors by the stability of their past residuals (predictability vs a latent-factor model) rather than majority-agreement",
     "consistently-informative contributors get greater influence even when they disagree; improves out-of-sample predictive performance"
    ],
    "not_visible": [
     "author h-indices (Borgs h-54, Chayes h-60) - bibliometric metadata not on the page"
    ],
    "quote": "We find evidence of strategic conformity: minority contributors' evaluations drift toward the majority ... weights contributors by the stability of their past residuals ... even when they disagree ... improves out-of-sample predictive performance while avoiding penalization of disagreement",
    "notes": "Every element of the claim is confirmed in the HTML full text, including the Sept 2022 eligibility-tied consensus auditing, the strategic-conformity/minority-drift finding on controversial items, and the two-stage residual-stability aggregation. Retrieved page is v2 (2026-05-15); the arXiv ID and source_date (2026-03-17) correspond to the v1 submission. The reported two-stage improvement (mean abs residual ~5.73%, median ~27.99% on one-week-ahead predictions) is additional detail beyond the claim.",
    "checked_at": "2026-07-15",
    "retrieval": "fulltext"
   },
   "source_retrieval_meta": {
    "retrieval": "fulltext",
    "fetched": [
     "https://arxiv.org/html/2603.18053"
    ],
    "resolved_title": "Auditing the Auditors: Does Community-based Moderation Get It Right?",
    "resolved_date": "arXiv:2603.18053v2 dated 2026-05-15 (v1 submission month March 2026 per the 2603 arXiv ID; source_date 2026-03-17)",
    "title_match": true
   },
   "used_on": [
    {
     "page": "index.html",
     "anchor": "problem",
     "label": "Overview - the problem"
    },
    {
     "page": "research.html",
     "anchor": "research-cards",
     "label": "Research - findings by decision weight"
    },
    {
     "page": "decisions.html",
     "anchor": "invariants",
     "label": "Decisions - settled constraints"
    }
   ],
   "decision_use": {
    "grade": "A",
    "basis": "human incentive result",
    "measured_on": "human-work",
    "note": "Consensus-rewarded auditing induces conformity (Community Notes). Incentive design; durable.",
    "models_measured": []
   }
  },
  {
   "id": "CG-17",
   "domain": "critic-and-gapfill",
   "area": "Content-moderation QA programs - the largest deployed analog of humans-reviewing-human-judgment at scale. Meta/TikTok/Google have run decade-old programs auditing moderator decisions: golden-set seeding into live queues, audit sampling rates, overturn/appeal-rate monitoring, QA-of-the-QA layers, measured reviewer base-rate and fatigue effects, and published transparency-report methodology plus academic studies (e.g., on moderator agreement and audit design). The industry-practice domain covered only AI-training-data vendors; this adjacent industry has already answered several taxonomy Q6/Q9 questions operationally (seeded known-verdict items, easy-case dilution, de-graduation triggers).",
   "claim": "Google's audit-quality method (ICLR 2023 Tiny Paper, Google LLC authors) decomposes inter-rater agreement per rubric question rather than per item - in their worked example one question had Fleiss kappa 0.094 against a 0.475 overall, isolating it as the ambiguous criterion - and identifies sequential non-blind review (a later reviewer seeing the earlier verdict) as a distinct, measurable source of audit risk, testable via A/B or difference-in-differences.",
   "load_bearing": false,
   "evidence": "Extracted full PDF text via pdftotext. Methods: per-question Fleiss kappa to locate high-disagreement criteria; chi-square to test criterion-verdict linkage; t-test/ANOVA to detect review teams systematically deviating from ground truth; binomial CIs to extrapolate sampled error rates. Caveats: 2-page tiny paper, synthetic dataset (3 reviewers, 9-question rubric, 1,528 products), no production deployment claimed. Code at https://github.com/xuanyang0607/openreviewpaper.",
   "implication": "For root cause #1 (subjective interpretation of axes): measure agreement per evaluation criterion, not per item - this converts 'reviewers disagree' into 'criterion 3 is ambiguous, rewrite it', which is the closed-loop the AutoQA needs to feed back into project instruction sets. Also: keep any second/QA review blind to the first verdict to avoid anchoring.",
   "source": {
    "raw": "Statistical Methods for Auditing the Quality of Manual Content Reviews (arXiv 2306.07466, ICLR 2023 Tiny Papers) | https://arxiv.org/abs/2306.07466 | 2023-06-12 | academic",
    "title": "Statistical Methods for Auditing the Quality of Manual Content Reviews (arXiv 2306.07466, ICLR 2023 Tiny Papers)",
    "url": "https://arxiv.org/abs/2306.07466",
    "date": "2023-06-12",
    "type": "academic"
   },
   "venue": {
    "peer_reviewed": true,
    "marker": "ICLR"
   },
   "same_source_claims": [],
   "in_corpus_verification": null,
   "external_verification": {
    "verdict": "confirmed",
    "checked": [
     "authors Xuan Yang, Andrew Smart, Daniel Theron; affiliation Google LLC",
     "Published as a Tiny Paper at ICLR 2023",
     "per-question Fleiss kappa locates the ambiguous criterion",
     "Table 1: Question 3 Fleiss kappa 0.093997 vs Overall Fleiss Kappa 0.475072",
     "Question 3 flagged as systematically higher disagreement",
     "sequential non-blind review (one reviewer sees previous reviewers' results) named as audit risk",
     "testable via A/B test or Difference-in-Difference",
     "synthetic dataset: three reviewers, nine-question rubric, 1,528 products",
     "chi-square for criterion-verdict linkage; t-test/ANOVA for deviating review teams",
     "binomial-distribution confidence interval of the error rate",
     "code at github.com/xuanyang0607/openreviewpaper"
    ],
    "not_visible": [],
    "quote": "Table 1 ... Question 3 0.093997 ... Overall Fleiss Kappa 0.475072 ... synthetic data set consisting of three reviewers and their answers regarding a nine question rubric on 1,528 products ... reviewers conduct reviews sequentially ... A/B test ... Difference-in-Difference",
    "notes": "Extracted full PDF text via pdftotext (WebFetch saved the binary; local extraction succeeded). Every element of the claim confirmed verbatim: 0.094 vs 0.475 (0.093997 / 0.475072), Google LLC, ICLR 2023 Tiny Paper, sequential-review/A-B/DiD, dataset dimensions, chi-square/t-test/ANOVA, binomial CI (line 283), and the GitHub URL (line 182).",
    "checked_at": "2026-07-15",
    "retrieval": "pdf"
   },
   "source_retrieval_meta": {
    "retrieval": "pdf",
    "fetched": [
     "https://arxiv.org/abs/2306.07466",
     "https://arxiv.org/pdf/2306.07466"
    ],
    "resolved_title": "Statistical Methods for Auditing the Quality of Manual Content Reviews",
    "resolved_date": "2023-06-12 (Published as a Tiny Paper at ICLR 2023)",
    "title_match": true
   },
   "used_on": [],
   "decision_use": {
    "grade": "A",
    "basis": "method",
    "measured_on": "human-work",
    "note": "Per-rubric-question agreement decomposition (Google). Method.",
    "models_measured": []
   }
  },
  {
   "id": "CG-18",
   "domain": "critic-and-gapfill",
   "area": "Content-moderation QA programs - the largest deployed analog of humans-reviewing-human-judgment at scale. Meta/TikTok/Google have run decade-old programs auditing moderator decisions: golden-set seeding into live queues, audit sampling rates, overturn/appeal-rate monitoring, QA-of-the-QA layers, measured reviewer base-rate and fatigue effects, and published transparency-report methodology plus academic studies (e.g., on moderator agreement and audit design). The industry-practice domain covered only AI-training-data vendors; this adjacent industry has already answered several taxonomy Q6/Q9 questions operationally (seeded known-verdict items, easy-case dilution, de-graduation triggers).",
   "claim": "A 2025 New Media & Society study (screen-share observation of commercial moderators in India) documents that throughput pressure distorts verdict distributions through the action-selection interface itself: moderators systematically chose a 2-click removal over a 4-click de-ranking regardless of policy fit, mechanically deleted tool-highlighted words without context assessment, and privately compressed broad guidelines into simplified dos/don'ts lists - i.e., queue and UI composition changed outcomes independent of moderator judgment quality.",
   "load_bearing": false,
   "evidence": "Fetched The Conversation summary (2025-07-22) by the study authors (Chatterjee IIT Delhi/UQ, Gupta, Thomas; New Media & Society, peer-reviewed). Evidence is qualitative (interviews + observed sessions), no quantitative error rates; an observed moderator said she 'would never recommend de-ranking content as it would take time.' Complements industry-standard per-item handling-time metrics (TSPA) that create the pressure.",
   "implication": "Reviewer 'harshness' variance is partly an artifact of per-item cost asymmetries, not belief: any AutoQA that infers attempter/reviewer quality from verdict distributions must first control for the click/effort cost of each verdict option, and the one-human-touch interaction in the hybrid loop should make the correct action the cheapest action. - DISAGREEMENT: Facebook's moderator accuracy target: The Verge's original Phoenix reporting (Feb 2019, fetched) states a 95% target with actual scores in the high-80s to 92; CNBC's Tampa follow-up (Jun 2019) and The Irish Times' Dublin reporting (Feb 2020) both state a 98% target. Likely site/contract variation, but the primary sources genuinely conflict on the number; all agree the target was chronically missed. - DISAGREEMENT: What low reviewer agreement means: production vendor practice (Facebook/Cognizant, Accenture/CPL) treats disagreement-with-auditor as individual reviewer error and scores/fires on it, while Musubi Labs (Nov 2025, practitioner), Pinterest's framework (2026), and the Google tiny paper treat low agreement primarily as a policy/criterion-ambiguity signal ('don't force agreement'); the Community Notes paper (Mar 2026) goes further, showing agreement-rewarded auditing actively corrupts the signal via strategic conformity. - DISAGREEMENT: Platform self-reported accuracy vs external measurement: TikTok's fifth DSA transparency report claims 99.2% moderation accuracy (self-defined, sample re-review), while academic audits of the DSA transparency database found up to ~50-percentage-point inconsistencies in TikTok's self-reported automation figures, and moderator testimony ('this number is fake - accuracy is only judged by agreement') argues agreement-based accuracy overstates true correctness. Self-reported accuracy figures from this industry should not be used as calibration anchors. - DISAGREEMENT: Value of majority vote: Pinterest measured 3x-human majority at only +3.6pp accuracy over a single non-expert human (arguing redundancy is a weak, expensive lever), whereas Roblox's published practice treats >=80% multi-moderator alignment as the gate for scaled consistency - different uses (measurement vs gating) but opposite implicit views on how much signal replication buys. - GAP: No public quantitative study of queue-composition effects on reviewer harshness in content moderation specifically (easy-case dilution shifting strictness, prevalence-induced criterion drift): the mechanism is well established in adjacent vigilance/low-prevalence-effect literature, but I found no moderation-industry measurement of it; the 2025 New Media & Society evidence is qualitative only. This taxonomy question remains empirically open even in the largest deployed analog. - GAP: Gold-injection density is nowhere disclosed: no platform, vendor, or paper states what fraction of a live queue is seeded known-verdict items, only that seeding exists and must be indistinguishable (TSPA) and that golden sets should oversample hard cases (Pinterest, Musubi). - GAP: Meta's current (2025-2026) internal audit sampling rates and Community Standards Enforcement Report reviewer-accuracy methodology are not publicly disclosed at mechanic level; the best internals remain 2019-2020 journalism. Kenya/Meta litigation (Majorel/Sama) disclosures center on labor harms, not QA mechanics - I found no litigation-produced QA-design documents newer than that reporting. - GAP: De-graduation thresholds are only known via journalism (miss accuracy target -> remedial program -> termination); no published numeric trigger (e.g., N misses in window, minimum sample before action) from any platform. - GAP: Pinterest's paper, the single best 2026 primary source, omits absolute golden-set size, SME adjudication protocol, and absolute (non-delta) accuracy values, so its numbers transfer as design patterns, not calibration constants.",
   "source": {
    "raw": "Hard labour conditions of online moderators directly affect how well the internet is policed (The Conversation, summarizing New Media & Society study) | https://theconversation.com/hard-labour-conditions-of-online-moderators-directly-affect-how-well-the-internet-is-policed-new-study-261386 | 2025-07-22 | academic",
    "title": "Hard labour conditions of online moderators directly affect how well the internet is policed (The Conversation, summarizing New Media & Society study)",
    "url": "https://theconversation.com/hard-labour-conditions-of-online-moderators-directly-affect-how-well-the-internet-is-policed-new-study-261386",
    "date": "2025-07-22",
    "type": "academic"
   },
   "venue": {
    "peer_reviewed": null,
    "marker": "academic"
   },
   "same_source_claims": [],
   "in_corpus_verification": null,
   "external_verification": {
    "verdict": "confirmed",
    "checked": [
     "removing posts required only two steps; de-ranking involved four steps",
     "moderator skipped/avoided de-ranking to save time ('Would never recommend de-ranking content as it would take time')",
     "removing tool-flagged words without evaluating context",
     "moderators develop a simplified list of 'dos and don'ts'",
     "commercial content moderators in India; screen-share observation",
     "study published in New Media & Society; author Tania Chatterjee"
    ],
    "not_visible": [
     "literal word 'clicks' (article says 'steps': two steps vs four steps)",
     "literal word 'throughput' (concept present via targets/time pressure)",
     "quantitative error rates (none - study is qualitative, as the claim acknowledges)"
    ],
    "quote": "removing posts required only two steps ... reducing the visibility of content (de-ranking) involved four steps ... Would never recommend de-ranking content as it would take time.",
    "notes": "Core mechanism confirmed: the action-selection UI (2-step removal vs 4-step de-ranking) plus throughput pressure shifted outcomes independent of policy fit ('To save time, she skipped the content flagged to be de-ranked'; 'removing flagged words without evaluating the context'; 'dos and don'ts'). Only wording nuance: article says 'steps' not 'clicks'.",
    "checked_at": "2026-07-15",
    "retrieval": "fulltext"
   },
   "source_retrieval_meta": {
    "retrieval": "fulltext",
    "fetched": [
     "https://theconversation.com/hard-labour-conditions-of-online-moderators-directly-affect-how-well-the-internet-is-policed-new-study-261386"
    ],
    "resolved_title": "Hard labour conditions of online moderators directly affect how well the internet is policed - new study",
    "resolved_date": "2025-07-22",
    "title_match": true
   },
   "used_on": [],
   "decision_use": {
    "grade": "A",
    "basis": "human science",
    "measured_on": "adjacent-domain",
    "note": "Throughput pressure degrades moderator judgment (2025 observational). Human factors; durable.",
    "models_measured": []
   }
  },
  {
   "id": "CG-19",
   "domain": "critic-and-gapfill",
   "area": "Feedback-efficacy science from education and organizational psychology. Every domain independently reports 'no evidence that QA feedback changes attempter behavior' as a gap, yet feedback-intervention research is a mature field: Kluger & DeNisi's feedback intervention theory (a third of feedback interventions REDUCE performance, with known moderators), formative-assessment literature on feedback specificity/timing, and workplace studies on feedback under pay-linked evaluation (which predicts gaming/monoculture, taxonomy Q8). The sweep searched only for annotation-specific longitudinal studies and found none - the general literature was never consulted.",
   "claim": "Feedback intervention theory's core empirical result stands: feedback improves performance on average (K&D 1996: 131 studies, ~12,000+ participants, mean d approximately 0.38-0.41) but MORE THAN ONE THIRD of feedback interventions REDUCE performance, and effectiveness declines as feedback cues move attention from the task toward the self (praise, person-level evaluation, normative comparison).",
   "load_bearing": false,
   "evidence": "Verified in the authors' own summary (Kluger & DeNisi 1998, Current Directions in Psychological Science): 'although FIs improve performance on average, they reduce performance in more than one third of the cases.' Independently corroborated by Wisniewski et al. 2020, which cites K&D as 131 studies, >12,000 participants, average effect 0.38 with roughly a third of effects negative. FIT's mechanism: feedback that directs attention to meta-task/self processes (threat to self, praise, social comparison) depletes task attention and backfires; task- and process-focused cues help.",
   "implication": "The AutoQA feedback 4-tuple must be strictly task/criterion-referenced and evidence-anchored, never person-referenced or rank-referenced; a feedback channel is not presumptively net-positive, so the design should treat 'feedback reduces this attempter's subsequent quality' as an expected outcome for a substantial minority and instrument for it (per-attempter pre/post error-rate deltas), not assume monotone benefit.",
   "source": {
    "raw": "Feedback Interventions (Kluger & DeNisi 1998), summarizing K&D 1996 Psychological Bulletin meta-analysis | https://doi.org/10.1111/1467-8721.ep10772989 | 1998-06 (meta-analysis 1996-03; foundational) | academic",
    "title": "Feedback Interventions (Kluger & DeNisi 1998), summarizing K&D 1996 Psychological Bulletin meta-analysis",
    "url": "https://doi.org/10.1111/1467-8721.ep10772989",
    "date": "1998-06 (meta-analysis 1996-03; foundational)",
    "type": "academic"
   },
   "venue": {
    "peer_reviewed": null,
    "marker": "academic"
   },
   "same_source_claims": [],
   "in_corpus_verification": null,
   "external_verification": {
    "verdict": "partially_confirmed",
    "checked": [
     "Kluger & DeNisi 1998 article in Current Directions in Psychological Science (Vol 7, Issue 3, pp 67-72, June 1998)",
     "cites the 1996 Psychological Bulletin meta-analysis (Kluger & DeNisi 1996, vol 119, pp 254-284)"
    ],
    "not_visible": [
     "131 studies",
     "~12,000+ participants",
     "mean d approximately 0.38-0.41",
     "phrase 'reduce performance in more than one third of the cases'",
     "attention shifting task->self mechanism (praise, normative comparison)"
    ],
    "quote": "Feedback Interventions: Toward the Understanding of a Double-Edged Sword",
    "notes": "DOI 302-redirects (server-side) to journals.sagepub.com, followed per allowed-variant rule. Page is restricted access: only title/authors/metadata and the 1996 citation are visible; no abstract or body text. The article's identity and the 1996 meta-analysis citation are confirmed, but every headline number and the 'more than one third reduce performance' phrase are behind the paywall and not visible in retrieved text.",
    "checked_at": "2026-07-15",
    "retrieval": "landing"
   },
   "source_retrieval_meta": {
    "retrieval": "landing",
    "fetched": [
     "https://doi.org/10.1111/1467-8721.ep10772989",
     "https://journals.sagepub.com/doi/10.1111/1467-8721.ep10772989"
    ],
    "resolved_title": "Feedback Interventions: Toward the Understanding of a Double-Edged Sword",
    "resolved_date": "Current Directions in Psychological Science, Vol. 7(3), pp. 67-72, June 1998",
    "title_match": true
   },
   "used_on": [
    {
     "page": "decisions.html",
     "anchor": "d1",
     "label": "Decisions - enforcement weight"
    }
   ],
   "decision_use": {
    "grade": "A",
    "basis": "human science",
    "measured_on": "human-work",
    "note": "Kluger & DeNisi FIT: a third of feedback interventions backfire; task-referenced helps. Foundational; durable.",
    "models_measured": []
   }
  },
  {
   "id": "CG-20",
   "domain": "critic-and-gapfill",
   "area": "Feedback-efficacy science from education and organizational psychology. Every domain independently reports 'no evidence that QA feedback changes attempter behavior' as a gap, yet feedback-intervention research is a mature field: Kluger & DeNisi's feedback intervention theory (a third of feedback interventions REDUCE performance, with known moderators), formative-assessment literature on feedback specificity/timing, and workplace studies on feedback under pay-linked evaluation (which predicts gaming/monoculture, taxonomy Q8). The sweep searched only for annotation-specific longitudinal studies and found none - the general literature was never consulted.",
   "claim": "The March 2025 Cochrane update on audit-and-feedback (292 studies, 678 arms, healthcare professionals) finds median absolute improvement in desired practice of only 2.7% (IQR 0.0 to 8.6; weighted meta-analytic mean +6.2%, 95% CI 4.1-8.2, moderate certainty), with effects larger for low baseline performers, individual-level (not team-level) data, comparison to TOP peers or a benchmark (comparison to peer average showed no significant effect), a trusted local source, and action plans with specific advice - while repeated delivery was associated with LOWER effect size.",
   "load_bearing": false,
   "evidence": "Fetched the Cochrane summary (updated review of CD000259, published 2025-03-25): '292 studies with 678 arms'; median absolute improvement '2.7%, with an IQR of 0.0 to 8.6'; weighted mean '6.2% (95% CI 4.1 to 8.2; moderate-certainty evidence)'; OR 1.47. Moderator list quoted directly, including the counterintuitive 'repeated delivery was associated with lower effect size' and 'comparison to top-peers or a benchmark increased effects; comparing against the average of all peers did not.'",
   "implication": "This is the largest causal evidence base that verdict-plus-feedback changes skilled professionals' behavior: expect a real but modest median effect with a fat right tail, concentrated in currently-low performers. Design levers with evidence: target feedback at low-baseline attempters first, deliver individual-level data, pair every failed criterion with a specific corrective action, and benchmark against top-quality exemplars rather than cohort averages. Do NOT assume higher feedback frequency improves uptake.",
   "source": {
    "raw": "Audit and feedback: effects on professional practice (Cochrane review update, Ivers et al.) | https://www.cochranelibrary.com/cdsr/doi/10.1002/14651858.CD000259.pub4/full | 2025-03-25 | academic",
    "title": "Audit and feedback: effects on professional practice (Cochrane review update, Ivers et al.)",
    "url": "https://www.cochranelibrary.com/cdsr/doi/10.1002/14651858.CD000259.pub4/full",
    "date": "2025-03-25",
    "type": "academic"
   },
   "venue": {
    "peer_reviewed": true,
    "marker": "Cochrane"
   },
   "same_source_claims": [],
   "in_corpus_verification": null,
   "external_verification": {
    "verdict": "confirmed",
    "checked": [
     "292 studies with 678 arms",
     "median absolute improvement in desired practice of 2.7%, IQR 0.0 to 8.6 (177 studies)",
     "mean absolute increase 6.2% (95% CI 4.1 to 8.2; moderate-certainty)",
     "OR 1.47 (95% CI 1.31 to 1.64)",
     "Lower baseline performance associated with larger intervention effects",
     "individual-recipient-level data rather than team-level data",
     "compares performance to top peers or a benchmark",
     "comparison to average performance of all peers did NOT find significant effects",
     "local champion with existing relationship (trusted local source)",
     "actionable plan with specific advice for improvement",
     "Contrary to expectations, repeated delivery was associated with lower effect size",
     "Version published 25 March 2025"
    ],
    "not_visible": [],
    "quote": "mean absolute increase in desired practice of 6.2% (95% confidence interval (CI) 4.1 to 8.2; moderate-certainty evidence) and an OR of 1.47 (95% CI 1.31 to 1.64; moderate-certainty evidence)",
    "notes": "Every number and every moderator in the claim confirmed verbatim from the full-text abstract and results, including the counterintuitive repeated-delivery finding and the peer-average null. Page required a cookie-jar two-step (WebFetch 403; curl 412) to bypass the cookie gate.",
    "checked_at": "2026-07-15",
    "retrieval": "fulltext"
   },
   "source_retrieval_meta": {
    "retrieval": "fulltext",
    "fetched": [
     "https://www.cochranelibrary.com/cdsr/doi/10.1002/14651858.CD000259.pub4/full",
     "https://www.cochranelibrary.com/cdsr/doi/10.1002/14651858.CD000259.pub4/full?cookiesEnabled"
    ],
    "resolved_title": "Audit and feedback: effects on professional practice (Ivers, N - 2025)",
    "resolved_date": "2025-03-25 (Version published 25 March 2025)",
    "title_match": true
   },
   "used_on": [],
   "decision_use": {
    "grade": "A",
    "basis": "human science",
    "measured_on": "adjacent-domain",
    "note": "Cochrane 2025 audit-and-feedback update: modest median effects, concentrated in low performers. Durable.",
    "models_measured": []
   }
  },
  {
   "id": "CG-21",
   "domain": "critic-and-gapfill",
   "area": "Feedback-efficacy science from education and organizational psychology. Every domain independently reports 'no evidence that QA feedback changes attempter behavior' as a gap, yet feedback-intervention research is a mature field: Kluger & DeNisi's feedback intervention theory (a third of feedback interventions REDUCE performance, with known moderators), formative-assessment literature on feedback specificity/timing, and workplace studies on feedback under pay-linked evaluation (which predicts gaming/monoculture, taxonomy Q8). The sweep searched only for annotation-specific longitudinal studies and found none - the general literature was never consulted.",
   "claim": "In the closest annotation-analog RCT (105 analyzed Mechanical Turk workers writing product reviews), both timely external expert feedback and rubric-based self-assessment significantly improved work quality vs no feedback (expert ratings 6.01 and 6.35 vs 5.69 on a 9-point scale, p<0.05) with NO quality difference between external and self-assessment; external feedback uniquely drove revision behavior (56.5% revised vs 24.8% for self-assessment) and more output, and self-assessors over-rated their own work by 1.8 points (7.9 self vs 6.1 expert).",
   "load_bearing": false,
   "evidence": "Read the full Dow, Kulkarni, Klemmer & Hartmann CSCW 2012 PDF. Between-subjects, blind-to-condition expert grading; self-assessment condition showed significant learning over the task series (slope 0.25, p=0.001) vs borderline for external (0.10, p=0.08) and null for none. Also: crowd-peer assessors had low agreement with the expert (Kappa=0.20 aggregated), and the paper explicitly did not measure long-term learning. Attrition was higher under assessment (Self 78%, External 61%, vs None 47% incompleteness), so gains partly reflect weak performers dropping out plus learning.",
   "implication": "Directly refutes 'no evidence QA feedback changes attempter behavior' for paid micro-task workers: rubric-mediated, task-specific, synchronous feedback is the minimal effective unit. A concrete per-criterion rubric surfaced to attempters may capture most of the quality gain of expensive external feedback (making the AutoQA-generated feedback the 'external expert' at near-zero marginal cost), but attempter self-ratings cannot serve as measurement, and part of any observed 'improvement' will be selection (weak attempters exiting) - the design's efficacy telemetry must separate within-attempter learning from attrition.",
   "source": {
    "raw": "Shepherding the Crowd Yields Better Work (Dow, Kulkarni, Klemmer, Hartmann; CSCW 2012) | https://www.cs.cmu.edu/~spdow/files/Crowds-Shepherd-CSCW12.pdf | 2012-02-11 (foundational; only direct crowdwork feedback RCT found) | academic",
    "title": "Shepherding the Crowd Yields Better Work (Dow, Kulkarni, Klemmer, Hartmann; CSCW 2012)",
    "url": "https://www.cs.cmu.edu/~spdow/files/Crowds-Shepherd-CSCW12.pdf",
    "date": "2012-02-11 (foundational; only direct crowdwork feedback RCT found)",
    "type": "academic"
   },
   "venue": {
    "peer_reviewed": true,
    "marker": "CSCW"
   },
   "same_source_claims": [],
   "in_corpus_verification": null,
   "external_verification": {
    "verdict": "confirmed",
    "checked": [
     "105 participants analyzed (538 consumer reviews from 105 participants)",
     "expert 9-point ratings: External mu=6.01, Self mu=6.35, None mu=5.69; F(2,102)=3.02, p<0.05; pairwise None-vs-External and None-vs-Self both p<0.05",
     "NO significant difference between External and Self assessment (t(150)=1.40, p=0.18)",
     "revision rates: 56.5% External vs 24.8% Self changed their review",
     "self-assessors over-rated own work by 1.8 points (mu=7.9 self vs mu=6.1 expert)",
     "crowd-as-aggregate vs expert agreement Cohen's Kappa=0.20",
     "higher attrition in Self (78%) and External (61%) than None (47%)",
     "learning: Self slope 0.25 (p=0.001) significant; External 0.10 (p=0.08) borderline; None null"
    ],
    "not_visible": [],
    "quote": "The External condition (mu=6.01, SD=1.38) and the Self condition (mu=6.35, SD=1.63) outperformed the None condition (mu=5.69, SD=1.19) (F(2,102)=3.02, p<0.05) ... There was no significant difference between External and Self assessment (t(150)=1.40, p=0.18)",
    "notes": "Every number in the claim confirmed verbatim from the extracted PDF text (downloaded via curl, converted with pdftotext). Self over-rating '1.8 points higher (mu=7.9 ... mu=6.1)', revision '56.5% ... 24.8%', Kappa=0.20, attrition 78/61/47%, and learning slopes 0.25/0.10 all present exactly as stated.",
    "checked_at": "2026-07-15",
    "retrieval": "pdf"
   },
   "source_retrieval_meta": {
    "retrieval": "pdf",
    "fetched": [
     "https://www.cs.cmu.edu/~spdow/files/Crowds-Shepherd-CSCW12.pdf"
    ],
    "resolved_title": "Shepherding the Crowd Yields Better Work (Dow, Kulkarni, Klemmer, Hartmann)",
    "resolved_date": "CSCW'12, February 11-15, 2012",
    "title_match": true
   },
   "used_on": [],
   "decision_use": {
    "grade": "A",
    "basis": "human science",
    "measured_on": "human-work",
    "note": "Dow CSCW 2012: task-specific feedback improves paid microtask quality; self-ratings inflate. Closest crowdwork RCT; durable.",
    "models_measured": []
   }
  },
  {
   "id": "CG-22",
   "domain": "critic-and-gapfill",
   "area": "Feedback-efficacy science from education and organizational psychology. Every domain independently reports 'no evidence that QA feedback changes attempter behavior' as a gap, yet feedback-intervention research is a mature field: Kluger & DeNisi's feedback intervention theory (a third of feedback interventions REDUCE performance, with known moderators), formative-assessment literature on feedback specificity/timing, and workplace studies on feedback under pay-linked evaluation (which predicts gaming/monoculture, taxonomy Q8). The sweep searched only for annotation-specific longitudinal studies and found none - the general literature was never consulted.",
   "claim": "In the largest educational feedback meta-analysis (435 studies, k=994, N>61,000), information content is the dominant moderator: high-information feedback (task + process + self-regulation content) yields d=0.99 [0.82-1.15] versus d=0.46 for corrective right/wrong feedback and d=0.24 for bare reinforcement/punishment - with overall d=0.48 masking huge heterogeneity (I2=83%) and 17% of raw effects negative.",
   "load_bearing": false,
   "evidence": "Fetched the open-access Frontiers article (Wisniewski, Zierer & Hattie 2020). Feedback-type moderator significant (QB=41.52, p<0.0001); outcome moderator significant: cognitive d=0.51 vs motivational d=0.33; of negative motivational effects, 86% came from uninformative (reward/punishment-style) feedback. Authors conclude feedback 'cannot be understood as a single consistent form of treatment.'",
   "implication": "This is the direct evidential anchor for taxonomy Q7's feedback-vs-verdict-only branch: a pass/fail verdict is the reinforcement/corrective tier (expected d approximately 0.24-0.46), while explaining WHAT is wrong against the criterion, WHY (process), and HOW to self-check next time (self-regulation) roughly doubles the expected effect. The minimal feedback unit worth building is therefore criterion-cited + evidence-grounded + process-level; verdict-only QA forfeits most of the achievable behavior change, and 'feedback is decoration' should only be concluded if high-information feedback (not verdicts) fails to move within-attempter error rates.",
   "source": {
    "raw": "The Power of Feedback Revisited: A Meta-Analysis of Educational Feedback Research | https://www.frontiersin.org/journals/psychology/articles/10.3389/fpsyg.2019.03087/full | 2020-01-22 | academic",
    "title": "The Power of Feedback Revisited: A Meta-Analysis of Educational Feedback Research",
    "url": "https://www.frontiersin.org/journals/psychology/articles/10.3389/fpsyg.2019.03087/full",
    "date": "2020-01-22",
    "type": "academic"
   },
   "venue": {
    "peer_reviewed": null,
    "marker": "academic"
   },
   "same_source_claims": [],
   "in_corpus_verification": null,
   "external_verification": {
    "verdict": "confirmed",
    "checked": [
     "435 studies, k = 994 effect sizes, N > 61,000",
     "high-information feedback d = 0.99 [0.82-1.15] (k=42)",
     "corrective right/wrong feedback d = 0.46 [0.39-0.55] (k=238)",
     "reinforcement/punishment d = 0.24 [0.06-0.43] (k=39)",
     "overall weighted d = 0.48 (outlier-cleaned; CL 0.44-0.51), I2 = 83.40%",
     "17% of raw effects negative (full set before outlier removal)",
     "feedback-type moderator QB = 41.52, df=2, p < 0.0001",
     "cognitive outcomes d = 0.51 vs motivational d = 0.33",
     "86% of negative motivational effects from uninformative (reward/punishment) feedback",
     "'feedback cannot be understood as a single consistent form of treatment' (verbatim)"
    ],
    "not_visible": [],
    "quote": "feedback cannot be understood as a single consistent form of treatment",
    "notes": "Every headline number matches. Minor caveat surfaced in text: d=0.48 is the outlier-adjusted estimate (initial integration d=0.55 with 17% negative); the claim uses the 0.48 cleaned figure and the 17%-negative full-set figure, both consistent with the paper.",
    "checked_at": "2026-07-15",
    "retrieval": "fulltext"
   },
   "source_retrieval_meta": {
    "retrieval": "fulltext",
    "fetched": [
     "https://www.frontiersin.org/journals/psychology/articles/10.3389/fpsyg.2019.03087/full"
    ],
    "resolved_title": "The Power of Feedback Revisited: A Meta-Analysis of Educational Feedback Research",
    "resolved_date": "2020-01-22 (Frontiers in Psychology, Vol. 10-2019)",
    "title_match": true
   },
   "used_on": [],
   "decision_use": {
    "grade": "A",
    "basis": "human science",
    "measured_on": "adjacent-domain",
    "note": "Information content dominates feedback efficacy (d~0.99 high-information vs 0.46 bare verdicts). Durable.",
    "models_measured": []
   }
  },
  {
   "id": "CG-23",
   "domain": "critic-and-gapfill",
   "area": "Feedback-efficacy science from education and organizational psychology. Every domain independently reports 'no evidence that QA feedback changes attempter behavior' as a gap, yet feedback-intervention research is a mature field: Kluger & DeNisi's feedback intervention theory (a third of feedback interventions REDUCE performance, with known moderators), formative-assessment literature on feedback specificity/timing, and workplace studies on feedback under pay-linked evaluation (which predicts gaming/monoculture, taxonomy Q8). The sweep searched only for annotation-specific longitudinal studies and found none - the general literature was never consulted.",
   "claim": "The 2025 Annual Review of Organizational Psychology's 25-year retrospective concludes the science of workplace feedback 'is not yet a story of coherent and cumulative progress': definitions are generic, assumptions diverge across six disconnected research substreams, and simple universal rules about feedback effectiveness do not survive contact with organizational reality.",
   "load_bearing": false,
   "evidence": "Crawled the full open-access review (Anseel & Sherf, Annu. Rev. Organ. Psychol. Organ. Behav. 12:19-43, published 2025-01-21). Abstract states insights 'often appear disconnected from the way feedback is practiced and experienced in organizations'; the review calls for explicated assumptions and paradigms mirroring complex realities. The companion 2025 systematic review (Heine, Stouten & Liden, J. Organ. Behav., 2025-10-26) reaches the same verdict for supervisor performance feedback: most studies fail even to specify feedback valence, and feedback quality/accuracy findings rest on inconsistent constructs.",
   "implication": "Tempering prior for the whole Q7 branch: the general literature supplies directional moderators (task-focus, specificity, information content, source credibility) but NO validated plug-in recipe, and effect heterogeneity is the norm. The project-agnostic foundation should therefore ship feedback design as parameterized hypotheses with built-in efficacy measurement (per-project A/B of feedback tiers against repeat-error rate), not as fixed doctrine imported from any single meta-analysis.",
   "source": {
    "raw": "A 25-Year Review of Research on Feedback in Organizations: From Simple Rules to Complex Realities (Anseel & Sherf) | https://www.annualreviews.org/content/journals/10.1146/annurev-orgpsych-110622-031927 | 2025-01-21 | academic",
    "title": "A 25-Year Review of Research on Feedback in Organizations: From Simple Rules to Complex Realities (Anseel & Sherf)",
    "url": "https://www.annualreviews.org/content/journals/10.1146/annurev-orgpsych-110622-031927",
    "date": "2025-01-21",
    "type": "academic"
   },
   "venue": {
    "peer_reviewed": null,
    "marker": "academic"
   },
   "same_source_claims": [],
   "in_corpus_verification": null,
   "external_verification": {
    "verdict": "not_retrievable",
    "checked": [],
    "not_visible": [
     "'is not yet a story of coherent and cumulative progress'",
     "insights 'often appear disconnected from the way feedback is practiced and experienced in organizations'",
     "diverging assumptions across six research substreams",
     "citation Annu. Rev. Organ. Psychol. Organ. Behav. 12:19-43, published 2025-01-21"
    ],
    "quote": "",
    "notes": "Page is protected by a Cloudflare managed challenge that requires JavaScript execution. WebFetch returned 403; curl with a browser user-agent received only the 'Just a moment...' challenge HTML; the Playwright browser loaded the exact same authorized URL but the managed challenge never cleared (HTTP 403) within ~14s. No article text retrieved; no substitute source permitted, so the claim cannot be verified.",
    "checked_at": "2026-07-15",
    "retrieval": "failed"
   },
   "source_retrieval_meta": {
    "retrieval": "failed",
    "fetched": [
     "https://www.annualreviews.org/content/journals/10.1146/annurev-orgpsych-110622-031927 (WebFetch: HTTP 403; curl browser-UA: Cloudflare managed JS challenge 'Just a moment...'; Playwright same URL: HTTP 403, challenge unresolved after ~14s)"
    ],
    "resolved_title": "not retrievable",
    "resolved_date": "not retrievable",
    "title_match": false
   },
   "used_on": [],
   "decision_use": {
    "grade": "A",
    "basis": "human science",
    "measured_on": "adjacent-domain",
    "note": "25-year workplace-feedback retrospective: field immature; effects heterogeneous. Calibrates expectations.",
    "models_measured": []
   }
  },
  {
   "id": "CG-24",
   "domain": "critic-and-gapfill",
   "area": "Feedback-efficacy science from education and organizational psychology. Every domain independently reports 'no evidence that QA feedback changes attempter behavior' as a gap, yet feedback-intervention research is a mature field: Kluger & DeNisi's feedback intervention theory (a third of feedback interventions REDUCE performance, with known moderators), formative-assessment literature on feedback specificity/timing, and workplace studies on feedback under pay-linked evaluation (which predicts gaming/monoculture, taxonomy Q8). The sweep searched only for annotation-specific longitudinal studies and found none - the general literature was never consulted.",
   "claim": "Under incentives, relative-rank feedback is a double-edged lever: lab and field economics find rank feedback raises output in flat-wage settings (Charness et al. 2014; Tafkov 2013) but induces costly sabotage and cheating to improve rank that offsets the gains, and in at least one field experiment (Barankay 2012) REMOVING rank feedback improved performance.",
   "load_bearing": false,
   "evidence": "Confirmed via the literature synthesis in a peer-reviewed Leadership Quarterly article ('Feedback quality and performance in organisations', 2021), which states: Charness et al. (2014) found 'offering relative rank feedback increases output... and subjects are willing to engage in costly sabotage and cheating activities to improve their relative rank, thus offsetting the positive effects'; Barankay (2012) is cited as the exception where rank feedback hurt. Primary Charness/Barankay texts not independently opened, so treated as well-sourced secondary evidence.",
   "implication": "For taxonomy Q8 (gaming/monoculture under pay-linked evaluation): if AutoQA outputs become visible rank or pass-rate leaderboards tied to pay, the literature predicts optimization of the metric (score-hacking, mimicry of known-passing templates) rather than quality. Keep attempter-facing feedback private, criterion-referenced, and decoupled from visible peer ranking; note the tension with Cochrane's top-peer-benchmark moderator (see disagreements) - benchmark against exemplar WORK, not against ranked PEOPLE. - DISAGREEMENT: Feedback frequency: pre-2025 audit-and-feedback guidance (Ivers 2012 Cochrane, Hysong 2006) held that repeated/more frequent delivery increases effect; the 2025 Cochrane update finds repeated delivery associated with LOWER effect size. Unresolved - could be confounding (repeated A&F deployed where problems persist) or genuine habituation. - DISAGREEMENT: Normative comparison: the 2025 Cochrane update finds comparison to top peers or a benchmark INCREASES behavior change in healthcare professionals, while FIT (Kluger & DeNisi) predicts normative comparison shifts attention to self and degrades performance, and incentive economics (Charness 2014, Barankay 2012) finds rank feedback triggers gaming/sabotage under competitive stakes. Plausible reconciliation: comparison to an exemplar standard helps when stakes are professional-norm-based; comparison as interpersonal rank hurts when pay/status is on the line - but no study directly adjudicates this. - DISAGREEMENT: Self-assessment vs external feedback: Dow 2012 found rubric self-assessment equal to external expert feedback for quality improvement (arguing feedback machinery could be replaced by surfaced rubrics), but the same study found self-ratings inflated by 1.8/9 points and education literature (Winstone 2016 recipience work) holds self-assessment only works when later external verification is believed to occur. External QA may be load-bearing as a credibility backstop even if the information could be self-generated. - DISAGREEMENT: Effect magnitude: education meta-analyses report medium standardized effects (d approximately 0.48), while the healthcare A&F median is a small 2.7% absolute improvement on already-trained professionals. For skilled adult annotators the healthcare prior (small median, heterogeneous, concentrated in low performers) is likely the better calibration than the education prior. - GAP: Still no longitudinal study of QA feedback effects on paid ANNOTATION workers' repeat-error rates: Dow 2012 explicitly did not measure long-term learning, and no 2024-2026 annotation-platform study of feedback efficacy was found (searched Exa, Tavily, OpenAlex/Crossref). The original sweep's gap is real; what changed is that adjacent causal literature (Cochrane A&F 2025, crowdwork RCT) transfers with stated caveats. - GAP: No study found on AI-GENERATED feedback to human annotators under pay-linked evaluation - the exact AutoQA deployment condition. Closest analogs are AI-feedback-to-employees work (e.g., Tong et al. 2021 SMJ, disclosure reduces effect) which was not deep-dived here. - GAP: Charness et al. 2014 and Barankay 2012 primaries were not opened (claims verified only through a peer-reviewed secondary synthesis); exact effect sizes for gaming-offset were not extracted. - GAP: Feedback-specificity tradeoff (Goodman & Wood 2004/2011: high specificity aids immediate performance but can impair exploration and transfer to novel cases) was identified in the Anseel & Sherf reference base but not independently verified - relevant to whether highly prescriptive AutoQA feedback creates template-following monoculture (Q8) and worth one follow-up read. - GAP: The 2025 individual-differences meta-analysis (Condrea & Iliescu, EJWOP, 2025-12-16) on reactions to feedback was located but is too new to have accessible full text; could sharpen per-attempter moderation (feedback orientation, self-esteem) but was not extractable.",
   "source": {
    "raw": "Feedback quality and performance in organisations (Leadership Quarterly; synthesizing Charness et al. 2014, Tafkov 2013, Barankay 2012) | https://www.sciencedirect.com/science/article/abs/pii/S1048984321000394 | 2021 (synthesizing 2012-2016 primaries) | academic",
    "title": "Feedback quality and performance in organisations (Leadership Quarterly; synthesizing Charness et al. 2014, Tafkov 2013, Barankay 2012)",
    "url": "https://www.sciencedirect.com/science/article/abs/pii/S1048984321000394",
    "date": "2021 (synthesizing 2012-2016 primaries)",
    "type": "academic"
   },
   "venue": {
    "peer_reviewed": null,
    "marker": "academic"
   },
   "same_source_claims": [],
   "in_corpus_verification": null,
   "external_verification": {
    "verdict": "not_retrievable",
    "checked": [],
    "not_visible": [
     "feedback quality and performance in organisations",
     "relative rank feedback raising output (Charness et al. 2014; Tafkov 2013)",
     "sabotage/cheating offsetting positive effects",
     "Barankay 2012 - removing rank feedback improving performance"
    ],
    "quote": "",
    "notes": "WebFetch returned HTTP 403 Forbidden. curl -sL with a browser user-agent returned a Cloudflare bot-challenge shell (title 'ScienceDirect'; contains __cf_chl_tk tokens, 'Captcha', 'Enable JavaScript', meta http-equiv refresh 360, noscript) with zero article content - no abstract, authors, journal, or any mention of rank feedback / Charness / Barankay / Tafkov / Leadership Quarterly. Claim could not be verified against the source per the WebFetch->curl->not_retrievable fallback chain.",
    "checked_at": "2026-07-15",
    "retrieval": "failed"
   },
   "source_retrieval_meta": {
    "retrieval": "failed",
    "fetched": [
     "https://www.sciencedirect.com/science/article/abs/pii/S1048984321000394"
    ],
    "resolved_title": null,
    "resolved_date": null,
    "title_match": false
   },
   "used_on": [],
   "decision_use": {
    "grade": "A",
    "basis": "human incentive result",
    "measured_on": "adjacent-domain",
    "note": "Rank feedback under incentives induces gaming/sabotage. Benchmark against exemplar work, never ranked people.",
    "models_measured": []
   }
  },
  {
   "id": "DA-01",
   "domain": "durable-anchors",
   "area": null,
   "claim": "NYC Local Law 144 of 2021, enforced from July 5, 2023, prohibits employers and employment agencies from using an automated employment decision tool unless it has undergone an independent bias audit within one year of use, with a summary of results publicly posted, and requires notice to candidates and employees.",
   "load_bearing": true,
   "evidence": "Provenance: proposed from model knowledge during the 2026-07-15 validity adjudication (the corpus's legal floor was EU-only); admitted only after source verification against the NYC DCWP's official page/rules. Decision relevance: if AutoQA verdicts influence pay, standing, or continued engagement of NYC-based workers and function as an employment decision tool, bias-audit and notice obligations attach - the D5 workforce census must be two-sided (EU and US), not EU-only.",
   "implication": "D5 gates on a two-sided workforce census; US-side counsel review joins EU counsel review as a launch precondition where scope conditions may hold.",
   "source": {
    "title": "NYC Local Law 144 - Automated Employment Decision Tools (DCWP)",
    "url": "https://www.nyc.gov/site/dca/about/automated-employment-decision-tools.page",
    "date": "2023-07-05 (enforcement start)",
    "type": "law"
   },
   "venue": {
    "peer_reviewed": null,
    "marker": "law"
   },
   "same_source_claims": [],
   "in_corpus_verification": null,
   "external_verification": {
    "verdict": "partially_confirmed",
    "checked": [
     "Named as 'Local Law 144 of 2021' - confirmed verbatim (Statement of Basis and Purpose).",
     "Prohibits employers AND employment agencies from using an AEDT unless conditions met - confirmed verbatim.",
     "Bias audit required 'within one year of the use of the tool' - confirmed verbatim.",
     "Independent bias audit - confirmed: rules define 'Independent Auditor' and require independence ('an \"independent auditor\" may not be employed or have a financial interest in an employer').",
     "Public posting of a summary of results - confirmed: Sec. 5-303 Published Results requires making 'publicly available on the employment section of their website ... The date of the most recent bias audit of the AEDT and a summary of the results'.",
     "Notice to candidates/employees - confirmed: Sec. 5-304 'Notice to Candidates and Employees' (per Sec. 20-871(b) of the Code), at least 10 business days before use."
    ],
    "not_visible": [
     "Enforcement start date 'July 5, 2023' - NOT present in the retrieved text. This document is the rule adoption (references only 2022 proposal dates and Jan 23, 2023 hearing); the July 5, 2023 enforcement date is not stated here and the primary nyc.gov page that would carry it returned 403.",
     "'penalties attach per violation' - NOT visible. No penalty/civil-penalty/fine language appears in the retrieved rule text (penalties are set in Administrative Code Sec. 20-872, which is not quoted in this document)."
    ],
    "quote": "Local Law 144 of 2021 prohibits employers and employment agencies from using an automated employment decision tool unless the tool has been subject to a bias audit within one year of the use of the tool, information about the bias audit is publicly available, and certain notices have been provided to employees or job candidates.",
    "notes": "Rule-adoption document confirms the substance of the prohibition, the one-year bias-audit condition, independent-auditor requirement, public posting (Sec. 5-303), and candidate/employee notice (Sec. 5-304). The two date/penalty specifics in the claim are outside this document's retrieved text. Verdict partially_confirmed on those two specifics only.",
    "checked_at": "2026-07-15",
    "retrieval": "landing"
   },
   "source_retrieval_meta": {
    "fetched": [
     "https://www.nyc.gov/site/dca/about/automated-employment-decision-tools.page (HTTP 403 Forbidden - not retrieved)",
     "https://rules.cityofnewyork.us/wp-content/uploads/2023/04/DCWP-NOA-for-Use-of-Automated-Employment-Decisionmaking-Tools-2.pdf (HTTP 200; 451 KB PDF; text extracted via pdftotext)"
    ],
    "resolved_title": "New York City Department of Consumer and Worker Protection - Notice of Adoption of Final Rule (rules implementing Local Law 144 of 2021 / Automated Employment Decision Tools; amends Title 6 RCNY, adding subchapter 5 Sec. Sec. 5-300 to 5-304)"
   },
   "used_on": [
    {
     "page": "research.html",
     "anchor": "research-cards",
     "label": "Research - findings by decision weight"
    },
    {
     "page": "decisions.html",
     "anchor": "d5",
     "label": "Decisions - where it runs and which law binds"
    }
   ],
   "decision_use": {
    "grade": "A",
    "basis": "law",
    "measured_on": "structural",
    "note": "Round-2 anchor: proposed from model knowledge 2026-07-15 under the 'relevant and certain' bar, then source-verified the same day. Not part of the original archive corpus; provenance disclosed by design.",
    "models_measured": []
   }
  },
  {
   "id": "DA-02",
   "domain": "durable-anchors",
   "area": null,
   "claim": "The EEOC Uniform Guidelines on Employee Selection Procedures (1978, codified at 29 CFR Part 1607) require that selection procedures with adverse impact be validated (criterion, content, or construct strategies), apply to any procedure used as a basis for an employment decision, and adopt the four-fifths rule as the practical test of adverse impact.",
   "load_bearing": true,
   "evidence": "Provenance: proposed from model knowledge during the 2026-07-15 validity adjudication; admitted only after source verification against the eCFR text. Decision relevance: consequential AutoQA verdicts over US workers are selection-procedure-shaped; validation obligations and adverse-impact monitoring are standards-grade constraints on D1 enforcement weight and A4 authority boundaries, independent of EU law.",
   "implication": "Where US workers face consequential verdicts, validity evidence and adverse-impact monitoring are compliance requirements, not research niceties - reinforcing measurement-first sequencing.",
   "source": {
    "title": "29 CFR Part 1607 - Uniform Guidelines on Employee Selection Procedures",
    "url": "https://www.ecfr.gov/current/title-29/subtitle-B/chapter-XIV/part-1607",
    "date": "1978 (current codification)",
    "type": "law"
   },
   "venue": {
    "peer_reviewed": null,
    "marker": "law"
   },
   "same_source_claims": [],
   "in_corpus_verification": null,
   "external_verification": {
    "verdict": "confirmed",
    "checked": [
     "Uniform Guidelines, 1978, codified at 29 CFR 1607 - confirmed: heading 'PART 1607 - UNIFORM GUIDELINES ON EMPLOYEE SELECTION PROCEDURES (1978)'.",
     "Adverse-impact procedures must be validated - confirmed: Sec. 1607.3(A) a procedure with adverse impact is 'considered to be discriminatory and inconsistent with these guidelines, unless the procedure has been validated'.",
     "Criterion / content / construct validity - confirmed: Sec. 1607.5(A) 'users may rely upon criterion-related validity studies, content validity studies or construct validity studies'.",
     "Scope - any employment decision - confirmed verbatim: Sec. 1607.2(B).",
     "Four-fifths (80%) rule as practical test of adverse impact - confirmed: Sec. 1607.4(D)."
    ],
    "not_visible": [],
    "quote": "Sec. 1607.2(B): 'These guidelines apply to tests and other selection procedures which are used as a basis for any employment decision.' Sec. 1607.4(D): a rate 'less than four-fifths (4/5) (or eighty percent) of the rate for the group with the highest rate' will 'generally be regarded by the Federal enforcement agencies as evidence of adverse impact.'",
    "notes": "All four sub-claims confirmed verbatim from Sec. 1607.2(B), Sec. 1607.3(A), Sec. 1607.4(D), and Sec. 1607.5(A). Fallback is the 2011 CFR edition; the guidelines themselves date to 1978 and the retrieved text is materially identical to the current codification. Sec. 1607.4(D) also notes smaller/greater differences may deviate from the four-fifths test where statistically significant or samples are small - consistent with 'practical test' framing in the claim.",
    "checked_at": "2026-07-15",
    "retrieval": "landing"
   },
   "source_retrieval_meta": {
    "fetched": [
     "https://www.ecfr.gov/current/title-29/subtitle-B/chapter-XIV/part-1607 (302 redirect to https://unblock.federalregister.gov/ - bot-block host, no content)",
     "https://www.govinfo.gov/content/pkg/CFR-2011-title29-vol4/xml/CFR-2011-title29-vol4-part1607.xml (HTTP 200; regulatory text retrieved)"
    ],
    "resolved_title": "PART 1607 - UNIFORM GUIDELINES ON EMPLOYEE SELECTION PROCEDURES (1978) (29 CFR Part 1607)"
   },
   "used_on": [
    {
     "page": "research.html",
     "anchor": "research-cards",
     "label": "Research - findings by decision weight"
    },
    {
     "page": "pilot.html",
     "anchor": "main",
     "label": "Pilot - gates and objectives"
    },
    {
     "page": "decisions.html",
     "anchor": "a4",
     "label": "Decisions - authority boundaries"
    },
    {
     "page": "decisions.html",
     "anchor": "d5",
     "label": "Decisions - where it runs and which law binds"
    }
   ],
   "decision_use": {
    "grade": "A",
    "basis": "law",
    "measured_on": "structural",
    "note": "Round-2 anchor: proposed from model knowledge 2026-07-15 under the 'relevant and certain' bar, then source-verified the same day. Not part of the original archive corpus; provenance disclosed by design.",
    "models_measured": []
   }
  },
  {
   "id": "DA-03",
   "domain": "durable-anchors",
   "area": null,
   "claim": "The multitask principal-agent model (Holmstrom & Milgrom, 1991) shows that when some task dimensions are measurable and others are not, high-powered incentives on the measured dimensions divert agent effort away from the unmeasured ones - and can make low-powered or no incentives optimal.",
   "load_bearing": true,
   "evidence": "Provenance: proposed from model knowledge during the 2026-07-15 validity adjudication; admitted only after source verification. Decision relevance: this is the theorem behind the feedback-monoculture and teaching-to-the-judge concerns - it bounds enforcement weight wherever quality is only partially measured (D1), and grounds the ban on stylistic guidance and rank-referenced feedback as incentive design rather than taste.",
   "implication": "Enforcement weight on partially-measured quality is bounded by theorem, not prudence; the kernel's K9/T7/T11 derivations cite this result as their economic formalization.",
   "source": {
    "title": "Multitask Principal-Agent Analyses: Incentive Contracts, Asset Ownership, and Job Design (JLEO 7)",
    "url": "https://academic.oup.com/jleo/article/7/special_issue/24/861573",
    "date": "1991",
    "type": "academic"
   },
   "venue": {
    "peer_reviewed": true,
    "marker": "academic"
   },
   "same_source_claims": [],
   "in_corpus_verification": null,
   "external_verification": {
    "verdict": "partially_confirmed",
    "checked": [
     "Author/title/journal/year identity of the source - confirmed: Holmstrom & Milgrom, 'Multitask Principal-Agent Analyses', JLEO, 1991 (vol 7, special issue, pp 24-52).",
     "The anchor points to the correct paper cited in the claim - confirmed."
    ],
    "not_visible": [
     "The substantive core result - that when some tasks/dimensions are measurable and others are not, high-powered incentives on measured dimensions divert effort away from unmeasured dimensions, and low-powered or no incentives can be optimal - is NOT present in any retrieved text. The OUP landing page has no abstract ('only available as a PDF') and the JSTOR fallback returned 403. Per the retrieved-text-only standard, the claim's economic content cannot be verified from what was retrieved (though it is the well-known finding of this paper)."
    ],
    "quote": "Title: 'Multitask Principal-Agent Analyses: Incentive Contracts, Asset Ownership, and Job Design.' Authors: Bengt Holmstrom (Yale University) and Paul Milgrom (Stanford University). The Journal of Law, Economics, and Organization, Volume 7, special_issue, Pages 24-52, published 01 January 1991. The page states: 'This content is only available as a PDF.'",
    "notes": "Bibliographic anchor is solid and correctly identifies the cited work. The claim's substantive proposition is behind the paywall on both allowed surfaces and could not be quoted; verdict partially_confirmed strictly because the result text is not_visible, not because of any contradiction.",
    "checked_at": "2026-07-15",
    "retrieval": "landing"
   },
   "source_retrieval_meta": {
    "fetched": [
     "https://academic.oup.com/jleo/article/7/special_issue/24/861573 (HTTP 200; landing page with metadata; no abstract or body text)",
     "https://www.jstor.org/stable/764957 (HTTP 403 Forbidden - not retrieved)"
    ],
    "resolved_title": "Multitask Principal-Agent Analyses: Incentive Contracts, Asset Ownership, and Job Design - Bengt Holmstrom & Paul Milgrom, The Journal of Law, Economics, and Organization, Vol. 7, Issue special_issue, pp. 24-52 (1991); DOI 10.1093/jleo/7.special_issue.24"
   },
   "used_on": [
    {
     "page": "research.html",
     "anchor": "research-cards",
     "label": "Research - findings by decision weight"
    },
    {
     "page": "decisions.html",
     "anchor": "d1",
     "label": "Decisions - enforcement weight"
    }
   ],
   "decision_use": {
    "grade": "A",
    "basis": "theorem (canonical economics)",
    "measured_on": "structural",
    "note": "Round-2 anchor: proposed from model knowledge 2026-07-15 under the 'relevant and certain' bar, then source-verified the same day. Not part of the original archive corpus; provenance disclosed by design.",
    "models_measured": []
   }
  }
 ]
}