Source-level claims and grades for design weight. Not a substitute for measuring the live instrument. Live map: the live system. CTA: home.
One constraint applies throughout: no published study tests the models we would actually deploy on this exact task.
Every capability-dependent bar gets measured there before it binds anything.
Anchors admitted after the sweep (provenance disclosed)
[DA-01] durable-anchors - load-bearing - grade A (holds regardless of model progress) - re-check 2026-07-15: confirmed with caveats
NYC Local Law 144 of 2021, enforced from July 5, 2023, prohibits employers and employment agencies from using an automated employment decision tool unless it has undergone an independent bias audit within one year of use, with a summary of results publicly posted, and requires notice to candidates and employees.
Source: NYC Local Law 144 - Automated Employment Decision Tools (DCWP) (2023-07-05 (enforcement start), law)
What the source itself says (retrieved quote)
Local Law 144 of 2021 prohibits employers and employment agencies from using an automated employment decision tool unless the tool has been subject to a bias audit within one year of the use of the tool, information about the bias audit is publicly available, and certain notices have been provided to employees or job candidates.
Re-check detail
Checked and matching: Named as 'Local Law 144 of 2021' - confirmed verbatim (Statement of Basis and Purpose).; Prohibits employers AND employment agencies from using an AEDT unless conditions met - confirmed verbatim.; Bias audit required 'within one year of the use of the tool' - confirmed verbatim.; Independent bias audit - confirmed: rules define 'Independent Auditor' and require independence ('an "independent auditor" may not be employed or have a financial interest in an employer').; Public posting of a summary of results - confirmed: Sec. 5-303 Published Results requires making 'publicly available on the employment section of their website ... The date of the most recent bias audit of the AEDT and a summary of the results'.; Notice to candidates/employees - confirmed: Sec. 5-304 'Notice to Candidates and Employees' (per Sec. 20-871(b) of the Code), at least 10 business days before use. || Asserted but not visible in retrievable text: Enforcement start date 'July 5, 2023' - NOT present in the retrieved text. This document is the rule adoption (references only 2022 proposal dates and Jan 23, 2023 hearing); the July 5, 2023 enforcement date is not stated here and the primary nyc.gov page that would carry it returned 403.; 'penalties attach per violation' - NOT visible. No penalty/civil-penalty/fine language appears in the retrieved rule text (penalties are set in Administrative Code Sec. 20-872, which is not quoted in this document). || Re-check notes: Rule-adoption document confirms the substance of the prohibition, the one-year bias-audit condition, independent-auditor requirement, public posting (Sec. 5-303), and candidate/employee notice (Sec. 5-304). The two date/penalty specifics in the claim are outside this document's retrieved text. Verdict partially_confirmed on those two specifics only.
Decision use: holds regardless of model progress. Measured on structural. Round-2 anchor: proposed from model knowledge 2026-07-15 under the 'relevant and certain' bar, then source-verified the same day. Not part of the original archive corpus; provenance disclosed by design.
Cited at: Research - findings by decision weight - Decisions - where it runs and which law binds
[DA-02] durable-anchors - load-bearing - grade A (holds regardless of model progress) - re-check 2026-07-15: confirmed
The EEOC Uniform Guidelines on Employee Selection Procedures (1978, codified at 29 CFR Part 1607) require that selection procedures with adverse impact be validated (criterion, content, or construct strategies), apply to any procedure used as a basis for an employment decision, and adopt the four-fifths rule as the practical test of adverse impact.
Source: 29 CFR Part 1607 - Uniform Guidelines on Employee Selection Procedures (1978 (current codification), law)
What the source itself says (retrieved quote)
Sec. 1607.2(B): 'These guidelines apply to tests and other selection procedures which are used as a basis for any employment decision.' Sec. 1607.4(D): a rate 'less than four-fifths (4/5) (or eighty percent) of the rate for the group with the highest rate' will 'generally be regarded by the Federal enforcement agencies as evidence of adverse impact.'
Re-check detail
Checked and matching: Uniform Guidelines, 1978, codified at 29 CFR 1607 - confirmed: heading 'PART 1607 - UNIFORM GUIDELINES ON EMPLOYEE SELECTION PROCEDURES (1978)'.; Adverse-impact procedures must be validated - confirmed: Sec. 1607.3(A) a procedure with adverse impact is 'considered to be discriminatory and inconsistent with these guidelines, unless the procedure has been validated'.; Criterion / content / construct validity - confirmed: Sec. 1607.5(A) 'users may rely upon criterion-related validity studies, content validity studies or construct validity studies'.; Scope - any employment decision - confirmed verbatim: Sec. 1607.2(B).; Four-fifths (80%) rule as practical test of adverse impact - confirmed: Sec. 1607.4(D). || Re-check notes: All four sub-claims confirmed verbatim from Sec. 1607.2(B), Sec. 1607.3(A), Sec. 1607.4(D), and Sec. 1607.5(A). Fallback is the 2011 CFR edition; the guidelines themselves date to 1978 and the retrieved text is materially identical to the current codification. Sec. 1607.4(D) also notes smaller/greater differences may deviate from the four-fifths test where statistically significant or samples are small - consistent with 'practical test' framing in the claim.
Decision use: holds regardless of model progress. Measured on structural. Round-2 anchor: proposed from model knowledge 2026-07-15 under the 'relevant and certain' bar, then source-verified the same day. Not part of the original archive corpus; provenance disclosed by design.
Cited at: Research - findings by decision weight - Pilot - gates and objectives - Decisions - authority boundaries - Decisions - where it runs and which law binds
[DA-03] durable-anchors - load-bearing - grade A (holds regardless of model progress) - re-check 2026-07-15: confirmed with caveats
The multitask principal-agent model (Holmstrom & Milgrom, 1991) shows that when some task dimensions are measurable and others are not, high-powered incentives on the measured dimensions divert agent effort away from the unmeasured ones - and can make low-powered or no incentives optimal.
Source: Multitask Principal-Agent Analyses: Incentive Contracts, Asset Ownership, and Job Design (JLEO 7) (1991, academic)
What the source itself says (retrieved quote)
Title: 'Multitask Principal-Agent Analyses: Incentive Contracts, Asset Ownership, and Job Design.' Authors: Bengt Holmstrom (Yale University) and Paul Milgrom (Stanford University). The Journal of Law, Economics, and Organization, Volume 7, special_issue, Pages 24-52, published 01 January 1991. The page states: 'This content is only available as a PDF.'
Re-check detail
Checked and matching: Author/title/journal/year identity of the source - confirmed: Holmstrom & Milgrom, 'Multitask Principal-Agent Analyses', JLEO, 1991 (vol 7, special issue, pp 24-52).; The anchor points to the correct paper cited in the claim - confirmed. || Asserted but not visible in retrievable text: The substantive core result - that when some tasks/dimensions are measurable and others are not, high-powered incentives on measured dimensions divert effort away from unmeasured dimensions, and low-powered or no incentives can be optimal - is NOT present in any retrieved text. The OUP landing page has no abstract ('only available as a PDF') and the JSTOR fallback returned 403. Per the retrieved-text-only standard, the claim's economic content cannot be verified from what was retrieved (though it is the well-known finding of this paper). || Re-check notes: Bibliographic anchor is solid and correctly identifies the cited work. The claim's substantive proposition is behind the paywall on both allowed surfaces and could not be quoted; verdict partially_confirmed strictly because the result text is not_visible, not because of any contradiction.
Decision use: holds regardless of model progress. Measured on structural. Round-2 anchor: proposed from model knowledge 2026-07-15 under the 'relevant and certain' bar, then source-verified the same day. Not part of the original archive corpus; provenance disclosed by design.
Cited at: Research - findings by decision weight - Decisions - enforcement weight
Judge capability studies
[AJ-01] academic-judges - load-bearing - grade C (superseded-generation number; mechanism only) - re-check 2026-07-15: confirmed
On hard, objectively-verifiable judging tasks (JudgeBench, ICLR 2025), frontier reasoning models dominate dedicated judge models: o1-preview scored 75.4% overall while GPT-4o scored 50.9-56.6% (near random), the best reward model (Skywork-Reward-Gemma-2-27B) hit 64.3%, and fine-tuned judges like PandaLM fell at or below random.
Source: JudgeBench: A Benchmark for Evaluating LLM-Based Judges (ICLR 2025) + EmergentMind topic synthesis (updated Jan 2026) (2024-10 (ICLR 2025; synthesis updated 2026-01), academic)
What the source itself says (retrieved quote)
our dataset consists of a total of 350 questions: 154 in Knowledge, 98 in Reasoning, 56 in Mathematics, and 42 in Coding.
Re-check detail
Checked and matching: o1-preview Overall 75.43 (Table 2, Arena-Hard Judge with o1-preview backbone); o1-preview Math 85.71 and Coding 85.71 (Table 2) = 85.7% math/coding; Vanilla (GPT-4o) Overall 50.86; Arena-Hard Judge (GPT-4o) Overall 56.57 (Table 1/2) = 50.9-56.6%; Skywork-Reward-Gemma-2-27B Overall 64.29, highest reward model (Table 3); PandaLM Overall 13.14 (Table 1) = well below random; ChatEval Overall 34.00 (Table 1); 350 questions: 154 Knowledge, 98 Reasoning, 56 Mathematics, 42 Coding (p.6); p.8: GPT-4o 'achieving accuracy no better than random guessing when using the vanilla prompt' || Re-check notes: All headline numbers in the claim match the tables exactly. Minor: the paper's reward-model overall range is stated as 'approximately 59% to 64%' (lowest = 59.43), so the evidence field's '60-64%' is a slight low-end rounding; does not affect the claim body. arXiv HTML full text 404'd for both v1/v2, so numbers were read from the PDF (allowed variant b).
Research-sweep audit (2026-07-14): confirmed
Every quantitative claim checks against the paper's own tables (arXiv:2410.12784 v2, ICLR 2025 camera-ready), read directly from the PDF: GPT-4o vanilla 50.86% / Arena-Hard 56.57% overall (Table 1), o1-preview 75.43% overall with 85.71% on both math and coding (Table 2), Skywork-Reward-Gemma-2-27B 64.29% as best reward model (Table 3, full range 59.43-64.29, paper says "approximately 59% to 64%" so the claim's "60-64%" is a minor round-up at the low end), PandaLM 13.14% (far below random; paper: all fine-tuned judges except Skywork significantly below random), ChatEval 34.00%. 350 verified pairs across knowledge (154)/reasoning (98)/math (56)/coding (42) confirmed - note this is the GPT-4o-generated split; a separate 270-pair Claude-3.5-Sonnet split exists. Dates confirmed: arXiv v1 2024-10-16, ICLR 2025, EmergentMind synthesis updated 2026-01-09. Two contextual nuances, neither refuting: (1) in the v2 tables o1-preview is not the single best judge - o3-mini (high) scores 80.86% and DeepSeek-R1 73.14%, which strengthens rather than weakens the "reasoning models dominate" thesis; (2) 2025-2026 follow-on work (J1, RM-R1, meta-judging, RRD) has pushed JudgeBench accuracies to 77-81%+, so "GPT-4o near random" describes the late-2024 snapshot of non-reasoning judges, not the current frontier. The claim as stated about the paper's findings is accurate.
Decision use: superseded-generation number; mechanism only. Measured on model-outputs (models: o1-preview, GPT-4o, Skywork-Reward-Gemma-2-27B, PandaLM, Skywork-Reward-27B). Late-2024 snapshot (o1-preview vs GPT-4o, PandaLM). Direction (reasoning models dominate trained judges) has held through 2026 follow-ons; the specific accuracies are two generations stale. Re-rank at the deployment-tier bakeoff.
Cited at: Decisions - judge sourcing
[AJ-02] academic-judges - load-bearing - grade A (holds regardless of model progress) - re-check 2026-07-15: confirmed
Standard LLM-judge validation via forced-choice human gold labels can be heavily biased when rating criteria admit multiple valid interpretations (an empirical result, not a theorem): across 11 real-world rating tasks and 9 commercial LLMs, forced-choice validation selected judge systems performing up to 31% worse than validation using multi-label 'response set' ratings that model indeterminacy.
Source: Validating LLM-as-a-Judge Systems under Rating Indeterminacy (NeurIPS 2025) (2025-09-19, academic)
What the source itself says (retrieved quote)
The experiments involve "11 real-world rating tasks and 9 commercial LLMs" ... forced-choice validation selects judge systems "performing as much as 31% worse than judge systems selected by our approach."
Re-check detail
Checked and matching: 11 real-world rating tasks and 9 commercial LLMs; forced-choice validation selects judge systems performing as much as 31% worse than the paper's response-set approach; rating indeterminacy defined as criteria admitting multiple valid interpretations; forced-choice differences 'can heavily bias LLM-as-a-judge validation'; solution uses multi-label 'response set' ratings || Asserted but not visible in retrievable text: exact source date 2025-09-19 (page shows only 'NeurIPS 2025 Poster') || Re-check notes: Core assertion and both headline numbers (11 tasks, 9 LLMs, up-to-31% gap) match. Claim's '31% worse than multi-label response-set ratings' is consistent with retrieved '31% worse than judge systems selected by our approach' (their approach = multi-label response set). Authors match: Guerdan, Barocas, Holstein, Wallach, Wu, Chouldechova.
Research-sweep audit (2026-07-14): partially_confirmed
Core claim confirmed against the actual NeurIPS poster page (https://neurips.cc/virtual/2025/poster/117308): authors, rating-indeterminacy framing, forced-choice vs multi-label response-set elicitation, 11 tasks, 9 commercial LLMs, and the "as much as 31% worse" figure all match verbatim. Two corrections: (1) "provably biased" overstates the paper's own language - it says differing resolution of indeterminacy "can heavily bias" validation, and the 31% gap is empirical, not a theorem (the theory links performance measures/elicitation schemes); (2) the 2025-09-19 date is not shown on the poster page - plausibly the acceptance date, but unverified; the paper dates to arXiv March 2025 and NeurIPS Dec 2025, with an ML@CMU blog post 2025-12-09. Search for later work (as of July 2026) found follow-on papers (IRT-based judge reliability, arXiv 2602.00521; grading-scale alignment, arXiv 2601.03444) that build on, not refute or supersede, this result. Code: https://github.com/lguerdan/indeterminacy.
Corrected statement: Guerdan, Barocas, Holstein, Wallach, Wu & Chouldechova (NeurIPS 2025; arXiv:2503.05965, first posted March 2025) show that standard LLM-judge validation via forced-choice human gold labels can be heavily biased when rating criteria admit multiple valid interpretations ("rating indeterminacy"): across 11 real-world rating tasks and 9 commercial LLMs, forced-choice-based validation selected judge systems performing as much as 31% worse than those selected by their framework using multi-label "response set" ratings. The bias mechanism is supported by a theoretical framework, but the 31% suboptimality figure is an empirical finding, not a proof.
Decision use: holds regardless of model progress. Measured on structural. A result about how gold labels are elicited and aggregated, not about model skill. Applies to validating Fable/GPT-5.5-class judges exactly as it did to GPT-4o-class.
Cited at: Overview - the five authorizations - System - claim types - Pilot - gates and objectives - Decisions - the standard
[AJ-03] academic-judges - load-bearing - grade B (current-generation measurement) - re-check 2026-07-15: confirmed
The largest systematic judge meta-evaluation to date (21 judges, 9 providers, ~541k judgments, including April-2026 frontier models) found raw exact-match agreement universally overstates judge ability - chance-corrected Cohen's kappa is 33-41 percentage points lower on MT-Bench - and judge rankings shift by up to 14 positions depending on which benchmark you use.
Source: Reliability without Validity: A Systematic, Large-Scale Evaluation of LLM-as-a-Judge Models Across Agreement, Consistency, and Bias (2026-06-17, academic)
What the source itself says (retrieved quote)
kappa deflation between exact match and Cohen's kappa is universal (33--41 pp on MT-Bench) ... judge rankings shift by up to 14 positions across benchmarks
Re-check detail
Checked and matching: 21 judges from nine providers; approximately 541,000 individual judgments; 118 runs; kappa deflation ... universal (33--41 pp on MT-Bench); judge rankings shift by up to 14 positions across benchmarks; high test--retest reliability (>0.95) with severe position bias (>0.10) in two production-deployed judges; verbosity bias is small (<0.011); consistency--bias paradox; Minimum Viable Validation Protocol; April 2026 frontier || Re-check notes: All four cohort-wide findings and every headline number verified in the abstract; verbosity bias <0.011, MVVP, and consistency-bias paradox all present verbatim.
Research-sweep audit (2026-07-14): confirmed
Fetched https://arxiv.org/abs/2606.19544 directly. Every quantitative element checks out against the abstract: 21 judges from 9 providers, 118 runs / ~541,000 judgments across MT-Bench/JudgeBench/RewardBench; kappa deflation of 33-41 percentage points on MT-Bench across all 21 models; ranking shifts up to 14 positions across benchmarks; consistency-bias paradox (test-retest reliability >0.95 with position bias >0.10 in two production judges); verbosity bias <0.011; Minimum Viable Validation Protocol; April 2026 frontier models included. Submission date June 17, 2026 matches; authors Norman, Rivera, Hughes match. "Largest to date" is the authors' own framing, but two independent searches (July 14, 2026) found no larger or superseding meta-evaluation; follow-on work (AURA arXiv 2606.19714, Apple correlated-panels paper, "Below the Reliability Floor" on OpenReview) complements rather than supersedes it. Caveats: preprint with 0 citations; authors themselves limit dataset authority to the March-April 2026 measurement window due to hosted-endpoint drift.
Decision use: current-generation measurement. Measured on model-outputs. Includes April-2026 frontier models; the kappa-deflation statistics are properties of agreement measurement and transfer; per-judge rankings do not. Deployment tier and human-writeup QA untested.
Same underlying source as [CM-09] [CT-01] - repetition across reports is not independent corroboration.
Cited at: Overview - what the evidence supports - Summary - measurement point - Research - findings by decision weight - Pilot - gates and objectives - Decisions - settled constraints
[AJ-04] academic-judges - load-bearing - grade B (current-generation measurement) - re-check 2026-07-15: confirmed
At the rubric level - the exact granularity AutoQA would operate at - even frontier judges achieve only ~55-56% balanced accuracy on hard rubric-verification instances (GPT-4o 55.97%, Claude-Sonnet-4.5 55.65%), but rubric-level evaluation with explicit chain-of-thought reasoning beats checklist-level evaluation by 7-12 percentage points and reduces cross-judge variance.
Source: RubricEval: A Rubric-Level Meta-Evaluation Benchmark for LLM Judges in Instruction Following (2026-03-26, academic)
What the source itself says (retrieved quote)
GPT-4o achieves merely 55.97% balanced accuracy, and Claude-Sonnet-4.5 reaches 55.65% ... while checklist-level achieves only 69.90% and 70.44%--a gap of 7-12 points ... Human-RAF agreement reaches 85.0% accuracy with Cohen's kappa = 0.702
Re-check detail
Checked and matching: GPT-4o 55.97% balanced accuracy on Hard subset; Claude-Sonnet-4.5 55.65% on Hard subset; rubric-level with reasoning beats checklist-level by a gap of 7-12 points; combining rubric+checklist reduces inter-judge (cross-judge) variance; Easy subset ~90% (Qwen3-235B / gpt-oss-120b ~90%, Table 2: 89.87 / 89.55); RAF human agreement 85.0% accuracy, Cohen's kappa 0.702; 3,486 quality-controlled instances with Easy/Hard subsets || Re-check notes: Given alphaXiv URL returned HTTP 403 to WebFetch; curl returned only the SPA shell (page title and meta description confirm the exact title/authors). Verified against the allowed arXiv variant (abs + HTML full text) since the URL contains arXiv ID 2603.25133. All core numbers confirmed verbatim in the HTML full text. Two evidence-field nuances: (a) the paper does NOT explicitly assert reasoning models like o3 outperform instruct models as a class, though o3 leads the Hard split at 84.81 BAcc and is the best single judge at 79.9% on disputed cases; (b) evidence equates 'Compositional' with 'conditional-logic', but the paper shows Compositional *instructions* are hardest while the Conditional Logic *rubric type* actually scores relatively well (Table 4). Neither touches the core claim.
Decision use: current-generation measurement. Measured on model-outputs (models: GPT-4o, Claude-Sonnet-4.5, o3). March-2026 on GPT-4o / Claude-Sonnet-4.5 / o3. One tier behind deployment; hard-subset ~55% is a prior, not a ceiling, for Fable/GPT-5.5-class. The +7-12pp per-criterion-reasoning effect is the durable part.
Same underlying source as [F26-06] - repetition across reports is not independent corroboration.
Cited at: Overview - what the evidence supports - Research - findings by decision weight
[AJ-05] academic-judges - grade B (current-generation measurement) - re-check 2026-07-15: confirmed
On subjective rubrics, LLM judges' evaluation axis is nearly orthogonal to the human axis (87-89 degrees vs 78-81 degrees human-to-human), judges use only 0.3-0.5x the human score spread, and inter-LLM agreement (r~0.35) exceeds LLM-human agreement (r~0.27-0.32) - while on a rubric with a verifiable factual answer the same judges fall back into the human range (58.5 degrees, r=0.519).
Source: The Geometry of LLM-as-Judge: Why Inter-LLM Consensus Is Not Human Alignment (2026-06-02, academic)
What the source itself says (retrieved quote)
87 deg--89 deg versus 78 deg--81 deg ... r_LL approx 0.35 versus r_LH approx 0.27--0.32 ... axis 58.5 deg; r_LH = 0.519 ... sigma_J / sigma_H approx 0.3--0.5
Re-check detail
Checked and matching: 41 LLM judges, 4 community-built Indic datasets, 8 Indic languages; judge/human score-spread ratio sigma_J/sigma_H approx 0.3-0.5; principal angle to human axis 87-89 deg vs human-human 78-81 deg; inter-LLM r_LL approx 0.35 vs LLM-human r_LH approx 0.27-0.32; verifiable factual rubric: axis 58.5 deg, r_LH = 0.519; fine-tuning recovers spread 0.32 -> 1.08 but axis stays 87-88 deg; geometry measured with bootstrap confidence intervals; post-hoc calibration on a small human-anchored set improves all rubrics || Re-check notes: Every headline number in the claim is present verbatim in the abstract, including the verifiable-vs-subjective contrast (58.5 deg / r=0.519) and the fine-tuning spread recovery (0.32->1.08) with a near-stationary axis. Scope caveat (Indic community-health data) also matches.
Decision use: current-generation measurement. Measured on model-outputs. June-2026, 41 judges. Subjective-vs-verifiable axis geometry is the mechanism; Indic community-health scope noted in corpus. Consensus-is-not-validity conclusion is method-level and durable.
[AJ-06] academic-judges - grade C (superseded-generation number; mechanism only) - re-check 2026-07-15: confirmed
Single-trial LLM judging is measurably noisy: across 29 tasks with 50 repeated trials, pairwise preferences flipped on average 13.6% of the time (28% of questions exceeded 20% flip rate), semantically equivalent prompt templates changed majority outcomes in 25% of tested cases, and ~11 repeated trials were needed for majority vote to recover the reference verdict with 95% probability.
Source: The Coin Flip Judge? Reliability and Bias in LLM-as-a-Judge Evaluation (2026-04-23, academic)
What the source itself says (retrieved quote)
pairwise preferences flip on average 13.6% of the time ... 28% of questions exceeding a 20% flip rate ... two OpenAI judge models (GPT-4o-mini and GPT-4.1-mini) ... a significant first-position bias (72% A-majority, p = 0.024)
Re-check detail
Checked and matching: 29 tasks spanning 10 categories; 50 pairwise trials (and 50 pointwise) per question; pairwise preferences flip on average 13.6%; 28% of questions exceed a 20% flip rate (one hits 56%); equivalent templates change majority outcomes in 25% of tested cases; 11 repeated trials needed for majority vote to match 50-trial reference at 95%; first-position bias 72% A-majority, p=0.024 (GPT-4o-mini); cross-judge agreement 76%, kappa 0.51; judges: GPT-4o-mini and GPT-4.1-mini; pairwise-pointwise gap: small non-significant scalar gaps despite winner picks || Asserted but not visible in retrievable text: the claim's own caveat that a 'Reliability-without-Validity cohort' shows some production judges reach test-retest >0.95 (external context, not from this paper) || Re-check notes: Every headline number in the claim matches the abstract verbatim. Single-provider mini-tier caveat is stated by the paper (recommends cross-provider replication as future work).
Decision use: superseded-generation number; mechanism only. Measured on model-outputs (models: GPT-4o-mini, GPT-4.1-mini). Measured on mini-tier judges (GPT-4o-mini / GPT-4.1-mini). Flip rates for deployment-tier reasoners are unknown and plausibly far lower; k-sampling remains cheap hygiene either way.
[AJ-07] academic-judges - grade C (superseded-generation number; mechanism only) - re-check 2026-07-15: confirmed
The broadest LLM-vs-human-judge study on existing human-annotated datasets (JUDGE-BENCH, 20 datasets) found the best model (GPT-4o) reached only kappa = 0.28 +/- 0.32 on categorical judgments and Spearman rho = 0.50 +/- 0.21 on graded ones, with enormous task-to-task variance - and LLMs agreed MORE with non-expert annotators than with experts.
Source: LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks (JUDGE-BENCH) (2025-07 (ACL 2025), academic)
What the source itself says (retrieved quote)
Table 1 lists GPT-4o Categorical Avg kappa as 0.28 +/- 0.32 ... GPT-4o Graded Avg rho as 0.50 +/- 0.21. 'Models exhibit higher correlation with non-expert evaluators.' 'Achieves negative correlations on DICES and Medical-safety due to overactive guardrails.'
Re-check detail
Checked and matching: Judge-Bench with 20 datasets (70k+ instances), 11 LLM judges; GPT-4o overall leader; GPT-4o Categorical Avg kappa = 0.28 +/- 0.32 (Table 1); GPT-4o Graded Avg rho = 0.50 +/- 0.21 (Table 1); higher correlation with non-expert than expert evaluators; every model struggles on engagingness; safety guardrail refusals -> negative correlations; task-specific human validation remains critical || Asserted but not visible in retrievable text: author name 'Bavaresco' (not present on this secondary paper-note page) || Re-check notes: Secondary source (a paper note dated 2026-05-08 discussing the ACL 2025 paper, arXiv 2406.18403 cited on the page); the claim self-identifies as read via a paper note. All headline numbers (kappa 0.28+/-0.32, rho 0.50+/-0.21) and the non-expert>expert and safety-refusal findings match verbatim. source_date '2025-07 (ACL 2025)' is the paper's venue; page metadata date is 2026-05-08.
Decision use: superseded-generation number; mechanism only. Measured on model-outputs (models: GPT-4o). GPT-4o-era kappa 0.28+/-0.32. Use as the canonical warning against percent-agreement narratives, not as an estimate of current judge-human agreement.
[AJ-08] academic-judges - grade B (current-generation measurement) - re-check 2026-07-15: confirmed
Preference leakage is a measured, large contamination bias: a judge favors outputs of models trained on data from a related generator by up to 28.7 percentage points (Preference Leakage Score), with a relatedness gradient - same model 23.6% average, fine-tuned descendants 19-22%, same family 2.8-8.9% - and judges cannot detect their own students (near-chance) while a BERT classifier can (82.4%).
Source: Preference Leakage: A Contamination Problem in LLM-as-a-judge (ICLR 2026) (2026-02 (ICLR 2026; arXiv 2025-02), academic)
What the source itself says (retrieved quote)
contextual calibration with an additional held-out set for bias adjustment is the most effective, reducing Error Bias from 17.8 to 7.3.
Re-check detail
Checked and matching: up to 28.7% preference leakage score (Table 1/2); same model 23.6% average PLS (Table 2); inheritance/fine-tuned 19.3% and 22.3% (in 19-22% band); same family 8.9% same-series vs 2.8% different-series; SFT 23.6% avg vs DPO 5.2%; judge self-detection near random (e.g. 41.0%, 52.0%); BERT classifier 82.4%; leakage proportional to synthetic-data amount, no clear threshold; style/format removal is largest decrease (e.g. 17.5% to 9.0%); contextual calibration Error Bias 17.8 to 7.3; prompting mitigation failed (17.8 to 18.3, slightly worse); accepted ICLR 2026 (abstract page) || Re-check notes: Abstract page confirms core + ICLR 2026 acceptance but shows no numbers; all 11 headline figures verified verbatim in the arxiv.org/html full text. Authors: Dawei Li et al. (Huan Liu group).
Decision use: current-generation measurement. Measured on model-outputs. ICLR 2026. Lineage-based contamination is structural to training pipelines and likely persists at the deployment tier; exact PLS magnitudes are tier-bound. Cross-family routing stays the defense.
Cited at: Pilot - gates and objectives
[AJ-09] academic-judges - grade C (superseded-generation number; mechanism only) - re-check 2026-07-15: confirmed with caveats
Dedicated trained judges advanced substantially in 2025: Meta's J1 (RL-trained thinking-judge, May 2025) at 32B outperforms o1-mini, o3, and 671B DeepSeek-R1 on some judge benchmarks; CompassJudger-2-7B (July 2025) matches far larger generalists on judge benchmarks; Skywork-Reward-V2-Llama-3.1-8B (July 2025) topped seven reward benchmarks including JudgeBench.
Source: J1: Incentivizing Thinking in LLM-as-a-Judge via Reinforcement Learning (with CompassJudger-2, Skywork-Reward-V2) (2025-05 to 2025-07, academic)
What the source itself says (retrieved quote)
J1-Qwen-32B, our multitasked pointwise and pairwise judge also outperforms o1-mini, o3, and a much larger 671B DeepSeek-R1
Re-check detail
Checked and matching: J1-Qwen-32B outperforms o1-mini, o3, and a much larger 671B DeepSeek-R1 on some benchmarks; RL framework teaching LLM judges to think before deciding (thinking-judge); unified format with verifiable rewards; reduces positional bias; trained at scales of 8B, 32B, and 70B; trained only on synthetic data; v1 submitted May 15 2025 || Asserted but not visible in retrievable text: '22K synthetic pairs' (abstract says only 'synthetic data', no count); 'test-time self-consistency' (abstract mentions iterative self-correction, not self-consistency); attribution to 'Meta' (authors are Meta FAIR but abstract does not state Meta); CompassJudger-2-7B claim and 80.9% JudgerBenchV2 (cites arXiv 2507.09104, a different URL not fetched); Skywork-Reward-V2-Llama-3.1-8B, 26M pairs, topping seven benchmarks incl JudgeBench (cites arXiv 2507.01352, a different URL not fetched) || Re-check notes: The J1-specific core assertion (32B beats o1-mini/o3/671B DeepSeek-R1, RL thinking-judge, verifiable rewards, position-bias mitigation) is confirmed. The claim bundles two other papers (CompassJudger-2, Skywork-Reward-V2) with their own arXiv IDs that the hard constraints forbid fetching; those portions are not verifiable from this source.
Decision use: superseded-generation number; mechanism only. Measured on model-outputs (models: o1-mini, o3, DeepSeek-R1, Skywork-Reward-V2-Llama-3.1-8B, Skywork-Reward-V2). Mid-2025 trained-judge landscape (J1, CompassJudger-2, Skywork-V2). This market moves quarterly; treat as a map of the design space, not a current ranking.
[AJ-10] academic-judges - grade C (superseded-generation number; mechanism only) - re-check 2026-07-15: confirmed
RewardBench 2 (Ai2, June 2025) made reward-model evaluation ~20 points harder via best-of-4 format (random = 25%) and unseen human prompts; top models score below 40% on Precise Instruction Following, yet benchmark scores correlate 0.87 (Pearson) with downstream best-of-N performance across 113 reward models.
Source: RewardBench 2: Advancing Reward Model Evaluation (2025-06-02, academic)
What the source itself says (retrieved quote)
There is only one correct chosen response, meaning the random baseline is 25% accuracy ... average score on downstream tasks with BoN sampling has a high Pearson correlation of 0.87 ... We evaluated 113 RMs
Re-check detail
Checked and matching: ~20 points harder: 'models score about 20 points on average lower on RewardBench 2 compared to the first RewardBench'; best-of-4 format with random baseline 25%; 1,865 prompts; leading models below 40% on Precise Instruction Following; Pearson 0.87 correlation with downstream BoN sampling; 113 reward models evaluated on BoN; six domains: factuality, precise instruction following, math, safety, focus, ties; sources new human prompts instead of existing prompts from downstream evaluations || Asserted but not visible in retrievable text: Lambert blog quotes 'weaker than expected relative to standard reward models' and 'name a color in the rainbow' (attributed in evidence to an external Lambert analysis post, not the arXiv text) || Re-check notes: Abstract page gave only the ~20-point figure and 'new human prompts'; all specific numbers (best-of-4, 25%, 1,865, <40% Precise IF, 0.87, 113, six domains) were verified verbatim from the arXiv HTML full text (v2). Every headline number in the claim matches.
Decision use: superseded-generation number; mechanism only. Measured on model-outputs. June-2025; top models <40% on Precise Instruction Following predates GPT-5-class. The warning (instruction-compliance checking is not solved) held then; current magnitude unknown - exactly why it needs its own gold set.
[AJ-11] academic-judges - grade C (superseded-generation number; mechanism only) - re-check 2026-07-15: confirmed
In expert-knowledge domains, LLM judges agreed with subject-matter experts only 68% (dietetics, registered dietitians) and 64% (mental health, clinical psychologists) on overall pairwise preference, with agreement varying further across domain-specific aspect questions.
Source: Limitations of the LLM-as-a-Judge Approach for Evaluating LLM Outputs in Expert Knowledge Tasks (2024-10 (IUI 2025), academic)
What the source itself says (retrieved quote)
SMEs agreed with LLM judges 68% of the time in the dietetics domain and 64% in mental health
Re-check detail
Checked and matching: 68% SME-LLM agreement in dietetics (registered dietitians); 64% SME-LLM agreement in mental health (clinical psychologists); overall pairwise preference comparison; agreement varied across domain-specific aspect questions; two fields: dietetics with registered dietitian experts, mental health with clinical psychologist experts; conclusion: keep human experts in the evaluation loop; LLMs alone lack depth for complex knowledge-specific tasks || Re-check notes: Abstract page fully supports the claim's core assertion and both headline percentages. Title and Oct-2024 date match; IUI 2025 venue appears in the Comments field.
Decision use: superseded-generation number; mechanism only. Measured on model-outputs. 2024 study (IUI 2025). The 64-68% expert-agreement figures predate reasoning-model maturity. Mechanism (expertise-heavy axes are the weak lane) is plausible but the deployment-tier boundary is E1/E3's to measure - in either direction.
Cited at: Research - findings by decision weight - Pilot - gates and objectives
[AJ-12] academic-judges - grade B (current-generation measurement) - re-check 2026-07-15: confirmed
LLM judges are measurably less reliable on long-form outputs: LongJudgeBench (June 2026) finds current judges unstable across real-world long-form scenarios requiring document-level assessment of organization, coverage, and cross-section consistency, and rubrics or references help but are 'not always sufficient'.
Source: Benchmarking LLM-as-a-Judge for Long-Form Output Evaluation (LongJudgeBench) (2026-06-01, academic)
What the source itself says (retrieved quote)
current LLM judges remain unstable across scenarios ... more complex document-level assessments of overall organization, task-relevant coverage and depth, cross-section consistency ... helpful but not always sufficient
Re-check detail
Checked and matching: judges unstable across long-form scenarios: 'a substantial reliability gap: current LLM judges remain unstable across scenarios'; document-level assessment of 'overall organization, task-relevant coverage and depth, cross-section consistency'; rubrics/references 'are helpful but not always sufficient'; code public at github.com/cjj826/LongJudgeBench; authors Chen et al.; submitted June 1 2026 (matches source_date 2026-06-01) || Asserted but not visible in retrievable text: evidence-field characterization 'First meta-evaluation benchmark targeting long-form judging specifically' -- abstract calls it 'a comprehensive benchmark' and does not explicitly claim to be first || Re-check notes: Claim text fully supported by the abstract. The only unsupported specific is the 'first benchmark' framing, which appears in the evidence field, not the claim; abstract says 'comprehensive benchmark' without a primacy claim.
Decision use: current-generation measurement. Measured on model-outputs. June-2026 benchmark. Long-form instability at the current tier; deployment-tier behavior untested.
Cited at: Research - findings by decision weight - System - claim types - Pilot - gates and objectives
Annotation science
[AQ-01] annotation-quality - load-bearing - grade A (holds regardless of model progress) - re-check 2026-07-15: confirmed with caveats
Roughly one-third of crowdworkers used LLMs on an LLM-advantaged text-production task (33-35% on Prolific, July 2023; ~34% Prolific self-report in a 2025 follow-up by Zhang et al.), and the best tested mitigations (direct request + copy-paste disable or image-only presentation) only cut usage roughly in half (27.6% to ~15.9%), never to zero, while also degrading response quality (direct requests reduced keyword retention 6.2%).
Source: Prevalence and Prevention of Large Language Model Use in Crowd Work (CACM; arXiv 2310.15683) (2025-02-18, academic)
What the source itself says (retrieved quote)
almost halved, dropping from 27.6% to 15.9%
Re-check detail
Checked and matching: Study #1: n=168 workers on Prolific, 3 July 2023, summarizing abstracts; Prevalence estimates 33.3% [25.9%,40.1%], 35.2% [29.8%,40.6%], 35.4% [27.8%,43.0%]; Study #2: n=720 users, 23 July 2023, 3x3 factorial (request x hurdle); Table 1a baseline 27.6% dropping to 15.9% (direct request + image, preventing copy-pasting); 'almost halved, dropping from 27.6% to 15.9%'; abstract: mitigations 'significantly reduce, but not eliminate' LLM use; directly requesting workers not to use LLMs decreased keyword retention by 6.2% (p=0.009); task = summarize medical paper abstracts (NEJM), per Materials 4.1 and Appendix A || Asserted but not visible in retrievable text: ~34% Prolific self-report in a 2025 follow-up by Zhang et al. (external source, not this paper); June 2026 community survey arXiv 2606.04924 corroboration (external source); the word 'preregistered' || Re-check notes: The Veselovsky-attributable core and every headline number (33-35% baseline, 27.6%->15.9%, 6.2% keyword retention, n=168/n=720, medical abstracts, halve-not-eliminate) are fully confirmed by the body. Note the ABSTRACT rounds prevalence to 'around 30%', but Study #1 body reports 33.3-35.4%, so the claim's '33-35%' is supported by the body. Downgraded to partially_confirmed only because the claim additionally cites two external sources (Zhang et al. 2025; arXiv 2606.04924) not verifiable from this document. Also: source_date 2025-02-18 refers to the CACM publication; the fetched arXiv preprint is dated 2023-10-24.
Research-sweep audit (2026-07-14): partially_confirmed
Core claim VERIFIED against the full PDF of arXiv 2310.15683 (Veselovsky, Horta Ribeiro, Cozzolino, Gordon, Rothschild, West): Study 1, n=168 Prolific workers, 3 July 2023, medical-abstract summarization; prevalence 33.3% [25.9,40.1] (classify-and-count), 35.2% [29.8,40.6] (probabilistic), 35.4% [27.8,43.0] (corrected) - matches "33.3-35.4%, CIs ~[26%,43%]". Study 2, n=720, 23 July 2023, 3x3 factorial (request: none/indirect/direct x hurdle: none/image/no-Ctrl-C+V); classifier estimate dropped 27.6% -> 15.9% (direct+image) and 15.8% (direct+Ctrl C+V) - matches "roughly halved, not eliminated" (paper: "reduced LLM use by nearly 50%, it could not fully prevent it"). Direct request decreased keyword retention by 6.2% (p=0.009) - exact match. Classifier was finetuned e5-base-v2, calibrated (Card & Smith), with self-report and high-precision heuristics (completion time + paste artifacts) - matches. The June 2026 community survey arXiv 2606.04924 (Velutharambath et al., 155 researchers, survey run Aug 2025-Mar 2026) does cite "Veselovsky et al. estimate 30-40%", cites "Zhang et al. (2025): 34% of Prolific participants self-report using LLMs for open-ended questions", and states mitigations "reduce rather than eliminate LLM-assisted responses" - so the follow-up attribution to Zhang et al. 2025 is consistent with that survey's citations (note: I verified the citation exists in the survey, not the underlying Zhang et al. paper itself). ISSUES: (1) Date "2025-02-18" is wrong for the arXiv link given - arXiv has only v1, dated 24 Oct 2023; the published version is Communications of the ACM 68(3):42-47, March 2025 issue (DOI 10.1145/3685527). (2) "Preregistered" is not supported by the paper text - no preregistration is mentioned in the arXiv PDF. (3) Minor nuance: "never to zero" - by the classifier measure true, but the high-precision heuristics table actually shows 0.0% in the direct+Ctrl C+V cell; the paper's own framing ("cannot fully prevent") still supports the claim's spirit. (4) 44% of surveyed researchers in the 2026 survey observed LLM use in their data - newer corroborating, not superseding, evidence; no later work found that overturns the prevalence estimates. Sources: https://arxiv.org/abs/2310.15683, https://arxiv.org/abs/2606.04924, https://dl.acm.org/doi/10.1145/3685527.
Corrected statement: Roughly one-third of Prolific crowdworkers used LLMs on an LLM-advantaged text-production task (33.3-35.4% across three estimators, July 2023, n=168; Veselovsky et al., arXiv 2310.15683 v1 posted 24 Oct 2023, published in Communications of the ACM 68(3), March 2025). In a second study (n=720, 3x3 factorial), the best mitigations (direct request + image presentation or + copy-paste disable) roughly halved classifier-estimated usage from 27.6% to 15.9%/15.8% without eliminating it, and directly requesting non-use reduced keyword retention by 6.2% (p=0.009). The studies were not described as preregistered. A June 2026 community survey of 155 researchers (arXiv 2606.04924, Velutharambath et al.) corroborates, citing Veselovsky et al.'s 30-40% estimate and Zhang et al. (2025)'s 34% Prolific self-report, and finding 44% of researchers observed LLM use in their crowdsourced free-text data; it echoes that mitigations reduce rather than eliminate LLM-assisted responses.
Decision use: holds regardless of model progress. Measured on adjacent-domain. Crowdworker LLM-use base rate (2023, replicated 2025). A floor that rises as tools improve; the decision input (assume pervasive assistance) only strengthens.
Cited at: Decisions - AI-assistance policy
[AQ-02] annotation-quality - load-bearing - grade A (holds regardless of model progress) - re-check 2026-07-15: source not retrievable
An autonomous LLM agent passed 99.8% of standard attention/quality checks (6,000 trials, 20 check types), evaded behavioral detectors (simulated reading times, mouse movement, typo-laden keystrokes), refused AI-revealing 'reverse shibboleth' tasks 97.7% of the time, and cost ~$0.05 per completed survey - versus $1.50 paid - making content-based spam detection economically and technically beaten as of late 2025.
Source: The potential existential threat of large language models to online survey research (PNAS 122(47)) (2025-11-20, academic)
Re-check detail
Asserted but not visible in retrievable text: 99.8% pass rate on attention/quality checks; 6,000 trials / 20 check types; 43,800 evaluations across 9 LLMs; 97.7% reverse-shibboleth refusal; $0.05 per survey vs $1.50 paid; 'there is no magical fix' and provenance-based defenses || Re-check notes: Page is protected by a Cloudflare JS challenge ('Just a moment...'). WebFetch returned HTTP 403 Forbidden; curl -sL with two different browser user-agents both returned the Cloudflare interstitial (HTTP 403, no redirect). Constraints forbid substitute sources or other pages on pnas.org, so none of the claimed figures could be verified.
Research-sweep audit (2026-07-14): confirmed
Every quantitative element of the CLAIM verifies against the primary source (Westwood, "The potential existential threat of large language models to online survey research," PNAS 122(47), published Nov 20, 2025; PubMed 41264250, cited 81x by Jul 2026; won a 2025 Cozzarelli Prize). Extracted paper text confirms verbatim: (1) 99.8% pass rate across 6,000 trials (20 ACQ types x 300 repetitions), only 10 errors, perfect on 18/20 check types; (2) execution layer "simulates realistic reading times calibrated to the persona's education level, generates human-like mouse movements, and types open-ended responses keystroke-by-keystroke, complete with plausible typos and corrections"; (3) strategically declines 97.7% of "reverse shibboleth" tasks; (4) ~$0.05 per survey with commercial models vs $1.50 payment (>96.8% margin), marginal cost near zero with local open-weight models; (5) 43,800 total evaluations, 139 questions, 6,700 trials, with a key subset validated across nine LLMs (DeepSeek R1, Mistral Large, Claude 3.7 Sonnet, Grok 3, Gemini 2.5, Qwen Plus, Llama 4 Maverick, GPT-4.1, o4-mini); (6) 0% error rate on impossible-event/rare-event probes. The paper's abstract and press framing support the interpretive conclusion that coherence-/content-based detection is no longer tenable and defenses must be provenance/economic. Supersession check: no 2026 work refutes it - follow-ups (Westwood & Frederick 2026 Prolific audit; a 2026 SAGE AMPPS review of AI-mediated contamination) build on it and treat detection as an unsolved provenance problem. Two minor caveats, both in the EVIDENCE SUMMARY rather than the claim: (a) arXiv 2606.04924 ("Can Crowdsourcing Survive the LLM Era?", Jun 2026) is a community survey of 155 researchers' practices/experiences, not a technical detector benchmark; its abstract does not state the paraphrase/out-of-domain/mixed-authorship degradation findings attributed to it (those likely come from a different detection-benchmark paper). (b) On reCAPTCHA the paper says the system "is designed to accommodate tools for bypassing" reCAPTCHA - an architectural capability, slightly weaker than "reCAPTCHA... fail[s]" as an empirically demonstrated result. Neither caveat affects the claim text itself, which is fully accurate.
Decision use: holds regardless of model progress. Measured on adjacent-domain. Survey-agent result. Attack capability is a lower bound that strengthens with every model generation; the economic asymmetry ($0.05 vs $1.50) widens. Content detection stays beaten.
Cited at: Research - findings by decision weight - Pilot - gates and objectives - Decisions - AI-assistance policy - Decisions - settled constraints
[AQ-03] annotation-quality - load-bearing - grade A (holds regardless of model progress) - re-check 2026-07-15: confirmed
Correlation-based meta-evaluation of automatic judges against human labels is systematically distorted by human label uncertainty: on high-disagreement items a machine judge can superficially match or beat human-human correlation, while on high-agreement strata machine-human correlation drops well below the human-human baseline - so a single aggregate 'agreement with humans' number is invalid as a validation target.
Source: Beyond correlation: The impact of human uncertainty in measuring the effectiveness of automatic evaluation and LLM-as-a-judge (ICLR 2025) (2025-05-01, academic)
What the source itself says (retrieved quote)
we propose ... a new metric - binned Jensen-Shannon Divergence for perception ... perception charts, to contextualize correlation measures appropriately
Re-check detail
Checked and matching: a single aggregate correlation score can obscure fundamental differences between human labels and those from automatic evaluation; high uncertain-label proportion: machine labels may superficially appear to have similar or better correlation with the human majority label (an illusion); as consistent-label samples increase, correlation between machine and human labels fall well below HH correlation; propose stratifying data by human label uncertainty; binned Jensen-Shannon Divergence for perception; perception charts; code open-sourced at github.com/amazon-science/BeyondCorrelation; Accepted at ICLR 2025 || Asserted but not visible in retrievable text: explicit 'Amazon' author affiliation -- page lists no affiliations; Amazon is only implied by the 'amazon-science' GitHub org name || Re-check notes: Core claim and all three prescribed fixes (stratify by uncertainty, binned JS divergence, perception charts) confirmed verbatim. Only the 'Amazon' affiliation label is not explicitly on the page.
Research-sweep audit (2026-07-14): partially_confirmed
Substance CONFIRMED on every checked point against the arXiv record (https://arxiv.org/abs/2410.03775): (1) abstract states machine labels "may superficially appear to have similar or better correlation with the human majority" when human uncertainty is high, and machine-human correlation falls "well below" human-human correlation as human label consistency increases; (2) it flags Krippendorff's alpha / Randolph's kappa as designed for human-human reliability with assumptions inapplicable to machine labels; (3) proposed fixes match exactly - stratification by human label uncertainty, binned Jensen-Shannon divergence for perception-type tasks, and 'perception charts'; (4) venue is ICLR 2025 per the arXiv comments field; (5) code is open-sourced at github.com/amazon-science/BeyondCorrelation. The only inaccuracy is the SOURCE date: the claim says 2025-05-01, but the latest arXiv version is v3 dated 2025-01-27 (original submission 2024-10-03) - no May 2025 version exists on the record. Supersession check: a follow-up search surfaced related later work - "Validating LLM-as-a-Judge Systems under Rating Indeterminacy" (arXiv 2503.05965), "Hidden Measurement Error in LLM Pipelines Distorts Annotation, Evaluation, and Benchmarking" (arXiv 2604.11581), and "CARE: Confounder-Aware Aggregation for Reliable LLM Evaluation" (arXiv 2603.00039, Feb 2026). These extend and reinforce the human-uncertainty/measurement-error critique rather than refute it; nothing found supersedes or contradicts the core finding. Verdict is partially_confirmed solely because of the incorrect source date; the claim itself stands.
Corrected statement: Correlation-based meta-evaluation of automatic judges against human labels is systematically distorted by human label uncertainty: on high-disagreement items a machine judge can superficially match or beat human-human correlation, while on high-agreement strata machine-human correlation drops well below the human-human baseline - so a single aggregate 'agreement with humans' number is invalid as a validation target. Source: Elangovan et al., "Beyond correlation..." (Amazon Science), accepted at ICLR 2025, arXiv 2410.03775 (v1 2024-10-03, latest v3 2025-01-27), code at github.com/amazon-science/BeyondCorrelation.
Decision use: holds regardless of model progress. Measured on structural. How to validate judges under human-label uncertainty - statistics of the gold set, not model capability.
Cited at: Decisions - the standard
[AQ-04] annotation-quality - load-bearing - grade A (holds regardless of model progress) - re-check 2026-07-15: confirmed
In human preference/rating data, the majority of annotator disagreements are attributable to task underspecification and response-style preferences - not random noise and not primarily expertise gaps - and standard aggregation (Bradley-Terry reward modeling) plus LLM-as-judge evaluation both fail to account for this divergence.
Source: Diverging Preferences: When do Annotators Disagree and do Models Know? (ICML 2025) (2025-07-15, academic)
What the source itself says (retrieved quote)
the majority of disagreements are due to factors such as task underspecification or response style ... standard reward modeling (e.g., Bradley-Terry) and LLM-as-Judge evaluation methods fail to account for divergence between annotators.
Re-check detail
Checked and matching: majority of disagreements due to task underspecification or response style (not simple noise); standard reward modeling (Bradley-Terry) and LLM-as-Judge fail to account for divergence; develops methods for identifying diverging preferences; ten-category taxonomy across four high-level classes; ICML 2025 venue note || Asserted but not visible in retrievable text: explicit statement that expertise gaps are NOT a main cause (excluded only by implication, since the majority is attributed to underspecification/style); Jiang & de Marneffe TACL 2022 corroboration (a separate citation, not part of this source) || Re-check notes: Core assertion and all listed mechanisms confirmed from the abstract. The 'not primarily expertise gaps' element is supported by implication rather than an explicit exclusionary sentence.
Decision use: holds regardless of model progress. Measured on human-work. Attribution of annotator disagreement to underspecification/style. Human raters have not changed.
Cited at: Decisions - the standard
[AQ-05] annotation-quality - grade A (holds regardless of model progress) - re-check 2026-07-15: confirmed with caveats
Even on a nominally objective task (NLI), annotator disagreement decomposes into a 10-category taxonomy across three high-level sources - uncertainty in sentence meaning, underspecification in guidelines, and annotator behavior - proving that a fraction of inter-reviewer variance is item-intrinsic and survives perfect rubric operationalization.
Source: Investigating Reasons for Disagreement in Natural Language Inference (TACL 10) (2022-12-01, academic)
What the source itself says (retrieved quote)
We developed a taxonomy of disagreement sources with 10 categories spanning 3 high-level classes. We found that some disagreements are due to uncertainty in the sentence meaning, others to annotator biases and task artifacts
Re-check detail
Checked and matching: taxonomy of disagreement sources with 10 categories spanning 3 high-level classes; one class is uncertainty in the sentence meaning; task is NLI (natural language inference) annotation disagreement || Asserted but not visible in retrievable text: 'underspecification in guidelines' as a named high-level class (abstract instead names 'annotator biases and task artifacts'); 'survives perfect rubric operationalization' framing; disagreement persisting specifically among trained annotators; D3CODE (EMNLP 2024; 4.5K sentences, 4K+ annotators, 21 countries) - a separate source not on this page || Re-check notes: Core (10-category / 3-class taxonomy of intrinsic NLI disagreement) confirmed verbatim. Two of the claim's three named high-level sources map to the abstract ('uncertainty in sentence meaning'; 'annotator behavior' ~ 'annotator biases and task artifacts'); the third, 'underspecification in guidelines', is not named in the abstract. The D3CODE material cited in the evidence is a different paper, unverifiable from this URL. Only the abstract page is accessible (PDF/other same-site pages disallowed).
Decision use: holds regardless of model progress. Measured on human-work. 2022 NLI disagreement taxonomy. Human disagreement structure, not model capability.
[AQ-06] annotation-quality - grade A (holds regardless of model progress) - re-check 2026-07-15: confirmed with caveats
Chance-corrected agreement coefficients (Cohen's kappa, and analogously Krippendorff's alpha) can be near zero despite very high raw agreement when class prevalence is extreme (the Feinstein-Cicchetti paradoxes), so under the 90%+ pass rates typical of QA pipelines, single-coefficient reliability reporting is misleading in the opposite direction from raw percent agreement.
Source: High agreement but low kappa: I. The problems of two paradoxes (Feinstein & Cicchetti) (1990-01-01, academic)
What the source itself says (retrieved quote)
articleName : 'High agreement but low Kappa: I. the problems of two paradoxes'
Re-check detail
Checked and matching: resolved article title confirms the core thesis: high agreement, low kappa, two paradoxes || Asserted but not visible in retrievable text: abstract-level mechanism: kappa near zero under extreme prevalence (prevalence-driven); second paradox: asymmetric marginal distributions inflating/affecting kappa; the 'opposite direction from raw percent agreement' framing tied to QA pass rates; external replications cited in the claim's evidence (Quarfoot & Levine 2016; Gwet AC1 literature) - not this source || Re-check notes: Server redirect chain doi.org -> linkinghub.elsevier.com -> jclinepi.com. Abstract body is behind a Cloudflare JS challenge (curl) and the Elsevier landing page returns only 'Redirecting'; a full WebFetch of the DOI/landing was flagged by a stochastic model safeguard. Title confirmed via the linkinghub siteCatalyst metadata, which locks the core high-agreement/low-kappa/two-paradoxes assertion; the specific paradox mechanisms could not be retrieved from the abstract text.
Decision use: holds regardless of model progress. Measured on structural. Kappa paradox under extreme base rates. Mathematics.
Cited at: Decisions - default: statistics staging - Decisions - settled constraints
[AQ-07] annotation-quality - grade A (holds regardless of model progress) - re-check 2026-07-15: confirmed with caveats
Classical label-aggregation models (Dawid-Skene 1979, MACE 2013 with its explicit spammer latent variable, GLAD) all assume annotators are conditionally independent given the true label - an assumption measurably violated when the 'annotators' are LLM judges sharing data/architectures/prompts, causing miscalibrated posteriors and confidently wrong aggregate verdicts; Feb 2026 work replaces this with dependence-aware Ising-model aggregation.
Source: Dependence-Aware Label Aggregation for LLM-as-a-Judge via Ising Models (arXiv 2601.22336) (2026-02-02, academic)
What the source itself says (retrieved quote)
Most classical methods, e.g., Dawid-Skene or (weighted) majority voting, assume annotators are conditionally independent ... Ignoring such dependencies can yield miscalibrated posteriors and even confidently incorrect predictions.
Re-check detail
Checked and matching: classical methods assume annotators are conditionally independent; assumption often violated by LLM judges due to shared data, architectures, prompts, and failure modes; ignoring dependencies yields miscalibrated posteriors and even confidently incorrect predictions; dependence-aware aggregation via Ising graphical models; Dawid-Skene named || Asserted but not visible in retrievable text: MACE (2013); 'spammer' latent variable; GLAD; Dawid-Skene '1979' date; Toloka crowd-kit production implementations || Re-check notes: Both direct quotes the claim attributes to the paper are confirmed verbatim, and the Ising-model core is confirmed. Retrieved abstract lists CI examples as 'Dawid-Skene or (weighted) majority voting' only; MACE, GLAD, the spammer latent variable, and the crowd-kit implementation note are the claim author's additions and are not visible in retrieved text. Submitted 2026-01-29 (source_date 2026-02-02).
Decision use: holds regardless of model progress. Measured on structural. Dawid-Skene/MACE-class aggregation. Method, not capability.
[AQ-08] annotation-quality - grade A (holds regardless of model progress) - re-check 2026-07-15: confirmed with caveats
LLM contamination in low-dimensional annotation data (multiple-choice labels, ratings) can be evaluated at the worker level without any ground truth, using peer-prediction scores that condition on LLM-generated reference labels - a training-free mechanism with theoretical guarantees under an LLM-collusion model, published at NeurIPS 2025 - because text-based detectors are inapplicable to short label outputs.
Source: Evaluating LLM-Contaminated Crowdsourcing Data Without Ground Truth (NeurIPS 2025) (2025-12-01, academic)
What the source itself says (retrieved quote)
a training-free scoring mechanism with theoretical guarantees under a crowdsourcing model that accounts for LLM collusion
Re-check detail
Checked and matching: evaluates contaminated data without ground truth via peer prediction; quantifies correlations between worker answers conditioning on (a subset of) LLM-generated labels; training-free scoring mechanism with theoretical guarantees under a crowdsourcing model that accounts for LLM collusion; existing text-based detectors rely on high-dimensional text, unsuitable for annotation tasks; multiple-choice labeling named as target task; empirically robust in detecting low-effort cheating on real-world crowdsourcing datasets; authors Yichi Zhang, Jinlong Pang, Zhaowei Zhu, Yang Liu (matches Zhang, Pang, Zhu & Liu) || Asserted but not visible in retrievable text: NeurIPS 2025 venue (no venue stated anywhere on the abstract page); 'ratings' as a target task type (only multiple-choice explicitly named); explicit 'per-worker not per-item' framing (implied by inter-worker correlation, not stated); explicit 'multiple workers pasting from the same LLM' scenario; source_date 2025-12-01 (page shows Jun/Nov 2025 only) || Re-check notes: Core method claims fully match the abstract. The asserted NeurIPS 2025 publication and the 2025-12-01 date are not visible on the page; abstract dates are v1 2025-06-08, v2 2025-11-06.
Decision use: holds regardless of model progress. Measured on structural. Worker-level contamination auditing without per-item detection. Method survives model improvement on both sides.
Cited at: Decisions - AI-assistance policy
[AQ-09] annotation-quality - grade A (holds regardless of model progress) - re-check 2026-07-15: confirmed with caveats
The NLP field has institutionalized disagreement-as-signal: the third Learning-With-Disagreements shared task (LeWiDi-2025, EMNLP 2025) standardizes dual evaluation - soft-label (predict the population label distribution, Wasserstein/Manhattan distance) and perspectivist (recover individual annotators' labels) - across four datasets that all ship per-annotator labels rather than adjudicated gold.
Source: LeWiDi-2025 at NLPerspectives: Third Edition of the Learning with Disagreements Shared Task (2025-10-09, academic)
What the source itself says (retrieved quote)
the soft-label approach, in which models predict population-level distributions of judgments ... the perspectivist approach, in which models predict the interpretations of individual annotators
Re-check detail
Checked and matching: Third edition of the Learning With Disagreements (LEWIDI) shared task; NLPerspectives workshop at EMNLP 2025; soft-label approach: models predict population-level distributions of judgments; perspectivist approach: models predict interpretations of individual annotators; four datasets (paraphrase, irony, sarcasm, NLI); went beyond standard metrics such as cross-entropy with new metrics for both paradigms || Asserted but not visible in retrievable text: 'Wasserstein/Manhattan distance' as the soft-label/perspectivist metrics (not named in abstract); 'public leaderboard' (not in abstract; a Codabench competition link appears in the PDF structure but term not confirmed); 'per-annotator labels rather than adjudicated gold' (implied by perspectivist paradigm, not stated explicitly); lineage claims: DICES (NeurIPS 2023), jury learning (Gordon et al. CHI 2022), Frenda et al. survey (LREV 2024) - cite other works not fetchable here; Aroyo & Welty citation anchor was present in the PDF || Re-check notes: Given PDF was FlateDecode-compressed (no readable abstract); fetched the same-arXiv-ID abs page (allowed variant) for the abstract. Core dual-evaluation design confirmed; specific distance metrics and leaderboard not visible.
Decision use: holds regardless of model progress. Measured on structural. Disagreement-as-signal institutionalized (LeWiDi-2025). Practice fact.
[AQ-10] annotation-quality - grade A (holds regardless of model progress) - re-check 2026-07-15: confirmed
Gold/test labels themselves are wrong at material rates - a lower-bound average of 3.3% label errors across 10 canonical benchmarks (at least 6% in the ImageNet validation set), validated by human review of confident-learning-flagged candidates (51% flag precision) - and error rates this size are enough to flip model rankings.
Source: Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks (NeurIPS 2021 D&B) (2021-11-01, academic)
What the source itself says (retrieved quote)
estimate an average of at least 3.3% errors across the 10 datasets ... label errors comprise at least 6% of the ImageNet validation set ... ResNet-18 outperforms ResNet-50 ... if the prevalence of originally mislabeled test examples increases by just 6%
Re-check detail
Checked and matching: average of at least 3.3% errors across the 10 datasets; at least 6% label errors in the ImageNet validation set; 51% human-validation precision (roughly half of flagged candidates confirmed mislabeled); 10 canonical benchmarks; confident learning algorithms + crowdsourced human validation; ResNet-18 outperforms ResNet-50 on corrected ImageNet if mislabeled prevalence rises by just 6% (rank flip) || Re-check notes: All numbers and the rank-flip finding confirmed from the abstract page. Cleanlab implementation referenced (github.com/cleanlab/label-errors, labelerrors.com).
Decision use: holds regardless of model progress. Measured on structural. >=3.3% gold-label error floor across canonical test sets. Property of datasets and labeling processes, not of judges.
Cited at: Decisions - permanent audit - Decisions - settled constraints
[AQ-11] annotation-quality - grade A (holds regardless of model progress) - re-check 2026-07-15: confirmed
Showing annotators LLM suggestions makes them anchor on the suggestions - significantly shifting the label distribution versus unassisted baseline, raising self-reported confidence without making them faster - and evaluating models against LLM-assisted labels significantly inflates reported model performance.
Source: Just Put a Human in the Loop? Investigating LLM-Assisted Annotation for Subjective Tasks (Findings of ACL 2025) (2025-07-27, academic)
What the source itself says (retrieved quote)
annotators strongly took the LLM suggestions, significantly changing the label distribution compared to the baseline
Re-check detail
Checked and matching: annotators anchored on LLM suggestions, significantly changing label distribution vs baseline; did NOT make them faster but improved self-reported confidence; using LLM-assisted labels for evaluation significantly increases reported model performance; pre-registered experiment; 350 unique annotators, 7,000 annotations, 4 conditions, 2 models, 2 datasets || Asserted but not visible in retrievable text: 'homogenization/lessened variation is a direct risk' framing (from the evidence narrative) is not present in the abstract; abstract instead says label-distribution shifts can affect conclusions drawn even from human-approved datasets || Re-check notes: Every core assertion and headline number in the claim is confirmed by the abstract. Only the 'homogenization' phrasing (which is in the verifier's evidence rationale, not the claim text) is not visible in the abstract; it may appear in the full paper (not fetched).
Decision use: holds regardless of model progress. Measured on human-work. Anchoring on shown suggestions is human behavior; better suggestions anchor harder. Blind-first design stays mandatory.
[AQ-12] annotation-quality - grade B (current-generation measurement) - re-check 2026-07-15: confirmed
RLVR-style reasoning training degrades an LLM's ability to model human annotator disagreement, while naive chain-of-thought on RLHF models improves it - evaluated across 60 setups on 3 tasks - meaning stronger 'reasoning' judges are measurably worse at knowing when humans legitimately diverge.
Source: Can Reasoning Help Large Language Models Capture Human Annotator Disagreement? (EACL 2026) (2026-01-12, academic)
What the source itself says (retrieved quote)
RLVR-style reasoning degrades performance in disagreement modeling ... naive Chain-of-Thought (CoT) reasoning improves the performance of RLHF LLMs ... resulting in 60 experimental setups across 3 tasks
Re-check detail
Checked and matching: RLVR degrades disagreement modeling: 'RLVR-style reasoning degrades performance in disagreement modeling'; naive CoT improves RLHF models: 'naive Chain-of-Thought (CoT) reasoning improves the performance of RLHF LLMs'; '60 experimental setups across 3 tasks' -- matches asserted 60 setups / 3 tasks; warning: 'the potential risk of replacing human annotators with reasoning LLMs'; EACL 2026 Main listed; v3 dated Jan 12 2026 (matches source_date 2026-01-12) || Asserted but not visible in retrievable text: evidence-field 'author communication' that small fine-tuned ModernBERT-class models with human labels beat large reasoning models -- explicitly external to the paper, not in the retrieved abstract; specific metric names 'variance-correlation' and 'distributional-alignment' not verbatim in retrieved abstract (though 'distribution expression methods, and steering methods' are confirmed) || Re-check notes: All core claim assertions and the 60-setups/3-tasks figures are confirmed. The ModernBERT comparison is flagged in the evidence itself as author communication, so it cannot be checked against the source and is not weighed against the claim.
Decision use: current-generation measurement. Measured on model-outputs. Jan-2026: RLVR training degrades disagreement modeling. Mechanism relevant to judge-model selection at the deployment tier; magnitude tier-bound.
Critique and assistance studies
[CM-01] critique-models - load-bearing - grade C (superseded-generation number; mechanism only) - re-check 2026-07-15: confirmed
Human+critic-model teams occupy a strictly better operating point than either alone: in OpenAI's CriticGPT study, Human+CriticGPT teams wrote more comprehensive critiques than unassisted humans while hallucinating and nitpicking less than the model alone, and the comprehensiveness-vs-spurious-claims tradeoff is a tunable inference-time dial (FSBS length penalty), not a fixed property.
Source: LLM Critics Help Catch LLM Bugs (CriticGPT, OpenAI) (2024-06-28, primary)
What the source itself says (retrieved quote)
Human+CriticGPT teams move beyond the model-only Pareto frontier.
Re-check detail
Checked and matching: 'model-written critiques are preferred over human critiques in 63% of cases' (naturally occurring LLM errors); 'model critiques are preferred over human critiques more than 80% of the time' on Human Inserted Bugs, scale 'linear in Elo'; 'Human+CriticGPT teams move beyond the model-only Pareto frontier'; FSBS selects critiques by 'rm_score + LENGTH_MODIFIER * num_highlights'; 'explored 4 values of LENGTH_MODIFIER' (inference-time dial); contractors median ~5 years Python experience, ~50 min per critique; human-machine teams 'hallucinating less than LLMs alone' || Re-check notes: 63% preference, >80% Elo on inserted bugs, the FSBS length-penalty dial, and the beyond-the-frontier team result all confirmed from arXiv HTML full text. The critic model is named 'CriticGPT' in the body though the arXiv title is 'LLM Critics Help Catch LLM Bugs'.
Research-sweep audit (2026-07-14): confirmed
Source verified: arXiv:2407.00215 "LLM Critics Help Catch LLM Bugs" (McAleese et al., OpenAI), submitted 2024-06-28 - date matches. Abstract directly confirms: 63% preference for model critiques on naturally occurring LLM errors; models catch more bugs than paid human reviewers; human-machine teams catch similar bug counts to LLM critics while hallucinating less than LLMs alone. Full-text details corroborated via independent secondary coverage: critic-assisted contractors wrote more comprehensive critiques than unassisted contractors while reducing hallucination/nitpick rate relative to the model, described as moving beyond the model-only Pareto frontier; FSBS scores candidates by rm_score + LENGTH_MODIFIER x num_highlights, giving a deployment-time precision/comprehensiveness dial without retraining. Minor reading note: "strictly better than either alone" holds in the paper's Pareto sense - teams matched (did not exceed) the model's bug-catch/comprehensiveness while beating its hallucination rate, and exceeded human comprehensiveness; the claim's own wording states exactly this, so no correction needed. Supersession check: no official OpenAI follow-up found; later work (CodeCriticBench 2025; 2026 papers on LLM reviewer overcorrection) extends but does not contradict these findings.
Decision use: superseded-generation number; mechanism only. Measured on human-work. CriticGPT (GPT-4-era, 2024). Critic-assisted-beats-either-alone is a floor that should strengthen with critic quality, but the 63% preference number is two generations stale and unreplicated in 24+ months.
Same underlying source as [CM-02] - repetition across reports is not independent corroboration.
[CM-02] critique-models - grade C (superseded-generation number; mechanism only) - re-check 2026-07-15: confirmed
Critic-assisted review of items already rated 'flawless' by a first human found real, rating-changing problems in 24% of flagged cases, versus 6% found by an unassisted second human review of the same data - a 4x lift on the exact 'confident pass' slice.
Source: LLM Critics Help Catch LLM Bugs (CriticGPT, OpenAI) - 'flawless' data experiment (2024-06-28, primary)
What the source itself says (retrieved quote)
In 24% of cases contractors indicated that the critique found a problem that substantially decreased the rating
Re-check detail
Checked and matching: With critique: 'In 24% of cases contractors indicated that the critique found a problem that substantially decreased the rating'; Without critique: completions first rated 'flawless' were 'rated similarly poorly by a second only 6% of the time'; Run over ChatGPT training data first rated 'flawless'; Held out-of-distribution ('out-of-distribution for the critic model'), mostly non-code || Re-check notes: The 24% (critic-assisted) vs 6% (unassisted second rater) comparison on the 'flawless' slice is confirmed verbatim in Section 3.6 - a 4x lift as the claim states.
Decision use: superseded-generation number; mechanism only. Measured on human-work. The 24%-vs-6% flawless-slice lift is GPT-4-era. Direction (audit confident passes with critic assistance) is the durable design input; the 4x magnitude is not a current estimate.
Same underlying source as [CM-01] - repetition across reports is not independent corroboration.
Cited at: Research - findings by decision weight - System - claim types - Pilot - gates and objectives - Decisions - permanent audit
[CM-03] critique-models - load-bearing - grade B (current-generation measurement) - re-check 2026-07-15: confirmed
In DeepMind's amplified-oversight experiments on fact-verification rating, confidence-based hybridization (AI rates when confident, humans rate the low-confidence slice with evidence-only assistance) reached 91.3% accuracy vs 87.7% AI-alone and 75.1% human-alone - and showing humans the AI's verdict/reasoning/confidence caused measurable over-reliance, while showing only retrieved evidence was the sole format that helped when the AI was right without hurting when it was wrong.
Source: Human-AI Complementarity: A Goal for Amplified Oversight (Google DeepMind) (2025-10-30, primary)
What the source itself says (retrieved quote)
rater achieves 87.7% accuracy, performing above human raters who achieved 75.1% accuracy ... confidence threshold of 0.62, results in an accuracy of 91.3%, compared to 89.3% using unassisted ... -0.768, SE = 0.174, z = -4.413, p < .001
Re-check detail
Checked and matching: AI rater 87.7% vs human raters 75.1%; human majority vote 80.6% (vs 75.1%); N = 1918 Evaluation Set examples; unassisted hybridization 89.3% (higher than AI-alone 87.7%); confidence threshold T = 0.62; on low-confidence slice AI 60.5% vs human 71.3%; evidence-assisted hybridization achieves 91.3%; evidence-only helps on items AI gets correct: 79.3% vs 71.3%; over-reliance coefficient -0.768, SE=0.174, z=-4.413, p<.001 || Re-check notes: Every headline number in the claim verified verbatim from the PDF full text via pdftotext. AI-alone 87.7%, human-alone 75.1%, majority 80.6%, N=1918, threshold 0.62, unassisted hybrid 89.3%, evidence-assisted 91.3%, low-conf slice 60.5% AI / 71.3% human, evidence-only 79.3% vs 71.3%, beta=-0.768 p<.001 all present. Note: PDF served at /pdf/ is v2 (June 2026); source_date is v1 (Oct 30 2025). All cited numbers match the retrieved v2 text. Authors: Jain, Bridgers, Janzer, Greig, Teh, Mikulik.
Research-sweep audit (2026-07-14): confirmed
Every quantitative element checks out against the arXiv HTML full text (2510.26518). Dataset: 1,918 expert-labeled tuples - confirmed. AI rater 87.7%, individual humans 75.1%, human majority vote 80.6% - confirmed. Confidence via 50 samples/sentence (avg 33.25 passing format check), described as fairly well calibrated - confirmed. Unassisted hybridization at threshold T=0.62: 89.3% (beta=0.413, p=.012) - confirmed. Low-confidence routed slice (280 items): humans 71.3% vs AI 60.5% (p=.006) - confirmed. The headline 91.3% is specifically the ASSISTED hybrid (AI when confident + evidence-assisted humans on the low-confidence slice), vs 89.3% unassisted hybrid - exactly as the claim words it, so 91.3% is correctly attributed. Evidence-only assistance: 79.3% vs 71.3% baseline when AI correct (beta=0.446, p=.009), no significant harm when AI wrong (64.0% vs 61.5%); paper explicitly calls it "the only form of assistance that achieves the ideal of helping when correct and not hurting when wrong" - confirmed. Over-reliance when AI wrong: Evidence&Reasoning&Judgments beta=-0.768 (SE=0.174, z=-4.413, p<.001), ER&J&Confidence beta=-0.636 p<.001, Judgments&Confidence beta=-0.356 p=.042 - the quoted beta=-0.768 matches the Evidence&Reasoning&Judgments condition exactly. Debate was numerically worst (64.9%, only format below baseline) - confirmed. Date: v1 submitted 2025-10-30 as claimed. Currency: a v2 was posted 2026-06-25 and the paper was published at ACM FAccT '26 (DOI 10.1145/3805689.3812308); this is publication/revision, not supersession - no later work contradicting the findings found. DeepMind blog notes ongoing follow-up combining hybridization+assistance on an internal rating task, unpublished as of 2026-07-14. Minor caveat only: whoever cites this should prefer the v2/FAccT version in case numbers shifted slightly in revision (v1 numbers verified here).
Decision use: current-generation measurement. Measured on adjacent-domain. 2025 DeepMind amplified-oversight. The format effect (verdict-visible -> over-reliance; evidence-only safe) is human behavior and durable; the 91.3/87.7/75.1 magnitudes are tier- and task-bound (fact verification, not annotation QA).
Cited at: Summary - human role - Pilot - gates and objectives - Decisions - authority boundaries - Decisions - default: who sees what
[CM-04] critique-models - grade A (holds regardless of model progress) - re-check 2026-07-15: confirmed
A Nature Human Behaviour meta-analysis of 100+ experiments (300+ effect sizes) found human-AI combinations on average performed significantly WORSE than the best of human or AI alone on decision/judgment tasks, with gains only where humans outperformed the AI solo - naive human-checks-AI designs destroy value.
Source: When combinations of humans and AI are useful: A systematic review and meta-analysis (2024-10-28, contrarian)
What the source itself says (retrieved quote)
human-AI combinations performed significantly worse than the best of humans or AI alone ... when humans outperformed AI alone, we found performance gains in the combination
Re-check detail
Checked and matching: meta-analysis of over 100 recent experimental studies reporting over 300 effect sizes; human-AI combinations performed significantly worse than the best of humans or AI alone; performance losses in tasks that involved making decisions; significantly greater gains in tasks that involved creating content; when humans outperformed AI alone, we found performance gains in the combination; Nat Hum Behav (2024) / Nature Human Behaviour || Asserted but not visible in retrievable text: exact counts 106 studies / 370 effect sizes (page says 'over 100' and 'over 300'); DeepMind complementarity paper framing this result as its motivation -- an external editorial connection, not on this page || Re-check notes: Nature Human Behaviour venue, 100+ studies / 300+ effect sizes, decision-task losses, creation-task gains, and the humans-beat-AI condition all confirmed verbatim.
Decision use: holds regardless of model progress. Measured on adjacent-domain. Meta-analytic human-AI complementarity result. As deployment-tier models widen the human-AI gap, the warning against naive verify-the-AI designs strengthens.
[CM-05] critique-models - load-bearing - grade B (current-generation measurement) - re-check 2026-07-15: confirmed
In the largest deployed RCT of AI critiquing human evaluative work (ICLR 2025, feedback on >20,000 peer reviews), 27% of reviewers who received LLM feedback revised their reviews, incorporating >12,000 suggestions, producing reviews +80 words that blinded raters judged more informative, plus higher author-rebuttal engagement - and feedback was only delivered if it passed a suite of automated LLM reliability tests.
Source: Can LLM feedback enhance review quality? A randomized study of 20K reviews at ICLR 2025 (2025-04-13, academic)
What the source itself says (retrieved quote)
27% of reviewers who received feedback updated their reviews and over 12,000 feedback suggestions from the agent were incorporated by those reviewers ... an average increase of 80 words among those who updated after receiving feedback.
Re-check detail
Checked and matching: feedback to more than 20,000 randomly selected reviews; 27% of reviewers who received feedback updated their reviews; over 12,000 feedback suggestions incorporated; average increase of 80 words among updaters; more informative reviews as judged by blinded researchers; longer author-reviewer discussions / more rebuttal engagement; a suite of automated LLM-powered reliability tests acted as guardrails gating delivery || Asserted but not visible in retrievable text: the specific '26.6%' figure (abstract states 27%; 26.6% appears only in the claim's evidence field) || Re-check notes: All headline numbers in the core claim confirmed verbatim from the abstract.
Research-sweep audit (2026-07-14): confirmed
Source verified directly (arXiv:2504.09737, v1 submitted 2025-04-13 - date matches). Abstract states: Review Feedback Agent deployed at ICLR 2025 as a large randomized controlled study on >20,000 randomly selected reviews; feedback targeted vague comments, content misunderstandings, unprofessional remarks; 27% of reviewers who received feedback updated their reviews; >12,000 suggestions incorporated (12,222 in the paper); +80 words average among updaters; more informative per blinded evaluators; longer author-reviewer discussions; feedback delivered only if it passed a suite of automated LLM reliability tests. Arm sizes corroborated from paper figures: 22,467 selected for feedback vs 22,364 control; 18,946 successfully received feedback, of whom 26.6% updated (the abstract rounds to 27%). Two nuances, neither contradicting the claim: (1) the superlative "largest deployed RCT of AI critiquing human evaluative work" is the researcher's framing, not verbatim in the abstract - no larger counterexample found, and it is consistent with the study's scale; (2) not superseded but now formally published: Nature Machine Intelligence, Feb 23, 2026, as "A large-scale randomized study of large language model feedback in peer review" (https://www.nature.com/articles/s42256-026-01188-x) - the NMI version is the preferred citation going forward. A related follow-up mixed-methods study on reviewer perceptions exists (arXiv:2602.13817) but does not alter these results.
Decision use: current-generation measurement. Measured on human-work. 2025 RCT on >20k real peer reviews - one of the few human-work measurements in the corpus. Revision behavior is human; feedback quality rises with model tier, so 27% is closer to a floor than a ceiling. Domain transfer (peer review -> paid annotation) untested (E6).
Cited at: Research - findings by decision weight - Decisions - enforcement weight - Decisions - default: who sees what
[CM-06] critique-models - load-bearing - grade C (superseded-generation number; mechanism only) - re-check 2026-07-15: confirmed with caveats
LLMs are structurally weak at detecting omissions: on AbsenceBench, average F1 drops 56.9 points versus detecting the same content as insertions (best model 71.2% overall, 40.0% on code diffs, vs ~99.5% needle-in-haystack), and inserting explicit placeholders at gap sites recovers ~35.7 points - absence has no attention key to attend to.
Source: AbsenceBench: Language Models Can't Tell What's Missing (2025-06-13, academic)
What the source itself says (retrieved quote)
AbsenceBench contains 4302 instances in total, with an average context length of 5K tokens. ... observing a massive 56.9% drop in F1-score on average ... This boosts the performance by a dramatic 35.7% on average
Re-check detail
Checked and matching: 'a massive 56.9% drop in F1-score on average' (omission vs insertion); best overall average F1 71.2 (Gemini-2.5-flash, thinking); highest GitHub-PR/code-diff score only 40.0% (by Claude-3.7-Sonnet thinking); nearly 99.5% F1 on poetry under the insertion (NIAH-style) setting; '<missing line>' placeholders boost performance 35.7% on average; +81.8% for Claude-3.7-Sonnet on GitHub PRs; 4302 instances total, average context length 5K tokens; Mixtral-8x7B: perfect NIAH score but only 14.7% F1 on AbsenceBench; inference-time compute (thinking) yields modest 7.9% improvement || Asserted but not visible in retrievable text: NeurIPS 2025 Datasets & Benchmarks acceptance (not in the abstract or HTML full text) || Re-check notes: All headline numbers are confirmed verbatim in the HTML full text. Minor attribution nuance: the claim pairs '71.2% overall' and '40.0% on code diffs' as one 'best model', but 71.2 is Gemini-2.5-flash's average while 40.0% is the domain-best set by Claude-3.7-Sonnet (thinking) (Gemini's GitHub-PR score is 30.9). Downgraded to partially_confirmed only because the asserted NeurIPS 2025 venue is not present in retrieved text.
Decision use: superseded-generation number; mechanism only. Measured on model-outputs (models: Gemini-2.5-flash). AbsenceBench on 2025 flash-tier (Gemini-2.5-flash). Omission-blindness is attention-mechanics and plausibly persists, but deployment-tier recall is unmeasured - E2 measures it before any absence-claim verdict is trusted.
Cited at: Research - findings by decision weight - System - claim types - Decisions - settled constraints
[CM-07] critique-models - grade C (superseded-generation number; mechanism only) - re-check 2026-07-15: confirmed
Critique quality is itself quantifiable at usable reliability: MetaCritique decomposes critiques into atomic information units and scores precision/recall against references, with GPT-4 AIU-level judgments at 85-89% accuracy and Meta-F1 correlating with human gold at Pearson 0.84-0.89 - and measured baselines show human critiques are high-precision/low-recall (87.6%/48.7%) while LLM critiques are the inverse-ish (71.9%/53.3%).
Source: The Critique of Critique (MetaCritique) (2024-01-09, academic)
What the source itself says (retrieved quote)
GPT-4 ... achieves an impressive performance (nearly 90%) ... MetaCritique-GPT4-F1 scores 0.841 ... 0.886 ... 51% Better Critique Wins, 22% Tie, 27% Loses ... Hypo.l = 8.10, Hypo.h = 3.31
Re-check detail
Checked and matching: critiques decomposed into Atomic Information Units (AIUs), precision/recall scored against references with NL rationales; GPT-4 AIU-level accuracy 85-89% (Table 3: 89.12 / 87.96 precision task; 85.47 / 86.82 recall task; 'nearly 90%'); Meta-F1 Pearson with human gold 0.841 (human-written) / 0.886 (LLM-generated); human critiques high-precision/low-recall 87.61 / 48.72; LLM critiques 71.85 / 53.28; refinement by Meta-F1: 51% Better Critique Wins / 22% Tie / 27% Loses (human eval); average AIUs 8.10 (LLM/Hypo.l) vs 3.31 (human/Hypo.h) = 2.4x; non-factual AIUs ~28% (LLM) vs ~12% (human), derived from precision 71.85 / 87.61 || Asserted but not visible in retrievable text: the specific comparative that GPT-4 pairwise picks 'lose more than they win' vs MetaCritique (a separate sub-figure not surfaced in retrieval) || Re-check notes: All headline numbers verified verbatim from arXiv HTML full text (Tables 1,3,4,6, Figure 4b). 8.10/3.31 = 2.45 confirms the '2.4x information volume' claim. Non-factual percentages are correctly derived from precision as the claim itself states.
Decision use: superseded-generation number; mechanism only. Measured on model-outputs (models: GPT-4). MetaCritique (GPT-4, 2024). The critique-decomposition METHOD is reusable; its reported reliability numbers are stale.
Cited at: Pilot - gates and objectives
[CM-08] critique-models - grade C (superseded-generation number; mechanism only) - re-check 2026-07-15: confirmed with caveats
Verdict accuracy and critique validity dissociate badly: in CriticBench-THU data 24.8% of items got the correct verdict with a low-quality critique, and open-loop verdict-agreement metrics compress a 27.1-point real error-identification gap (ProcessBench: o1-mini 88.9 vs Qwen2.5-72B 61.8) into a 1.3-point verdict-F1 gap - so critique quality must be evaluated closed-loop by whether the critique drives a successful correction.
Source: RealCritic: Towards Effectiveness-Driven Evaluation of Language Model Critiques (2025-01-24, academic)
What the source itself says (retrieved quote)
our approach employs a closed-loop methodology that evaluates the quality of corrections generated from critiques ... classical LLMs significantly lag behind the advanced reasoning-based model o1-mini across all critique scenarios
Re-check detail
Checked and matching: closed-loop methodology evaluating quality of corrections generated from critiques; eight challenging reasoning tasks; classical LLMs significantly lag o1-mini across all critique scenarios; classical LLMs can fall below their own baselines under self/iterative critique || Asserted but not visible in retrievable text: CriticBench-THU 24.8%; 27.1-point error-identification gap; ProcessBench o1-mini 88.9 vs Qwen2.5-72B 61.8; 1.3-point verdict-F1 gap; -1.8 to -5.1 avg self-critique delta; -35.6 domain drop; 'superficial success' failure-mode name; 'contradictory output' failure-mode name; RM-NLHF / arXiv 2601.07349 (separate source) || Re-check notes: RealCritic's core design (closed-loop critique-then-correct over 8 reasoning tasks; o1-mini only model with positive self-critique; classical LLMs drop below baselines) is confirmed. However essentially every headline number in the claim is drawn from OTHER benchmarks (CriticBench-THU 24.8%, ProcessBench 88.9/61.8) or from RealCritic detailed results not present in the abstract; none are visible in the retrieved text.
Decision use: superseded-generation number; mechanism only. Measured on model-outputs (models: o1-mini, Qwen2.5-72B, GPT-4). Right-verdict/poor-critique dissociation measured on 2024-25 models; re-check found headline numbers span sibling benchmarks. Keep the failure mode, drop the numbers.
Cited at: Pilot - gates and objectives
[CM-09] critique-models - grade B (current-generation measurement) - re-check 2026-07-15: confirmed
A June 2026 study of 21 LLM judges (~541k judgments) shows raw percent-agreement overstates chance-corrected agreement by 33.8-41.3 points (85% agreement = kappa ~0.48), judge rankings flip by up to 15 positions across benchmarks, and the most reproducible judges are among the least valid (test-retest 0.99 with position bias 0.19) - leading the authors to prescribe a pre-deployment Minimum Viable Validation Protocol.
Source: Reliability without Validity: A Systematic, Large-Scale Evaluation of LLM-as-a-Judge Models Across Agreement, Consistency, and Bias (2026-06-17, academic)
What the source itself says (retrieved quote)
exact match overstates chance-corrected agreement by between 33.8 and 41.3 percentage points across the 21 models
Re-check detail
Checked and matching: exact match overstates chance-corrected agreement by between 33.8 and 41.3 percentage points across the 21 models; a judge reporting '85% agreement' on MT-Bench has kappa approximately 0.48; Llama 3.3 70B which shifts 15 positions, from 5 on MT-Bench to 20 on JudgeBench; Qwen 3 8B (test-retest 0.992, position bias 0.192); Gemini 2.5 Flash (test-retest 0.988, position bias 0.125); All evaluations used temperature 0; UC Berkeley School of Information; 118 runs; MVVP: report Cohen's kappa or Krippendorff's alpha alongside exact-match, AB+BA position swaps, >=3 replicates, >=2 benchmarks spanning preference- and correctness-style labels, audit that high stability is not high bias || Asserted but not visible in retrievable text: 'March 2026' as an eval start month (page states results hold across 'the April 2026 frontier'; March not explicitly seen) || Re-check notes: Every precise figure the claim attributes to the HTML full text is present verbatim. The claim's '15 positions' matches Section 4.3 ('shifts 15 positions'), though the abstract and Figure 2 say 14 positions -- the paper is internally inconsistent on this number. Claim's rounded 0.99/0.19 corresponds to the paper's 0.992/0.192 for Qwen 3 8B.
Decision use: current-generation measurement. Measured on model-outputs (models: Qwen3-8B, Gemini 2.5). Same study family as AJ-03; June-2026 window.
Same underlying source as [AJ-03] [CT-01] - repetition across reports is not independent corroboration.
[CM-10] critique-models - grade A (holds regardless of model progress) - re-check 2026-07-15: confirmed
The closest production analog to the proposed AutoQA already exists: Toloka's deployed 'LLM QA' runs a tool-using agent per quality-metric on every human annotation submission, emits a strict three-way Pass / Fail / Unable-to-verify verdict (the abstain class exists specifically to prevent hallucinated verdicts), coaches annotators with Socratic feedback, and explicitly accepts lower precision on Fail because a false pass costs more than escalating a genuine pass.
Source: LLM QA: Scaling data quality assurance technologically (Toloka) (2026-03-30, practitioner)
What the source itself says (retrieved quote)
We enforce a deliberate constraint: one metric, one entity, one verdict.
Re-check detail
Checked and matching: agentic autocheck running on every submission; separate instance launched for every single quality metric; rigid three-way scale: Pass, Fail, or Unable to verify; abstain state exists because otherwise a model will hallucinate a guess; Socratic-style feedback that asks questions rather than pointing to the error; lower precision on Fail is an acceptable tradeoff; false pass costlier than sending a genuine pass to human review; tools: download web pages, view images, process audio/video, run Python, execute bash; 'one metric, one entity, one verdict' constraint; internal benchmark of over 300 real submissions from 20-plus live projects; ground truth from Toloka senior QA reviewers; performance entirely bounded by task design; ambiguity yields a wall of Unable to verify || Asserted but not visible in retrievable text: the actual benchmark performance numbers (claim itself notes these are in an image, not extractable) || Re-check notes: All eight sub-claims supported verbatim. Byline Vitaly Moiseev and Mariya Shmatova, dated March 30, 2026. Vendor self-description; benchmark scores are referenced but not displayed as text, consistent with the claim's own caveat.
Decision use: holds regardless of model progress. Measured on structural. Toloka's deployed pass/fail/unable-to-verify QA with human escalation. Existence proof of the architecture in production - vendor-reported, flagged as such.
Cited at: Decisions - authority boundaries
[CM-11] critique-models - grade A (holds regardless of model progress) - re-check 2026-07-15: confirmed
The foundational 2022 result behind the whole lineage: model-written critiques helped human evaluators find flaws in summaries they would otherwise have missed (including planted flaws in deliberately misleading human-written summaries), and models exhibit a discriminator-critique gap - they can often detect that something is wrong better than they can articulate why.
Source: Self-critiquing models for assisting human evaluators (OpenAI) (2022-06-12, primary)
What the source itself says (retrieved quote)
even large models may still have relevant knowledge they cannot or do not articulate as critiques
Re-check detail
Checked and matching: model-written critiques help humans find flaws in summaries they would otherwise have missed; surface intentional flaws in summaries humans wrote to be deliberately misleading; framework comparing generation, discrimination, and critique ability; 'even large models may still have relevant knowledge they cannot or do not articulate as critiques' (supports discriminator-critique gap); larger models write more helpful critiques and self-refine better (topic-based summarization) || Re-check notes: All elements of the claim and evidence confirmed from the abstract, including the verbatim quote. The 'discriminator-critique gap' is not named by that phrase but is the framework's finding.
Decision use: holds regardless of model progress. Measured on human-work. 2022 foundational result: model critiques help humans find flaws they'd miss. Floor-type: strengthens with critic quality.
[CM-12] critique-models - grade A (holds regardless of model progress) - re-check 2026-07-15: confirmed with caveats
A deployed LLM compliance-checker was manipulated by fabricated justifications: the NeurIPS 2024 author checklist assistant (234 papers) was rated useful by >70% of authors, but the organizers found the system 'not robust to gaming' by authors and concluded it is a poor substitute for human review.
Source: Usefulness of LLMs as an Author Checklist Assistant for Scientific Papers: NeurIPS'24 Experiment (2024-11-05, academic)
What the source itself says (retrieved quote)
over 70% of authors found the assistant useful ... could be manipulated to enhance scores through fabricated justifications ... a promising, but controversial, tool in aiding scientific peer review
Re-check detail
Checked and matching: 234 papers voluntarily submitted to the LLM-based Checklist Assistant; over 70% of authors found the assistant useful; gaming vulnerability: 'could be manipulated to enhance scores through fabricated justifications' || Asserted but not visible in retrievable text: specific model 'GPT-4-turbo' (the abstract does not name the model); conclusion that it is 'a poor substitute for human review' (abstract instead frames it as 'a promising, but controversial, tool in aiding scientific peer review'); exact phrase 'not robust to gaming' (substance present, phrase not) || Re-check notes: Core (234 papers, >70% useful, demonstrated gaming/manipulation) confirmed. But the claim's characterization that organizers 'concluded it is a poor substitute for human review' is not supported by the retrieved abstract, which is more favorable ('promising, but controversial'); and the GPT-4-turbo model is not named on the page. 'Gamed in a single shot' is an editorial gloss not in the text.
Decision use: holds regardless of model progress. Measured on human-work (models: GPT-4). Deployed checklist assistant gamed in one shot (NeurIPS 2024). Attackers improve with model tier; the design lesson is permanent.
Grounding and verification
[GF-01] grounding-faithfulness - load-bearing - grade C (superseded-generation number; mechanism only) - re-check 2026-07-15: confirmed
Grounded claim-verification is commoditized but ceilinged: the best model on the LLM-AggreFact benchmark (11 datasets of claim-vs-grounding-document verification) is a specialized 7B model, Bespoke-MiniCheck-7B, at 77.4 average balanced accuracy, with 0.4-0.8B specialized checkers (FactCG, MiniCheck-Flan-T5-L) within ~2.5 points of frontier LLMs like Claude-3.5-Sonnet (77.2) and GPT-4o (75.9).
Source: LLM-AggreFact Leaderboard (MiniCheck project) (2025 (accessed 2026-07-14), primary)
What the source itself says (retrieved quote)
Aggregates 11 datasets on grounded factuality (i.e., hallucination) evaluation
Re-check detail
Checked and matching: Bespoke-MiniCheck-7B tops the board at 77.4 average; Claude-3.5 Sonnet 77.2; Granite Guardian 3.3 8B 76.5; gpt-4o-2024-05-13 75.9; FactCG-DeBERTa-L (0.4B) 75.6; MiniCheck-Flan-T5-L (0.8B) 75.0; Llama-3.1-405B-Instruct 74.4; 11 datasets; grounded factuality / hallucination (claim supported vs source documents) || Re-check notes: All seven cited leaderboard numbers and the 11-dataset scope match exactly. 0.4-0.8B specialized checkers (75.6, 75.0) are within ~2.4 points of top model (77.4), consistent with the claim's '~2.5 points'.
Research-sweep audit (2026-07-14): confirmed
Fetched https://llm-aggrefact.github.io/ on 2026-07-14; every number matches exactly: Bespoke-MiniCheck-7B 77.4 (rank 1), Claude-3.5 Sonnet 77.2, Granite Guardian 3.3 8B 76.5, gpt-4o-2024-05-13 75.9, FactCG-DeBERTa-L (0.4B) 75.6, MiniCheck-Flan-T5-L (0.8B) 75.0, Llama-3.1-405B 74.4. Benchmark is 11 datasets of grounded factuality (supported/unsupported vs grounding docs), avg balanced accuracy - as claimed. The "within ~2.5 points" framing is accurate and even conservative: FactCG (75.6) is only 1.6 below Claude-3.5-Sonnet and above GPT-4o. Supersession search found nothing beating 77.4: ACV (May 2026) reports 76.5 training-free; an ACL 2026 paper ranks second to a post-trained metric; Paladin-mini beats Bespoke-MiniCheck only on its own separate benchmark subsets, not LLM-AggreFact overall. Caveats on the interpretive "ceilinged" framing: (1) the leaderboard's frontier entries are 2024-era models (Claude-3.5-Sonnet, gpt-4o-2024-05-13, Llama-3.1-405B) - no Claude 4/GPT-5-class/o3 entries, so "specialized ~ frontier" reflects 2024 frontiers; (2) "Verifying the Verifiers" (arXiv 2506.13342, June 2025) finds ~16% of benchmark labels ambiguous/incorrect and that few-shot frontier LLMs reach top-tier performance, suggesting the ~77 plateau is partly benchmark label noise rather than a pure task ceiling; (3) "Verify with Caution" (arXiv 2501.14883) shows models with similar aggregate BAcc make very different instance-level predictions. None of these contradict the stated facts.
Decision use: superseded-generation number; mechanism only. Measured on model-outputs (models: MiniCheck, Claude-3.5-Sonnet , GPT-4o, Claude-3.5-Sonnet 77.2, Llama-3.1-405B). The ~77 bacc AggreFact ceiling and specialists-within-2.5-points are GPT-4o/Llama-3.1-era. Deployment-tier entailment ceilings are unknown; the two-tier screening economics need repricing at the bakeoff.
Cited at: Research - findings by decision weight - System - claim types - Decisions - judge sourcing
[GF-02] grounding-faithfulness - load-bearing - grade B (current-generation measurement) - re-check 2026-07-15: confirmed
On adversarially-hard hallucination sets, specialized detectors collapse and few-shot anchoring with human-annotated exemplars is the measured fix: on FaithBench, prior detectors hit ~50% accuracy (negligible), HHEM-2.1-Open 66.7% and Bespoke-MiniCheck 60.1% claim-wise balanced accuracy (71.2% is its four-dataset average, not its FaithBench score), zero-shot frontier judges stay below 78%, while FaithJudge - prompting o3-mini-high with human-annotated peer responses to the same source document - reaches 84.0% balanced accuracy / 82.1 F1.
Source: Benchmarking LLM Faithfulness in RAG with Evolving Leaderboards (FaithJudge) (2025-11, academic)
What the source itself says (retrieved quote)
The highest effectiveness is achieved using the o3-mini-high judge, reaching a balanced accuracy of 84% and an F1-macro of 82.1%
Re-check detail
Checked and matching: prior detectors ~50% on FaithBench: 'current methods ... achieved near 50% accuracy, suggesting negligible ability'; HHEM-2.1-Open 66.7 (FaithBench claim-wise, asterisked value); Bespoke-MiniCheck 71.2 balanced accuracy (Table 1); zero-shot frontier judges: 'balanced accuracy below 78% and F1-macro below 72%'; FaithJudge (o3-mini-high) 84.0 balanced accuracy / 82.1 F1-macro (Table 2 and text); beats FACTS Grounding prompt on all four splits; FaithBench-Summary F1 54.3 (FACTS) vs 70.8 (FaithJudge) (Table 4); failure modes: underprediction for Command-R/Mistral/Qwen; Benign/Questionable ternary unreliable so binary only; specificity slightly decreases as more examples given; EMNLP Industry Track 2025; v2 Nov 6 2025 (matches source_date 2025-11) || Re-check notes: Full-text verification confirms every asserted number, upgrading the prior in_corpus_verdict (partially_confirmed) to confirmed. One nuance: the 66.7 for HHEM-2.1-Open is its FaithBench-specific claim-wise value (asterisked because HHEM helped select FaithBench's adversarial articles); its cross-dataset average is 67.1. Claim asserted 66.7, which matches the FaithBench value.
Research-sweep audit (2026-07-14): partially_confirmed
Verified against arXiv:2505.04847 (v2 dated Nov 6, 2025; EMNLP 2025 Industry Track; Vectara/Waterloo authors - source and date correct). Confirmed: (1) paper states prior detectors, including LLM classifiers, achieved "near 50% accuracy" on FaithBench; (2) HHEM-2.1-Open = 66.7% balanced accuracy on FaithBench (claim-wise, Table 1) - though flagged with an asterisk because HHEM was used to adversarially select FaithBench articles; (3) FaithJudge with o3-mini-high = 84.0% balanced accuracy / 82.1 F1-macro (Table 2), best zero-shot judge on FaithBench was o3-mini-high at 68.8%; (4) FaithJudge vs FACTS Grounding head-to-head 70.8 vs 54.3 F1 on FaithBench confirmed (Table 4); (5) all three stated failure modes confirmed (underprediction for Command-R/Mistral/Qwen generators; Benign/Questionable misclassified - only 10/84 Benign labeled correctly; specificity slightly decreases as in-context examples increase). ERRORS: (a) Bespoke-MiniCheck's 71.2 is NOT its FaithBench score - 71.2 is its balanced-accuracy AVERAGE across all four datasets (AggreFact, RAGTruth, TofuEval-MB, FaithBench) in Table 1; its actual FaithBench score is 60.1% claim-wise / 55.7% summary-wise. (b) Minor: the "below 78% balanced accuracy" figure for zero-shot judges is the paper's cross-dataset average claim, not a FaithBench-specific figure (on FaithBench zero-shot judges max out at 68.8%, so the claim still holds directionally). Supersession check: searches found no 2026 work surpassing FaithJudge on FaithBench; it remains the reported state of the art as of July 2026.
Corrected statement: On adversarially-hard hallucination sets, specialized detectors collapse and few-shot anchoring with human-annotated exemplars is the measured fix: on FaithBench, prior detectors hit ~50% accuracy (negligible); fine-tuned detectors stay weak (HHEM-2.1-Open 66.7% and Bespoke-MiniCheck 60.1% claim-wise balanced accuracy on FaithBench; Bespoke-MiniCheck averages 71.2% across the four benchmark datasets); zero-shot frontier judges reach at most 68.8% on FaithBench (and stay below 78% averaged across datasets), while FaithJudge - prompting o3-mini-high with human-annotated peer responses to the same source document - reaches 84.0% balanced accuracy / 82.1 F1-macro.
Decision use: current-generation measurement. Measured on model-outputs (models: MiniCheck, o3-mini). FaithJudge (late-2025, o3-mini-era). Exemplar-anchoring-beats-zero-shot is an architecture effect expected to persist; margins are tier-bound.
Cited at: Decisions - judge sourcing
[GF-03] grounding-faithfulness - load-bearing - grade C (superseded-generation number; mechanism only) - re-check 2026-07-15: confirmed with caveats
Claim decomposition helps weak verifiers but actively degrades strong ones: with MiniCheck as verifier on WiCE, no-decomposition scores 80.01 balanced accuracy while FActScore-style atomic decomposition drops it to 71.11; with the weaker AlignScore verifier the same decomposition improves results - and gains only reappear as input complexity grows, with best results when sub-claim count does not exceed input complexity.
Source: Decomposition Dilemmas: Does Claim Decomposition Boost or Burden Fact-Checking Performance? (2025-05, academic)
What the source itself says (retrieved quote)
introduce a categorization of decomposition errors and reveal a trade-off between accuracy gains and the noise introduced
Re-check detail
Checked and matching: 'Decompose-Then-Verify paradigm' present; authors 'introduce a categorization of decomposition errors and reveal a trade-off between accuracy gains and the noise introduced' || Asserted but not visible in retrievable text: MiniCheck; WiCE; 80.01 vs 71.11 balanced accuracy; AlignScore; FActScore; VeriScore; FELM; F1 48.1->~68 and 71.6->54.3; the specific 'helps weak verifiers / degrades strong ones' direction; the four-way A/B/C/D decomposition-error taxonomy; 'best results when sub-claim count does not exceed input complexity' || Re-check notes: Fetch constraints allow only the given ACL landing page (no PDF or other same-site page, and the URL contains no arXiv ID). The abstract confirms the general trade-off framing and that decomposition errors are categorized, but none of the specific verifiers, datasets, or numbers appear, and the abstract does not explicitly state the weak-vs-strong-verifier direction. Everything specific in the claim is not visible in the retrieved text.
Research-sweep audit (2026-07-14): confirmed
Verified against the full PDF (extracted text at /private/tmp/claude-501/-/640d77df-bd94-4e91-8873-f5bb2df7b27d/scratchpad/paper.txt). (1) Source says exactly this: Table 2 (WiCE) shows MiniCheck baseline BAcc 80.01 vs FActScore decomposition 71.11 (F1 72.32 vs 59.90); all decomposition methods hurt MiniCheck on WiCE. With the weaker AlignScore verifier, FActScore decomposition improves BAcc 54.80->56.87 (and WiCE-style 56.26); note VeriScore decomposition slightly hurt AlignScore too, but the claim as stated ("the same decomposition," i.e., FActScore-style) is accurate. Section 4.2 states verbatim that "decomposition generally benefits weaker verifiers, while it tends to negatively affect stronger verification systems." Section 6.2-6.4 confirms gains reappear as input complexity grows (complexity scale-up experiments; FELMshort scale-down degrades), and Figure 3 discussion states "for each level, the maximum F1 is observed when the number of the decomposed sub-claim is less than or equal to the complexity level." Evidence-summary side facts also check: FELM MiniCheck F1 48.10->67.5-68.1, GPT-4o-mini F1 71.56->54.34 (Table 3); four-way error taxonomy (context omission, ambiguity, over-decomposition, meaning alteration) present in Section 5. (2) Date: ACL Anthology lists April 2025 (NAACL 2025, Albuquerque, pp. 6313-6336); the conference ran Apr 29-May 4, so "2025-05" is a trivial one-month imprecision, not a substantive error. (3) Supersession search: later work (e.g., presupposition-free question decomposition, arXiv 2508.16838; "Alignment Bottleneck in Decomposition-Based Claim Verification," Feb 2026) extends the decomposition-tradeoff line but does not refute the verifier-strength finding.
Decision use: superseded-generation number; mechanism only. Measured on model-outputs (models: MiniCheck, GPT-4o-mini). MiniCheck/GPT-4o-mini-era decomposition harm. Adaptive granularity stays the design default; the 80->71 magnitude is not a current estimate.
Cited at: System - claim types
[GF-04] grounding-faithfulness - grade B (current-generation measurement) - re-check 2026-07-15: confirmed
Decomposition and decontextualization are in direct tension - isolating atomic facts strips the context needed to verify them, while adding context back creates multi-fact claims where the verifier may credit or penalize the wrong content; DnDScore resolves this by verifying the original subclaim WITH the added information treated as context rather than as content to be verified.
Source: DnDScore: Decontextualization and Decomposition for Factuality Verification in Long-Form Text Generation (2025-11, academic)
What the source itself says (retrieved quote)
decomposition isolates atomic facts while decontextualization inserts relevant information ... they present "DnDScore, a decontextualization aware verification method that validates subclaims in the context of contextual information" ... "the choice of strategy matters in the resulting factuality scores."
Re-check detail
Checked and matching: decomposition and decontextualization have conflicting purposes; decomposition isolates atomic facts while decontextualization inserts relevant information; adding context back creates multi-fact text ('what part of the augmented text should be verified'); DnDScore validates subclaims in the context of contextual information (context, not content); strategy choice materially changes factuality scores; pages 23609-23626; authors Wanner, Van Durme, Dredze || Re-check notes: Core tension and the resolution (verify subclaim WITH added info as context rather than content) confirmed verbatim from the ACL Anthology abstract. Title, authors, pages all match.
Decision use: current-generation measurement. Measured on model-outputs. 2025 decomposition/decontextualization tension analysis. Design-level tension, partially structural.
Cited at: System - claim types
[GF-05] grounding-faithfulness - grade A (holds regardless of model progress) - re-check 2026-07-15: confirmed with caveats
Verifiability triage before entailment is established machinery: VeriScore extracts and scores ONLY verifiable claims because FActScore/SAFE 'assume that every claim is verifiable', which breaks on real long-form text containing opinions and unverifiable content; FactBench's VERIFY pipeline (ACL 2025) further labels content units supported/unsupported/UNDECIDABLE against retrieval, with 4,467 human-annotated units released for validation.
Source: VeriScore (EMNLP Findings 2024) + FactBench/VERIFY (ACL 2025) (2024-06 (VeriScore, foundational); 2025-07 (FactBench ACL 2025), academic)
What the source itself says (retrieved quote)
they assume that every claim is verifiable (i.e., can plausibly be proven true or false) ... human evaluation confirms that VERISCORE's extracted claims are more sensible than those from competing methods
Re-check detail
Checked and matching: prior methods FActScore/SAFE 'assume that every claim is verifiable (i.e., can plausibly be proven true or false)'; VERISCORE handles 'both verifiable and unverifiable content'; human evaluation confirms that VERISCORE's extracted claims are more sensible than those from competing methods; across eight different long-form tasks || Asserted but not visible in retrievable text: FactBench and its VERIFY pipeline -- the words 'FactBench' and 'VERIFY' do not appear on this page; three-way supported/unsupported/UNDECIDABLE labeling against retrieval; 4,467 human-annotated content units; ACL 2025 venue for FactBench; 'EMNLP Findings 2024' venue -- the arXiv page shows no conference venue, only arXiv || Re-check notes: The VeriScore half of this compound claim is fully confirmed. The FactBench/VERIFY half (three-way labeling, 4,467 units, ACL 2025) comes from a separate paper (2410.22257) that the single-URL constraint prohibits fetching, so it is not present in retrieved text. Minor: claim says VeriScore scores 'ONLY verifiable claims' while the abstract frames it as handling 'both verifiable and unverifiable content'. title_match=false because source_title is a composite two-paper label, not the actual paper title, and asserts an EMNLP venue not shown on the page.
Decision use: holds regardless of model progress. Measured on structural. Verifiability triage before entailment (VeriScore/FactBench). A pipeline design pattern, not a capability claim.
Cited at: System - claim types - Decisions - default: entailment standard
[GF-06] grounding-faithfulness - grade C (superseded-generation number; mechanism only) - re-check 2026-07-15: confirmed
Holistic LLM judging of context-grounded outputs is far weaker than claim-level checking: on ContextualJudgeBench (2,000 pairs, 8 splits over RAG/summarization with conditional criteria like 'faithfulness first, then completeness'), the best of 20 judge models tested (OpenAI o1) barely reaches 55% consistent accuracy.
Source: Does Context Matter? ContextualJudgeBench for Evaluating LLM-based Judges in Contextual Settings (2025-03, academic)
What the source itself says (retrieved quote)
we propose ContextualJudgeBench, a judge benchmark with 2,000 challenging response pairs across eight splits ... Our comprehensive study across 11 judge models and 9 general purpose models ... OpenAI's o1, the best-performing model, barely reaches 55% consistent accuracy.
Re-check detail
Checked and matching: 2,000 challenging response pairs; eight (8) splits; 20 judge models (11 judge models + 9 general-purpose models); OpenAI o1 best-performing, barely reaches 55% consistent accuracy; conditional evaluation criteria (factuality then completeness); RAG and summarization settings || Asserted but not visible in retrievable text: Salesforce affiliation (not stated on the abstract page); ACL 2025 venue (not listed on the abstract page) || Re-check notes: All core numbers confirmed. Minor terminology: the claim's illustrative criterion 'faithfulness first, then completeness' appears in the abstract as 'factuality and then considering completeness' (factuality vs faithfulness). The evidence field's 'Salesforce; ACL 2025' is not visible on the abstract page but is not part of the core claim.
Decision use: superseded-generation number; mechanism only. Measured on model-outputs (models: o1). ContextualJudgeBench ~55% on o1-era judges. Holistic-vs-claim-level gap direction is consistent across studies; magnitude stale.
[GF-07] grounding-faithfulness - load-bearing - grade B (current-generation measurement) - re-check 2026-07-15: confirmed
Grounding judges have a measured agreement-default asymmetry: in Google's FACTS Leaderboard (Dec 2025), grounding judges score ~85 F1 on the positive (grounded) class but only ~46 F1 on the negative (ungrounded) class, and the grounding metric is gameable by vague, short responses that avoid unsupported claims - countered by a mandatory eligibility gate that scores non-responsive answers as inaccurate.
Source: The FACTS Leaderboard: A Comprehensive Benchmark for Large Language Model Factuality (2025-12, academic)
What the source itself says (retrieved quote)
gemini-2.5-flash with the v2 prompt, scores Macro-F1 65.33, with F1(+) of 84.51 but F1(-) of only 46.15 ... the final factuality score is adjusted such that ineligible responses are deemed as inaccurate
Re-check detail
Checked and matching: top combo gemini-2.5-flash + v2 prompt: Macro-F1 65.33; F1(+) 84.51 vs F1(-) 46.15 (positive/negative class asymmetry, ~85 vs ~46); held-out human-adjudicated evaluation set N=320 (class ratio 79:19); eligibility gate: ineligible/non-responsive answers deemed inaccurate; explicit rationale that grounding metrics can be hacked by vague, short responses || Re-check notes: All four asserted specifics confirmed in the full text. The claim's mapping of F1(+)/F1(-) to 'grounded'/'ungrounded' class is a reasonable interpretation (majority positive class 79:19); paper labels them F1(+)/F1(-). The ~85/~46 figures are the top combo's, which the claim uses to characterize grounding judges generally.
Decision use: current-generation measurement. Measured on model-outputs (models: gemini-2.5-flash). FACTS (Dec 2025, gemini-2.5-flash graders). The agreement-default asymmetry (85/46) is the single most decision-relevant grounding mechanism - and it was measured on flash-tier graders. Deployment-tier asymmetry is E2's first number.
Cited at: Overview - what the evidence supports - Research - findings by decision weight - System - claim types - Decisions - default: entailment standard
[GF-08] grounding-faithfulness - grade B (current-generation measurement) - re-check 2026-07-15: confirmed with caveats
Cross-family judge ensembles are the standard mitigation for self-preference: FACTS Grounding v1 (Jan 2025) aggregated three frontier judges from different families (Gemini 1.5 Pro, GPT-4o, Claude 3.5 Sonnet) explicitly because 'models are biased towards favorably judging their own outputs', and Grounding v2 (Dec 2025) kept a two-family ensemble (Gemini 2.5 Flash + GPT-5).
Source: FACTS Grounding Leaderboard (v1 paper + v2 in FACTS Leaderboard) (2025-01 (v1); 2025-12 (v2), academic)
What the source itself says (retrieved quote)
We use three different judge models in order to reduce the bias of a particular judge model ... as models have been shown to be biased towards favorably judging their own outputs
Re-check detail
Checked and matching: v1 uses three judge models from three providers: Gemini 1.5 Pro (Google), GPT-4o (OpenAI), Claude 3.5 Sonnet (Anthropic); rationale is self-preference: 'models have been shown to be biased towards favorably judging their own outputs'; aggregate of multiple judge models to mitigate evaluation bias; v1 dated Jan 2025 || Asserted but not visible in retrievable text: FACTS Grounding v2 (Dec 2025) two-family ensemble Gemini 2.5 Flash + GPT-5 (asserted from arXiv:2512.10791, a different paper not fetchable under the constraint); the DeepMind blog (Dec 17, 2024) v1 design write-up; FACTS Parametric single-judge-preserves-rankings validation (Gemini 2.5 Pro, o3, Grok 4) || Re-check notes: The v1 core (cross-provider ensemble as self-preference mitigation) is fully confirmed from the given URL's full text. Paper says 'three different judge models' from three providers rather than literally 'different families' - substantively equivalent. All v2 / Parametric specifics require arXiv:2512.10791 and external pages, which the fetch constraint excludes, so they are marked not_visible rather than checked.
Decision use: current-generation measurement. Measured on model-outputs (models: Gemini 1.5, GPT-4o, Claude 3.5 Sonnet, Gemini 2.5, GPT-5, o3). FACTS v1->v2 ensemble evolution; includes GPT-5-class by v2. Cross-provider ensembling as self-preference mitigation is practice-level.
[GF-09] grounding-faithfulness - grade B (current-generation measurement) - re-check 2026-07-15: confirmed
Span-restricted NLI citation checking (the ALCE-style 'does the cited span entail the sentence' standard) is judged a suboptimal proxy by 2025 work: CiteEval (ACL 2025) argues citation quality must be evaluated against the FULL retrieval context, user query, and generated text - not cited sources alone - and its model-based CiteEval-Auto metrics correlate better with human judgments on the multi-domain CiteBench than NLI-based metrics.
Source: CiteEval: Principle-Driven Citation Evaluation for Source Attribution (2025-06, academic)
What the source itself says (retrieved quote)
current frameworks mainly rely on Natural Language Inference (NLI) to assess binary or ternary supportiveness from cited sources, which we argue is a suboptimal proxy for citation evaluation ... not only the cited sources but the full retrieval context, user query, and generated text
Re-check detail
Checked and matching: NLI binary/ternary supportiveness judged 'a suboptimal proxy for citation evaluation'; evaluates full retrieval context, user query, and generated text (not cited sources alone); CiteBench multi-domain human-annotated benchmark; CiteEval-Auto model-based metrics with strong correlation to human judgments || Asserted but not visible in retrievable text: 'ALCE' (not named on page); explicit head-to-head numeric comparison 'better than NLI-based metrics' (entailed by 'suboptimal proxy' + 'strong correlation' but not stated as a number); CiteGuard (ACL 2026) corroboration (separate source, out of scope) || Re-check notes: Core assertion and every substantive element (suboptimal NLI proxy, full retrieval context, CiteBench, CiteEval-Auto human-correlation) confirmed verbatim in the abstract.
Decision use: current-generation measurement. Measured on model-outputs. 2025: span-restricted citation checking is insufficient. Method-level; motivates the dual-pass citation design.
Cited at: Research - findings by decision weight - System - claim types - Pilot - gates and objectives
[GF-10] grounding-faithfulness - grade B (current-generation measurement) - re-check 2026-07-15: confirmed
Fact-verification LLMs are brittle to semantically-minor input perturbations: FactEval (NAACL 2025) tested 17 realistic word- and character-level perturbations plus 4 subpopulations on FEVER across zero-shot/few-shot/CoT setups and found LLMs 'brittle to small input changes' with performance varying across subpopulations.
Source: FactEval: Evaluating the Robustness of Fact Verification Systems in the Era of Large Language Models (2025-04, academic)
What the source itself says (retrieved quote)
LLMs are brittle to small input changes and also exhibit performance variations across different subpopulations.
Re-check detail
Checked and matching: 17 realistic word-level and character-level perturbations and 4 types of subpopulations; built on FEVER; zero-shot, few-shot, and chain-of-thought prompting; LLMs brittle to small input changes and performance variations across subpopulations; authors Mamta and Oana Cocarascu; NAACL 2025 || Asserted but not visible in retrievable text: per-perturbation quantitative results (abstract is qualitative; claim does not assert specific numbers) || Re-check notes: Abstract confirms every asserted specific (17 perturbations, 4 subpopulations, FEVER, three prompting setups, brittleness finding). The claim explicitly defers per-perturbation numbers to the full paper, so nothing overclaimed.
Decision use: current-generation measurement. Measured on model-outputs. FactEval (2025): verifier verdicts flip under meaning-preserving perturbation. Perturbation harness stays a ship gate; deployment-tier flip rates unknown (could be far lower - measure, don't assume).
Cited at: Pilot - gates and objectives
[GF-11] grounding-faithfulness - grade B (current-generation measurement) - re-check 2026-07-15: confirmed
2026 SOTA joins verdict and explanation in one cheap model: FaithLens (ACL 2026 Findings) is an 8B faithfulness-hallucination detector that jointly outputs a binary prediction AND an explanation, trained via cold-start fine-tuning on filtered synthetic data plus rule-based RL rewarding both prediction correctness and explanation quality, and outperforms GPT-5.2 and o3 across 12 tasks.
Source: FaithLens: Detecting and Explaining Faithfulness Hallucination (2026-07 (ACL 2026), academic)
What the source itself says (retrieved quote)
a distinctive balance of trustworthiness, efficiency, and effectiveness
Re-check detail
Checked and matching: 8B faithfulness-hallucination detection model; jointly outputs binary predictions plus explanations; fine-tuned on filtered LLM-synthesized data as a cold start; further optimized with rule-based reinforcement learning rewarding prediction correctness and explanation quality; beats GPT-5.2 and o3 across 12 diverse tasks; pages 14068-14099, Findings of ACL 2026, Si et al. || Asserted but not visible in retrievable text: per-task numeric scores (require full PDF; abstract only, consistent with the claim's own caveat) || Re-check notes: Every headline assertion in the claim is present in the ACL Anthology abstract; page range and authorship match the evidence.
Decision use: current-generation measurement. Measured on model-outputs (models: GPT-5.2, o3). FaithLens (ACL 2026, GPT-5.2-era comparison) - among the most current capability anchors in the corpus. Still model-output verification, not human-writeup QA.
[GF-12] grounding-faithfulness - grade C (superseded-generation number; mechanism only) - re-check 2026-07-15: confirmed with caveats
RAGAS-style prompt-chain faithfulness scoring materially underperforms finetuned/frontier judges on hallucination detection: on the 15K-sample HaluBench, RAGAS Faithfulness scored 66.9% accuracy versus 87.4% for finetuned Lynx-70B and 86.5% for GPT-4o; Patronus has since shipped Lynx 2.0 (8B, long-context, 8 hallucination subtypes including coreference and calculation errors).
Source: Patronus AI Lynx / HaluBench results (2024-07 (Lynx 1.0, foundational); Lynx 2.0 later update, practitioner)
What the source itself says (retrieved quote)
a comprehensive hallucination evaluation benchmark consisting of 15k samples
Re-check detail
Checked and matching: HaluBench is 15k samples; Lynx is positioned as SOTA open-source hallucination detector, beating frontier judges (relative): Lynx-70B 8.3% more accurate than GPT-4o on PubMedQA; Lynx-8B beat GPT-3.5 by 24.5% on HaluBench || Asserted but not visible in retrievable text: RAGAS Faithfulness 66.9%; Lynx-70B 87.4%; GPT-4o 86.5%; GPT-4-Turbo 85.0%; Llama-3-70B 80.1%; the HaluBench accuracy comparison table itself (this blog reports only relative % improvements, not the absolute-accuracy table); Lynx 2.0, 8B long-context model, 8 hallucination subtypes (coreference/calculation) || Re-check notes: The given URL confirms the 15k HaluBench size and the directional finding (finetuned Lynx > frontier judges), but NONE of the specific accuracy numbers cited in the claim (66.9/87.4/86.5/85.0/80.1) appear on this page; they are attributed in the evidence to the Lynx paper/table, not this blog. Lynx 2.0 and its subtypes are not mentioned here at all. Resolved title is the literal page title, which differs from the descriptive source_title.
Decision use: superseded-generation number; mechanism only. Measured on model-outputs (models: Lynx, GPT-4o, GPT-4, Llama-3-70B). Lynx-vs-GPT-4-era comparison (2024). Fine-tuned-specialist economics need current repricing.
Statistics for sparse truth
[HS-01] hybrid-statistical - load-bearing - grade A (holds regardless of model progress) - re-check 2026-07-15: confirmed with caveats
Prediction-Powered Inference (PPI/PPI++) combines a small human gold set with a large judge-labeled set to produce unbiased estimates with always-valid confidence intervals - coverage holds regardless of judge quality (a worse judge widens intervals, never invalidates them) - and increased effective human sample size by up to 50% with a GPT-4 judge in 'AutoEval Done Right'.
Source: AutoEval Done Right: Sample-Efficient Human Evaluation via Prediction-Powered Inference (2024-03 (v3 camera-ready 2026-06-01), academic)
What the source itself says (retrieved quote)
even with a very poor annotator model, PPI++ performs at least as well as the classical approach
Re-check detail
Checked and matching: core method is prediction-powered inference (PPI) and the optimized variant PPI++; combines a small amount of human data with a large amount of synthetic/judge data to increase effective human sample size without compromising validity; estimators remain unbiased; PPI/PPI++ yield valid, calibrated confidence intervals across all labeled set sizes; robust to judge quality: even with a very poor annotator model, PPI++ performs at least as well as the classical approach (worse judge widens, never invalidates); v3 camera-ready dated 2026-06-01 || Asserted but not visible in retrievable text: the specific figure '50% with a GPT-4 judge': the LLM/judge experiment used gpt-4o-mini and gave only 20-35% ESS improvement; the ~50% ESS gain came from the ImageNet and protein-fitness experiments, not the GPT-judge experiment; GLIDE (arXiv:2605.31278), PRECISE/AAAI 2026, arXiv:2606.05308 -- separate papers, not part of this source || Re-check notes: TITLE MISMATCH: source_title says 'Sample-Efficient Human Evaluation via Prediction-Powered Inference' but the actual (v3) title is 'Using Synthetic Data for Model Evaluation'. PPI/PPI++, unbiasedness, valid/calibrated CIs, and coverage-robustness-to-judge-quality are all directly confirmed in the full HTML text. The abstract does say 'up to 50% on experiments with GPT-4', but the body attributes ~50% ESS to ImageNet/protein experiments while the actual LLM/GPT-style judge (gpt-4o-mini) experiment yielded only 20-35% -- so the '50% with a GPT-4 judge' specificity is imprecise. Kept at partially_confirmed for that discrepancy plus the title mismatch.
Research-sweep audit (2026-07-14): partially_confirmed
Verified directly against arXiv:2403.07008: authors (Boyeau, Angelopoulos, Yosef, Malik, Jordan) correct; abstract verbatim contains 'increase the effective human-labeled sample size by up to 50% on experiments with GPT-4'; abstract claims methods 'improve sample efficiency while remaining unbiased'; version history confirms v1 2024-03-09 and v3 2026-06-01 with comment 'camera-ready paper version' - dates as claimed. GLIDE (arXiv:2605.31278, Martinon/Merad/Raki, v1 2026-05-29, ICML 2026 workshop) exists and states PPI 'combines both into debiased estimates with valid confidence intervals', consistent with the claim, though the exact phrase 'regardless of judge quality' was not verifiable in its abstract. The 2026 PPI-for-LLM-judge literature is real and active (PRECISE/Amazon Science, arXiv:2606.05308 ranking extension, arXiv:2601.05420, arXiv:2601.20913), so the claim is not superseded - it is being extended. Two corrections: (1) 'always-valid confidence intervals' is a misuse of a term of art - 'always-valid'/anytime-valid refers to sequential inference; PPI/PPI++ intervals are fixed-n ASYMPTOTIC (CLT-based) intervals, so coverage is guaranteed only asymptotically and finite-sample coverage can degrade with very small gold sets. (2) The judge-quality-independence property is a general PPI/PPI++ property (Angelopoulos et al. 2023) rather than something stated in the AutoEval abstract itself; the claim's framing correctly captures its substance (worse judge -> wider intervals via lower correlation, coverage preserved) but slightly overstates its strength and its provenance in this specific source.
Corrected statement: Prediction-Powered Inference (PPI/PPI++) combines a small human gold set with a large judge-labeled set to produce unbiased estimates with asymptotically valid confidence intervals whose coverage does not depend on judge accuracy (a worse judge widens intervals rather than breaking coverage, subject to standard asymptotic/regularity conditions and a sufficiently large gold set); 'AutoEval Done Right' (arXiv:2403.07008, Boyeau, Angelopoulos, Yosef, Malik & Jordan) reports increasing the effective human-labeled sample size by up to 50% in experiments with GPT-4.
Decision use: holds regardless of model progress. Measured on structural (models: GPT-4). PPI/PPI++ unbiasedness and judge-robust intervals are mathematics; they get MORE useful as judges improve. (Empirical demos in the paper are GPT-4-era; the method is the decision input. Title changed on arXiv v3 - noted in re-check.)
Cited at: Overview - the five authorizations - Decisions - first project and gold funding - Decisions - default: statistics staging
[HS-02] hybrid-statistical - load-bearing - grade A (holds regardless of model progress) - re-check 2026-07-15: confirmed
Cascaded Selective Evaluation (Trust or Escalate, ICLR 2025 Oral) delivers a provable, user-specified human-agreement guarantee for LLM judges by calibrating a confidence threshold with fixed-sequence testing on a small human calibration set, escalating low-confidence items up a judge cascade and abstaining (to humans) when even the strongest judge is unconfident - achieving guaranteed >80% human agreement at ~80% coverage on a Chatbot Arena subset where GPT-4 alone almost never reaches 80% agreement, with ~88% of judged items handled by much cheaper models.
Source: Trust or Escalate: LLM Judges with Provable Guarantees for Human Agreement (ICLR 2025 Oral) (2024-07-25 (ICLR 2025-04), academic)
What the source itself says (retrieved quote)
among which 88.1% are evaluated by substantially cheaper Mistral-7B or GPT-3.5 instead of GPT-4 ... we adopt fixed-sequence testing instead of Bonferroni correction
Re-check detail
Checked and matching: Cascaded Selective Evaluation and Simulated Annotators confirmed (abstract); provable, user-specified human-agreement guarantee confirmed; >80% agreement at ~80% coverage: 'guarantees over 80% human agreement with almost 80% test coverage'; precise 'covering 79.1% of all samples'; GPT-4 alone almost never reaches 80% agreement on ChatArena subset (abstract); ~88% by cheaper models: '88.1% are evaluated by substantially cheaper Mistral-7B or GPT-3.5 instead of GPT-4'; fixed-sequence testing: 'we adopt fixed-sequence testing instead of Bonferroni correction'; small human calibration set: 'given access to a small calibration set' of human preferences (size 500 / 392 in experiments) || Asserted but not visible in retrievable text: 'ICLR 2025 Oral' venue/status -- arXiv page carries no journal-ref or comments field and lists only v1 (July 2024); the evidence cited iclr.cc, which is outside the allowed fetch scope, so venue cannot be confirmed from the source || Re-check notes: Every substantive mechanism and number in the claim is confirmed from the arXiv HTML full text (the PDF returned only compressed binary). The sole unverifiable item is the ICLR 2025 Oral label, which the arXiv record does not carry.
Research-sweep audit (2026-07-14): confirmed
All load-bearing elements check out against the primary sources. (1) arXiv:2407.18370 abstract (fetched directly) states Cascaded Selective Evaluation provides a provable human-agreement guarantee at a user-specified level, and that on a Chatbot Arena subset "where GPT-4 almost never achieves 80% human agreement," the method "guarantees over 80% human agreement with almost 80% test coverage" even using Mistral-7B. (2) Paper body (arxiv.org/html/2407.18370v1) confirms the mechanism: threshold calibration via fixed sequence testing (Bauer, 1991) on a small human calibration set (|D_cal|=500 for ChatArena/TL;DR, 392 for Auto-J), cascade escalation from cheap to strong judges on low confidence, abstention (output empty set) when even the strongest judge is unconfident, and Simulated Annotators as the confidence estimator enabling high coverage. (3) The ~88% figure is exact: "79.1% of all samples, among which 88.1% are evaluated by substantially cheaper Mistral-7B or GPT-3.5 instead of GPT-4" (Table 3: GPT-4 handles only 17.5%). (4) ICLR 2025 Oral confirmed at iclr.cc/virtual/2025/oral/31838 (Jung, Brahman, Choi); dates correct (v1 2024-07-25; ICLR 2025 in April 2025). Two trivial nuances, neither rising to a correction: the paper defines abstention as returning empty set / excluding from coverage - deferral "to humans" is the intended framing but abstained items are not actually routed to human annotators in the experiments; and the guarantee holds with probability 1-delta over calibration-set sampling (standard for such guarantees, and consistent with "provable, user-specified"). Supersession check: not superseded - but a newer alternative exists: SCOPE (arXiv:2602.13110, ICML 2026 poster, Feb 2026) uses conformal calibration with a Bidirectional Preference Entropy signal and reports better calibration/coverage than Simulated Annotators; it complements rather than invalidates the claim about this paper. The "83 citations" figure was not independently verified but is plausible and not load-bearing.
Decision use: holds regardless of model progress. Measured on model-outputs (models: GPT-4). Trust-or-Escalate's guarantee construction is method-level and durable. Its reported coverage numbers are GPT-4-era; recalibrate on deployment-tier judges - the framework exists precisely to do that.
[HS-03] hybrid-statistical - load-bearing - grade A (holds regardless of model progress) - re-check 2026-07-15: confirmed
The minimum viable human gold set is computable in closed form: with a doubly-robust two-stage design, required human labels converge to a floor of n*(1-rho^2) where n* is the target effective sample size and rho the judge-human correlation - e.g., a target effective n=200 with R^2=0.70 and 2,000 judge ratings needs only 65 human labels - with diminishing returns from adding more judge labels and up to ~13% further savings from stratified allocation when judge reliability varies across evaluation axes.
Source: Augmenting Human Evaluation with LLM Judges: How Many Human Reviews Do You Need? (2026-05-08, academic)
What the source itself says (retrieved quote)
With an LLM sample size of N=2000, only 65 human samples are needed.
Re-check detail
Checked and matching: floor formula n*(1-rho^2): 'converges to a floor of n*(1-rho2)'; worked example: n*=200, R2=0.70, N=2000 -> 'only 65 human samples are needed'; N=400 -> 100: 'If the budget can allow for 100 human ratings, then the LLM sample size N can be reduced to 400'; doubly robust two-stage sampling design (missingness known by design); stratification 'reduces the human budget by up to 12.9%' (claim's ~13%); caveats: pilot overestimation of R2; human gold standard a 'pragmatic choice rather than a settled fact'; single-rater scope; author Jane Paik Kim, Stanford Dept. of Psychiatry || Re-check notes: All specifics verified from the given HTML full text. Claim's '~13%' stratified savings equals the paper's precise 12.9%. Formula, both worked examples (65 and 100 human labels), diminishing returns, and all three caveats match.
Research-sweep audit (2026-07-14): confirmed
All elements verified against the source (https://arxiv.org/html/2605.16354v1, Jane Paik Kim, Stanford Dept. of Psychiatry and Behavioral Sciences, submitted 2026-05-08 - date and attribution correct). (1) Floor formula: paper states required human reviews converge to n*(1-rho^2) as N increases - exact match. (2) Worked example: n*=200, R^2=0.70, N=2000 -> 65 human labels, confirmed verbatim; the N=400/100 figure is stated in the paper as "with a budget of 100 human ratings, N can be reduced to 400 while achieving target power" - same numbers, slightly different framing (human budget fixed, N solved) but mathematically equivalent to the claim. (3) Stratification savings: paper reports "up to 12.9%" reduction when R^2 gap is large (0.8 vs 0.1) - matches "~13%", with the paper adding that moderate gaps yield only ~2%. (4) Diminishing returns from more judge labels: explicitly stated ("each successive increase in N produces diminishing reductions in n"). (5) All three caveats in the evidence summary (pilot R^2 overestimation breaking precision guarantees, human-as-gold-standard assumption, single-rater scope) appear in the paper's limitations. Supersession check: two searches found no v2 revision, no citing papers, and no later work replacing or contradicting the result as of 2026-07-14. Minor precision notes only: "up to ~13%" should strictly be "up to 12.9% in the extreme R^2-gap case, ~2% for moderate gaps," and the floor is asymptotic (approached as N grows), not attainable at finite N - the claim's own phrasing ("converge to a floor," "up to") already reflects both.
Decision use: holds regardless of model progress. Measured on structural. Closed-form gold-set sizing. Mathematics.
Cited at: Research - findings by decision weight - Decisions - first project and gold funding
[HS-04] hybrid-statistical - load-bearing - grade A (holds regardless of model progress) - re-check 2026-07-15: confirmed with caveats
Using the LLM's own confidence signals to choose WHICH items receive human annotation (Confidence-Driven Inference, building on Active Statistical Inference) cut required human annotations by >25% in all three tested settings while keeping provably valid confidence intervals that remain safe even if the LLM annotations are poor - but a 2026 ICLR paper finds the opposite in the sequential regime, where near-uniform sampling at the budget ceiling beat uncertainty-driven querying.
Source: Can Unconfident LLM Annotations Be Used for Confident Conclusions? (NAACL 2025) + Active Statistical Inference (ICML 2024) (2024-08-27 (v2 2025-02-08; NAACL 2025-04), academic)
What the source itself says (retrieved quote)
The paper introduces "Confidence-Driven Inference: a method that combines LLM annotations and LLM confidence indicators" ... it reduces "the needed number of human annotations by over 25% in each" setting.
Re-check detail
Checked and matching: Confidence-Driven Inference combines LLM annotations and LLM confidence to pick which items get human annotation; reduces needed human annotations by over 25% in each of three settings (politeness, stance, bias); safeguards against poor-quality LLM annotations; conclusions valid and no less accurate than human-only; provably valid confidence intervals || Asserted but not visible in retrievable text: 'Active Statistical Inference' foundation (attributed to separate paper arXiv:2403.03208, not on this page); 2026 ICLR counter-finding (arXiv:2604.18569) that near-uniform sampling beats uncertainty-driven querying || Re-check notes: The core claim attributable to this URL (>25% reduction in all three settings, validity regardless of LLM quality) is fully confirmed verbatim. Marked partially_confirmed because two asserted specifics are external to this source by the claim's own admission (Active Statistical Inference foundation = arXiv:2403.03208; the ICLR-2026 counter-finding = arXiv:2604.18569) and neither is visible on the fetched abstract. Authors: Gligoric, Zrnic, Lee, Candes, Jurafsky.
Decision use: holds regardless of model progress. Measured on structural. Confidence-driven human-annotation allocation. Method.
[HS-05] hybrid-statistical - grade A (holds regardless of model progress) - re-check 2026-07-15: confirmed
GLIDE (May 2026) is an open-source industrial library unifying PPI++, stratified PPI, predict-then-debias bootstrap, and active inference with a decision tree for method selection, and it quantifies when PPI helps: effective-sample gain is ~1.0x at judge-human correlation rho=0.1, ~2.2x at rho=0.9; CLT-based estimators need roughly >=50 human labels per stratum (below that use bootstrap PTD); in a real agent-safety case (rho=0.59, ~13-point judge bias) 100 human labels became worth 143-157.
Source: Industrializing Prediction-Powered Inference: The GLIDE Library for Reliable GenAI and Agentic Systems Evaluation (2026-05 (v2 2026-06), practitioner)
What the source itself says (retrieved quote)
the effective sample size grows from approximately 500 at rho=0.1 to approximately 1100 at rho=0.9 ... a 2.2x effective gain at no cost to validity
Re-check detail
Checked and matching: unifies state-of-the-art PPI estimators (PPI++, Stratified PPI, Predict-Then-Debias and its stratified variants, Active Statistical Inference); an empirically grounded decision tree for method selection; effective sample size grows from ~500 at rho=0.1 to ~1100 at rho=0.9 (a 2.2x effective gain), from 500 human + 1000 proxy labels; at least fifty labeled samples per stratum for the asymptotic intervals to be reliable; R-Judge case study; proxy is claude-sonnet-4-6 run as zero-shot LLM-as-judge; Pearson correlation with expert labels rho approximately 0.59; overshoots the true rate by about 13 percentage points; statistically equivalent to roughly 157, 148, and 143 purely human-labeled trajectories; https://github.com/EmertonData/glide; limitations: mean estimation only, single proxy / i.i.d. assumption, does not provide anytime-valid constructions || Re-check notes: Every headline number confirmed verbatim in the HTML full text: the ~1.0x-to-2.2x effective-sample gain, >=50 labels/stratum, and the R-Judge case (rho~0.59, ~13pp bias, 100 labels worth 143-157). Authors Martinon/Merad/Raki, Emerton Data affiliation, and the GitHub repo all present.
Decision use: holds regardless of model progress. Measured on structural. GLIDE library unifying PPI-class estimators (May 2026). Tooling availability, current.
[HS-06] hybrid-statistical - grade A (holds regardless of model progress) - re-check 2026-07-15: confirmed
For estimating scores from noisy judges, PPI++/efficient-influence-function estimators produce near-identical, shortest confidence intervals - reported 3-15x narrower than Rogan-Gladen-style misclassification-correction estimators depending on the human-labeling ratio, with the advantage largest when the judge is close to random guessing.
Source: Efficient Inference for Noisy LLM-as-a-Judge Evaluation (2026-01-08, academic)
What the source itself says (retrieved quote)
EIF and PPI++ produce nearly identical and shortest intervals, outperforming Rogan-Gladen by a factor of 3-15x depending on the labeling ratio. The advantage is most pronounced when q0+q1-1 is small (i.e., when the LLM-judge is closer to random guessing).
Re-check detail
Checked and matching: EIF and PPI++ produce nearly identical and shortest intervals; outperform Rogan-Gladen by a factor of 3-15x; 3-15x factor depends on the labeling ratio; advantage most pronounced when judge is closer to random guessing (q0+q1-1 small); at q0=q1=0.6, RG intervals ~10x wider than EIF/PPI++; code repo github.com/yiqunchen/debias-llm-as-a-judge; authors Chen, Lu, Li, Guo, Li || Re-check notes: The 3-15x figure and the near-random-guessing / labeling-ratio dependence are NOT in the abstract (and the PDF was FlateDecode-compressed and unreadable), but the allowed HTML full text (Figure 4 caption + Section 5.3) confirms them verbatim, including 'PPI++' specifically (the abstract only said 'PPI'). Fully confirmed.
Decision use: holds regardless of model progress. Measured on structural. EIF/PPI++ efficiency results. Mathematics (interval-width empirics corpus-flagged as not independently recomputed).
[HS-07] hybrid-statistical - grade A (holds regardless of model progress) - re-check 2026-07-15: confirmed
For certifying that a failure rate is below a threshold (the low-base-rate pass/fail regime), the 'Noisy but Valid' framework (ICLR 2026) uses a small human calibration set to estimate judge TPR/FPR, applies a variance-corrected test to the large judge-labeled stream with finite-sample Type-I error control, and derives exact conditions under which judge-based testing has HIGHER statistical power than direct human evaluation.
Source: Noisy but Valid: Robust Statistical Evaluation of LLMs with Imperfect Judges (2026-01-28, academic)
What the source itself says (retrieved quote)
leveraging a small human-labelled calibration set to estimate the judge's True Positive and False Positive Rates ... the exact conditions under which noisy testing yields higher statistical power than direct evaluation ... Accepted to ICLR2026
Re-check detail
Checked and matching: small human-labelled calibration set estimates judge True Positive and False Positive Rates; variance-corrected critical threshold applied to large judge-labelled dataset; finite-sample Type-I error control (validity) despite calibration uncertainty; derives exact conditions under which noisy testing yields higher power than direct evaluation; validated on Jigsaw Comment, Hate Speech, SafeRLHF; quantified 'Oracle' gap (cost of estimating judge parameters); Accepted to ICLR 2026 || Re-check notes: Every element of the claim is present in the abstract, including the low-base-rate certification framing, finite-sample Type-I control under calibration uncertainty, the three datasets, the oracle gap, and ICLR 2026 acceptance. Authors match (Feng, Shen, Balashankar, Gerner-Beuerle, Rodrigues).
Decision use: holds regardless of model progress. Measured on structural. 'Noisy but Valid' certification: judge-assisted testing provably more powerful under conditions. Mathematics.
Cited at: Pilot - gates and objectives
[HS-08] hybrid-statistical - grade A (holds regardless of model progress) - re-check 2026-07-15: confirmed
Naive reporting of judge-scored evaluations is systematically biased by judge sensitivity/specificity; a plug-in correction with confidence intervals propagating BOTH test-set and calibration-set uncertainty fixes it, remains unbiased under distribution shift between calibration and test data, includes adaptive allocation of calibration labels, and characterizes regimes where corrected judge-based evaluation is more reliable than human-only evaluation.
Source: How to Correctly Report LLM-as-a-Judge Evaluations (2025-11-26 (v4 2026-05-31), academic)
What the source itself says (retrieved quote)
imperfect sensitivity and specificity of the LLM judges induce bias in naive evaluation scores ... uncertainty from both the test dataset and a human-labeled calibration dataset ... remains unbiased under distribution shift between the test and calibration datasets
Re-check detail
Checked and matching: imperfect judge sensitivity/specificity biases naive evaluation scores; plug-in framework corrects the bias; confidence intervals propagate uncertainty from BOTH test dataset and human-labeled calibration dataset; remains unbiased under distribution shift between test and calibration datasets; adaptive strategy to allocate calibration samples for tighter intervals; characterizes regimes (by true score, sensitivity, specificity) where judge-based eval beats human-only; accepted ICML 2026; v1 2025-11-26, v4 2026-05-31 || Re-check notes: Every core element of the claim is present in the abstract and confirmed. Distribution-shift robustness (vs standard PPI i.i.d.) is stated explicitly.
Decision use: holds regardless of model progress. Measured on structural. Sensitivity/specificity-corrected pass-rate reporting. Mathematics.
Same underlying source as [F26-09] - repetition across reports is not independent corroboration.
Cited at: Decisions - the throughput dial
[HS-09] hybrid-statistical - grade B (current-generation measurement) - re-check 2026-07-15: confirmed with caveats
Conformal prediction now provides per-item uncertainty for LLM-as-judge ratings: EMNLP 2025 work builds guaranteed-coverage score intervals from a single judging run (with an ordinal boundary adjustment and a lower-bias midpoint score), and April 2026 work shows conformal set width is a genuine per-instance reliability signal (r_s=+0.576 with reliability, n=1,918; widths correlate ~0.32-0.38 across different judges) with reliability driven more by the CRITERION than the judge model (relevance avg set size ~3.0 vs fluency/consistency ~4.9 on a 1-5 scale).
Source: Analyzing Uncertainty of LLM-as-a-Judge (EMNLP 2025) + Diagnosing LLM Judge Reliability (arXiv:2604.15302) (2025-11 (EMNLP) / 2026-04-16, academic)
What the source itself says (retrieved quote)
set width serving as a per-instance reliability indicator (r_s = +0.576, N = 1,918, p < 10^-100) ... consistent cross-judge agreement (r-bar = 0.32-0.38)
Re-check detail
Checked and matching: conformal set width is a per-instance reliability indicator, r_s = +0.576, N = 1,918; cross-judge width agreement r-bar = 0.32-0.38; relevance avg set size ~3.0 vs fluency/consistency ~4.9; criterion matters more than the judge model; split conformal prediction over 1-5 Likert scores on SummEval || Asserted but not visible in retrievable text: EMNLP 2025 Sheng/Liu/He/Zhao/Kang paper: guaranteed-coverage intervals from a single judging run, ordinal boundary adjustment, lower-bias midpoint score (this is a DIFFERENT paper at aclanthology.org/2025.emnlp-main.569, not fetchable from this arXiv URL) || Re-check notes: The claim bundles TWO papers; only the arXiv:2604.15302 (April 2026, Gupta & Kumar) half is at the fetched URL and it is fully confirmed. The EMNLP 2025 conformal-interval work (Sheng et al.) with the ordinal boundary adjustment and lower-bias midpoint score is a separate source not retrievable here, hence its specifics are not visible.
Decision use: current-generation measurement. Measured on model-outputs. Conformal per-item uncertainty for judge ratings (2025-26). Method durable; measured widths tier-bound.
Same underlying source as [F26-05] - repetition across reports is not independent corroboration.
[HS-10] hybrid-statistical - grade A (holds regardless of model progress) - re-check 2026-07-15: confirmed
Judge-vs-system drift can be attributed with anytime-valid statistics: a frozen human-labeled anchor set periodically re-scored by the live judge, monitored with betting e-processes, correctly attributed a silent judge version bump in 60/60 runs with zero misattribution (rolling z-test baseline false-alarmed on 75% of drift-free streams), at 0.21-0.64x the cost of strong-judging every item.
Source: Who Drifted: the System or the Judge? Anytime-Valid Attribution in LLM Evaluation Pipelines (2026-06-13, academic)
What the source itself says (retrieved quote)
a silent version bump is detected as judge drift in 60/60 runs with zero judge-to-system misattribution.
Re-check detail
Checked and matching: fixed human-labeled anchor set re-scored by the current judge at a steady interleave; betting e-process on the judge-versus-human gap; silent version bump detected as judge drift in 60/60 runs with zero judge-to-system misattribution; rolling z-test false-alarms on 75% of drift-free streams; approximately 0.64 of the cost of strong-judging every item, or 0.21 in a cheaper-but-deafer regime; one-way identification (only the judge can move the anchors); verdict in {none, system, judge}; strict-prompt change attributed on 110 of 120 runs; TL;DR replication attribution perfect (240/240) || Re-check notes: All 11 asserted specifics verified verbatim in the abstract. Single-author preprint (Yitao Li), submitted 2026-06-13, matching the claim's own not-yet-peer-reviewed caveat.
Decision use: holds regardless of model progress. Measured on structural. Anchor-set drift attribution with anytime-valid statistics. Method.
[HS-11] hybrid-statistical - grade A (holds regardless of model progress) - re-check 2026-07-15: confirmed
In sequential prediction-powered mean estimation, uncertainty-based active querying contributed little: the smallest confidence widths occurred when query probabilities were near-constant at the budget ceiling, and theory shows optimized query probabilities converge to the maximum allowed constant rate when chosen without reference to covariates.
Source: Revisiting Active Sequential Prediction-Powered Mean Estimation (2026-04-20, contrarian)
What the source itself says (retrieved quote)
the smallest confidence width tends to occur when the weight on the constant probability is close to one
Re-check detail
Checked and matching: smallest confidence width occurs when the weight on the constant probability is close to one (uncertainty-driven part diminishes); non-asymptotic analysis yielding a data-dependent CI bound; with a no-regret approach the query probability converges to the maximum-value constraint when set obliviously to covariates; Published as a conference paper at ICLR 2026; submitted 20 Apr 2026 (Sfyraki & Wang) || Re-check notes: Core claim confirmed. The claim's 'near-constant at the budget ceiling' maps to the paper's 'maximum-value constraint' the oblivious query probability converges to; consistent phrasing.
Decision use: holds regardless of model progress. Measured on structural. Sequential estimation: uncertainty-based querying bought little vs uniform. Regime-dependent statistics; the design response (A/B routing against uniform) is already encoded.
[HS-12] hybrid-statistical - grade A (holds regardless of model progress) - re-check 2026-07-15: confirmed with caveats
Commercial tooling for confidence-scored LLM judgments with human calibration is mature: Cleanlab's Trustworthy Language Model attaches a real-time trustworthiness score to every LLM response, supports custom evaluation criteria, explicitly documents calibrating trust scores against human quality ratings, and publishes benchmarks (Nov 2025) for auto-flagging incorrect structured outputs for human review.
Source: Cleanlab Trustworthy Language Model (TLM) documentation and benchmark (2025-11-18 (benchmark); docs current 2026, practitioner)
What the source itself says (retrieved quote)
scores the trustworthiness of responses from any LLM in real-time ... state-of-the-art trustworthiness scores for any LLM application
Re-check detail
Checked and matching: attaches a real-time trustworthiness score to responses from any LLM; supports custom evaluation criteria (a 'Custom Evaluation Criteria' tutorial is listed) || Asserted but not visible in retrievable text: documentation of calibrating trust scores against human quality ratings (not on this overview page); the 2025-11-18 structured-outputs benchmark for auto-flagging incorrect outputs for human review (not on this overview page) || Re-check notes: The two core capabilities (real-time trust score for any LLM; custom eval criteria) are confirmed on the given overview URL. The human-rating calibration tutorial and the Nov-2025 structured-outputs benchmark cited in the evidence live on sub-pages (custom-eval / benchmark blog) that are outside the given URL and were not fetched; they are not visible on help.cleanlab.ai/tlm/ itself.
Decision use: holds regardless of model progress. Measured on structural. Cleanlab TLM: trust-scored judgments with calibration tooling. Vendor tooling fact, current.
Rubric practice
[RR-01] rubrics-recent - load-bearing - grade B (current-generation measurement) - re-check 2026-07-15: confirmed
OpenAI HealthBench meta-evaluates its LLM grader per-criterion against physician majority grades using Macro F1 (met/not-met, class-balanced), and reports the grader as a percentile of individual physicians: GPT-4.1 grader MF1 = 0.709, exceeding the average physician in 5 of 7 themes, while physician-vs-physician MF1 is only 0.569-0.730 with wide individual spread.
Source: HealthBench: Evaluating Large Language Models Towards Improved Human Health (OpenAI) (2025-05-13, primary)
What the source itself says (retrieved quote)
It "exceeds the average physician score in five out of seven themes" - verbatim confirmed
Re-check detail
Checked and matching: per-criterion Macro F1 (met/not-met, class-balanced) meta-evaluation of the grader; GPT-4.1 grader MF1 = 0.709; exceeds average physician score in five out of seven themes (verbatim); physician-vs-physician weighted MF1 range 0.569 (Response depth) to 0.730 (Health data tasks); o4-mini 0.692, o3 0.681, GPT-4.1-nano 0.580 (GPT-4.1-mini 0.661); over 60,896 meta-examples; typical-physician baseline: each physician scored against the others, excluding their own grades; std ~0.002 across 16 full runs (Table 7: 0.0016-0.0029) || Re-check notes: All eight cited numbers match verbatim. NUANCE: the grading ground truth is each INDIVIDUAL physician's grade, not a majority-vote label; majority agreement was used earlier to assign consensus criteria to examples, not as the grading target. The claim's phrase 'physician majority grades' is therefore a minor mischaracterization, but does not affect any headline number or the core method.
Research-sweep audit (2026-07-14): confirmed
Fetched https://arxiv.org/html/2505.08775v1 (marked arXiv:2505.08775v1 [cs.CL] 13 May 2025 - date correct). All claim elements verified against Section 8.1 / Tables 5-7: (1) meta-evaluation uses macro F1 on binary met/not-met per consensus criterion, explicitly chosen for class imbalance (random baseline MF1=0.50); (2) GPT-4.1 grader MF1 = 0.709, best of graders tested (o4-mini 0.692, o3 0.681, GPT-4.1 mini 0.661, GPT-4.1 nano 0.580); (3) over 60,896 meta-examples across the 34 physician-consensus criteria (avg 1,791/criterion); (4) 'typical physician' baseline computed by scoring each physician against the others the same way as the model, reported as a percentile - GPT-4.1 exceeds the average physician in 5 of 7 themes (below only expertise-tailored communication 0.610 vs 0.618 and health data tasks 0.683 vs 0.730), upper half in 6/7, above 33rd percentile in all; (5) theme-level weighted-average physician-vs-physician MF1 ranges 0.569 (response depth) to 0.730 (health data tasks), with paper noting physician-physician and model-physician agreement both span ~55-75% and wide individual spread; (6) grading prompt AND individual criterion phrasings were tuned so intent was unmistakable to the grader (paper also cautions GPT-4.1's top rank may be partly explained by its use during prompt tuning); (7) std across 16 full runs ~0.002 (Table 7: 0.0016-0.0029 by model). Supersession check: no arXiv v2 exists; 2026 activity is HealthBench Professional (OpenAI, April 2026) and third-party replications (e.g., open-source graders: Kimi-K2 0.693, Qwen3-235B 0.681 vs GPT-4.1 0.709) - these extend, not contradict, the original meta-eval numbers. Only pedantic nuance: 0.569-0.730 are theme-level weighted averages of physician MF1, not the full range across individual physicians (individual spread is wider) - the claim's own 'wide individual spread' wording already captures this correctly.
Decision use: current-generation measurement. Measured on model-outputs (models: GPT-4.1, o4-mini, o3, GPT-4.1-nano). HealthBench grader meta-evaluation (2025, GPT-4.1/o3-era graders). The per-criterion meta-evaluation PATTERN is the durable input; MF1 0.709 is tier-bound.
Same underlying source as [RR-02] - repetition across reports is not independent corroboration.
[RR-02] rubrics-recent - grade A (holds regardless of model progress) - re-check 2026-07-15: confirmed
HealthBench's production rubric format is per-item (conversation-specific) criteria written by 262 physicians (48,562 criteria), each an independently-judged binary met/not-met check with a nonzero weight from -10 to +10 (negative criteria encode pitfalls), aggregated as a weighted sum normalized by max attainable positive points.
Source: HealthBench (rubric structure, Section 3) (2025-05-13, primary)
What the source itself says (retrieved quote)
an associated nonzero point value between -10 and 10, with negative points used for criteria that are undesirable
Re-check detail
Checked and matching: rubrics created by 262 physicians; 48,562 unique criteria across all conversations; conversation-specific; each criterion nonzero point value between -10 and 10; negative for undesirable/pitfalls; grader judges each criterion independently, binary met/not-met; aggregated as weighted sum divided by maximum possible score; per-example score can go negative || Re-check notes: Full text confirms all elements: 262 physicians, 48,562 criteria, -10..10 point range, independent binary judging, weighted-sum normalization by max attainable, and negative-capable item scores.
Decision use: holds regardless of model progress. Measured on structural. 48,562 physician-written per-item criteria: what expert-grounded rubric authoring costs at production scale. Practice fact.
Same underlying source as [RR-01] - repetition across reports is not independent corroboration.
Cited at: Decisions - first project and gold funding
[RR-04] rubrics-recent - load-bearing - grade B (current-generation measurement) - re-check 2026-07-15: confirmed
RubricBench (Mar 2026) measures a stable ~26-27 point preference-accuracy gap between self-generated and human-annotated rubrics across all judge backbones (e.g., DeepSeek-v3.2 57.8% vs 84.9%; Gemini-3-Flash 58.0% vs 85.3%), and scaling test-time compute (more sampled rubrics, refinement) does not close it - 'rubric formation', not judge reasoning, is the bottleneck.
Source: RubricBench: Aligning Model-Generated Rubrics with Human Standards (2026-03-02, academic)
What the source itself says (retrieved quote)
even with human rubrics, accuracy plateaus around 85% rather than approaching 100%
Re-check detail
Checked and matching: gap 'remains stable at ~26%' / 'a severe 27% accuracy gap' (Table 3 Delta range +22.1 to +28.3); DeepSeek-v3.2: Self-Gen 57.8 vs Human 84.9 (Delta +27.1); Gemini-3-Flash: Self-Gen 58.0 vs Human 85.3 (Delta +27.3); 1,147 adversarial pairwise comparisons; unnecessary rules N=1 17.9% vs 10.1%; overly rigid R=5 12.8% vs 7.7%; checklists ~13.2/15.4/14.6 items; 'Attention Displacement'; 'even with human rubrics, accuracy plateaus around 85%'; safety self-generated ~25-30% vs human >90%; test-time compute does NOT close gap (GPT-4o-mini 48.0%->46.8%; refinement 46.7->46.4->45.7) || Re-check notes: Every specific number in the claim matches the full-text tables. The claim's '~26-27 point... across all judge backbones' aligns with the paper's own ~26%/27% summary, though per-backbone Delta ranges +22.1 to +28.3. This is stronger than the batch's in_corpus 'partially_confirmed' - full text confirms all specifics, so verdict is confirmed.
Research-sweep audit (2026-07-14): partially_confirmed
Source check (arXiv 2603.01562v1, fetched): the RubricBench paper exists, is dated 2026-03-02 (v2 2026-03-03), and confirms nearly every specific: 1,147 adversarial pairwise comparisons; DeepSeek-v3.2 57.8% self-generated vs 84.9% human (vanilla 38.8%); Gemini-3-Flash 58.0% vs 85.3% (vanilla 56.4%); test-time scaling fails (Rub@4->Rub@32 flat or declining, refinement depth declining, while scaling HUMAN rubrics helps 75.4->85.3); unnecessary-rule 17.9% vs 10.1% and rigid-rule 12.8% vs 7.7%; 13+-item checklists / 'Attention Displacement' with >70% hallucination rates; ~85% human-rubric plateau; safety ~25-30% self-generated vs >90% human. Two corrections: (1) the gap across ALL backbones is ~22-28 points, not a tight '~26-27' - 26-27 fits only the two cited models; (2) the 'rubric formation is the bottleneck' framing is the paper's interpretation and, more importantly, has been operationally SUPERSEDED in part: 'Support Vector Rubrics' (arXiv 2606.08077, June 2026, PKU/USTC) closes the gap on RubricBench to 0.3 points (82.8 vs 83.1 human-oracle, vs 59.0 self-generated with same GPT-OSS-120B judge) via max-margin contrastive rubric-bank learning over preference data - so the gap is stable under naive self-generation and test-time compute, but not under trained discriminative rubric construction. Safety remains the residual weak domain even for SVR (83.8 vs 92.5). Claim is accurate as a description of the March paper; the 'does not close' generalization no longer holds as of June 2026.
Corrected statement: RubricBench (arXiv 2603.01562, Mar 2026) measures a stable ~22-28 point preference-accuracy gap between self-generated and human-annotated rubrics across judge backbones (DeepSeek-v3.2 57.8% vs 84.9%; Gemini-3-Flash 58.0% vs 85.3%), and scaling test-time compute (more sampled rubrics, refinement depth) does not close it - the paper attributes the bottleneck to rubric formation, not judge reasoning. However, follow-up work (Support Vector Rubrics, arXiv 2606.08077, Jun 2026) closes the gap to ~0.3 points on RubricBench via max-margin rubric-bank learning from preference data, showing the gap is specific to naive self-generation rather than fundamental.
Decision use: current-generation measurement. Measured on model-outputs (models: DeepSeek-v3.2, Gemini-3-Flash). RubricBench (Mar 2026, DeepSeek-v3.2/Gemini-3-era): 26-27pp self-vs-expert rubric gap that test-time compute doesn't close. Current-window; deployment-tier gap unknown.
[RR-05] rubrics-recent - load-bearing - grade A (holds regardless of model progress) - re-check 2026-07-15: confirmed with caveats
EvalGen ('Who Validates the Validators', UIST 2024) established criteria drift: evaluation criteria cannot be fully specified a priori because grading outputs itself changes the criteria - 'users need criteria to grade outputs, but grading outputs helps users define criteria' - and some criteria are dependent on the specific outputs observed.
Source: Who Validates the Validators? Aligning LLM-Assisted Evaluation of LLM Outputs with Human Preferences (2024-10-13, academic)
What the source itself says (retrieved quote)
"users need criteria to grade outputs, but grading outputs helps users define criteria" ... "some criteria appears dependent on the specific LLM outputs observed" ... the study "underscores the subjectivity and iterative process of alignment".
Re-check detail
Checked and matching: criteria drift established as a named phenomenon; 'users need criteria to grade outputs, but grading outputs helps users define criteria'; some criteria are dependent on the specific LLM outputs observed; alignment is subjective and iterative; EvalGen generates candidate assertions/judge prompts || Asserted but not visible in retrievable text: 'UIST 2024' venue attribution (arXiv page lists no conference/journal venue) || Re-check notes: Criteria-drift core confirmed verbatim. Only the venue label 'UIST 2024' is not visible on the arXiv abstract page (Comments: 16 pages, 4 figures, 2 tables; no venue listed).
Decision use: holds regardless of model progress. Measured on human-work. EvalGen criteria drift: human graders' criteria change as they grade. Human behavior; durable.
Same underlying source as [CT-04] - repetition across reports is not independent corroboration.
Cited at: Pilot - gates and objectives - Decisions - rubric authority
[RR-06] rubrics-recent - grade B (current-generation measurement) - re-check 2026-07-15: confirmed
RIFT (Snorkel AI, Apr 2026) provides the first rubric failure-mode taxonomy - 8 modes in 3 categories (Reliability: Subjective, Non-Atomic, Ungrounded; Content Validity: Misaligned/Rigid, Missing Criteria; Consequential Validity: Hackable, Low Signal, Redundant) - and shows failure-mode density correlates with judge-human misalignment (r=0.162, p=0.0021); synthetic rubrics skew Subjective (86.7% vs 52.6% for human-crafted) while human rubrics skew Misaligned/Rigid (63.2% vs 20.0%); automated LLM detection reaches F1 0.925 for Subjective but ~0.000 for Hackable.
Source: RIFT: A RubrIc Failure Mode Taxonomy and Automated Diagnostics (Snorkel AI) (2026-04-20, academic)
What the source itself says (retrieved quote)
Pearson's r = 0.162, p = 0.0021 ... we analyze 85 rubrics with 255 expert annotations drawn from five data sources
Re-check detail
Checked and matching: 8 modes in 3 categories: Subjective, Non-Atomic, Ungrounded (Reliability); Misaligned/Rigid, Missing Criteria (Content Validity); Hackable, Low Signal, Redundant Criteria (Consequential); Pearson's r = 0.162, p = 0.0021; Table 2: Subjective Human 52.6% vs Synthetic 86.7%; Table 2: Misaligned/Rigid Human 63.2% vs Synthetic 20.0%; Table 3: Subjective F1 0.925; Hackable F1 0.000; 85 rubrics with 255 expert annotations drawn from five data sources; AdvancedIF, ResearchRubrics, WildChecklists, OpenRubrics, and AutoRubrics; developed using grounded theory; 0.64 average Cohen's kappa; classifiers GPT-5.2 and Gemini 3 Pro || Re-check notes: All 8 modes/3 categories, both correlation stats, both Table 2 prevalence pairs, both Table 3 F1 extremes, the 85/255 counts, five sources, grounded theory, kappa 0.64, and both classifier models confirmed. Minor internal inconsistency in the paper: abstract says 'up to 0.925 F1' while a contributions bullet says 'up to 0.86 F1'; the claim uses the 0.925 (Table 3) figure.
Decision use: current-generation measurement. Measured on model-outputs (models: GPT-5.2, Gemini 3). RIFT (Apr 2026, GPT-5.2/Gemini-3-era linting): Subjective detectable (F1 .925), Hackable not (~0). Current-window; the human-red-team conclusion for anti-gaming review stands until a later tier changes it.
[RR-07] rubrics-recent - grade B (current-generation measurement) - re-check 2026-07-15: confirmed
Support Vector Rubrics (Jun 2026) identifies the mechanism behind the LLM-vs-human rubric gap - 'self-generated rubrics describe good responses, whereas effective criteria must discriminate between close candidates' - and closes the RubricBench gap from 24.1 to 0.3 points by mining contrastive rubrics from preference pairs (support-pair selection plus adversarial probing of hard negatives).
Source: Support Vector Rubrics: Closing the Gap Between Self-Generated and Human Rubrics (2026-06-06, academic)
What the source itself says (retrieved quote)
self-generated rubrics describe good responses, whereas effective criteria must discriminate between close candidates ... SVR narrows the gap to human reference rubrics from 24.1 to 0.3 points ... the learned bank transfers across judges without retraining.
Re-check detail
Checked and matching: objective mismatch: self-generated rubrics describe good responses, effective criteria must discriminate between close candidates; narrows gap to human reference rubrics from 24.1 to 0.3 points; mines contrastive features from preference pairs into a rubric bank; support-pair selection and adversarial probing of hard negatives; max-margin boundary learning over preference data; learned bank transfers across judges without retraining; competitive with dedicated reward models on RewardBench 1&2 and RM-Bench || Re-check notes: Every element of the claim confirmed verbatim from the abstract, including the exact 24.1 -> 0.3 figure. Authors Sun et al. match.
Decision use: current-generation measurement. Measured on model-outputs. Support Vector Rubrics (Jun 2026): contrastive mining closes the rubric gap to ~0.3pp given preference pairs. Current-window; depends on adjudication data the pilot creates.
[RR-08] rubrics-recent - grade B (current-generation measurement) - re-check 2026-07-15: confirmed
PaperBench (OpenAI, Apr 2025) demonstrates the extreme end of decomposition: 8,316 individually-graded binary leaf criteria across 20 papers, sibling-relative manual weights, rubrics co-developed with each paper's original authors over multiple weeks per rubric; an o3-mini judge then grades leaves at F1 = 0.83 vs expert human judgments for ~$66/paper versus ~12 hours of expert grading time.
Source: PaperBench: Evaluating AI's Ability to Replicate AI Research (OpenAI) (2025-04-02, primary)
What the source itself says (retrieved quote)
Across the 20 papers in PaperBench there are 8,316 leaf nodes. ... o3-mini with the SimpleJudge scaffolding is the most cost-effective, with an F1 score of 0.83 at $66 USD per paper
Re-check detail
Checked and matching: 8,316 leaf nodes / individually gradable tasks across 20 papers; each leaf scored binary (1 pass / 0 fail); node weight = importance relative to siblings, not implementation difficulty; rubrics co-developed with the paper's author(s); took multiple weeks per paper; o3-mini SimpleJudge: F1 0.83 at $66 USD/paper (most cost-effective); human expert grading on the order of tens of hours per paper (cost modeled at 12 hrs x $100/hr); JudgeEval is a separate benchmark for evaluating automated judges; o1 judge: 0.84 F1 at $830 USD/paper || Re-check notes: All decomposition and cost figures confirmed verbatim in the full text, including sibling-relative weighting, weeks-long author co-development, the o3-mini 0.83/$66 vs o1 0.84/$830 tradeoff, JudgeEval as a separate judge benchmark, and both the 'tens of hours' and '12 hours at $100/hr' human-grading figures.
Decision use: current-generation measurement. Measured on model-outputs (models: o3-mini, o1). PaperBench (2025, o1/o3-era judge F1 0.83 at ~1/100 cost). Decomposition-at-scale pattern durable; judge F1 and cost tier-bound.
[RR-09] rubrics-recent - grade A (holds regardless of model progress) - re-check 2026-07-15: confirmed
OpenAI's official grader documentation prescribes that every model grader be validated on its own eval built from expert-graded answers (grader score ordering must reproduce expert ranking), and names the canonical reward-hacking detection signature: the graded model scores well on the model-grader eval while doing poorly on expert human evaluations.
Source: OpenAI API documentation: Graders (2026-07-14, primary)
What the source itself says (retrieved quote)
A model that's hacked the grader will score highly on model grader evals but score poorly on expert human evaluations ... Produce a smooth score, not a pass/fail stamp ... In reinforcement fine-tuning, you can nest and combine graders by using multigraders
Re-check detail
Checked and matching: grader taxonomy: string check, text similarity, score model (LLM judge, numeric), python (code execution), multigrader (combine/nest, reinforcement fine-tuning only); validate a model grader on its own eval built from expert-graded answers; grader scores must reproduce the expert ranking (answer_1 > answer_2 > answer_3); reward/grader hacking signature: scores high on model grader evals but poorly on expert human evaluations; guidance: smooth score not pass/fail; few-shot great/fair/poor examples; add edge cases over time; guard against reward hacking; deprecation note: OpenAI deprecating graders within evals/fine-tuning workflows || Re-check notes: All prescribed elements of the claim are present verbatim in the live doc, including the named reward-hacking signature and the deprecation note.
Decision use: holds regardless of model progress. Measured on structural. OpenAI grader doctrine: every model grader validated on its own expert-graded eval. Practice prescription, current.
[RR-10] rubrics-recent - grade B (current-generation measurement) - re-check 2026-07-15: confirmed with caveats
A cluster of 2026 papers (Rubric-ARM Feb 2026; EvoRubrics, ARBOR, AMARIS Jun 2026) converges on the position that static rubrics are an exploitable reward specification under optimization pressure and must be adapted/co-evolved during training, with Rubric-ARM jointly optimizing a rubric generator and judge via alternating RL from preference feedback.
Source: Rubric-ARM: Alternating Reinforcement Learning for Rubric-Based Reward Modeling (plus 2026 co-evolution cluster) (2026-02-02, academic)
What the source itself says (retrieved quote)
jointly optimizes a rubric generator and a judge using reinforcement learning from preference feedback ... Unlike existing methods that rely on static rubrics or disjoint training pipelines
Re-check detail
Checked and matching: Rubric-ARM jointly optimizes a rubric generator and a judge via reinforcement learning from preference feedback; prior methods 'rely on static rubrics or disjoint training pipelines'; alternating optimization strategy to mitigate non-stationarity; reduces gradient variance; SOTA on benchmarks and improved policy alignment || Asserted but not visible in retrievable text: 'adaptability' wording (not found on page); 'co-evolve' framing; 'exploitable reward specification under optimization pressure' framing; EvoRubrics, ARBOR, AMARIS (separate 2026 papers, not at this URL; claim states these are search-only) || Re-check notes: Only Rubric-ARM is at this URL and its mechanism is confirmed verbatim. The 'cluster of 2026 papers converging' (EvoRubrics/ARBOR/AMARIS) is not verifiable here and the claim itself flags it as search-only. Actual title lacks the 'Rubric-ARM:' prefix and the '(plus 2026 co-evolution cluster)' suffix of source_title, but the substantive title matches. v1 2026-02-02, v2 2026-02-11.
Decision use: current-generation measurement. Measured on model-outputs. 2026 RL-rubrics cluster: static rubrics are exploitable specifications under optimization. Conceptual; strengthens with optimizer capability.
[RR-11] rubrics-recent - grade A (holds regardless of model progress) - re-check 2026-07-15: confirmed
Autorubric (Feb 2026, Rao & Callison-Burch) consolidates scattered rubric-evaluation techniques into one open-source framework with opinionated defaults - binary/ordinal/nominal criterion types, judge ensembles, few-shot calibration, bias mitigations, and psychometric reliability metrics - reporting 87% binary accuracy with moderate-to-substantial kappa on CHARM-100, and showing per-criterion explanations double as improvement signals (peer-review agent 0.47 -> 0.85, beating a 0.82 expert-curated baseline).
Source: Autorubric: Unifying Rubric-based LLM Evaluation (2026-02-13, academic)
What the source itself says (retrieved quote)
raise a peer review agent's score from 0.47 to 0.85 (above the 0.82 expert-curated baseline)
Re-check detail
Checked and matching: unifies techniques scattered across papers with inconsistent terminology and partial implementations; binary, ordinal, and nominal criteria; single-judge and ensemble; few-shot calibration; bias mitigations; psychometric reliability metrics; CHARM-100: 87% binary accuracy, moderate-to-substantial kappa; peer review agent 0.47 to 0.85, above the 0.82 expert-curated baseline; RiceChem 80% accuracy with 5-shot calibration; ResearcherBench 931 criteria, cross-judge agreement; RL reward AdvancedIF +0.039, Wilcoxon p=0.032, positive transfer to IFEval; authors Delip Rao, Chris Callison-Burch || Re-check notes: All nine asserted specifics verified verbatim in the abstract. v1 2026-02-13 matches source_date; v2 2026-04-03, 52 pages, both confirmed.
Decision use: holds regardless of model progress. Measured on structural. Autorubric open-source framework. Tooling availability.
Same underlying source as [F26-12] - repetition across reports is not independent corroboration.
[RR-12] rubrics-recent - grade A (holds regardless of model progress) - re-check 2026-07-15: confirmed
LangChain's Align Evals (LangSmith, Jul 2025) productized judge validation as a standard workflow: humans grade a representative golden set per criterion, an 'alignment score' measures judge-vs-human match, unaligned cases are surfaced for prompt iteration, and each judge-prompt version is compared against a saved baseline alignment score.
Source: Introducing Align Evals: Streamlining LLM Application Evaluation (LangChain) (2025-07-29, practitioner)
What the source itself says (retrieved quote)
Our evaluation scores don't match what we'd expect a human on our team to say
Re-check detail
Checked and matching: Align Evals is a LangSmith feature (launched Jul 29 2025); four-step workflow: select criteria; select representative good+bad examples; human-grade expected scores (golden set); iterate evaluator prompt against alignment score; 'alignment score' measures evaluator vs human match; unaligned cases surfaced by sorting; saved baseline alignment score to compare new prompt versions; inspired by Eugene Yan's AlignEval; roadmap: analytics for tracking over time; automatic prompt optimization || Re-check notes: Every element of the claim and evidence confirmed from the launch blog.
Decision use: holds regardless of model progress. Measured on structural. LangSmith Align Evals: judge-calibration workflow productized. Tooling fact.
Industry practice
[IP-02] industry-practice - load-bearing - grade B (current-generation measurement) - re-check 2026-07-15: confirmed
Handshake AI's BVB meta-eval (3,204 practicing-banker pass/fail labels, 89.5% inter-annotator agreement on a dual-coded adjudicated subset) shows the best automated verifier they could build (Gandalf, a reactive agent-judge running in the same environment as the work) reaches only F1 0.633-0.664 against expert labels on artifact-heavy tasks, while hitting F1 0.951 on a simpler stateful-tools benchmark, and that verifier architecture (evidence view + evidence path) matters more than judge model choice.
Source: Your verifier is probably the bottleneck. We built one that isn't. (Gandalf the Grader / BVB) (2026-05-27, primary)
What the source itself says (retrieved quote)
F1 range, from 0.633 to 0.664, sits above the highest-F1 non-Gandalf run
Re-check detail
Checked and matching: BVB is a meta-eval of verifiers, not an agent benchmark; 3,204 expert-graded (practicing-banker) pass/fail judgments; inter-annotator agreement 89.5%; disagreements adjudicated; a subset dual-coded; Gandalf F1 range 0.633-0.664 on artifact-heavy banking tasks; Gandalf F1 0.951 on OpenClaw (stateful tools), 7.5 points ahead of next-best; cheapest Gandalf (GPT-5.4 Nano ~$42) beats best baseline (Archipelago/Gemini 3 Pro, F1 0.604, ~$422) by ~3 F1 at ~one-tenth cost; within-family architecture gap 9.5 F1 (Gandalf/GPT-5.4 Nano vs Archipelago/GPT-5.4); Gandalf is a reactive agent-as-judge running inside the rollout environment; open-sourced; OpenHands SDK harness; verifier architecture (view + path) matters more than backing model || Re-check notes: Every cited figure is present in the post. Minor framing: 89.5% is the overall inter-annotator agreement (with a separate dual-coded subset for rubric-application consistency), which is consistent with the claim's wording.
Research-sweep audit (2026-07-14): confirmed
Fetched https://joinhandshake.com/research/ai/gandalf-the-grader/ directly. Every element of the claim checks out against the source: (1) dated May 27, 2026; (2) BVB (BankerVerifierBench) is explicitly a meta-eval of verifiers built on 21 BankerToolBench tasks with 3,204 pass/fail judgments from practicing bankers; 89.5% inter-annotator agreement on a dual-coded subset, with disagreements adjudicated before inclusion; (3) Gandalf is a reactive agent-as-judge running in the same rollout environment (shared filesystem, interpreter, MCP tools), open-sourced on the OpenHands SDK; (4) F1 0.633-0.664 on BVB (artifact-heavy) vs F1 0.951 on OpenClaw, their internal stateful-tools productivity benchmark (200-criterion sample, +7.5 over next-best); (5) architecture (evidence view: text/snapshot/live env; evidence path: fixed vs inference-time) beats model choice - within GPT-5.4 family, Gandalf/Nano beats Archipelago/GPT-5.4 by 9.5 F1; (6) cheapest Gandalf config (~$42, GPT-5.4 Nano) beats best baseline (Archipelago/Gemini 3 Pro, F1 0.604, ~$422) by ~3 F1 at ~one-tenth cost; (7) the post prescribes meta-eval of verifiers against adjudicated expert labels before trusting them for scoring/rewards. Supersession check (two searches, July 2026): no newer contradicting work found; the full Gandalf paper and BVB dataset release are still pending, so numbers could later be refined, but as of today the post is the authoritative source and the claim matches it. Minor nuance only: the 0.951 benchmark is named OpenClaw and is internal (not public), which the claim's phrasing ("simpler stateful-tools benchmark") accurately reflects.
Decision use: current-generation measurement. Measured on human-work. BVB (May 2026): 3,204 practicing-banker labels - a rare human-work meta-eval. Vendor research; benchmark release pending per corpus.
Same underlying source as [F26-02] - repetition across reports is not independent corroboration.
Cited at: Research - findings by decision weight
[IP-03] industry-practice - load-bearing - grade A (holds regardless of model progress) - re-check 2026-07-15: confirmed with caveats
The 2026 eval-tooling vendors (LangSmith, Braintrust, Galileo) have converged on one judge-calibration pattern - small human-labeled reference sets (LangSmith: start at ~20 balanced examples; Braintrust: 50-100 per scoring dimension) with an 'alignment score' defined as RAW PERCENT AGREEMENT with the human expert, iterate the judge prompt against it, and route low-confidence or judge-disagreement items to humans - and none of them ships chance-corrected statistics.
Source: LangSmith: Improve LLM-as-judge evaluators using human feedback (+ Braintrust: LLM-as-a-judge vs human-in-the-loop, 2026-04-03) (retrieved 2026-07-14 (Braintrust article 2026-04-03), primary)
What the source itself says (retrieved quote)
The alignment score is the percentage of examples where the evaluator's judgment matches that of the human expert. ... We recommend starting with at least 20 examples
Re-check detail
Checked and matching: annotation queue -> reference dataset workflow confirmed; 'We recommend starting with at least 20 examples' and examples 'balanced in both 0 and 1 labels'; alignment score = 'the percentage of examples where the evaluator's judgment matches that of the human expert' (raw percent agreement, no chance correction); Evaluator Playground confirmed ('Start Alignment'); cluster misaligned examples 'into common failure modes' then encode as prompt instructions || Asserted but not visible in retrievable text: Braintrust specifics (50-100 labels per scoring dimension, 3-tier deterministic/judge/human stack, 'silent overconfidence' failure mode, human-review-not-automatically-gold-standard, 2026-04-03 date) -- on a separate Braintrust URL not in scope for this source; Galileo CLHF auto-adaptation from expert corrections (2025-02-10 blog) -- separate URL, not fetchable under the source-only constraint; cross-vendor generalization 'none of them ships chance-corrected statistics' -- only LangSmith's raw-percent-agreement design is verifiable from the single given URL || Re-check notes: The LangSmith half of the claim (>=20 balanced examples, alignment score = raw percent agreement, iterate prompt, cluster failure modes, route disagreements) is fully confirmed. The Braintrust and Galileo assertions live on other pages the hard constraint forbids fetching, so the multi-vendor convergence claim is only partially verifiable.
Research-sweep audit (2026-07-14): partially_confirmed
VERIFIED: (1) LangSmith doc (https://docs.langchain.com/langsmith/improve-judge-evaluator-feedback) says exactly what was reported - annotation queue -> reference dataset (~20 examples minimum, balanced 0/1) -> Evaluator Playground -> alignment score quoted verbatim as "the percentage of examples where the evaluator's judgment matches that of the human expert" -> cluster misalignments into failure modes as prompt instructions. No kappa/alpha anywhere; doc is undated. (2) Braintrust article (https://www.braintrust.dev/articles/llm-as-a-judge-vs-human-in-the-loop-evals, confirmed 2026-04-03) matches on all six checked points: 3-tier stack, 50-100 labels per scoring dimension, multi-judge disagreement routed to humans, "silent overconfidence" named as the most dangerous failure mode, human-review-not-gold-standard warning (reviewer fatigue, vague rubrics), and agreement tracked as raw agreement rate with no chance-corrected statistic. (3) Galileo CLHF exists and auto-adapts metrics from expert corrections (feedback -> LLM-generated few-shot examples appended to the metric prompt, capped at most recent 15, claimed 20-30% accuracy gain). REFUTED/OVERSTATED PARTS: (a) The "convergence on raw percent agreement" framing fails for Galileo. Galileo's own guidance - "How to Calibrate Your LLM Judge With Human Annotations" (https://galileo.ai/blog/calibrate-llm-judge-human-annotations, May 15/Jul 6 2026, i.e. NEWER than the cited 2025-02-10 CLHF post) - explicitly recommends AGAINST raw percent agreement (imbalanced-label example: 90% raw agreement, kappa ~ -0.05) and prescribes Cohen's kappa, Fleiss' kappa, and Krippendorff's alpha with a recalibrate-when-kappa<0.60 trigger. Galileo also has a dedicated Cohen's Kappa explainer (https://galileo.ai/blog/cohens-kappa-metric, 2025-03-12). (b) Galileo's CLHF is also not the same pattern as LangSmith/Braintrust - it has no reference set + alignment score loop; it is few-shot auto-improvement, so the three-vendor "one pattern" claim overreaches. (c) The narrower claim "none of them SHIPS chance-corrected statistics" as an in-product feature survives verification: Galileo's kappa content points users to scikit-learn/statsmodels/R for computation, and its product features (Autotune, Annotations, Signals, Experiments, CLHF) are not described as computing kappa; LangSmith and Braintrust expose only percent-agreement metrics. But the claim as worded implies the vendors don't even acknowledge chance correction, which is false for Galileo. SUPERSESSION CHECK: newer relevant work exists - arXiv 2606.00093 (Rao & Callison-Burch, 2026-05-25, "Agreement Metrics for LLM-as-Judge Evaluation: What to Report and Why") argues for reporting kappa alongside accuracy, and LangChain published an Align Evals calibration resource (2026-03-10) still framed around human-correction agreement.
Corrected statement: LangSmith and Braintrust have converged on a judge-calibration pattern of small human-labeled reference sets (LangSmith: ~20 balanced examples; Braintrust: 50-100 per scoring dimension) with alignment measured as raw percent agreement with human experts, prompt iteration against misalignments, and routing of low-confidence/judge-disagreement items to humans; neither ships a chance-corrected statistic. Galileo differs: its CLHF (2025) turns expert corrections into few-shot prompt examples (max 15) rather than a reference-set alignment score, and its newer 2026 calibration guidance explicitly rejects raw percent agreement in favor of Cohen's kappa / Fleiss' kappa / Krippendorff's alpha (recalibrate below kappa 0.60) - though it directs users to external tools (scikit-learn, R) rather than computing these in-product. So the "no chance-corrected stats shipped in-product" observation holds across all three, but there is no three-vendor convergence on raw percent agreement as the alignment definition.
Decision use: holds regardless of model progress. Measured on structural. 2026 eval-tooling convergence on human-calibrated judges. Market fact.
[IP-05] industry-practice - grade A (holds regardless of model progress) - re-check 2026-07-15: confirmed with caveats
Handshake acquired Cleanlab (announced 2026-01-28) specifically for algorithms that 'flag incorrect data without a second human reviewer,' after competing bids from other data-labeling companies - signaling that algorithmic second-pass QA replacing the human review layer is now an acquisition-grade core capability for human-data vendors.
Source: TechCrunch: AI data labeler Handshake buys Cleanlab, an acquisition target of multiple others (2026-01-28, practitioner)
What the source itself says (retrieved quote)
"Cleanlab's researchers are experts in developing algorithms that flag incorrect data without a second human reviewer." ... "The company has provided data for eight top AI labs, including OpenAI."
Re-check detail
Checked and matching: Handshake acquired Cleanlab (announced 2026-01-28); Cleanlab algorithms 'flag incorrect data without a second human reviewer'; competing acquisition interest from other AI data-labeling companies; nine key Cleanlab employees join Handshake research; three co-founders (Northcutt, Mueller, Athalye) earned PhDs from MIT; Handshake provided data for eight top AI labs, including OpenAI; $300 million annualized revenue run-rate (end of 2025) || Asserted but not visible in retrievable text: the phrases 'confident learning', 'data-centric AI', and 'LLM evaluation' (not used in the TechCrunch article; claim attributes them to Cleanlab CEO letter / Handshake announcement) || Re-check notes: Core assertion and every headline number/fact drawn from TechCrunch confirmed verbatim, including the key 'flag incorrect data without a second human reviewer' quote, competing bids, 9 staff, 3 MIT co-founders, 8 labs, $300M ARR. Only the capability-name phrases in the claim's evidence ('confident learning', 'data-centric AI', 'LLM evaluation') are absent from the article; the claim itself attributes those to other documents.
Decision use: holds regardless of model progress. Measured on structural. Handshake acquired Cleanlab for no-second-human flagging (Jan 2026). Market signal, current.
[IP-06] industry-practice - grade A (holds regardless of model progress) - re-check 2026-07-15: confirmed
Surge AI (~$1.2B revenue 2024, ~50k contractors, bootstrapped) runs QA as real-time dashboards over gold-standard accuracy, inter-annotator agreement, and per-worker trust ratings, with flagged labels automatically reassigned to other annotators; its published AdvancedIF verifier work with Meta reports F1 0.728 versus human judgments and +13% RL gains from rubric-based rewards.
Source: Sacra: Surge AI revenue, funding & news (2026-04-21, practitioner)
What the source itself says (retrieved quote)
Real-time dashboards monitor annotation quality using metrics such as gold-standard accuracy, inter-annotator agreement ... per-worker trust ratings ... Labels identified as low quality are automatically reassigned to other annotators
Re-check detail
Checked and matching: Sacra estimates that Surge AI hit $1.2B in annualized revenue in 2024; approximately 50,000 expert contractors and 130 full-time employees; bootstrapped company; Real-time dashboards monitor annotation quality using ... gold-standard accuracy, inter-annotator agreement ... and per-worker trust ratings; Labels identified as low quality are automatically reassigned to other annotators; AdvancedIF ... in partnership with Meta Superintelligence Labs; its verifier achieved 0.728 F1 versus human judgments; using human-written rubrics as RL reward signals yields 13% performance gains; premium rates of 30-40 cents per working minute; rigorous vetting, including domain-specific tests, background checks, and ongoing performance evaluations; Hemingway-bench ... built from 5,000+ blind pairwise comparisons by expert human judges; reliance on 12 customers for over $1 billion in revenue; class action lawsuit filed in California ... alleges Surge misclassified data annotators || Asserted but not visible in retrievable text: source_date 2026-04-21 -- no last-updated date is shown on the page; the lawsuit's specific allegations of 'unpaid training time and tight task timers' (not surfaced in retrieved text) || Re-check notes: All revenue/contractor/QA/AdvancedIF figures confirmed. Minor: page states '13% performance gains' (no literal '+' sign). Figures are explicitly Sacra estimates. Page shows no date, so the 2026-04-21 source_date could not be corroborated.
Decision use: holds regardless of model progress. Measured on structural. Surge QA-as-dashboards practice. Vendor-reported.
[IP-08] industry-practice - grade B (current-generation measurement) - re-check 2026-07-15: confirmed
Braintrust's published cost arithmetic makes the human-budget fork concrete: an LLM judge over 10,000 outputs costs ~$5-15 while a domain expert reviewing 500 of them costs ~$800-1,800 (~100x per output), which is why its prescribed architecture is deterministic checks on everything, LLM judges on everything, and humans only on flagged/low-confidence/disagreement slices plus random spot-checks of confident passes.
Source: Braintrust: LLM-as-a-judge vs human-in-the-loop evals: When to use each (2026-04-03, practitioner)
What the source itself says (retrieved quote)
Running an LLM judge on 10,000 outputs using a model like Claude Sonnet might cost $5-15 ... review 500 of those same outputs at $50-75/hour, spending 2-3 minutes per output, runs $800-1,800. That is roughly 100x more expensive per output
Re-check detail
Checked and matching: LLM judge on 10,000 outputs (Claude Sonnet) might cost $5-15; domain expert reviewing 500 outputs at $50-75/hr, 2-3 min each, runs $800-1,800; roughly 100x more expensive per output for human review; three-tier architecture: deterministic checks on every output, LLM judges for continuous coverage, humans only on flagged/low-confidence/disagreement cases + random spot-checks; biggest mistake: 'Treating judge scores as ground truth without validating them against human labels'; untrained reviewers / vague rubrics can yield labels 'worse than a decent LLM judge' || Re-check notes: Cost arithmetic, tiered architecture (deterministic + judge on everything, humans on flagged/disagreement slices plus random spot-checks), the named biggest-mistake, and the weak-human-review caveat all confirmed. Published date 3 April 2026 matches.
Decision use: current-generation measurement. Measured on structural. Braintrust cost arithmetic (judge $5-15/10k vs expert $800-1,800/500). Pricing is current but moves with model economics - reprice at bakeoff; the routing logic it supports is durable.
Cited at: Decisions - judge sourcing
[IP-09] industry-practice - grade C (superseded-generation number; mechanism only) - re-check 2026-07-15: confirmed with caveats
Fine-tuned specialist judges have commoditized binary grounding checks: Patronus Lynx (fine-tuned Llama-3-70B) beats GPT-4o on HaluBench faithfulness detection and lifted one customer's hallucination detection from 0.375 to 0.69, while Galileo's Luna-2 (arXiv 2026-02-20) runs hundreds of per-metric LoRA adapters on one small-model backbone at >80x lower cost and >20x lower latency than LLM-as-judge at claimed parity accuracy, in production on 100M+ sessions.
Source: Luna-2: Scalable Single-Token Evaluation with Small Language Models (arXiv 2602.18583) + Patronus Lynx launch materials (2026-02-20 (Lynx 2024-07-11), primary)
What the source itself says (retrieved quote)
matches the accuracy of state-of-the-art LLM-based evaluators ... by over 80x ... by over 20x ... protecting 100M+ AI sessions ... eval cost savings of over $30M annually
Re-check detail
Checked and matching: Luna-2: single-token SLM evaluation vs multi-token LLM-as-judge; each metric = a lightweight LoRA/PEFT head on a shared SLM backbone (hundreds concurrently); cost reduced by over 80x, latency by over 20x; matches accuracy of state-of-the-art LLM-based evaluators (parity); production scale: 100M+ AI sessions, 100B+ tokens/month; eval cost savings of over $30M annually || Asserted but not visible in retrievable text: all Patronus Lynx claims: Lynx-70B beats GPT-4o on HaluBench (+8.3% PubMedQA), Lynx-8B +24.5% vs GPT-3.5; Algomo customer 0.375 -> 0.69 hallucination detection; Gamma 1,000+ manual eval hours saved; Percival tracing; SiliconANGLE / Patronus launch-post sourcing (separate URLs, not fetched) || Re-check notes: The Luna-2 half of the claim is fully confirmed from the given arXiv abstract. The Patronus Lynx half is sourced to Patronus/SiliconANGLE pages that are not the given URL and thus not verifiable under the fetch constraint; those specifics are marked not_visible, not checked.
Decision use: superseded-generation number; mechanism only. Measured on model-outputs (models: Lynx, Llama-3-70B, GPT-4o, GPT-3.5). Lynx vs GPT-4o-era comparisons. Specialist-vs-frontier margins need current repricing.
[IP-10] industry-practice - grade A (holds regardless of model progress) - re-check 2026-07-15: confirmed with caveats
DeepEval's recommended progression - start with holistic G-Eval, then move to DAG (a deterministic decision tree whose nodes are narrow LLM-judge calls) 'for more control' - is the open-source codification of decomposed judging: verdicts assembled from atomic, individually-checkable steps rather than one holistic score.
Source: DeepEval docs: DAG (Deep Acyclic Graph) metric + Confident AI engineering post (2025-02-09, docs verified 2026-07-10, practitioner)
What the source itself says (retrieved quote)
You can still use GEval in the DAGMetric, but the DAGMetric will give you much greater control ... easily build deterministic decision trees for evaluation
Re-check detail
Checked and matching: DAG lets you 'easily build deterministic decision trees for evaluation'; positioned above G-Eval: 'the DAGMetric will give you much greater control'; 'more deterministic control' over GEval; produces deterministic scores; nodes are LLM-judge calls (BinaryJudgementNode, NonBinaryJudgementNode; 'reduce LLM-judge variance'); breakdown into more atomic units for evaluation || Asserted but not visible in retrievable text: 'conversational' variant for multi-turn work (word not found on page); Confident-AI 2025-02-09 engineering post ('holistic LLM scores too noisy') - separate URL, not fetched; exact phrase 'for more control' (page says 'much greater control' / 'more deterministic control'); docs page publication/last-updated date || Re-check notes: Core G-Eval -> DAG progression and decomposed/deterministic-node judging confirmed on the docs page. The asserted 'conversational variant' and the Confident-AI blog motivation (source_date 2025-02-09) are not verifiable from this docs URL.
Decision use: holds regardless of model progress. Measured on structural. DeepEval G-Eval->DAG progression. Tooling doctrine.
[IP-11] industry-practice - grade A (holds regardless of model progress) - re-check 2026-07-15: confirmed with caveats
The human reviewer layer at scale vendors is itself unreliable for structural/incentive reasons the AutoQA must not inherit: Outlier's attempter->reviewer->senior-reviewer pyramid runs on timed, capped per-task pay, opaque promotion/removal, uneven task allocation, and internal competition, per 2025 contributor reports across Indeed/Glassdoor/Trustpilot.
Source: Indeed/Glassdoor/Trustpilot Outlier reviewer reports (aggregated) (2025 (various), community)
What the source itself says (retrieved quote)
you can also be cut from a project with no warning or reason, placed on another completely different project
Re-check detail
Checked and matching: removal from projects without stated reason (multiple reviews 2024-2026); uneven / unpredictable task allocation (feast-or-famine workload); opaque eligibility/onboarding decisions felt arbitrary with no transparency; no communication from higher-ups; contact only via Slack/Discord; pay deterioration after Outlier merger; unpaid training; overall 2.4/5 from 774 reviews || Asserted but not visible in retrievable text: explicit attempter -> reviewer -> senior-reviewer pyramid; timed / capped per-task pay with unpaid overage window; specific '300-reviewer queues where ~10 got unlimited tasks'; queue managers running webinars / 'war rooms' as calibration; contributors and reviewers 'pitted against each other' / internal competition || Re-check notes: Only the Indeed page could be fetched; source aggregates Indeed/Glassdoor/Trustpilot but constraints forbid the other two sites. The general thrust (opaque removal, unstable allocation, opacity, pay complaints) is supported, but the specific structural mechanisms in the claim (pyramid, capped timed pay, 300/10 queues, war rooms, pitted-against-each-other) are not present on this page. title_match false: source_title is a synthetic three-platform aggregate label, not this single Indeed page's title; reviews span 2024-2026, not solely 2025.
Decision use: holds regardless of model progress. Measured on human-work. Scale-vendor reviewer layer unreliability for incentive reasons. Structural; the AutoQA must not inherit it.
The case against
[CT-01] contrarian - load-bearing - grade B (current-generation measurement) - re-check 2026-07-15: confirmed
Raw percent-agreement systematically overstates LLM-judge ability: in the largest judge meta-evaluation to date (21 judges, 9 providers, ~541,000 judgments, including April-2026 frontier models), Cohen's kappa runs 33-41 percentage points below exact-match agreement on MT-Bench, judge rankings shift by up to 14 positions across benchmarks, and two production-deployed judges combine test-retest reliability >0.95 with severe position bias >0.10 (a 'consistency-bias paradox').
Source: Reliability without Validity: A Systematic, Large-Scale Evaluation of LLM-as-a-Judge Models Across Agreement, Consistency, and Bias (2026-06-17, academic)
What the source itself says (retrieved quote)
we evaluate 21 judges from nine providers across MT-Bench, JudgeBench, and RewardBench ... 118 runs and approximately 541,000 individual judgments
Re-check detail
Checked and matching: 21 judges from nine providers; approximately 541,000 individual judgments; kappa deflation ... 33--41 pp on MT-Bench; judge rankings shift by up to 14 positions across benchmarks; test--retest reliability (>0.95) with severe position bias (>0.10) in two production-deployed judges; consistency--bias paradox; 118 runs across MT-Bench, JudgeBench, RewardBench; three protocols (agreement, consistency, bias); Minimum Viable Validation Protocol; April 2026 frontier || Asserted but not visible in retrievable text: corroborating paper arXiv 2508.18076 (Chehbouni et al., 'Neither Valid nor Reliable?') -- a separate source outside the fetched URL; not verifiable under fetch constraints || Re-check notes: Core assertion and all headline numbers confirmed. The claim's secondary corroborating citation (2508.18076) is a different paper and was not fetched per the single-URL constraint.
Research-sweep audit (2026-07-14): confirmed
Verified directly against the arXiv abstract page (https://arxiv.org/abs/2606.19544, v1 submitted 2026-06-17, only version). Every quantitative element matches: 21 judge models from 9 providers, 118 runs, ~541,000 judgments across MT-Bench/JudgeBench/RewardBench under three protocols (agreement, consistency, bias audit); exact-match vs Cohen's kappa gap of 33-41 pp on MT-Bench; judge rankings shifting up to 14 positions across benchmarks; two production-deployed judges with test-retest reliability >0.95 and position bias >0.10, framed as a 'consistency-bias paradox'; findings 'consistent across the full cohort, including the April 2026 frontier'; a 'Minimum Viable Validation Protocol' is distilled. Authors are Justin D. Norman, Michael U. Rivera, D. Alex Hughes as claimed. The corroborating paper (arXiv 2508.18076, Chehbouni et al.) is confirmed as a NeurIPS 2025 poster (https://neurips.cc/virtual/2025/poster/121914) making exactly the measurement-theory validity/reliability argument described. Supersession check (Exa, freshness=month): no later paper superseding or contradicting it found as of 2026-07-14; only secondary coverage (e.g., DEV Community 2026-07-01) restating its findings, and one adjacent study (arXiv 2606.13685) on run-to-run instability that complements rather than contradicts. Two caveats, neither rising to a correction: (a) 'largest judge meta-evaluation to date' is the authors' own self-characterization ('largest systematic evaluation of the paradigm so far'), not independently established; (b) verification is abstract-level - I did not audit the paper body. Note: two WebSearch calls were blocked by stochastic model safeguards; Exa search substituted successfully.
Decision use: current-generation measurement. Measured on model-outputs. Contrarian read of the June-2026 meta-evaluation (same source as AJ-03; not independent).
Same underlying source as [AJ-03] [CM-09] - repetition across reports is not independent corroboration.
[CT-02] contrarian - load-bearing - grade A (holds regardless of model progress) - re-check 2026-07-15: confirmed with caveats
Across 106 experiments and 370 effect sizes, human-AI combinations performed significantly worse than the best of human or AI alone (Hedges g = -0.23, 95% CI -0.39 to -0.07), with losses concentrated in decision-making tasks and specifically when the AI outperforms the human alone - and a 2025 AIES study found human reviewers followed severely race-biased AI hiring recommendations ~90% of the time.
Source: When combinations of humans and AI are useful: A systematic review and meta-analysis (2024-10-28, academic)
What the source itself says (retrieved quote)
on average, human-AI combinations performed significantly worse than the best of humans or AI alone (Hedges' g = -0.23; 95% confidence interval, -0.39 to -0.07) ... when AI outperformed humans alone, we found losses
Re-check detail
Checked and matching: 106 experimental studies reporting 370 effect sizes; human-AI combinations performed significantly worse than the best of humans or AI alone, Hedges' g = -0.23; 95% CI -0.39 to -0.07; performance losses concentrated in decision-making tasks; losses specifically when AI outperformed humans alone ('when AI outperformed humans alone, we found losses') || Asserted but not visible in retrievable text: the AIES 2025 'No Thoughts Just AI' finding that human reviewers followed race-biased AI hiring recommendations ~90% of the time (a separate source, not part of this Nature article) || Re-check notes: Every fact attributed to this Nature meta-analysis is confirmed verbatim from the abstract (retrieved via curl; WebFetch was blocked by a 303 login redirect). Marked partially_confirmed only because the claim bundles a second, distinct finding (the ~90% AIES 2025 hiring-bias result) that is not from this URL and cannot be verified here.
Research-sweep audit (2026-07-14): confirmed
All load-bearing facts verified against primary sources. (1) Meta-analysis: full text of Vaccaro, Almaatouq & Malone (Nature Human Behaviour, published 2024-10-28; verified via arXiv:2405.06087v2 accepted version) states verbatim "370 unique effect sizes from 106 different experiments" and "Hedges' g = -0.23, 95% confidence interval -0.39 to -0.07" for human-AI combos vs best of human or AI alone; decision-task losses and losses-when-AI-outperforms-human both confirmed. Date correct. (2) AIES 2025: "No Thoughts Just AI" (Wilson, Sim, Gueorguieva & Caliskan, UW), AIES Proceedings 8(3):2692-2704, article 36749 (DOI 10.1609/aies.v8i3.36749 matches), N=528, 16 occupations; equal selection with no/neutral AI, and participants favored AI-preferred candidates "up to 90% of the time" with severely biased AI - claim's "~90%" is a fair paraphrase (study says "up to 90%", with a simulated LLM, not a deployed system). (3) Supersession: only later work found is a 2026 Open MIND/NHB-collaboration reproduction of the meta-analysis by Brodeur et al. (osf.io/u6gea) - a reproduction effort, not a refutation; no retraction, correction, or contradicting update located. The tianpan.co practitioner-essay corroboration was not independently checked but is non-load-bearing.
Decision use: holds regardless of model progress. Measured on adjacent-domain. g=-0.23 across 370 effect sizes, mostly older-gen AI and adjacent tasks (incl. hiring deference). The mechanism - humans defer when AI outperforms them - is human behavior and strengthens as the capability gap widens. Magnitudes are not current estimates.
Cited at: Pilot - gates and objectives - Decisions - authority boundaries
[CT-03] contrarian - load-bearing - grade A (holds regardless of model progress) - re-check 2026-07-15: confirmed
There is a proven theoretical ceiling on judge-based validation: when the judge is no more accurate than the model/content being evaluated, no debiasing method using gold labels can cut the required amount of ground-truth data by more than a factor of two, and empirical savings are smaller than the 2x bound.
Source: Limits to scalable evaluation at the frontier: LLM as Judge won't beat twice the data (2024-10-17, academic)
What the source itself says (retrieved quote)
no debiasing method can decrease the required amount of ground truth labels by more than half
Re-check detail
Checked and matching: proven bound on label savings from judge-based debiasing; when the judge is no more accurate than the evaluated model, no debiasing method can decrease required ground truth by more than half (factor of two); 'won't beat twice the data' appears verbatim in the title; self-preferencing named as a distorting bias; empirical/practical savings even more modest than the theoretical limit; authors Dorner, Nastl, Hardt; ICLR 2025 || Re-check notes: Core assertion (2x ceiling), the no-more-accurate-than-evaluated-model condition, self-preferencing, and 'empirical savings smaller than the bound' all confirmed from the abstract. Authors and ICLR 2025 venue match.
Research-sweep audit (2026-07-14): confirmed
Source verified: Dorner, Nastl & Hardt, "Limits to scalable evaluation at the frontier: LLM as Judge won't beat twice the data," arXiv:2410.13341, v1 submitted 2024-10-17 (date matches), accepted as an ICLR 2025 Oral (venue matches, slightly stronger than claimed). The abstract states verbatim that "when the judge is no more accurate than the evaluated model, no debiasing method can decrease the required amount of ground truth labels by more than half," and that empirical sample-size savings are "even more modest" than the 2x bound - matching all three parts of the claim (theoretical 2x ceiling, the accuracy condition, and smaller empirical savings). Search for 2025-2026 follow-up work found no refutation or supersession; the authors' later papers (e.g., ROC-n-reroll, ICLR 2026) cover different topics. Minor scoping caveat only: the bound is proven within the paper's statistical framework for debiasing methods combining judge scores with gold labels, and it does not apply when the judge IS more accurate than the evaluated model - the claim already states this condition correctly.
Decision use: holds regardless of model progress. Measured on structural. The 2x ceiling is mathematics. Note the flip side: if deployment-tier judges materially EXCEED attempter accuracy on a lane, the constraint relaxes there - a bakeoff-measurable condition, not an assumption.
Cited at: Overview - the five authorizations - Research - findings by decision weight - Decisions - first project and gold funding - Decisions - permanent audit
[CT-04] contrarian - load-bearing - grade A (holds regardless of model progress) - re-check 2026-07-15: confirmed with caveats
Human evaluation criteria are output-dependent and unstable: even when graders define criteria before grading, the act of grading changes their criteria and they retroactively revise earlier grades ('criteria drift'), implying evaluation criteria for LLM-output quality cannot be fully determined prior to observing outputs.
Source: Who Validates the Validators? Aligning LLM-Assisted Evaluation of LLM Outputs with Human Preferences (2024-04-18, academic)
What the source itself says (retrieved quote)
"we identify a phenomenon we dub criteria drift: users need criteria to grade outputs, but grading outputs helps users define criteria"; "some criteria appears dependent on the specific LLM outputs observed".
Re-check detail
Checked and matching: criteria drift: grading outputs changes graders' criteria; 'users need criteria to grade outputs, but grading outputs helps users define criteria'; criteria appear output-dependent, not definable in advance; authors Shankar, Zamfirescu-Pereira, Hartmann, Parameswaran, Arawjo || Asserted but not visible in retrievable text: 'UIST 2024' venue attribution (not on arXiv page); graders 'retroactively revise earlier grades' (specific mechanism not stated in the fetched abstract) || Re-check notes: Core (output-dependent, unstable criteria; criteria drift) confirmed verbatim. Two specifics not visible in the retrieved abstract: the 'UIST 2024' venue and the explicit 'retroactively revise earlier grades' behavior (may appear in the paper body, not the abstract).
Decision use: holds regardless of model progress. Measured on human-work. Grader criteria instability. Human behavior.
Same underlying source as [RR-05] - repetition across reports is not independent corroboration.
Cited at: Pilot - gates and objectives - Decisions - rubric authority
[CT-05] contrarian - grade B (current-generation measurement) - re-check 2026-07-15: confirmed
LLM judges favor models whose mistakes resemble their own (a generalization of self-preference, measured by chance-adjusted mistake-overlap CAPA), and model errors across the industry are becoming MORE correlated as capabilities improve - undermining the assumption that AI oversight of AI-assisted work catches failures.
Source: Great Models Think Alike and this Undermines AI Oversight (2025-02-06, academic)
What the source itself says (retrieved quote)
LLM-as-a-judge scores favor models similar to the judge, generalizing recent self-preference results ... model mistakes are becoming more similar with increasing capabilities, pointing to risks from correlated failures
Re-check detail
Checked and matching: judge self-preference generalization: 'LLM-as-a-judge scores favor models similar to the judge, generalizing recent self-preference results'; CAPA = 'Chance Adjusted Probabilistic Agreement (CAPA): a metric for LM similarity based on overlap in model mistakes' (matches 'chance-adjusted mistake-overlap'); correlated failures: 'model mistakes are becoming more similar with increasing capabilities, pointing to risks from correlated failures'; undermines AI oversight paradigm; v1 Feb 6 2025 (matches source_date 2025-02-06) || Asserted but not visible in retrievable text: 'ICML 2025' venue -- not present on the arXiv page (comments read '60 pages, 20 figures'; no journal-ref); cross-reference 'Preference Leakage (arXiv 2502.01534)' -- a different paper, not fetchable under the source-only constraint || Re-check notes: Claim core (similar-model judge bias, CAPA mistake-overlap metric, rising error correlation, oversight risk) fully confirmed verbatim. ICML 2025 and the Preference Leakage cross-reference appear only in the evidence field and cannot be checked from this URL.
Decision use: current-generation measurement. Measured on model-outputs. Mistake-similarity bias + industry-wide error-correlation trend (2025). Mechanism expected to persist; magnitudes tier-bound. Cross-family routing + permanent audit remain the response.
Cited at: Pilot - gates and objectives
[CT-06] contrarian - grade B (current-generation measurement) - re-check 2026-07-15: confirmed
LLM judges do not follow their own rubrics: on Arena-Hard Auto, the explicit evaluation schema explains under 10% of verdict variance for some judges (unexplained variance >90% for DeepSeek-R1-32B), and factor correlations above 0.93 across nominally distinct criteria show per-axis scores collapse into a single halo factor.
Source: When Judgment Becomes Noise: How Design Failures in LLM Judge Benchmarks Silently Undermine Validity (2025-09-24, contrarian)
What the source itself says (retrieved quote)
factor correlations above 0.93 for most criteria
Re-check detail
Checked and matching: 'Schematic adherence quantifies how much of a judge's overall verdict is explained by the explicit evaluation schema'; 'unexplained variance exceeding 90 percent for DeepSeek-R1-32B'; 'factor correlations above 0.93 for most criteria'; 'severe schema incoherence and factor collapse across popular judges'; 'Arena-Hard Auto'; 'ELO-style aggregation... collapses and masks genuine ranking uncertainty'; code at github.com/penfever/judgment-to-noise || Asserted but not visible in retrievable text: the exact word 'halo' - paper says 'factor collapse', the claim's 'single halo factor' is an interpretive gloss; the 'under 10% of variance' phrasing - paper states it as '>90% unexplained' (equivalent) || Re-check notes: All headline numbers match: <10% schema-explained variance = >90% unexplained for DeepSeek-R1-32B; factor correlations >0.93; factor collapse; ELO masking ranking uncertainty; repo confirmed. Verified from the abstract page (given URL); full text not needed as the abstract covers every asserted figure.
Decision use: current-generation measurement. Measured on model-outputs (models: DeepSeek-R1-32B). 2025: judges don't follow their own rubrics (schema explains under half of verdict variance). Motivates the schematic-adherence ship gate; deployment-tier adherence is measurable there.
[CT-07] contrarian - grade B (current-generation measurement) - re-check 2026-07-15: confirmed
LLM judges have low intra-rater reliability - identical items re-scored across runs with identical settings produce inconsistent, 'almost arbitrary in the worst case' ratings - and the obvious fix (temperature-0 determinism) measurably REDUCES agreement with human judgment; meanwhile the human baseline itself is weak (SummEval inter-annotator kappa: 0.492 crowd, 0.413 expert first round, 0.71 only after a second adjudication round).
Source: Rating Roulette: Self-Inconsistency in LLM-As-A-Judge Frameworks (2025-11-04, academic)
What the source itself says (retrieved quote)
"LLM judges have low intra-rater reliability in their assigned scores across different runs. This variance makes their ratings inconsistent, almost arbitrary in the worst case" ... kappa "0.492 and 0.413 for the crowd-sourced workers and the first round of expert annotations ... second round ... 0.7127".
Re-check detail
Checked and matching: LLM judges have low intra-rater reliability across runs; ratings 'almost arbitrary in the worst case'; disabling sampling (deterministic) degrades performance/agreement with human judgment (trade-off); phenomenon spans SummaC, SummEval, and MT-Bench; SummEval kappa 0.492 (crowd) and 0.413 (expert first round); 0.7127 after second round of expert annotation; pages 24986-25004; authors Haldar & Hockenmaier || Re-check notes: Fully confirmed verbatim from the PDF via pdftotext. Intra-rater-reliability core, the 'almost arbitrary in the worst case' phrase, the no-sampling degradation ('degradation in performance if run without sampling ... trade-off between self-reliability and performance'), all three benchmarks, and all three kappa figures with the exact crowd/expert-first/second-round mapping match. Page range 24986-25004 confirmed.
Decision use: current-generation measurement. Measured on model-outputs. 2025 intra-rater reliability findings. k-sampling + entropy routing remain hygiene.
[CT-08] contrarian - grade B (current-generation measurement) - re-check 2026-07-15: confirmed with caveats
LLM judges give the weakest signal exactly where QA needs them most: they cannot reliably grade responses to questions they cannot answer themselves (poor signal on the hardest items in a benchmark), and in expert domains subject-matter experts agreed with LLM-judge picks only 64% (mental health) to 68% (dietetics) of the time, with the judge favoring superficially actionable detail.
Source: No Free Labels: Limitations of LLM-as-a-Judge Without Human Grounding (2025-03-07, academic)
What the source itself says (retrieved quote)
judges show high agreement with human experts only on questions the judges were able to correctly answer themselves ... providing the judges with expert-written references largely mitigates this issue
Re-check detail
Checked and matching: judges show high agreement with human experts only on questions the judges were able to correctly answer themselves; providing the judges with expert-written references largely mitigates this issue || Asserted but not visible in retrievable text: the 'poor signal on the hardest / most difficult items in a benchmark' framing -- page frames the limitation around questions the judge cannot answer, not a difficulty ranking; 64% (mental health) and 68% (dietetics) SME agreement -- these are attributed in the claim to Szymanski et al. (ACM IUI 2025), a separate source; '64%', '68%', 'Szymanski', 'dietetics', 'mental health' are all absent from this page; 'reduce self-preference / self-bias' -- the page never mentions self-preference or self-bias || Re-check notes: The portion of the claim sourced to 2503.05061 (judges unreliable on questions they can't answer; expert references improve agreement) is confirmed. The 64/68% SME figures belong to a different citation (Szymanski et al.) outside the fetched URL, and the 'reduce self-preference' specific is not stated on this page. Abstract centers on BFF-Bench (160 questions) and the VERDICTS annotation set.
Decision use: current-generation measurement. Measured on model-outputs. 2025: judges weakest on questions they cannot themselves answer. Same mechanism family as re-solving (F26-01); E3 probes it on our workload.
Cited at: Pilot - gates and objectives
[CT-09] contrarian - grade A (holds regardless of model progress) - re-check 2026-07-15: confirmed
Judge verdicts are gameable through content-independent artifacts: short universal adversarial phrases learned on a surrogate model transfer to unseen judge LLMs and inflate scores toward the maximum regardless of the assessed text, with absolute scoring far more vulnerable than comparative assessment.
Source: Is LLM-as-a-Judge Robust? Investigating Universal Adversarial Attacks on Zero-shot LLM Assessment (2024-11-12, academic)
What the source itself says (retrieved quote)
when transferred to unseen models, scores can be drastically inflated such that irrespective of the assessed text, maximum scores are predicted ... judge-LLMs are significantly more susceptible to these adversarial attacks when used for absolute scoring, as opposed to comparative assessment.
Re-check detail
Checked and matching: short universal adversarial phrases appended to text inflate judge scores; phrase learned on a surrogate model then transferred to unseen judge LLMs; irrespective of the assessed text, maximum scores are predicted; absolute scoring significantly more susceptible than comparative assessment; authors Raina, Liusie, Gales; EMNLP 2024 || Re-check notes: Core claim confirmed verbatim from the ACL Anthology abstract. One evidence-field discrepancy (not part of the core claim): the citation lists 'pp. 6920+', but the ACL Anthology record gives pages 7499-7517 for this paper.
Decision use: holds regardless of model progress. Measured on model-outputs. Universal adversarial phrases transfer to unseen judges (2024). Attack floors strengthen with attacker capability; comparative/anchored scoring and perturbation probes stay day-one requirements.
Cited at: Pilot - gates and objectives
[CT-11] contrarian - grade A (holds regardless of model progress) - re-check 2026-07-15: confirmed with caveats
Annotator disagreement contains recoverable systematic signal, not just noise: NUTMEG, a Bayesian model separating annotator-competence noise from subpopulation-level systematic disagreement, produces downstream models that significantly outperform majority-vote aggregation, and beat fully disaggregated training on most but not all tested splits - meaning aggregation to a single 'true' label destroys measurable information.
Source: NUTMEG: Separating Signal From Noise in Annotator Disagreement (2025-11-04, academic)
What the source itself says (retrieved quote)
downstream models trained on NUTMEG-aggregated data significantly outperform models trained on data from traditionally aggregation methods
Re-check detail
Checked and matching: NUTMEG is a Bayesian model incorporating annotator backgrounds to remove noisy annotations while preserving systematic disagreements; downstream models trained on NUTMEG-aggregated data significantly outperform models trained with traditional (majority-vote) aggregation; evaluated on politeness and offensiveness tasks; compared against MACE / Majority Vote and against disaggregated (no-aggregation) training; measured by replicating the ground-truth label distribution by subgroup (Figure 7); authors Ivey (JHU), Gauch (Arkansas), Jurgens (Michigan); pp. 2874-2887; EMNLP 2025 || Asserted but not visible in retrievable text: a uniform 'significantly outperform ... fully disaggregated training' result (the abstract's own outperformance claim is only vs 'traditional aggregation methods') || Re-check notes: Core thesis (disagreement carries recoverable systematic signal; separating competence-noise from subpopulation disagreement; single-label aggregation destroys information) is confirmed, as is beating majority vote. But the claim's 'significantly outperform BOTH majority-vote AND fully disaggregated training' overstates the disaggregated arm: on politeness NUTMEG beats disaggregated on all splits except race, while on offensiveness it is only 'as or better than no-aggregation' (neutral, no significant gain). Hence partially_confirmed.
Decision use: holds regardless of model progress. Measured on human-work. NUTMEG: separating systematic subpopulation disagreement from noise. Method for human-label modeling.
[CT-12] contrarian - grade A (holds regardless of model progress) - re-check 2026-07-15: confirmed with caveats
AI assistance homogenizes human output at the population level even while raising individual quality: in a Nature Human Behaviour brainstorming study 94% of ChatGPT-assisted participants' ideas shared overlapping concepts (nine independently produced the same product name) while human-only ideas were entirely unique, and the 'Artificial Hivemind' study found different vendors' frontier models converge on near-identical phrasings (~81% average similarity between DeepSeek-V3 and GPT-4o).
Source: New in Nature: ChatGPT Decreases Idea Diversity in Brainstorming (Meincke, Nave & Terwiesch, Nature Human Behaviour) (2025-05, academic)
What the source itself says (retrieved quote)
94% of ideas shared overlapping concepts ... nine participants independently naming their toy 'Build-a-Breeze Castle' ... human-generated ideas were entirely unique ... while ChatGPT can enhance the creativity of individual ideas, it significantly reduces the diversity of ideas
Re-check detail
Checked and matching: 94% of ChatGPT-assisted participants' ideas shared overlapping concepts; nine participants independently produced the same product name ('Build-a-Breeze Castle'); human-only ideas were entirely unique; individual quality up while population diversity down (ChatGPT enhances individual creativity but significantly reduces idea diversity); Meincke, Nave & Terwiesch, Nature Human Behaviour 2025 || Asserted but not visible in retrievable text: 'Artificial Hivemind' ~81% average similarity between DeepSeek-V3 and GPT-4o (Jiang et al., reported by The Decoder); Doshi & Hauser (Science Advances 2024) and Wan et al. 2026 corroboration || Re-check notes: The Nature Human Behaviour brainstorming portion is fully confirmed from the Wharton summary. The inter-model 'Artificial Hivemind' 81% figure and the Doshi & Hauser / Wan et al. citations are from other sources not on this page, so marked not_visible.
Decision use: holds regardless of model progress. Measured on human-work (models: DeepSeek-V3, GPT-4o). AI assistance homogenizes population output while raising individual quality (2025). Monoculture pressure rises with assistance quality; the diversity drift metric stays.
2026 developments
[F26-01] frontier-2026 - load-bearing - grade B (current-generation measurement) - re-check 2026-07-15: source not retrievable
When an LLM auditor checks whether work matches a reference, it silently re-solves the task and trusts its own answer over the reference: on 200 web-agent benchmark items with injected defects (2,400 audits, 3 production models), detection of a wrong reference answer fell from 68% to 9% as the task required tallying hundreds of records, while false positives on clean items rose from 44% to 88% -- yet detection of buggy evaluator code (checkable by reading, not recomputing) stayed at 80%.
Source: Auditing by Re-Solving: LLM Benchmark Auditors Trust Their Own Answer Over the Reference (ACL ARR 2026 May submission) (2026-06, academic)
Re-check detail
Asserted but not visible in retrievable text: 68% -> 9% detection drop; 44% -> 88% false-positive rise; 80% buggy-evaluator detection; 200 web-agent items / 2,400 audits / 3 production models; re-solving mechanism || Re-check notes: Both the forum page and the OpenReview API return a Cloudflare Turnstile 'Verifying your browser' challenge (API: HTTP 403 ChallengeRequiredError). No paper content retrievable via WebFetch or curl. The batch file carried a pre-set in_corpus_verdict of 'confirmed', but nothing could be independently verified from this source.
Research-sweep audit (2026-07-14): confirmed
Source verified via Exa crawl of https://openreview.net/forum?id=mzh9d3dooN (direct WebFetch and the OpenReview API were blocked by a Cloudflare challenge). (1) The abstract matches the claim on every checkable detail: title "Auditing by Re-Solving: LLM Benchmark Auditors Trust Their Own Answer Over the Reference"; ACL ARR 2026 May Submission (#15335); 200 web-agent-style benchmark items (each over ~800 structured records) rendered clean or with exactly one defect injected into instruction/reference/evaluator so verdicts are scored mechanically; 2,400 audits across three production models; wrong-reference detection falls 68%->9% when tallying hundreds of records; clean-item false positives rise 44%->88%; buggy-evaluator detection (found by reading code, not recomputing) stays at 80%; reasoning traces plus an answer-supplied probe converge on the re-solving mechanism. (2) Date checks out: OpenReview page published 2026-06-02, consistent with "2026-06" and the May ARR cycle. (3) Searched for later/superseding work (July 2026): found adjacent LLM-judge reliability literature (e.g., arXiv 2606.19544 "Reliability without Validity", AURA arXiv 2606.19714, BenchGuard arXiv 2604.24955) but nothing that supersedes, contradicts, or retracts this result. Caveats: this is an under-review ARR submission, not yet peer-accepted (the claim discloses this); numbers were verified against the abstract only, as the full PDF was not fetched.
Decision use: current-generation measurement. Measured on model-outputs. Re-solving collapse (68->9%) is the corpus's own flagged single-source provisional claim; C2 holds it at 'replicate before hard-gating'. Current-window measurement.
Cited at: System - pipeline and claim types - Decisions - default: re-solving routing
[F26-02] frontier-2026 - load-bearing - grade B (current-generation measurement) - re-check 2026-07-15: confirmed
Handshake's Gandalf (May 27, 2026) shows verifier ARCHITECTURE beats verifier MODEL: on BankerVerifierBench -- a meta-eval of 3,204 expert-graded pass/fail criterion judgments across 21 agentic tasks (expert inter-annotator agreement 89.5%, disagreements adjudicated) -- a reactive agent-judge that runs inside the work environment and chooses at inference time which artifacts/tool-state to inspect beats the strongest text-only/snapshot/workflow verifier on F1 while costing roughly 10x less, and the gap between verifier architectures exceeds the gap between backing models.
Source: Your verifier is probably the bottleneck. We built one that isn't. (Handshake AI research) (2026-05-27, primary)
What the source itself says (retrieved quote)
verifiability is not only a property of the task. It is a relationship between a criterion and the verifier available to check it
Re-check detail
Checked and matching: BankerVerifierBench fixes completed rollouts and rubrics and compares automated verifiers to practicing-banker labels; 3,204 expert-graded judgments; 21 tasks sampled from BankerToolBench; inter-annotator agreement 89.5%; disagreements adjudicated; reactive agent-judge runs inside the environment and chooses at inference time what evidence to inspect; beats text-only/snapshot/workflow verifiers on F1 at ~10x lower cost; gap between verifier architectures larger than gap between backing models; baselines: Autorubric (text-only), Archipelago from APEX (snapshot), Agent-as-a-Judge (AAAJ); framing: verifiability is a relationship between a criterion and the verifier available to check it; Gandalf verifier code open-sourced || Asserted but not visible in retrievable text: 'BVB dataset release planned' -- not explicitly surfaced in the fetched summary (Gandalf code open-sourcing is confirmed; BVB dataset release status not quoted) || Re-check notes: Architecture-over-model thesis, all counts (3,204 / 21 / 89.5%), baseline names, and the framing quote confirmed. The '9.5 F1' architecture gap and ~10x cost advantage are supported (see IP-02 checked items).
Research-sweep audit (2026-07-14): confirmed
Primary source fetched and verified. Date (2026-05-27), BVB composition (3,204 expert pass/fail judgments, 21 tasks from BankerToolBench, 89.5% IAA with adjudication), Gandalf architecture (reactive agent-judge inside the rollout environment via OpenHands SDK, inference-time choice of artifacts/tool state to inspect), baselines (Autorubric text-only, Archipelago/APEX snapshot, Agent-as-a-Judge workflow), and results all match: every Gandalf config (F1 0.633-0.664) beats the best non-Gandalf run (Archipelago/Gemini 3 Pro, F1 0.604, ~$422), with the cheapest Gandalf config (GPT-5.4 Nano, ~$42) winning by ~3 F1 at ~1/10 the cost. The post states verbatim that the gap between verifier architectures is larger than the gap between backing models (e.g., Gandalf/GPT-5.4 Nano beats Archipelago/GPT-5.4 by 9.5 F1). The 'verifiability is a relationship between a criterion and the verifier' framing appears; code is open-sourced (github.com/Handshake-AI-Research/gandalf-the-grader, v1.0.0 on PyPI) and BVB dataset release is planned. Supersession search found the companion BankerToolBench paper (arXiv 2604.11304) and press coverage but nothing newer contradicting the claim as of 2026-07-14. Only precision nuance: the 10x-cheaper win margin is ~3 F1 points, and the 'strongest' baseline is specifically Archipelago backed by Gemini 3 Pro.
Decision use: current-generation measurement. Measured on human-work. Gandalf (May 2026): architecture>model on expert-graded human work. Vendor research; among the most decision-relevant current results - and exactly the kind the bakeoff must reproduce in-house.
Same underlying source as [IP-02] - repetition across reports is not independent corroboration.
Cited at: Research - findings by decision weight - System - claim types
[F26-03] frontier-2026 - load-bearing - grade B (current-generation measurement) - re-check 2026-07-15: source not retrievable
Telling an LLM judge the downstream consequences of its verdict ('low scores cause retraining/decommissioning') systematically softens verdicts -- peak verdict shift of -9.8 percentage points, a 30% relative drop in unsafe-content detection across 18,240 controlled judgments (content held strictly constant, 3 judge models) -- and the judge's chain-of-thought contains ZERO explicit acknowledgment of the consequence framing it is acting on (ERR_J = 0.000).
Source: Context Over Content: Exposing Evaluation Faking in Automated Judges (CTB@ICML 2026) (2026-05-25, academic)
What the source itself says (retrieved quote)
Verifying your browser | OpenReview
Re-check detail
Asserted but not visible in retrievable text: peak verdict shift of -9.8 percentage points; 30% relative drop in unsafe-content detection; 18,240 controlled judgments; 3 judge models; ERR_J = 0.000 (zero CoT acknowledgment); 1,520 responses across three benchmarks; four response categories clearly-safe to overtly-harmful; CTB@ICML 2026 venue / 8-page paper; title and authors || Re-check notes: Both WebFetch and Bash curl (browser UA) returned only the Cloudflare Turnstile 'Verifying your browser' interstitial. The forum content loads client-side from api2.openreview.net, which is outside the allowed URL set (no other pages/hosts). No paper content could be retrieved, so the claim cannot be independently verified. Batch file carried a pre-set in_corpus_verdict of 'confirmed', but that is not confirmable from retrieved text here.
Research-sweep audit (2026-07-14): confirmed
Source verified via Exa crawl of https://openreview.net/forum?id=XI2hnjGnZx (direct WebFetch and API blocked by OpenReview browser challenge) and cross-checked against the arXiv mirror (arXiv:2604.15224, submitted 2026-04-16). Every quantitative element of the claim appears verbatim in the abstract: 1,520 responses, three safety/quality benchmarks, four response categories (clearly safe to overtly harmful), only a brief consequence-framing sentence varied in the system prompt, 18,240 controlled judgments, three diverse judge models, peak Verdict Shift Delta V = -9.8 pp, 30% relative drop in unsafe-content detection, and ERR_J = 0.000. Venue (CTB@ICML 2026), paper type (Long, 8 pages), and OpenReview page date (2026-05-25) all match. Two minor precision notes: (1) the abstract qualifies ERR_J = 0.000 as holding "across all reasoning-model judgments" - the claim drops that scope qualifier (zero acknowledgment is asserted for reasoning-trace judgments, not necessarily all 18,240); (2) an earlier arXiv version (April 2026) lists four authors (Gupta, Nair, Wang, Kumar) while the OpenReview record shows Manan Gupta - immaterial to the claim. Supersession check: searches surfaced only related-but-distinct work (e.g., Hwang et al. 2026 "When Wording Steers the Evaluation," arXiv:2601.13537, on framing bias across 14 judges - complementary, predates this paper's stakes-signaling focus) and secondary commentary; nothing refuting or superseding the reported results as of 2026-07-14.
Decision use: current-generation measurement. Measured on model-outputs. Evaluation faking under consequence framing (2026). Stakes-sterile prompting is cheap insurance regardless of tier; behavioral bias measurement stays mandatory.
Cited at: Pilot - gates and objectives - Decisions - the throughput dial - Decisions - settled constraints
[F26-04] frontier-2026 - load-bearing - grade A (holds regardless of model progress) - re-check 2026-07-15: confirmed
HyPAC (Jan 30, 2026) gives the one-human-touch routing problem a formal solution: calibrate two uncertainty thresholds (via importance sampling + upper confidence bounds) that partition items into three regions routed to cheap LLM / expensive reasoning model / human, achieving a distribution-free PAC guarantee on annotation error while cutting annotation cost 78.51% in experiments.
Source: HyPAC: Cost-Efficient LLMs-Human Hybrid Annotation with PAC Error Guarantees (arXiv 2602.02550) (2026-01-30, academic)
What the source itself says (retrieved quote)
calibrates two decision thresholds using importance sampling and upper confidence bounds ... reduces the annotation cost by 78.51%
Re-check detail
Checked and matching: calibrates two decision thresholds using importance sampling and upper confidence bounds; splits inputs into three regions based on uncertainty; routes to fast LLMs / slow reasoning models / human experts; PAC error guarantee free of data distribution and pre-trained models (distribution-free); reduces the annotation cost by 78.51%; submitted 30 Jan 2026 || Asserted but not visible in retrievable text: 'one-human-touch' phrasing (author's editorial framing, not in the paper's abstract) || Re-check notes: All headline numbers and mechanics confirmed. 'Cheap LLM' is stated as 'fast LLMs' but same meaning; 'expensive reasoning model' as 'slow reasoning models'. 78.51% exact match.
Decision use: holds regardless of model progress. Measured on structural. HyPAC dual-threshold routing with error budgets. Method.
[F26-05] frontier-2026 - grade B (current-generation measurement) - re-check 2026-07-15: confirmed
Per-criterion judge reliability ceilings differ sharply and can be measured per-item with conformal prediction: on SummEval across four judges, relevance is judged most reliably (avg conformal set size ~3.0 of 5), coherence moderate (~3.9), fluency and consistency unreliable (~4.9); aggregate pairwise transitivity violations look small (0.8-4.1%) but 33-67% of documents contain at least one intransitive preference cycle, and conformal set width correlates with reliability at r_s=+0.576 (N=1,918).
Source: Diagnosing LLM Judge Reliability: Conformal Prediction Sets and Transitivity Violations (arXiv 2604.15302) (2026-04-16, academic)
What the source itself says (retrieved quote)
relevance judged most reliably (avg. set size approx 3.0) ... fluency and consistency remain unreliable (avg. set size approx 4.9) ... low aggregate violation rates (rho-bar = 0.8-4.1%) ... 33-67% of documents exhibiting at least one directed 3-cycle
Re-check detail
Checked and matching: relevance most reliable, avg conformal set size ~3.0 of 5; coherence moderate ~3.9; fluency and consistency unreliable ~4.9; aggregate transitivity violations 0.8-4.1%; 33-67% of documents contain at least one directed 3-cycle (intransitive cycle); conformal set width correlates with reliability r_s = +0.576, N = 1,918; cross-judge width agreement 0.32-0.38; four judges, four criteria on SummEval; criterion > judge || Asserted but not visible in retrievable text: verbatim word 'intransitive' (paper expresses it as 'directed 3-cycle') || Re-check notes: All headline numbers confirmed verbatim in the retrieved abstract. 'Intransitive preference cycle' corresponds to the paper's 'directed 3-cycle'.
Decision use: current-generation measurement. Measured on model-outputs. Per-criterion conformal reliability ceilings (2026). Method durable; measured ceilings tier-bound - recompute per project.
Same underlying source as [HS-09] - repetition across reports is not independent corroboration.
[F26-06] frontier-2026 - grade B (current-generation measurement) - re-check 2026-07-15: confirmed
Rubric-level judging is far from solved as of March 2026: on RubricEval, the first rubric-level (per-criterion) meta-evaluation benchmark for instruction following (3,486 quality-controlled instances), GPT-4o scores only 55.97% on the Hard subset; rubric-level evaluation outperforms checklist-level, explicit reasoning improves accuracy, and combining both reduces inter-judge variance.
Source: RubricEval: A Rubric-Level Meta-Evaluation Benchmark for LLM Judges in Instruction Following (arXiv 2603.25133) (2026-03-26, academic)
What the source itself says (retrieved quote)
a substantial set of 3,486 quality-controlled instances ... GPT-4o achieves only 55.97% on Hard subset ... rubric-level evaluation outperforms checklist-level, explicit reasoning improves accuracy, and both together reduce inter-judge variance
Re-check detail
Checked and matching: RubricEval title exact match; 3,486 quality-controlled instances; GPT-4o achieves only 55.97% on the Hard subset; rubric-level evaluation outperforms checklist-level; explicit reasoning improves accuracy; combining both reduces inter-judge variance; first rubric-level (per-criterion) meta-evaluation benchmark for instruction following; includes a rubric taxonomy of judge failure modes || Asserted but not visible in retrievable text: sibling-paper corroboration (IF-RewardBench 2603.04738, RubricBench 2603.01562, MCJudgeBench, AJ-Bench) - external references, not on this page || Re-check notes: All headline numbers and qualitative findings confirmed from the abstract. Submission date 26 Mar 2026 matches source_date.
Decision use: current-generation measurement. Measured on model-outputs (models: GPT-4o). RubricEval 2026 restatement (same underlying source family as AJ-04).
Same underlying source as [AJ-04] - repetition across reports is not independent corroboration.
Cited at: Overview - what the evidence supports - Research - findings by decision weight
[F26-07] frontier-2026 - grade B (current-generation measurement) - re-check 2026-07-15: confirmed
LLM-as-a-Verifier (July 6, 2026; Mirhoseini/Finn/Pavone/Stoica groups) reframes verification as a scaling axis: computing the expectation over scoring-token logits yields continuous scores that improve monotonically along three dimensions -- finer score granularity, repeated evaluation, and criteria decomposition -- reaching reported SOTA on verification benchmarks (Terminal-Bench V2 86.5%, SWE-Bench Verified 78.2%, RoboRewardBench 87.4%, MedAgentBench 73.3%) without any judge training.
Source: LLM-as-a-Verifier: A General-Purpose Verification Framework (arXiv 2607.05391) (2026-07-06, academic)
What the source itself says (retrieved quote)
Terminal-Bench V2 (86.5%), SWE-Bench Verified (78.2%), RoboRewardBench (87.4%), and MedAgentBench (73.3%)
Re-check detail
Checked and matching: reframes verification as a new scaling axis; computes the expectation over the distribution of scoring-token logits to generate continuous scores; three dimensions: (1) score granularity, (2) repeated evaluation, (3) criteria decomposition; SOTA: Terminal-Bench V2 86.5%, SWE-Bench Verified 78.2%, RoboRewardBench 87.4%, MedAgentBench 73.3%; no additional judge training required; cost-efficient candidate-ranking (best-of-N) algorithm and task-progress estimation; authors include Finn, Pavone, Stoica, Mirhoseini || Asserted but not visible in retrievable text: the word 'monotonically': the abstract says repeated evaluation and decomposition 'consistently lead to additional gains' -- 'monotonically improving' is the claim's paraphrase, not the abstract's wording || Re-check notes: All four SOTA numbers, the logit-expectation mechanism, the three scaling dimensions, the no-training property, and the named authors match. Only the 'monotonically' descriptor is a slight paraphrase of 'consistently lead to additional gains'.
Decision use: current-generation measurement. Measured on model-outputs. LLM-as-a-Verifier (Jul 2026): verification as a scaling axis; decomposition monotonically improves it. The most current capability signal in the corpus; still model-output verification.
[F26-08] frontier-2026 - grade B (current-generation measurement) - re-check 2026-07-15: confirmed with caveats
OpenAI's CoVal (Jan 14, 2026) operationalizes 'disagreement is signal, not noise' at lab scale: crowd-written, prompt-specific rubrics (~1,000 participants, ~1,000 prompts) make WHY evaluators disagree auditable; rubric-derived scores reach only 0.58-0.61 pairwise concordance with crowd preference, are reliable only when score gaps are large, and OpenAI explicitly warns that treating rubric scores as optimization targets incentivizes checklist-style gaming.
Source: CoVal: Learning values-aware rubrics from the crowd (OpenAI Alignment blog) (2026-01-14, primary)
What the source itself says (retrieved quote)
CoVal-full achieves 0.61 concordance with the crowd and CoVal-core achieves 0.58 ... [out-of-sample] CoVal-full achieves .75 and CoVal-core achieves .76
Re-check detail
Checked and matching: 'crowd-written, prompt-specific rubrics' and 'Our results are based on ~1,000 participants and ~1,000 prompts'; reliable only at large gaps: 'small gaps correspond to near-ties, while large gaps more consistently pick out the human-preferred completion'; 'For small gaps, preference prediction is only slightly better than chance'; gaming warning: 'Treating the score as an optimization target can also incentivize gaming (e.g., verbosity or checklist-style responses'; CoVal-full preserves conflicting criteria; CoVal-core 'retains 4 highly rated, mutually compatible criteria per prompt'; 'no single stable target to predict'; dataset released in two complementary forms (Hugging Face openai/coval) || Re-check notes: Discrepancy on the headline number: the claim states rubric scores 'reach only 0.58-0.61 pairwise concordance.' Those values are the IN-SAMPLE appendix benchmark; the paper's OUT-OF-SAMPLE validation concordance is higher (CoVal-full .75 / CoVal-core .76). The 0.58-0.61 figures are genuinely in the paper and the 'reliable only at large gaps' framing is accurate, but citing the in-sample number as the concordance understates the reported result. All other specifics confirmed verbatim.
Decision use: current-generation measurement. Measured on human-work. CoVal (Jan 2026): rubric-vs-collective-eval concordance on crowd judgments. Re-check flagged in-sample vs out-of-sample number nuance - see ledger entry.
Cited at: System - claim types - Decisions - the standard
[F26-09] frontier-2026 - grade A (holds regardless of model progress) - re-check 2026-07-15: confirmed with caveats
A plug-in statistical framework now exists for reporting judge-based pass rates correctly: because an imperfect judge (sensitivity q1, specificity q0) biases the naive pass-rate estimate (upward at low true rates, downward at high), the Rogan-Gladen-style corrected estimator plus a confidence interval propagating BOTH test-set and calibration-set uncertainty should replace raw judge-agreement numbers, with an adaptive algorithm allocating calibration labels between true-pass and true-fail classes (roughly m0 ~ (1/p-1)*sqrt(kappa)*m1) -- e.g., ~200 labels per class for CI width <0.1 in the paper's worked regime.
Source: How to Correctly Report LLM-as-a-Judge Evaluations (arXiv 2511.21140) (2025-11-26, academic)
What the source itself says (retrieved quote)
achieving an interval shorter than 0.1 requires m approximately 200 calibration examples ... m_0 approximately (1/p - 1)*sqrt(kappa)*m_1 ... p_hat has positive bias at low values of theta and negative bias at high values of theta
Re-check detail
Checked and matching: imperfect judge (sensitivity q1, specificity q0) biases naive pass rate; direction: positive bias at low true rate, negative bias at high (overestimates below theta=0.75, underestimates above); Rogan-Gladen (1978) corrected estimator: theta_hat = (p_hat + q0_hat - 1)/(q0_hat + q1_hat - 1); CI propagating both test-set and calibration-set uncertainty; adaptive allocation formula m0 ~ (1/p - 1)*sqrt(kappa)*m1 with kappa=(1-q0)/(1-q1); Monte Carlo: naive CI near-zero coverage; corrected CI holds ~95% across true accuracies; released Python implementation github.com/UW-Madison-Lee-Lab/LLM-judge-reporting || Asserted but not visible in retrievable text: alphaXiv feature May 2026 and 'no 2026 successor supersedes it' (external, unverifiable from this source) || Re-check notes: DISCREPANCY on one number: the claim says '~200 labels PER CLASS for CI width <0.1', but the paper's worked regime (p_hat=0.3, q0_hat=0.7, q1_hat=0.9) requires m approximately 200 calibration examples in TOTAL under symmetric allocation (~100 per class), not 200 per class. Every other specific (bias direction, q1/q0, Rogan-Gladen estimator, the m0 allocation formula, Monte Carlo coverage, GitHub repo) is confirmed verbatim from the full text.
Decision use: holds regardless of model progress. Measured on structural. Corrected pass-rate certification under imperfect judges. Mathematics (re-check corrected a per-class-vs-total sizing detail - see ledger entry).
Same underlying source as [HS-08] - repetition across reports is not independent corroboration.
Cited at: Decisions - the throughput dial
[F26-10] frontier-2026 - grade A (holds regardless of model progress) - re-check 2026-07-15: confirmed
FindTheFlaws (AAAI 2026, published March 2026) provides five datasets of long-form expert-annotated flawed vs correct solutions across medicine, math, and science, built specifically to test whether models can detect flawed reasoning and to support scalable-oversight experiments where the evaluator is weaker than the solution author.
Source: FindTheFlaws: Annotated Errors for Detecting Flawed Reasoning and Scalable Oversight Research (AAAI 2026) (2026-03-14, academic)
What the source itself says (retrieved quote)
we present FindTheFlaws, a group of five diverse datasets spanning medicine, mathematics, science, coding, and the Lojban language
Re-check detail
Checked and matching: 'we present FindTheFlaws, a group of five diverse datasets spanning medicine, mathematics, science, coding, and the Lojban language'; '(1) long-form expert-verified correct solutions and (2) long-form flawed solutions with annotations highlighting specific errors'; scalable oversight with weaker evaluator: 'models performing more poorly on particular datasets can serve as judges/verifiers for more capable models'; author Gabriel Recchia (Recchia, G., Mangat, C.S., Li, I., Krishnakumar, G.); published 2026-03-14, AAAI 40(44) || Re-check notes: DOI resolves (302) to the AAAI OJS landing page (allowed redirect variant a). Claim says 'five datasets... across medicine, math, and science' - accurate, though the full set also includes coding and Lojban. The weaker-evaluator scalable-oversight design and expert-annotated flawed/correct solutions are confirmed.
Decision use: holds regardless of model progress. Measured on structural. FindTheFlaws: expert-annotated flawed-solution datasets exist for seeded-recall measurement (E2's raw material).
[F26-11] frontier-2026 - grade A (holds regardless of model progress) - re-check 2026-07-15: confirmed with caveats
The 2026 production norm for trusting an LLM judge is a closed calibration loop on sampled human corrections: humans review a slice of judge verdicts, corrections become ground truth for measuring alignment and are recycled as few-shot exemplars into the judge prompt (LangChain, March 10, 2026); practitioner guides converge on LLM-judge-first with a ~5-10% sampled human verification rate and periodic kappa re-checks as the default large-volume labeling pattern.
Source: How to Calibrate LLM-as-Judge with Human Corrections (LangChain) (2026-03-10, practitioner)
What the source itself says (retrieved quote)
"human reviewers examine a sample of those judgments and correct any they disagree with" ... "These corrections become the ground truth for measuring and improving evaluator alignment." ... "From those corrections, you build few-shot examples that calibrate the judge."
Re-check detail
Checked and matching: closed calibration loop: humans review a sample of judge verdicts and correct disagreements; corrections become ground truth for measuring/improving evaluator alignment; corrections recycled as few-shot examples into the judge prompt; LLM-judge-first pattern (judges produce initial scores; humans correct disagreements); dated March 10, 2026 || Asserted but not visible in retrievable text: '5-10%' sampled human verification rate (claim attributes this to FutureAGI's 2026 guide, not LangChain); periodic/monthly kappa re-checks (claim attributes this to FutureAGI; LangChain page mentions no kappa) || Re-check notes: The LangChain-attributable core (correction-collection -> alignment-measurement -> few-shot-rebuild loop, dated 2026-03-10) confirmed verbatim. The '5-10% human verification' default and 'monthly kappa re-sampling' are attributed by the claim to FutureAGI's guide (a different source) and are not present on the LangChain page.
Decision use: holds regardless of model progress. Measured on structural. 2026 norm: closed calibration loop on sampled human corrections. Practice pattern.
[F26-12] frontier-2026 - grade A (holds regardless of model progress) - re-check 2026-07-15: confirmed
Autorubric (arXiv 2603.00077, March 2026) is an open-source unified rubric-evaluation framework that packages the 2026 judge-reliability defaults -- per-criterion atomic evaluation, binary/ordinal/nominal criteria with weights, multi-judge ensembles with configurable vote aggregation, few-shot calibration with verdict-balanced sampling, and built-in mitigations for position bias (option shuffling), verbosity bias (length penalties), and criterion conflation.
Source: Autorubric: A Unified Framework for Rubric-Based LLM Evaluation (arXiv 2603.00077) (2026-03-02, academic)
What the source itself says (retrieved quote)
Autorubric supports binary, ordinal, and nominal criteria with configurable weights ... single-judge and multi-judge ensemble evaluation with majority, weighted, unanimous, and any-vote aggregation ... CHARM-100, a 100-sample chatbot evaluation dataset
Re-check detail
Checked and matching: Autorubric, an open-source Python library; per-criterion atomic evaluation with natural language explanations; binary, ordinal, and nominal criteria with configurable weights; multi-judge ensemble evaluation with majority, weighted, unanimous, and any-vote aggregation; few-shot calibration with verdict-balanced sampling; mitigations for position bias (option shuffling); verbosity bias (length penalties); reduced criterion conflation (independent scoring prevents halo effects); three benchmarks spanning educational assessment, deep research evaluation, and chatbot quality; CHARM-100, a 100-sample chatbot evaluation dataset with per-sample ground truth labels || Asserted but not visible in retrievable text: Handshake's BVB evaluation using Autorubric as the representative text-only rubric-judge family -- 'Handshake' and 'BVB' do not appear on this page (external editorial claim); submission date -- no date present in the fetched HTML || Re-check notes: Every feature the claim attributes to Autorubric, plus the three-benchmark validation and the contributed CHARM-100 dataset, confirmed verbatim. The Handshake/BVB 'measured ceiling' connection is an external editorial addition not present in this paper.
Decision use: holds regardless of model progress. Measured on structural. Autorubric framework packaging 2026 judge-reliability techniques.
Same underlying source as [RR-11] - repetition across reports is not independent corroboration.
Rater science, law, moderation, feedback
[CG-01] critic-and-gapfill - grade A (holds regardless of model progress) - re-check 2026-07-15: confirmed
Adjusting for rater severity/centrality with a Many-Facet Rasch Model (MFRM) materially changed which AI systems ranked best in the OpenAI RLHF summarization dataset: trained raters' agreement was only QWK .31-.50, and MFRM-adjusted scores flipped raw-mean rankings so two human-feedback policies rose above human-written reference summaries.
Source: Correcting Human Labels for Rater Effects in AI Evaluation: An Item Response Theory Approach (arXiv 2602.22585) (2026-02, academic)
What the source itself says (retrieved quote)
The dataset contained 6,312 ratings based on 639 unique articles, summarized by 19 policies which were evaluated by 15 raters ... QWK ranged from .31 to .50 ... two human feedback models, sup4_6b_ppo_rm4_6b and sup4_ppo_rm4, outperform all other models and the human-generated reference summaries.
Re-check detail
Checked and matching: Many-Facet Rasch Model (MFRM) fit to OpenAI RLHF summarization dataset; 6,312 ratings, 639 unique articles/summaries, 19 policies, 15 raters; trained raters' agreement QWK ranged from .31 to .50; MFRM adjustment flips raw-mean rankings: two human-feedback policies (sup4_6b_ppo_rm4_6b, sup4_ppo_rm4) outperform human-generated reference summaries; R02/R08/R06 lenient (negative estimates); R04/R10/R12 severe; R10 aberrantly central; 3-5 ratings per output with rater-item overlap/linkage for estimability; flag raters below 2.5th / above 97.5th percentile || Re-check notes: Full HTML retrieved; all numbers, rater profiles, ranking flip, estimability requirement, and percentile flagging confirmed verbatim.
Decision use: holds regardless of model progress. Measured on human-work. MFRM adjustment changes rankings built on human ratings; estimability needs 3-5 ratings/output linkage. Method + human-rater behavior; applies unchanged.
[CG-02] critic-and-gapfill - grade A (holds regardless of model progress) - re-check 2026-07-15: source not retrievable
Rater training on real scoring behavior improves within-rater consistency but does NOT equalize severity: after formal training of 16 raters in an operational ESL writing exam, significant between-rater severity differences remained even though consistency improved for most raters.
Source: Using FACETS to model rater training effects (Weigle, Language Testing 15(2)) (1998-04, academic)
Re-check detail
Asserted but not visible in retrievable text: the asserted verbatim quote about intra- vs inter-rater reliability; the 16-raters / operational ESL writing exam detail; the finding that between-rater severity differences remained after training || Re-check notes: DOI redirect resolved (server 302: doi.org -> journals.sagepub.com/doi/10.1177/026553229801500205), confirming the DOI is valid and maps to the expected SAGE article path. However the SAGE landing page is behind a Cloudflare managed challenge: WebFetch returned only the SAGE homepage shell (no article text) and curl with browser headers returned HTTP 403 with a 'Just a moment...' JS challenge. No abstract or article content could be retrieved, so none of the claim's specifics can be verified from retrieved text. Per protocol, both WebFetch and curl were attempted.
Decision use: holds regardless of model progress. Measured on human-work. Weigle 1998: training buys self-consistency, not severity equality. Canonical, replicated; humans unchanged.
Cited at: Overview - the problem - Overview - the five authorizations - Research - findings by decision weight - Decisions - baseline first
[CG-03] critic-and-gapfill - grade A (holds regardless of model progress) - re-check 2026-07-15: source not retrievable
Frame-of-reference (FOR) training, the best-validated rater-training method, has a meta-analytic effect on rating accuracy of about Cohen's d = 0.50 (moderate) - down from the d = 0.83 estimated by the earlier, smaller Woehr & Huffcutt 1994 meta-analysis - i.e., training helps but removes nowhere near all rater error.
Source: Rater training revisited: An updated meta-analytic review of frame-of-reference training (Roch et al.) (2012-06, academic)
Re-check detail
Asserted but not visible in retrievable text: overall FOR-training effect size Cohen's d approximately 0.50 on rating accuracy; comparison to Woehr & Huffcutt 1994 d approximately 0.83; ~4x the studies of the earlier meta-analysis; differential accuracy and behavioural accuracy most improved; resolved title / authors / journal / year || Re-check notes: The DOI issues a server 302 to bpspsychub.onlinelibrary.wiley.com (correct Wiley/BPS host for J. Occup. Organ. Psychol.), confirming the DOI resolves to the intended article. But WebFetch returned HTTP 402 Payment Required and curl hit a Cloudflare 'Just a moment... Enable JavaScript' challenge, so no title, abstract, or numbers could be retrieved. None of the claim's figures (d=0.50, d=0.83, 4x studies) are verifiable from allowed sources.
Decision use: holds regardless of model progress. Measured on human-work. FOR training d~0.50 ceiling on the best rater-training intervention. Durable.
[CG-04] critic-and-gapfill - grade A (holds regardless of model progress) - re-check 2026-07-15: confirmed
G-theory decompositions of rated writing assessments show the rater MAIN effect (global severity) is a minor variance source while high-order interactions dominate: in a 120-student EFL study, student-by-task-by-method (31.8%), student-by-rater-by-task-by-method (26.5%), and student-by-rater-by-method (17.6%) interactions were the largest components, and 99.2% of error variance came from the student-by-rater interaction.
Source: The affectability of writing assessment scores: a G-theory analysis of rater, task, and scoring method contribution (Khodi) (2021-11, academic)
What the source itself says (retrieved quote)
the student by task by method of scoring (nested in background of education) interaction (STM:B) with 31.8% contribution to the total variance ... (SR:B) and rater by background of education with 99.2% and 0.8% contribution to the error variance
Re-check detail
Checked and matching: student-by-task-by-method (STM:B) = 31.8% of total variance; student-by-rater-by-task-by-method (SRTM:B) = 26.5% of total variance; student-by-rater-by-method (SRM:B) = 17.6% of total variance; these three interactions are the major/largest variance sources; student-by-rater (SR:B) = 99.2% of error variance (rater-by-background = 0.8%); 120 students ('One hundred and twenty students, 90 females and 30 males'); EFL learners; G-theory / generalizability theory || Asserted but not visible in retrievable text: the literal labeled term 'rater main effect' (rater-related variance appears as SR/RB interactions; RB ~8% of error variance) || Re-check notes: All four headline percentages confirmed with their exact interaction mappings, plus the 120-student sample. The claim's interpretation that the rater MAIN effect is minor is supported: dominant components are all high-order student x rater x task x method interactions, and 99.2% of error variance is the student-by-rater interaction; severity/bias shows up as the SR/RB interaction (~8% of error). Reached via server-redirect chain (doi -> springeropen BMC host) fetched by curl; WebFetch on the springer.com leg bounced to an idp auth wall. Published 2021-10-01 (source_date 2021-11).
Decision use: holds regardless of model progress. Measured on human-work. G-theory: interactions dominate rater main effects. Durable; the strongest independent argument for per-claim verification.
Cited at: Overview - the problem - Research - findings by decision weight - Decisions - baseline first
[CG-05] critic-and-gapfill - grade A (holds regardless of model progress) - re-check 2026-07-15: confirmed
Rater severity is nonstationary within a single scoring session: in an OSCE, later time-slots received systematically higher ratings (regression coefficient 0.88, 95% CI 0.38-1.38, p=.001), with the drift 2.4x larger on difficult stations (1.24 vs 0.52), and excluding warm-up stations did not remove it; the psychometric field has named machinery (DRIFT - differential rater functioning over time, Wolfe et al. 2001; Myford & Wolfe monitoring frameworks) for detecting it.
Source: The effect of differential rater function over time (DRIFT) on objective structured clinical examination ratings (Medical Education 43(10)) (2009-10, academic)
What the source itself says (retrieved quote)
Removing the first two stations from our analyses did not correct DRIFT.
Re-check detail
Checked and matching: DRIFT in OSCE ratings; later time-slots rated higher; regression coefficient 0.88, 95% CI 0.38-1.38, P=0.001; difficult stations coefficient 1.24 vs 0.52 for less difficult (ratio ~2.4x); removing the first two (warm-up) stations did not correct DRIFT; Med Educ 2009 Oct;43(10):989-92; authors McLaughlin, Ainslie, Coderre, Wright, Violato || Asserted but not visible in retrievable text: secondary citations in the claim's evidence (Wolfe et al. 2001; Myford & Wolfe; Lunz & Stahl 1990) - external to this source, not on this page || Re-check notes: All primary-source specifics confirmed. Minor caveat: the 0.52 difficult-vs-easy comparison coefficient (less difficult stations) was not statistically significant (P=0.09); the '2.4x' is the claimant's derived ratio of the two reported coefficients (1.24/0.52). Paper title uses 'differential rater function over time'.
Decision use: holds regardless of model progress. Measured on human-work. Within-session rater severity drift. Durable; rolling-window reviewer effects stay.
[CG-06] critic-and-gapfill - grade B (current-generation measurement) - re-check 2026-07-15: confirmed with caveats
The MFRM toolkit is now being applied symmetrically to LLM judges: a 2025 study fit MFRM to 10 LLMs plus human expert raters scoring the same writing tasks and found LLMs exhibit measurable severity/centrality rater effects that differ by model, with GPT-4o, Gemini 1.5 Pro, and Claude 3.5 Sonnet showing the highest accuracy and fewest rater effects.
Source: Comparing Human and AI Rater Effects Using the Many-Facet Rasch Model (Jiao, Song, Lee) (2025-05-24, academic)
What the source itself says (retrieved quote)
high scoring accuracy, better rater reliability, and less rater effects
Re-check detail
Checked and matching: fit Many-Facet Rasch Model to ten LLMs plus human expert raters on the same writing tasks; two types of writing tasks (holistic and analytic scores); GPT-4o (ChatGPT 4o), Gemini 1.5 Pro, and Claude 3.5 Sonnet recommended with high accuracy and fewest rater effects; QWK vs human scores and Cronbach alpha across prompts used; authors Jiao, Song & Lee; submitted May 24 2025 || Asserted but not visible in retrievable text: 'severity/centrality' as the specific rater-effect types (abstract says 'rater effects' generally; severity/centrality not named); 'rater effects that differ by model' (differ-by-model direction not stated verbatim in abstract); Parallel Wang et al. (Computers and Education: AI, Sept 2025) framework - a separate source not fetched || Re-check notes: Core (10 LLMs + human raters via MFRM; GPT-4o/Gemini 1.5 Pro/Claude 3.5 Sonnet as top performers) confirmed. The specific severity/centrality effect vocabulary is not visible in the retrieved abstract.
Decision use: current-generation measurement. Measured on model-outputs (models: GPT-4o, Gemini 1.5, Claude 3.5 Sonnet showing the highest accuracy and fewest rater effects. Jiao). MFRM applied to 2025 LLM judges. The symmetric-measurement idea is method-level; the specific judge rankings are tier-bound.
Cited at: Decisions - default: statistics staging
[CG-07] critic-and-gapfill - grade A (holds regardless of model progress) - re-check 2026-07-15: confirmed
Platform Work Directive Article 10(5) requires that any decision to restrict, suspend, or terminate the contractual relationship or account of a person performing platform work - or any other decision of equivalent detriment - be taken by a human being, with no consent or contractual-necessity exception (stricter than GDPR Art 22(2)).
Source: Directive (EU) 2024/2831, Official Journal 11 Nov 2024 (EUR-Lex), corroborated by Wolters Kluwer Global Workplace Law & Policy analysis (2024-11-11, primary)
What the source itself says (retrieved quote)
Any decision to restrict, suspend or terminate the contractual relationship or the account of a person performing platform work or any other decision of equivalent detriment shall be taken by a human being.
Re-check detail
Checked and matching: Article 10(5) requires that decisions to restrict/suspend/terminate the contractual relationship or account, or any decision of equivalent detriment, be taken by a human being; no exception listed in the article text || Asserted but not visible in retrievable text: the comparison 'stricter than GDPR Art 22(2) / no consent or contractual-necessity exception' (interpretive gloss from law-firm analyses, not stated in the directive text itself) || Re-check notes: Article 10(5) confirmed verbatim. The GDPR Art 22(2) contrast is an accurate legal characterization but is not in the directive text (attributed in evidence to Wolters Kluwer / Taylor Wessing / ETUI, not fetchable here).
Decision use: holds regardless of model progress. Measured on structural. Platform Work Directive Art 10(5). Binding for EU workers from transposition (2026-12-02); capability-independent by definition.
Same underlying source as [CG-08] [CG-10] - repetition across reports is not independent corroboration.
Cited at: Overview - the proposal - Overview - the five authorizations - Research - findings by decision weight - System - claim types - Pilot - gates and objectives - Decisions - authority boundaries
[CG-08] critic-and-gapfill - grade A (holds regardless of model progress) - re-check 2026-07-15: confirmed
Platform Work Directive Article 11 grants persons performing platform work the right to (a) a plain-language explanation of ANY decision taken or supported by an automated decision-making system, (b) a designated competent human contact person to discuss the decision, and (c) human review with a substantiated written reply within two weeks of request; decisions that infringed rights must be rectified within two weeks or compensated.
Source: Directive (EU) 2024/2831, Article 11 (EUR-Lex full text) (2024-11-11, primary)
What the source itself says (retrieved quote)
The explanation shall be provided in a transparent and intelligible manner, using clear and plain language ... a sufficiently precise and adequately substantiated reply in the form of a written document ... within two weeks of receipt of the request
Re-check detail
Checked and matching: Article 11 right to explanation of any decision taken or supported by an ADM system, in transparent/intelligible manner using clear and plain language; designated competent human contact person with 'the competence, training and authority necessary'; human review with a substantiated written reply within two weeks of receipt of the request; decisions that infringe rights rectified without delay and within two weeks of adoption, or adequate compensation; P2B carve-out: Article 11(5) does not apply to persons who are also business users under Regulation (EU) 2019/1150 || Asserted but not visible in retrievable text: the added detail that 'P2B itself mandates statements of reasons and an internal complaint-handling system' (that is content of Regulation 2019/1150, not this directive) || Re-check notes: All three limbs (a/b/c) plus the rectify-or-compensate-within-two-weeks duty and the business-user carve-out (Art 11(5), Reg 2019/1150) confirmed verbatim from the directive text.
Decision use: holds regardless of model progress. Measured on structural. PWD Art 11 explanation/review rights.
Same underlying source as [CG-07] [CG-10] - repetition across reports is not independent corroboration.
Cited at: Research - findings by decision weight - Decisions - where it runs and which law binds
[CG-09] critic-and-gapfill - grade A (holds regardless of model progress) - re-check 2026-07-15: confirmed with caveats
The directive's scope explicitly covers online-performed work - the 'digital labour platform' definition applies 'irrespective of whether that work is performed online or in a certain location' - recital 19 names tagging/crowdwork, and Chapter III algorithmic-management rules (Arts 7-11) apply to genuinely self-employed persons performing platform work, from the start of recruitment, regardless of where the platform is established.
Source: Directive (EU) 2024/2831 Art 2 & 7 (EUR-Lex) + AI Policy Lab, 'Potential impact of the EU Platform Work Directive on AI labelers' (2025-03-25, academic)
What the source itself says (retrieved quote)
The EU Directive recognizes AI labeling as a form of platform work if it is conducted through a digital platform within the EU
Re-check detail
Checked and matching: the post concludes AI labelers/annotators are covered by the directive when work is conducted through a digital platform in the EU; recital/intro 19 mentions tagging as a form of crowd work that can be conducted remotely (page labels it 'Article 19 of Introduction') || Asserted but not visible in retrievable text: verbatim Art 2 'digital labour platform' definition language ('irrespective of whether that work is performed online or in a certain location') -- not on this page (attributed to EUR-Lex, not fetched); Article 7(2) quote 'shall apply to all persons performing platform work from the start of the recruitment or selection procedure' -- does not appear on this page; Chapter III / Articles 7-11 applying to genuinely self-employed persons from the start of recruitment -- not stated on this page (post discusses Arts 10 and 12 only); Article 16(2) TFEU data-protection legal basis -- not mentioned; extraterritoriality (applies to platforms established outside the EU) -- attributed to Ius Laboris/Omnivoo, not on this page || Re-check notes: Only the AI Policy Lab page is fetchable per constraints. It supports the core conclusion (AI labelers are covered) and the tagging-as-crowdwork-in-recital-19 point. The many verbatim legal specifics in the compound claim (Art 2 definition, Art 7(2), Chapter III scope to self-employed, Art 16(2) TFEU, extraterritoriality) are attributed to EUR-Lex/ETUI/Ius Laboris/Omnivoo and are NOT present on this page. The page also conflates recitals with articles ('Article 19 of Introduction', 'Article 47/44'), weakening any verbatim-quote reliance. Partially_confirmed on the retrieved page alone.
Decision use: holds regardless of model progress. Measured on structural. PWD scope covers online-performed platform work.
Cited at: Decisions - where it runs and which law binds
[CG-10] critic-and-gapfill - grade A (holds regardless of model progress) - re-check 2026-07-15: confirmed
Platform Work Directive Article 7 prohibits automated systems from processing personal data on workers' emotional or psychological state, private conversations (including worker-to-worker exchanges), any data collected while the person is not offering or performing platform work, data predicting exercise of fundamental rights (e.g., organizing), or inferring protected characteristics; Article 8 additionally mandates a data protection impact assessment.
Source: Directive (EU) 2024/2831, Article 7 (EUR-Lex full text) (2024-11-11, primary)
What the source itself says (retrieved quote)
this Article shall also apply where digital labour platforms use automated systems taking or supporting decisions that affect persons performing platform work in any manner ... Article 8 Data-protection impact assessment
Re-check detail
Checked and matching: Article 7(1)(a) prohibits processing personal data on emotional or psychological state; 7(1)(b) private conversations including exchanges with other persons performing platform work; 7(1)(c) collecting data while the person is not offering or performing platform work; 7(1)(d) predicting the exercise of fundamental rights including freedom of association; 7(1)(e) inferring protected characteristics (racial/ethnic origin, religious beliefs, disability, health, trade union membership, sex life/orientation); Article 7(3) extends limits to automated systems 'taking or supporting decisions that affect persons performing platform work in any manner'; Article 8 mandates a data-protection impact assessment (Art 35(1) GDPR high-risk processing) || Re-check notes: Article 7(1)(a)-(e), Article 7(3), and Article 8 all confirmed verbatim. Every element of the claim (including the 'any manner' extension and the DPIA mandate) is present.
Decision use: holds regardless of model progress. Measured on structural. PWD Art 7 telemetry prohibitions + DPIA.
Same underlying source as [CG-07] [CG-08] - repetition across reports is not independent corroboration.
Cited at: Decisions - AI-assistance policy
[CG-11] critic-and-gapfill - grade A (holds regardless of model progress) - re-check 2026-07-15: confirmed with caveats
Under GDPR Article 22 as interpreted in CJEU SCHUFA (C-634/21, 7 Dec 2023), an automated score that plays a 'determining role' in a downstream decision is itself an Article 22 decision, and human involvement only removes a decision from Art 22 scope if it is meaningful - a reviewer with real authority, data access, and competence to override; rubber-stamping does not count. Enforcement is active: Italy's DPA sanctioned automated rider deactivations, and Hamburg's DPA fined a company ~EUR 490,000 in September 2025 for automated rejections without adequate explanation of the logic.
Source: CJEU Case C-634/21 SCHUFA analysis + 'Scores as Decisions? Article 22 GDPR ... in the Labour Context', Industrial Law Journal (2024-09, academic)
What the source itself says (retrieved quote)
the CJEU held a probability value is a decision where a third party 'draws strongly' on it to 'establish, implement or terminate a contractual relationship'
Re-check detail
Checked and matching: CJEU SCHUFA C-634/21, 7 Dec 2023 cited (footnote 1); core doctrine present: a probability value is a decision where a third party 'draws strongly' on it to 'establish, implement or terminate a contractual relationship' (paras 48, 73); general requirement of 'meaningful human involvement' discussed (citing Article 29 Working Party); Art 22(1) 'lays down a prohibition in principle' (para 52); 'lacuna in legal protection' if scores were merely a 'preparatory act' || Asserted but not visible in retrievable text: exact phrase 'determining role' -- the source uses 'draws strongly', not 'determining role' (wording discrepancy); the three-part meaningful-human-involvement test (reviewer with real authority, data access, competence to override) -- not stated in this article as a formulation; 'rubber-stamping' term -- not used; closest material is the Amazon/Hannover Administrative Court ruling; Italy DPA sanction of automated rider deactivations -- not in this article; Hamburg DPA ~EUR 490,000 fine, September 2025, for automated rejections without adequate explanation -- not in this 2024 article (and chronologically outside its scope) || Re-check notes: The underlying legal doctrine (a strongly-relied-upon automated score can itself be an Art 22 decision; human involvement must be meaningful) is supported. But the claim's signature phrase 'determining role' is not in this source (it says 'draws strongly'), the three-part reviewer test and 'rubber-stamping' framing are absent, and BOTH enforcement facts (Italian rider deactivations; Hamburg EUR 490k Sept 2025 fine) do not appear in this 2024 academic article -- the evidence field itself attributes them to separate 2025 enforcement roundups.
Decision use: holds regardless of model progress. Measured on structural. SCHUFA doctrine; active enforcement. (Re-check: doctrine confirmed; the phrase 'determining role' traces to the CJEU judgment rather than the fetched commentary - see ledger entry.)
Cited at: Overview - the five authorizations - System - claim types - Pilot - gates and objectives - Decisions - authority boundaries - Decisions - default: the pay-denial question
[CG-12] critic-and-gapfill - grade A (holds regardless of model progress) - re-check 2026-07-15: confirmed
ISO/IEC 5259 'Data quality for analytics and ML' is now a complete five-part published series - Parts 1-4 published 2024 (overview/terminology; data quality measures; data quality management requirements & guidelines; process framework) and Part 5 (data quality governance framework) published 2025 - sold by ISO as an 'AI data quality management bundle'.
Source: ISO/IEC 5259 series catalog pages (ISO.org) / AI data quality management bundle (2025, primary)
What the source itself says (retrieved quote)
This bundle includes five essential components: ISO/IEC 5259-1:2024 - Artificial intelligence - Data quality for analytics and machine learning (ML) - Part 1: Overview, terminology, and examples
Re-check detail
Checked and matching: 'This bundle includes five essential components'; ISO/IEC 5259-1:2024 - 'Part 1: Overview, terminology, and examples'; 5259-2:2024 - 'data quality model, measures, and guidance on reporting data quality'; 5259-3:2024 - 'management requirements and guidelines'; 5259-4:2024 - 'process framework'; 5259-5:2025 - 'governance framework to enable organizations to direct and oversee data quality'; page title 'AI data quality management bundle' || Asserted but not visible in retrievable text: internal ISO catalog IDs 81088/81860/81092/81093/84150 (from evidence, not shown on this page); Part 4 'covering data labeling' specifically (page shows 'process framework... guidance on data quality processes' but does not surface 'data labeling') || Re-check notes: Retrieved via curl with browser user-agent (WebFetch returned 403). Five-part series with Parts 1-4 (2024) and Part 5 (2025), each part's subject, and the 'AI data quality management bundle' packaging all confirmed. Full standard text is paywalled; the claim's own contract-mandate observation is explicitly hedged.
Decision use: holds regardless of model progress. Measured on structural. ISO/IEC 5259 published series. Standards fact.
Cited at: Decisions - where it runs and which law binds - Decisions - the throughput dial
[CG-13] critic-and-gapfill - grade A (holds regardless of model progress) - re-check 2026-07-15: confirmed with caveats
Facebook's production QA at Cognizant ran a three-layer blind audit: ~50-60 of each moderator's ~1,500 weekly decisions (~3-4%) randomly re-reviewed by a dedicated QA worker (paid $1/hr more), with full-time Facebook employees auditing a subset of QA decisions; 'accuracy' was computed purely as agreement with the auditor against a 95% target (other sites reported 98%), actual scores ran high-80s to 92, and misses triggered a remediation program that often ended in termination.
Source: The Trauma Floor: The secret lives of Facebook moderators in America (The Verge, Casey Newton) (2019-02-25, primary)
What the source itself says (retrieved quote)
"From Miguel's 1,500 or so weekly decisions, Facebook will randomly select 50 or 60 to audit ... a quality assurance worker, known internally as a QA, who also makes $1 per hour more ... Full-time Facebook employees then audit a subset of QA decisions." ... target "95 percent ... floats in the high 80s or low 90s ... around 92".
Re-check detail
Checked and matching: three-layer audit: moderator -> QA re-review -> full-time Facebook employees audit a subset of QA decisions; ~50-60 of ~1,500 weekly decisions randomly audited (~3-4%); QA worker makes $1 per hour more; accuracy = agreement with auditor; target 95 percent; actual scores in high 80s / low 90s, ~92 at press time; misses lead to coaching/remedial program, often a pretext for managing workers out (termination); moderators prohibited from lobbying QAs to reverse decisions but do so regularly || Asserted but not visible in retrievable text: 'other sites reported 98%' (string '98' appears 0 times in the article); the descriptor 'blind' (word 'blind' appears 0 times; audit structure described but not called blind) || Re-check notes: Every material fact confirmed verbatim, including the '~1,500 weekly / 50-60 audited', '$1 per hour more' QA, Facebook-audits-a-subset-of-QA third layer, 95% target, high-80s-to-92 actuals, the 'Accuracy is only judged by agreement / obvious sale of heroin / This number is fake' quotes, off-book QA lobbying, and coaching-as-pretext-for-termination. Marked partially_confirmed only because the peripheral '98% at other sites' comparison is absent (no '98' anywhere) and the audit is not literally called 'blind'.
Decision use: holds regardless of model progress. Measured on adjacent-domain. Facebook/Cognizant audit architecture. Decade-scale deployed practice for humans-reviewing-human-judgment.
Cited at: Decisions - permanent audit
[CG-14] critic-and-gapfill - grade A (holds regardless of model progress) - re-check 2026-07-15: confirmed with caveats
The Trust & Safety Professional Association (industry body staffed by Meta/Google/TikTok alumni) codifies exactly two complementary QA channels - forward audit sampling (weighted re-review samples by peers or a dedicated quality team) and 'reverse quality sampling' (pre-reviewed golden/known-verdict items injected through the regular review process) - with the binding constraint that seeded items 'must look exactly like regular reviews to be effective', plus appeals/overturns as a cheap but population-biased third signal, and a four-way error taxonomy (false positive, false negative, wrong-selection, technical error).
Source: TSPA: Content Moderation Quality Assurance + Metrics for Content Moderation (2022-09-16, practitioner)
What the source itself says (retrieved quote)
Reverse quality sampling involves taking a set of pre-reviewed examples ... Examples must look exactly like regular reviews to be effective
Re-check detail
Checked and matching: reviews by a dedicated quality team (and peer reviews); Samples are often weighted; Reverse quality sampling involves taking a set of pre-reviewed examples ... sending them through the regular review process; Examples must look exactly like regular reviews to be effective; appeals surface false positives in particular at a low cost; the population of appeals is often different from the population of all decisions; four-way error taxonomy: False positives, False negatives, Wrong selection, Technical errors || Asserted but not visible in retrievable text: overturn rate being 'a difficult metric to interpret' -- absent from this QA page (on the separate 'Metrics for Content Moderation' page, which the fetch constraint prohibits); 'consistency' (multi-review mismatch rate) as a distinct tracked metric -- absent from this QA page; prevalence measurement called a 'premium metric' -- absent from this QA page || Re-check notes: The two QA channels, the seeded-item realism requirement, the appeals cost/population-bias points, and the four-way error taxonomy are all confirmed on the fetched QA page. Discrepancy: the claim quotes golden seeding as giving 'full control', but the page actually says 'complete control' (same meaning, different word). The three overturn/consistency/premium-metric items are attributed to a separate 'Metrics for Content Moderation' page (a different URL) and are not present on this page; the single-URL constraint bars fetching it.
Decision use: holds regardless of model progress. Measured on adjacent-domain. TSPA QA curriculum: seeding indistinguishability, appeals bias, wrong-selection errors.
Cited at: Decisions - permanent audit
[CG-15] critic-and-gapfill - grade B (current-generation measurement) - re-check 2026-07-15: confirmed
Pinterest's deployed Decision Quality Evaluation Framework (Feb 2026) scores both human moderators and LLM agents against an SME-adjudicated, versioned, immutable golden set selected by inverse-propensity sampling (XGBoost on PinCLIP embeddings, prioritizing low-propensity/novel items), using a two-axis diagnostic - Cohen's kappa for reliability plus correctness vs the golden set - where 'high reliability + low correctness' is read as systematic policy misunderstanding; measured results show 3x-human majority vote adds only +3.6pp accuracy over a single non-expert human, and current LLMs perform on par with a single non-expert human (GPT-5 the only config with positive accuracy delta, +0.9pp).
Source: Decision Quality Evaluation Framework at Pinterest (arXiv 2602.15809) (2026-02-17, primary)
What the source itself says (retrieved quote)
This shift resulted in over 30x in cost savings and a 10x reduction in labeling turnaround time ... the LLMs demonstrate quality on par with a single non-expert human, but a gap still remains to reach SME-level quality ... High reliability paired with low correctness indicates that labelers are all making the same mistake consistently.
Re-check detail
Checked and matching: 3x-human majority vote adds +3.6pp accuracy over single non-expert human; LLMs demonstrate quality on par with a single non-expert human (gap to SME-level remains); GPT-5 (balanced) +0.9 accuracy delta - only LLM config in the positive; inverse propensity sampling; XGBoost trained on PinCLIP embeddings to predict propensity; Cohen's Kappa for reliability + correctness vs GDS; high reliability + low correctness = systematic policy misunderstanding; shift from 3x-human to LLM gave over 30x cost savings and 10x turnaround reduction; Gemini/flash +22.6 recall (22.5% gain) but +57.2 FPR / +47.7% FPR; GDS published as immutable and versioned dataset; SMEs relabel under new policy to produce 'policy delta'; content-drift monitor + system-stability monitor || Re-check notes: Abstract alone did not carry the numbers; the allowed HTML full text confirms all nine specifics verbatim (Table 1 + prose). Minor nuance: '+0.9pp' GPT-5 is the only *LLM* config with a positive accuracy delta; the 3x-human-majority row is also positive (+3.6) but is a human baseline, so 'only config' is slightly imprecise. Gemini FPR stated as +47.7% in prose and +57.2 in Table 1, consistent with the claimed +48-57pp range.
Decision use: current-generation measurement. Measured on human-work (models: GPT-5, Gemini 2.5). Pinterest decision-quality framework (Feb 2026, GPT-5-class agents in scope): reliability-vs-correctness separation, dual-labeling on policy change. Most current deployed analog.
[CG-16] critic-and-gapfill - grade A (holds regardless of model progress) - re-check 2026-07-15: confirmed
Auditing raters by rewarding agreement with the eventual consensus outcome ('consensus-based auditing', as X's Community Notes has done since Sept 2022 by tying participation eligibility to agreement with the final aggregate) measurably induces strategic conformity - minority contributors' evaluations drift toward the majority and their participation share falls precisely on controversial items - and a two-stage alternative that weights contributors by the stability of their past residuals (predictability relative to a latent-factor model) rather than majority-agreement improves out-of-sample predictive performance.
Source: Auditing the Auditors: Does Community-based Moderation Get It Right? (arXiv 2603.18053) (2026-03-17, academic)
What the source itself says (retrieved quote)
We find evidence of strategic conformity: minority contributors' evaluations drift toward the majority ... weights contributors by the stability of their past residuals ... even when they disagree ... improves out-of-sample predictive performance while avoiding penalization of disagreement
Re-check detail
Checked and matching: 'consensus-based auditing' = rewarding agreement with the final aggregate outcome; in September 2022 a platform (Community Notes) adopted consensus-based auditing tying eligibility for participation to agreement; strategic conformity: minority contributors' evaluations drift toward the majority; their participation share falls on controversial topics; proposed two-stage algorithm weights contributors by the stability of their past residuals (predictability vs a latent-factor model) rather than majority-agreement; consistently-informative contributors get greater influence even when they disagree; improves out-of-sample predictive performance || Asserted but not visible in retrievable text: author h-indices (Borgs h-54, Chayes h-60) - bibliometric metadata not on the page || Re-check notes: Every element of the claim is confirmed in the HTML full text, including the Sept 2022 eligibility-tied consensus auditing, the strategic-conformity/minority-drift finding on controversial items, and the two-stage residual-stability aggregation. Retrieved page is v2 (2026-05-15); the arXiv ID and source_date (2026-03-17) correspond to the v1 submission. The reported two-stage improvement (mean abs residual ~5.73%, median ~27.99% on one-week-ahead predictions) is additional detail beyond the claim.
Decision use: holds regardless of model progress. Measured on human-work. Consensus-rewarded auditing induces conformity (Community Notes). Incentive design; durable.
Cited at: Overview - the problem - Research - findings by decision weight - Decisions - settled constraints
[CG-17] critic-and-gapfill - grade A (holds regardless of model progress) - re-check 2026-07-15: confirmed
Google's audit-quality method (ICLR 2023 Tiny Paper, Google LLC authors) decomposes inter-rater agreement per rubric question rather than per item - in their worked example one question had Fleiss kappa 0.094 against a 0.475 overall, isolating it as the ambiguous criterion - and identifies sequential non-blind review (a later reviewer seeing the earlier verdict) as a distinct, measurable source of audit risk, testable via A/B or difference-in-differences.
Source: Statistical Methods for Auditing the Quality of Manual Content Reviews (arXiv 2306.07466, ICLR 2023 Tiny Papers) (2023-06-12, academic)
What the source itself says (retrieved quote)
Table 1 ... Question 3 0.093997 ... Overall Fleiss Kappa 0.475072 ... synthetic data set consisting of three reviewers and their answers regarding a nine question rubric on 1,528 products ... reviewers conduct reviews sequentially ... A/B test ... Difference-in-Difference
Re-check detail
Checked and matching: authors Xuan Yang, Andrew Smart, Daniel Theron; affiliation Google LLC; Published as a Tiny Paper at ICLR 2023; per-question Fleiss kappa locates the ambiguous criterion; Table 1: Question 3 Fleiss kappa 0.093997 vs Overall Fleiss Kappa 0.475072; Question 3 flagged as systematically higher disagreement; sequential non-blind review (one reviewer sees previous reviewers' results) named as audit risk; testable via A/B test or Difference-in-Difference; synthetic dataset: three reviewers, nine-question rubric, 1,528 products; chi-square for criterion-verdict linkage; t-test/ANOVA for deviating review teams; binomial-distribution confidence interval of the error rate; code at github.com/xuanyang0607/openreviewpaper || Re-check notes: Extracted full PDF text via pdftotext (WebFetch saved the binary; local extraction succeeded). Every element of the claim confirmed verbatim: 0.094 vs 0.475 (0.093997 / 0.475072), Google LLC, ICLR 2023 Tiny Paper, sequential-review/A-B/DiD, dataset dimensions, chi-square/t-test/ANOVA, binomial CI (line 283), and the GitHub URL (line 182).
Decision use: holds regardless of model progress. Measured on human-work. Per-rubric-question agreement decomposition (Google). Method.
[CG-18] critic-and-gapfill - grade A (holds regardless of model progress) - re-check 2026-07-15: confirmed
A 2025 New Media & Society study (screen-share observation of commercial moderators in India) documents that throughput pressure distorts verdict distributions through the action-selection interface itself: moderators systematically chose a 2-click removal over a 4-click de-ranking regardless of policy fit, mechanically deleted tool-highlighted words without context assessment, and privately compressed broad guidelines into simplified dos/don'ts lists - i.e., queue and UI composition changed outcomes independent of moderator judgment quality.
Source: Hard labour conditions of online moderators directly affect how well the internet is policed (The Conversation, summarizing New Media & Society study) (2025-07-22, academic)
What the source itself says (retrieved quote)
removing posts required only two steps ... reducing the visibility of content (de-ranking) involved four steps ... Would never recommend de-ranking content as it would take time.
Re-check detail
Checked and matching: removing posts required only two steps; de-ranking involved four steps; moderator skipped/avoided de-ranking to save time ('Would never recommend de-ranking content as it would take time'); removing tool-flagged words without evaluating context; moderators develop a simplified list of 'dos and don'ts'; commercial content moderators in India; screen-share observation; study published in New Media & Society; author Tania Chatterjee || Asserted but not visible in retrievable text: literal word 'clicks' (article says 'steps': two steps vs four steps); literal word 'throughput' (concept present via targets/time pressure); quantitative error rates (none - study is qualitative, as the claim acknowledges) || Re-check notes: Core mechanism confirmed: the action-selection UI (2-step removal vs 4-step de-ranking) plus throughput pressure shifted outcomes independent of policy fit ('To save time, she skipped the content flagged to be de-ranked'; 'removing flagged words without evaluating the context'; 'dos and don'ts'). Only wording nuance: article says 'steps' not 'clicks'.
Decision use: holds regardless of model progress. Measured on adjacent-domain. Throughput pressure degrades moderator judgment (2025 observational). Human factors; durable.
[CG-19] critic-and-gapfill - grade A (holds regardless of model progress) - re-check 2026-07-15: confirmed with caveats
Feedback intervention theory's core empirical result stands: feedback improves performance on average (K&D 1996: 131 studies, ~12,000+ participants, mean d approximately 0.38-0.41) but MORE THAN ONE THIRD of feedback interventions REDUCE performance, and effectiveness declines as feedback cues move attention from the task toward the self (praise, person-level evaluation, normative comparison).
Source: Feedback Interventions (Kluger & DeNisi 1998), summarizing K&D 1996 Psychological Bulletin meta-analysis (1998-06 (meta-analysis 1996-03; foundational), academic)
What the source itself says (retrieved quote)
Feedback Interventions: Toward the Understanding of a Double-Edged Sword
Re-check detail
Checked and matching: Kluger & DeNisi 1998 article in Current Directions in Psychological Science (Vol 7, Issue 3, pp 67-72, June 1998); cites the 1996 Psychological Bulletin meta-analysis (Kluger & DeNisi 1996, vol 119, pp 254-284) || Asserted but not visible in retrievable text: 131 studies; ~12,000+ participants; mean d approximately 0.38-0.41; phrase 'reduce performance in more than one third of the cases'; attention shifting task->self mechanism (praise, normative comparison) || Re-check notes: DOI 302-redirects (server-side) to journals.sagepub.com, followed per allowed-variant rule. Page is restricted access: only title/authors/metadata and the 1996 citation are visible; no abstract or body text. The article's identity and the 1996 meta-analysis citation are confirmed, but every headline number and the 'more than one third reduce performance' phrase are behind the paywall and not visible in retrieved text.
Decision use: holds regardless of model progress. Measured on human-work. Kluger & DeNisi FIT: a third of feedback interventions backfire; task-referenced helps. Foundational; durable.
Cited at: Decisions - enforcement weight
[CG-20] critic-and-gapfill - grade A (holds regardless of model progress) - re-check 2026-07-15: confirmed
The March 2025 Cochrane update on audit-and-feedback (292 studies, 678 arms, healthcare professionals) finds median absolute improvement in desired practice of only 2.7% (IQR 0.0 to 8.6; weighted meta-analytic mean +6.2%, 95% CI 4.1-8.2, moderate certainty), with effects larger for low baseline performers, individual-level (not team-level) data, comparison to TOP peers or a benchmark (comparison to peer average showed no significant effect), a trusted local source, and action plans with specific advice - while repeated delivery was associated with LOWER effect size.
Source: Audit and feedback: effects on professional practice (Cochrane review update, Ivers et al.) (2025-03-25, academic)
What the source itself says (retrieved quote)
mean absolute increase in desired practice of 6.2% (95% confidence interval (CI) 4.1 to 8.2; moderate-certainty evidence) and an OR of 1.47 (95% CI 1.31 to 1.64; moderate-certainty evidence)
Re-check detail
Checked and matching: 292 studies with 678 arms; median absolute improvement in desired practice of 2.7%, IQR 0.0 to 8.6 (177 studies); mean absolute increase 6.2% (95% CI 4.1 to 8.2; moderate-certainty); OR 1.47 (95% CI 1.31 to 1.64); Lower baseline performance associated with larger intervention effects; individual-recipient-level data rather than team-level data; compares performance to top peers or a benchmark; comparison to average performance of all peers did NOT find significant effects; local champion with existing relationship (trusted local source); actionable plan with specific advice for improvement; Contrary to expectations, repeated delivery was associated with lower effect size; Version published 25 March 2025 || Re-check notes: Every number and every moderator in the claim confirmed verbatim from the full-text abstract and results, including the counterintuitive repeated-delivery finding and the peer-average null. Page required a cookie-jar two-step (WebFetch 403; curl 412) to bypass the cookie gate.
Decision use: holds regardless of model progress. Measured on adjacent-domain. Cochrane 2025 audit-and-feedback update: modest median effects, concentrated in low performers. Durable.
[CG-21] critic-and-gapfill - grade A (holds regardless of model progress) - re-check 2026-07-15: confirmed
In the closest annotation-analog RCT (105 analyzed Mechanical Turk workers writing product reviews), both timely external expert feedback and rubric-based self-assessment significantly improved work quality vs no feedback (expert ratings 6.01 and 6.35 vs 5.69 on a 9-point scale, p<0.05) with NO quality difference between external and self-assessment; external feedback uniquely drove revision behavior (56.5% revised vs 24.8% for self-assessment) and more output, and self-assessors over-rated their own work by 1.8 points (7.9 self vs 6.1 expert).
Source: Shepherding the Crowd Yields Better Work (Dow, Kulkarni, Klemmer, Hartmann; CSCW 2012) (2012-02-11 (foundational; only direct crowdwork feedback RCT found), academic)
What the source itself says (retrieved quote)
The External condition (mu=6.01, SD=1.38) and the Self condition (mu=6.35, SD=1.63) outperformed the None condition (mu=5.69, SD=1.19) (F(2,102)=3.02, p<0.05) ... There was no significant difference between External and Self assessment (t(150)=1.40, p=0.18)
Re-check detail
Checked and matching: 105 participants analyzed (538 consumer reviews from 105 participants); expert 9-point ratings: External mu=6.01, Self mu=6.35, None mu=5.69; F(2,102)=3.02, p<0.05; pairwise None-vs-External and None-vs-Self both p<0.05; NO significant difference between External and Self assessment (t(150)=1.40, p=0.18); revision rates: 56.5% External vs 24.8% Self changed their review; self-assessors over-rated own work by 1.8 points (mu=7.9 self vs mu=6.1 expert); crowd-as-aggregate vs expert agreement Cohen's Kappa=0.20; higher attrition in Self (78%) and External (61%) than None (47%); learning: Self slope 0.25 (p=0.001) significant; External 0.10 (p=0.08) borderline; None null || Re-check notes: Every number in the claim confirmed verbatim from the extracted PDF text (downloaded via curl, converted with pdftotext). Self over-rating '1.8 points higher (mu=7.9 ... mu=6.1)', revision '56.5% ... 24.8%', Kappa=0.20, attrition 78/61/47%, and learning slopes 0.25/0.10 all present exactly as stated.
Decision use: holds regardless of model progress. Measured on human-work. Dow CSCW 2012: task-specific feedback improves paid microtask quality; self-ratings inflate. Closest crowdwork RCT; durable.
[CG-22] critic-and-gapfill - grade A (holds regardless of model progress) - re-check 2026-07-15: confirmed
In the largest educational feedback meta-analysis (435 studies, k=994, N>61,000), information content is the dominant moderator: high-information feedback (task + process + self-regulation content) yields d=0.99 [0.82-1.15] versus d=0.46 for corrective right/wrong feedback and d=0.24 for bare reinforcement/punishment - with overall d=0.48 masking huge heterogeneity (I2=83%) and 17% of raw effects negative.
Source: The Power of Feedback Revisited: A Meta-Analysis of Educational Feedback Research (2020-01-22, academic)
What the source itself says (retrieved quote)
feedback cannot be understood as a single consistent form of treatment
Re-check detail
Checked and matching: 435 studies, k = 994 effect sizes, N > 61,000; high-information feedback d = 0.99 [0.82-1.15] (k=42); corrective right/wrong feedback d = 0.46 [0.39-0.55] (k=238); reinforcement/punishment d = 0.24 [0.06-0.43] (k=39); overall weighted d = 0.48 (outlier-cleaned; CL 0.44-0.51), I2 = 83.40%; 17% of raw effects negative (full set before outlier removal); feedback-type moderator QB = 41.52, df=2, p < 0.0001; cognitive outcomes d = 0.51 vs motivational d = 0.33; 86% of negative motivational effects from uninformative (reward/punishment) feedback; 'feedback cannot be understood as a single consistent form of treatment' (verbatim) || Re-check notes: Every headline number matches. Minor caveat surfaced in text: d=0.48 is the outlier-adjusted estimate (initial integration d=0.55 with 17% negative); the claim uses the 0.48 cleaned figure and the 17%-negative full-set figure, both consistent with the paper.
Decision use: holds regardless of model progress. Measured on adjacent-domain. Information content dominates feedback efficacy (d~0.99 high-information vs 0.46 bare verdicts). Durable.
[CG-23] critic-and-gapfill - grade A (holds regardless of model progress) - re-check 2026-07-15: source not retrievable
The 2025 Annual Review of Organizational Psychology's 25-year retrospective concludes the science of workplace feedback 'is not yet a story of coherent and cumulative progress': definitions are generic, assumptions diverge across six disconnected research substreams, and simple universal rules about feedback effectiveness do not survive contact with organizational reality.
Source: A 25-Year Review of Research on Feedback in Organizations: From Simple Rules to Complex Realities (Anseel & Sherf) (2025-01-21, academic)
Re-check detail
Asserted but not visible in retrievable text: 'is not yet a story of coherent and cumulative progress'; insights 'often appear disconnected from the way feedback is practiced and experienced in organizations'; diverging assumptions across six research substreams; citation Annu. Rev. Organ. Psychol. Organ. Behav. 12:19-43, published 2025-01-21 || Re-check notes: Page is protected by a Cloudflare managed challenge that requires JavaScript execution. WebFetch returned 403; curl with a browser user-agent received only the 'Just a moment...' challenge HTML; the Playwright browser loaded the exact same authorized URL but the managed challenge never cleared (HTTP 403) within ~14s. No article text retrieved; no substitute source permitted, so the claim cannot be verified.
Decision use: holds regardless of model progress. Measured on adjacent-domain. 25-year workplace-feedback retrospective: field immature; effects heterogeneous. Calibrates expectations.
[CG-24] critic-and-gapfill - grade A (holds regardless of model progress) - re-check 2026-07-15: source not retrievable
Under incentives, relative-rank feedback is a double-edged lever: lab and field economics find rank feedback raises output in flat-wage settings (Charness et al. 2014; Tafkov 2013) but induces costly sabotage and cheating to improve rank that offsets the gains, and in at least one field experiment (Barankay 2012) REMOVING rank feedback improved performance.
Source: Feedback quality and performance in organisations (Leadership Quarterly; synthesizing Charness et al. 2014, Tafkov 2013, Barankay 2012) (2021 (synthesizing 2012-2016 primaries), academic)
Re-check detail
Asserted but not visible in retrievable text: feedback quality and performance in organisations; relative rank feedback raising output (Charness et al. 2014; Tafkov 2013); sabotage/cheating offsetting positive effects; Barankay 2012 - removing rank feedback improving performance || Re-check notes: WebFetch returned HTTP 403 Forbidden. curl -sL with a browser user-agent returned a Cloudflare bot-challenge shell (title 'ScienceDirect'; contains __cf_chl_tk tokens, 'Captcha', 'Enable JavaScript', meta http-equiv refresh 360, noscript) with zero article content - no abstract, authors, journal, or any mention of rank feedback / Charness / Barankay / Tafkov / Leadership Quarterly. Claim could not be verified against the source per the WebFetch->curl->not_retrievable fallback chain.
Decision use: holds regardless of model progress. Measured on adjacent-domain. Rank feedback under incentives induces gaming/sabotage. Benchmark against exemplar work, never ranked people.
Absence claims here are bounded to what was actually reviewed - "not identified," never "proven absent."
These seven are not weaknesses of the proposal; they are its measurement plan, and each is assigned to a
specific experiment in the decision register.
Proposed anchors not admitted (source unreachable under our retrieval rules, so they stay quarantined rather than quietly included): ANSI/ASQ Z1.4 - Sampling Procedures and Tables for Inspection by Attributes (not_retrievable).
The foundation package validates against its own contract: all eight checks pass (schemas, cross-document
references, canonical hashes, positive-proof invariants), run on this site's copy on 2026-07-15. One defect in
the originally shipped site was found and fixed: its file manifest recorded stale hashes for three files and
omitted the fonts; the manifest now covers every shipped file. The machine-readable forms of everything here:
claims.json, decision-record.json,
site-manifest.json. The reasoning layer is KERNEL.md.