Domain: contrarian
Raw percent-agreement systematically overstates LLM-judge ability: in the largest judge meta-evaluation to date (21 judges, 9 providers, ~541,000 judgments, including April-2026 frontier models), Cohen's kappa runs 33-41 percentage points below exact-match agreement on MT-Bench, judge rankings shift by up to 14 positions across benchmarks, and two production-deployed judges combine test-retest reliability >0.95 with severe position bias >0.10 (a 'consistency-bias paradox'). Load-bearing
EVIDENCE: Norman, Rivera & Hughes ran 118 runs across MT-Bench, JudgeBench, RewardBench under three protocols (agreement, consistency, bias audit); kappa deflation was universal across the cohort; they distill a 'Minimum Viable Validation Protocol'. Verified from the arXiv abstract page. Corroborated by the NeurIPS 2025 position paper 'Neither Valid nor Reliable? Investigating the Use of LLMs as Judges' (arXiv 2508.18076, Chehbouni et al.), which argues via social-science measurement theory that LLJ adoption has outpaced scrutiny of validity/reliability and that human-agreement proxying is an unvalidated assumption. SOURCE: Reliability without Validity: A Systematic, Large-Scale Evaluation of LLM-as-a-Judge Models Across Agreement, Consistency, and Bias | https://arxiv.org/abs/2606.19544 | 2026-06-17 | academic IMPLICATION: Directly answers Q1's metric-standardization question: ban raw percent-agreement and single 'accuracy vs humans' numbers from all AutoQA meta-evaluation; require chance-corrected statistics per benchmark. And for Q5: run-to-run consistency must never be reported as evidence of validity - a judge can be perfectly repeatable and severely biased simultaneously, so a separate perturbation/bias audit is a mandatory ship gate independent of the consistency gate.
Across 106 experiments and 370 effect sizes, human-AI combinations performed significantly worse than the best of human or AI alone (Hedges g = -0.23, 95% CI -0.39 to -0.07), with losses concentrated in decision-making tasks and specifically when the AI outperforms the human alone - and a 2025 AIES study found human reviewers followed severely race-biased AI hiring recommendations ~90% of the time. Load-bearing
EVIDENCE: Vaccaro, Almaatouq & Malone (Nature Human Behaviour, preregistered meta-analysis; direction, decision-task losses, and AI-better-than-human losses verified from the arXiv abstract; g value corroborated across MIT Sloan, Nature's own summary, and ResearchGate). UW's 'No Thoughts Just AI' (AIES 2025, DOI 10.1609/aies.v8i3.36749, 528 participants, 16 job types) found participants matched biased AI picks even at moderate bias and ~90% at severe bias, vs equal selection rates with no/neutral AI. The 2026 practitioner literature (e.g., tianpan.co HITL rubber-stamp essay, 2026-04-15) reports the same mechanism in production review pipelines. SOURCE: When combinations of humans and AI are useful: A systematic review and meta-analysis | https://www.nature.com/articles/s41562-024-02024-1 | 2024-10-28 | academic IMPLICATION: Q6's fork is real and the evidence leans against the default: 'one shallow human touch verifying the AI verdict' is precisely the decision-task, AI-better-than-human configuration where the meta-analysis finds negative synergy and rubber-stamping. The foundation should treat zero-per-item-touch plus randomized BLIND deep audits (human judges never see the AI verdict before committing their own) as the null design to beat, and any verify-the-AI-verdict interaction must be measured against its own overturn rate to prove the human is adding information.
There is a proven theoretical ceiling on judge-based validation: when the judge is no more accurate than the model/content being evaluated, no debiasing method using gold labels can cut the required amount of ground-truth data by more than a factor of two, and empirical savings are smaller than the 2x bound. Load-bearing
EVIDENCE: Dorner, Nastl & Hardt (ICLR 2025) prove the bound for methods that combine cheap judge scores with a small gold-label set to correct judge biases such as self-preference; verified from the arXiv abstract page including the central quote 'when the judge is no more accurate than the evaluated model, no debiasing method can decrease the required amount of ground truth' by more than half. SOURCE: Limits to scalable evaluation at the frontier: LLM as Judge won't beat twice the data | https://arxiv.org/abs/2410.13341 | 2024-10-17 | academic IMPLICATION: Q1's gold-set arithmetic cannot be escaped by judge cleverness: on exactly the axes where the expertise-gap root cause bites (attempter competence near or above judge competence), certifying the AutoQA still costs at least half the gold labels a judge-free validation would - so hierarchical/pooled validation across projects and a permanent (not bootstrap) expert-audit channel are structural requirements, and any vendor claim that the judge 'validates itself' at scale should be treated as mathematically impossible in the regime that matters.
Human evaluation criteria are output-dependent and unstable: even when graders define criteria before grading, the act of grading changes their criteria and they retroactively revise earlier grades ('criteria drift'), implying evaluation criteria for LLM-output quality cannot be fully determined prior to observing outputs. Load-bearing
EVIDENCE: Shankar, Zamfirescu-Pereira, Hartmann, Parameswaran & Arawjo (UIST 2024, EvalGen system study); abstract verified via arXiv crawl: 'we identify a phenomenon we dub criteria drift: users need criteria to grade outputs, but grading outputs helps users define criteria'; the authors argue there is reason to believe criteria never fully settle because they adapt to the observed output distribution. SOURCE: Who Validates the Validators? Aligning LLM-Assisted Evaluation of LLM Outputs with Human Preferences | https://arxiv.org/abs/2404.12272 | 2024-04-18 | academic IMPLICATION: Falsifies Q4's implicit one-shot compilation contract: 'instructions compile once into artifacts, fresh judge hits target agreement with no conversation' is unstable because the project owner's own criteria will drift once they see real attempter submissions and judge verdicts. Rubric compilation must be designed as a versioned re-compilation loop with an explicit retroactive re-scoring policy, and owner sign-off must happen on graded real items, not on abstract criteria.
LLM judges favor models whose mistakes resemble their own (a generalization of self-preference, measured by chance-adjusted mistake-overlap CAPA), and model errors across the industry are becoming MORE correlated as capabilities improve - undermining the assumption that AI oversight of AI-assisted work catches failures.
EVIDENCE: Goel et al., ICML 2025; abstract verified via arXiv crawl: 'LLM-as-a-judge scores favor models similar to the judge' and 'model mistakes are becoming more similar with increasing capabilities, pointing to risks from correlated failures'. Related: 'Preference Leakage' (arXiv 2502.01534) documents contamination when generator and judge are related. SOURCE: Great Models Think Alike and this Undermines AI Oversight | https://arxiv.org/abs/2502.04313 | 2025-02-06 | academic IMPLICATION: Q5: judge model family must differ from the generator being critiqued AND from whatever assistant attempters plausibly used - per-project judge routing plus similarity reporting. Q1: the false-agreement rate (judge passes an LLM-assisted attempter because both share priors) is predicted to RISE over time, which converts the independent human expert-audit channel from a bootstrap phase into a permanent structural component.
LLM judges do not follow their own rubrics: on Arena-Hard Auto, the explicit evaluation schema explains under 10% of verdict variance for some judges (unexplained variance >90% for DeepSeek-R1-32B), and factor correlations above 0.93 across nominally distinct criteria show per-axis scores collapse into a single halo factor.
EVIDENCE: Feuer et al. introduce 'schematic adherence' (how much of the verdict the stated rubric explains) and psychometric validity checks; abstract verified via arXiv crawl: 'severe schema incoherence and factor collapse across popular judges'; ELO-style aggregation additionally masks genuine ranking uncertainty. Code at github.com/penfever/judgment-to-noise. SOURCE: When Judgment Becomes Noise: How Design Failures in LLM Judge Benchmarks Silently Undermine Validity | https://arxiv.org/abs/2509.20293 | 2025-09-24 | contrarian IMPLICATION: Q3/Q4: a per-axis verdict sheet can be theater - the judge may emit one holistic impression dressed up as N criterion scores. The foundation should require discriminant-validity and schematic-adherence tests per project config (do axis scores actually vary independently? does the rationale's schema explain the verdict?) before per-criterion outputs are exposed to attempters or used for specialist routing.
LLM judges have low intra-rater reliability - identical items re-scored across runs with identical settings produce inconsistent, 'almost arbitrary in the worst case' ratings - and the obvious fix (temperature-0 determinism) measurably REDUCES agreement with human judgment; meanwhile the human baseline itself is weak (SummEval inter-annotator kappa: 0.492 crowd, 0.413 expert first round, 0.71 only after a second adjudication round).
EVIDENCE: Haldar & Hockenmaier, Findings of EMNLP 2025 (pp. 24986-25004); verified by reading the paper PDF (pp. 1-3): contributions are (1) low agreement of LLM ratings across runs, (2) disabling sampling hurts human-agreement, (3) phenomenon persists across SummaC, SummEval, MT-Bench; the SummEval human kappa figures are quoted in their Section 3.1. SOURCE: Rating Roulette: Self-Inconsistency in LLM-As-A-Judge Frameworks | https://aclanthology.org/2025.findings-emnlp.1361.pdf | 2025-11-04 | academic IMPLICATION: Q5: verdict flip rate is intrinsic and cannot be silenced with greedy decoding without paying validity - so the build must budget k-sample voting AND treat residual flip clusters as underspecification signal routed to rubric revision. Q1: raw human labels sit below conventional reliability thresholds; validation targets must be adjudicated (multi-round) labels, which prices the gold set honestly.
LLM judges give the weakest signal exactly where QA needs them most: they cannot reliably grade responses to questions they cannot answer themselves (poor signal on the hardest items in a benchmark), and in expert domains subject-matter experts agreed with LLM-judge picks only 64% (mental health) to 68% (dietetics) of the time, with the judge favoring superficially actionable detail.
EVIDENCE: 'No Free Labels' (Kim et al., arXiv 2503.05061, Mar 2025) shows judge quality collapses on the most difficult items and that human-written reference answers improve agreement and reduce self-preference; Szymanski et al. (ACM IUI 2025, DOI 10.1145/3708359.3712091) measured the 64/68% SME agreement in pairwise comparisons on domain tasks. Both reported consistently across two independent search engines; abstracts inspected. SOURCE: No Free Labels: Limitations of LLM-as-a-Judge Without Human Grounding | https://arxiv.org/abs/2503.05061 | 2025-03-07 | academic IMPLICATION: Q5/Q2: the expertise-gap root cause is NOT automatically closed by an LLM judge - the judge's competence ceiling binds hardest on exactly the expert items where human reviewers also fail, so per-axis judge-vs-adjudicated-expert ceilings must be measured before assigning judge-autonomous lanes, and expert-authored reference answers/exemplars are the highest-leverage compile-time artifact (they measurably raise judge validity).
Judge verdicts are gameable through content-independent artifacts: short universal adversarial phrases learned on a surrogate model transfer to unseen judge LLMs and inflate scores toward the maximum regardless of the assessed text, with absolute scoring far more vulnerable than comparative assessment.
EVIDENCE: Raina, Liusie & Gales, EMNLP 2024 (main, pp. 6920+); abstract verified via OpenAlex/ACL record: attackers need no access to the deployed judge - surrogate attack then transfer; 'irrespective of the assessed text, maximum scores are predicted'. Complementary: 'Cheating Automatic LLM Benchmarks' (arXiv 2410.07137) showed null models emitting constant responses achieve top win rates on judged benchmarks. SOURCE: Is LLM-as-a-Judge Robust? Investigating Universal Adversarial Attacks on Zero-shot LLM Assessment | https://aclanthology.org/2024.emnlp-main.427/ | 2024-11-12 | academic IMPLICATION: Q8: pay-motivated attempters do not need to see the judge's rationale to Goodhart it - transfer attacks work blind, so filtering rationales is insufficient as the sole defense. Day-one requirements: comparative/anchored scoring formats over absolute Likert, rotating seeded probes with known verdicts, and evidence-perturbation tests (swap the cited span, verdict must flip) as a standing citation-theater detector.
Industrial-scale annotation QA has already lost to spam economics once: a large data vendor's 'Bulba Experts' program for Google (Mar 2023-Apr 2024) was flooded with contributors submitting 'writing gibberish, writing incorrect information, GPT-generated thought processes' who were paid anyway because per-item catching was infeasible and banned workers returned via VPN - while academic estimates put LLM use at 33-46% of crowdworkers on text tasks.
EVIDENCE: Futurism (2025-06-28) reporting Inc. Magazine's review of internal a large data vendor documents and former queue-manager interviews; article text read in full via crawl ('spammers could get away with just totally submitting garbage and there weren't enough people to track them down'; 'no background checks whatsoever'). Prevalence: Veselovsky et al. arXiv 2306.07899; NeurIPS 2025 'Evaluating LLM-Contaminated Crowdsourcing Data Without Ground Truth' proposes peer-prediction mechanisms precisely because text-detector approaches fail on structured annotations. SOURCE: The AI Company Zuckerberg Just Poured $14 Billion Into Is Reportedly a Clown Show of Ludicrous Incompetence (Inc. internal-document reporting) | https://futurism.com/scale-ai-zuckerberg-incompetence | 2025-06-28 | practitioner IMPLICATION: Q8: the strongest available field postmortem says per-item content classification is the losing side of the cost asymmetry - the foundation should weight identity/provenance gating, rate limits, work-history consistency, peer-prediction-style information scoring, and randomized deep audits over per-item spam detection, and should assume some fraction of 'human' annotations are LLM-assisted (making provenance policy per-project, with process telemetry as an explicit data-collection contract).
Annotator disagreement contains recoverable systematic signal, not just noise: NUTMEG, a Bayesian model separating annotator-competence noise from subpopulation-level systematic disagreement, produces downstream models that significantly outperform both majority-vote aggregation and fully disaggregated training - meaning aggregation to a single 'true' label destroys measurable information.
EVIDENCE: Ivey, Gauch & Jurgens, EMNLP 2025 main (pp. 2874-2887); PDF and poster abstract inspected: NUTMEG estimates true labels per subpopulation and beats MACE/majority-vote at replicating subgroup label distributions on politeness and offensiveness. Sits atop the perspectivist NLP literature (NLPerspectives workshops 2022-2025) which rejects single-gold-label resolution for subjective tasks. SOURCE: NUTMEG: Separating Signal From Noise in Annotator Disagreement | https://aclanthology.org/2025.emnlp-main.144.pdf | 2025-11-04 | academic IMPLICATION: Q2: inter-reviewer disagreement is partly legitimate signal, so an AutoQA that maximizes consistency-of-application on judgment-laden axes launders one interpretation into false objectivity and destroys exactly the pluralism the training data may need. The verdict ontology needs a measured 'systematic-disagreement' class (distinct from noise), triggered by subpopulation-conditioned disagreement, exempt from attempter penalty, and fed back as an instruction-gap report.
AI assistance homogenizes human output at the population level even while raising individual quality: in a Nature Human Behaviour brainstorming study 94% of ChatGPT-assisted participants' ideas shared overlapping concepts (nine independently produced the same product name) while human-only ideas were entirely unique, and the 'Artificial Hivemind' study found different vendors' frontier models converge on near-identical phrasings (~81% average similarity between DeepSeek-V3 and GPT-4o).
EVIDENCE: Meincke, Nave & Terwiesch (Nature Human Behaviour, 2025; Wharton Mack Institute summary opened and read); Jiang et al. (UW/CMU/AI2 'Artificial Hivemind', reported by The Decoder) documents intra-model repetition plus inter-model homogeneity. Consistent with Doshi & Hauser (Science Advances 2024) and Wan et al. 2026 (diverse AI personas partially mitigate homogenization). All are adjacent-domain (ideation/writing), not annotation-QA-feedback studies. SOURCE: New in Nature: ChatGPT Decreases Idea Diversity in Brainstorming (Meincke, Nave & Terwiesch, Nature Human Behaviour) | https://mackinstitute.wharton.upenn.edu/2025/new-in-nature-chatgpt-decreases-idea-diversity-in-brainstorming/ | 2025-05 | academic IMPLICATION: Q8's monoculture worry has strong analogical support: a single judge family delivering item-specific stylistic feedback to thousands of attempters is a homogenization pump aimed at the exact diversity the training data exists to capture. The foundation should ban stylistic guidance in feedback by default, keep feedback at principle level, and stand up a population-level output-diversity drift metric from day one - while flagging that direct evidence in the annotation-QA setting does not yet exist.
Tools & artifacts
- judgment-to-noise: Code and dataset for schematic-adherence and psychometric-validity diagnostics of LLM judges (Feuer et al. 2025) | https://github.com/penfever/judgment-to-noise | Directly reusable as a per-project ship gate: measures whether a judge's rubric actually explains its verdicts and whether per-axis scores have discriminant validity, before exposing per-criterion output.
- CAPA (Chance Adjusted Probabilistic Agreement): Metric for judge-model similarity based on chance-corrected overlap in mistakes, with website and code (Goel et al., ICML 2025) | https://model-similarity.github.io/ | Operationalizes per-project judge routing: quantifies how similar a candidate judge is to the generator/assistant models attempters likely used, bounding correlated-failure risk.
- LLM_contamination (peer-prediction detection): Code for NeurIPS 2025 'Evaluating LLM-Contaminated Crowdsourcing Data Without Ground Truth' - peer-prediction mechanisms that score information content of annotations without gold labels | https://github.com/yichiz97/LLM_contamination | An economics/information-based alternative to per-item LLM-text detection for the spam-filtering layer, designed for structured annotation tasks where text detectors fail.
- EvalGen: Mixed-initiative interface for aligning LLM-generated evaluators with human grades (Shankar et al., UIST 2024); the study that identified criteria drift | https://arxiv.org/abs/2404.12272 | Reference design for the rubric-compilation workshop: interleaves owner grading with criteria refinement rather than assuming criteria can be authored up front.
- Minimum Viable Validation Protocol: Reporting protocol distilled from the 541k-judgment meta-evaluation (Norman et al. 2026): chance-corrected agreement + consistency + bias audit as the minimum bar for claiming a judge works | https://arxiv.org/abs/2606.19544 | Candidate template for the AutoQA's own meta-evaluation reporting standard (Q1), replacing single-number accuracy claims.
- NUTMEG: Bayesian aggregation model separating annotator-competence noise from systematic subpopulation disagreement (EMNLP 2025) | https://aclanthology.org/2025.emnlp-main.144.pdf | Machinery for the variance-attribution pilot (Q1) and for the 'systematic-disagreement' verdict class (Q2): distinguishes items where reviewers disagree from noise vs. from legitimate perspective differences.
Disagreements
- Can LLM judges match human evaluators at all? Counter-evidence exists: some studies report GPT-4-class judges agreeing with human CONSENSUS at rates exceeding individual human raters (Chatbot Arena-style preference tasks), and one 2026 study reports inter-judge Krippendorff's alpha of 0.867 above the 0.80 threshold. The contrarian corpus (kappa deflation, 64-68% SME agreement in expert domains, schema incoherence) does not refute this; the reconciliation is regime-dependence - judges can beat the mean crowd rater on generic preference tasks while falling below adjudicated expert panels on expertise-heavy and subjective axes. Design consequence: per-axis ceilings, not a global verdict on judge viability.
- Human-in-the-loop value: Vaccaro et al. find human-AI combinations LOSE on decision tasks when AI outperforms the human, but the same meta-analysis finds GAINS when the human outperforms the AI - so the one-touch design is not universally bad; it is bad specifically where the judge is stronger than the reviewer, and potentially positive on expert axes where humans still beat judges (per the IUI 2025 expert-domain results). The two contrarian findings point in opposite directions depending on axis type.
- Claim decomposition: the decompose-then-verify literature is internally split. Proponents (FactScore lineage, LREC 2026 Japanese decomposition dataset) find atomic claims improve explainability and reduce annotator variability; critics ('Does Claim Decomposition Boost or Burden Fact-checking Performance?' OpenReview; Wanner et al. 2024; the ACL 2025 dynamic-decomposition paper's own literature review) find decomposition does not consistently improve verification across input lengths and verifier strengths, and that atomicity choice itself changes verdicts (accuracy swings of ~0.12 from decomposition policy alone). Neither side has settled the granularity question the AutoQA design poses.
- Determinism as a fix for judge inconsistency: common practitioner guidance says run judges at temperature 0 for reproducibility; Rating Roulette (EMNLP 2025) shows disabling sampling measurably REDUCES agreement with human judgment - reproducibility and validity trade off rather than align.
- Bias against AI-generated content: earlier work suggested LLM judges penalize (or favor) AI-generated text; Balog, Metzler & Qin (SIGIR 2025) report finding no evidence of bias against AI-generated content in their preliminary study, while confirming a significant bias TOWARD LLM-based rankers' outputs. The direction and existence of AI-content bias is unsettled; the similarity-based affinity bias (Great Models Think Alike) is the better-supported formulation.
Gaps
- No published longitudinal study of judge-feedback-induced homogenization of a paid annotator population (the monoculture finding rests on analogical evidence from ideation/creative-writing settings, not annotation QA feedback loops).
- No published measurement of the false-agreement rate (AutoQA judge and attempter agreeing while both wrong per expert ground truth) in a production annotation-QA setting; Great Models Think Alike predicts it rises with model similarity but does not measure it in this configuration.
- Found no practitioner postmortem of an LLM-based AutoQA layer for human annotation being Goodharted by attempters over feedback cycles - the a large data vendor postmortem is from the pre-LLM-judge QA era, and adversarial-judge evidence comes from benchmark/exam settings; the specific 'scores rise on AutoQA, flat on held-out human-graded goldens' divergence experiment appears not to have been run or published anywhere.
- Little direct evidence on positive-claim (praise) verification: no seeded-set study measuring whether judges pass plausible-but-ungrounded praise at higher rates than they fail grounded praise, per praise type - the leniency/sycophancy literature is about response evaluation generally, not verification of an evaluator's positive statements.
- Rating Roulette's per-benchmark intra-rater ICC/Krippendorff numbers were not extracted (only pages 1-3 read: abstract-level claims plus the SummEval human-kappa figures verified); the synthesis engine should not cite specific LLM intra-rater coefficients from this report.
- The 'Reliability without Validity' (2026) preprint is 0-citation and un-peer-reviewed as of 2026-07-14; its headline numbers are consistent with the older peer-reviewed literature (kappa deflation, position bias) but the specific 33-41pp figure rests on one team's protocol.
Verifications
- CLAIM: Raw percent-agreement systematically overstates LLM-judge ability: in the largest judge meta-evaluation to date (21 judges, 9 providers, ~541,000 judgments, including April-2026 frontier models), Cohen's kappa runs 33-41 percentage points below exact-match agreement on MT-Bench, judge rankings shift by up to 14 positions across benchmarks, and two production-deployed judges combine test-retest reliability >0.95 with severe position bias >0.10 (a 'consistency-bias paradox'). VERDICT: confirmed | Verified directly against the arXiv abstract page (https://arxiv.org/abs/2606.19544, v1 submitted 2026-06-17, only version). Every quantitative element matches: 21 judge models from 9 providers, 118 runs, ~541,000 judgments across MT-Bench/JudgeBench/RewardBench under three protocols (agreement, consistency, bias audit); exact-match vs Cohen's kappa gap of 33-41 pp on MT-Bench; judge rankings shifting up to 14 positions across benchmarks; two production-deployed judges with test-retest reliability >0.95 and position bias >0.10, framed as a 'consistency-bias paradox'; findings 'consistent across the full cohort, including the April 2026 frontier'; a 'Minimum Viable Validation Protocol' is distilled. Authors are Justin D. Norman, Michael U. Rivera, D. Alex Hughes as claimed. The corroborating paper (arXiv 2508.18076, Chehbouni et al.) is confirmed as a NeurIPS 2025 poster (https://neurips.cc/virtual/2025/poster/121914) making exactly the measurement-theory validity/reliability argument described. Supersession check (Exa, freshness=month): no later paper superseding or contradicting it found as of 2026-07-14; only secondary coverage (e.g., DEV Community 2026-07-01) restating its findings, and one adjacent study (arXiv 2606.13685) on run-to-run instability that complements rather than contradicts. Two caveats, neither rising to a correction: (a) 'largest judge meta-evaluation to date' is the authors' own self-characterization ('largest systematic evaluation of the paradigm so far'), not independently established; (b) verification is abstract-level - I did not audit the paper body. Note: two WebSearch calls were blocked by stochastic model safeguards; Exa search substituted successfully.
- CLAIM: Across 106 experiments and 370 effect sizes, human-AI combinations performed significantly worse than the best of human or AI alone (Hedges g = -0.23, 95% CI -0.39 to -0.07), with losses concentrated in decision-making tasks and specifically when the AI outperforms the human alone - and a 2025 AIES study found human reviewers followed severely race-biased AI hiring recommendations ~90% of the time. VERDICT: confirmed | All load-bearing facts verified against primary sources. (1) Meta-analysis: full text of Vaccaro, Almaatouq & Malone (Nature Human Behaviour, published 2024-10-28; verified via arXiv:2405.06087v2 accepted version) states verbatim "370 unique effect sizes from 106 different experiments" and "Hedges' g = -0.23, 95% confidence interval -0.39 to -0.07" for human-AI combos vs best of human or AI alone; decision-task losses and losses-when-AI-outperforms-human both confirmed. Date correct. (2) AIES 2025: "No Thoughts Just AI" (Wilson, Sim, Gueorguieva & Caliskan, UW), AIES Proceedings 8(3):2692-2704, article 36749 (DOI 10.1609/aies.v8i3.36749 matches), N=528, 16 occupations; equal selection with no/neutral AI, and participants favored AI-preferred candidates "up to 90% of the time" with severely biased AI - claim's "~90%" is a fair paraphrase (study says "up to 90%", with a simulated LLM, not a deployed system). (3) Supersession: only later work found is a 2026 Open MIND/NHB-collaboration reproduction of the meta-analysis by Brodeur et al. (osf.io/u6gea) - a reproduction effort, not a refutation; no retraction, correction, or contradicting update located. The tianpan.co practitioner-essay corroboration was not independently checked but is non-load-bearing.
- CLAIM: There is a proven theoretical ceiling on judge-based validation: when the judge is no more accurate than the model/content being evaluated, no debiasing method using gold labels can cut the required amount of ground-truth data by more than a factor of two, and empirical savings are smaller than the 2x bound. VERDICT: confirmed | Source verified: Dorner, Nastl & Hardt, "Limits to scalable evaluation at the frontier: LLM as Judge won't beat twice the data," arXiv:2410.13341, v1 submitted 2024-10-17 (date matches), accepted as an ICLR 2025 Oral (venue matches, slightly stronger than claimed). The abstract states verbatim that "when the judge is no more accurate than the evaluated model, no debiasing method can decrease the required amount of ground truth labels by more than half," and that empirical sample-size savings are "even more modest" than the 2x bound - matching all three parts of the claim (theoretical 2x ceiling, the accuracy condition, and smaller empirical savings). Search for 2025-2026 follow-up work found no refutation or supersession; the authors' later papers (e.g., ROC-n-reroll, ICLR 2026) cover different topics. Minor scoping caveat only: the bound is proven within the paper's statistical framework for debiasing methods combining judge scores with gold labels, and it does not apply when the judge IS more accurate than the evaluated model - the claim already states this condition correctly.