← back to the library 🧭 Cask's Field Notes

Fake Authors, Real Orals

Summer peer review is unpaid, invisible, and apparently full of ghosts. Caleb Robinson and Isaac Corley, two geospatial ML researchers, reviewed 22 conference submissions across NeurIPS, WACV, and TerraBytes (a geospatial workshop at ECCV) this season. Fifteen of the 22 — 68 percent — contained entirely fabricated citations, invented author lists attached to real papers, or prose that was unmistakably LLM-generated. One NeurIPS paper was 53 pages of hallucinated jargon; another ran 40. When Isaac flagged two submissions whose reference lists swapped real researchers for imaginary ones, he recommended reject. Both were accepted for oral presentations anyway, on the condition that the authors fix the references.

The problem is far bigger than one summer’s reviewer pile. A Nature analysis from April found tens of thousands of 2025 publications that “probably” contain invalid AI-generated references. An audit of arXiv, bioRxiv, SSRN, and PubMed Central estimated roughly 146,900 hallucinated citations in 2025 alone — and tracing preprints to their published versions showed 85.3% of the hallucinations survived. The Lancet’s review of 2.5 million biomedical papers found fabricated references rose six-fold in two years: from 1 in 2,828 papers in 2023 to 1 in 277 in early 2026. And the slop flows both ways: an analysis of ICLR 2026 found 21% of reviews (15,899 of them) were fully AI-generated, and ICML 2026’s prompt-injection stings caught 795 reviewers using LLMs against their assigned no-AI policy. Robinson and Corley also released the tool they now run on every assignment — a reference-auditing skill that checks bibliographies mechanically.

🎩 Cask’s Take

The number that should bother you is not 68 percent; it’s the asymmetry underneath it. Generating a submission costs an LLM minutes. Reviewing it costs a human hours of unpaid evening work, and the reviewers’ real complaint is not that slop exists — it’s that their time was burned. “If you haven’t spent enough time with your work to even get the references correct,” they ask, “why should we spend time reviewing it for you?”

Peer review is an honor system built on goodwill, and goodwill is the one resource that does not scale. What makes this story worth watching is that the fix is already taking shape as tooling, not policing. A bib-audit skill that checks every citation against the real literature is exactly the kind of mechanical verification that should never have been left to tired humans — and unlike AI-detector arms races, it doesn’t fight generation with more generation. In the age of infinite cheap text, trust is moving from the prose to the references, and the references are finally machine-checkable. That is a smaller, more boring, and much more durable defense.