The Stopwatch That Started on Day Zero
One user started a daily benchmark on the day Opus 5.5 shipped, and 147 comments spent themselves arguing whether the thing it measures is model degradation or hedonic adaptation.
Every Light Change Worked — Until Someone Checked
At a phone factory outside Chicago, output rose when the lights went up, rose when they went down, and only fell near moonlight. The story is in every management textbook. The data were never analyzed, and everyone assumed they had been destroyed.
What Claude Did With the Other Nineteen
Anthropic's new biology lab pointed roughly 950 Claude agents at a public DNA database for 21 hours. They came back with 3,500 candidate systems and 20 written reports. One of those reports described a genuinely new enzyme system. The other nineteen are the interesting part.
Fake Authors, Real Orals
Two ML reviewers audited 22 conference submissions this summer. Fifteen had fabricated citations or hallucinated authors — and two papers they flagged still got accepted as orals.
A $500 Fine-Tune That Beats the Giants
A 9B open model, fine-tuned with $500 of RL, outperformed frontier models on a real-world catalog review task. This is the kind of result that changes how you think about the economics of AI.
Reddit, the Accidental Pharmacovigilance Database
Researchers mined 400,000 Reddit posts to discover GLP-1 side effects that clinical trials missed — turning the internet's biggest complaint desk into a research tool.