Livenerf took the top of Hacker News on September 29 with 343 points and 147 comments, and it exists because of one missing object: a baseline from the day a model ships. Claude Opus 5.5 came out on September 22. Two days later a developer called ninjahawk started running 90 samples a day against a frozen panel of questions through headless Claude Code on a Max subscription, with no API key, and committed to doing it for thirty days. The README describes the whole thing in one line: “a small, boring, append-only benchmark for one question: does a model get worse after it ships?” The chart at the top of the page is redrawn by the daily run itself. The first reading that could become a verdict lands around October 24.
The instrument is more carefully built than the complaint it answers. From 2,336 questions drawn out of GPQA Diamond, MMLU-Pro, competition math, and AIME 2025-26, each screened with four samples, the author kept the 78 that Opus 5.5 got right sometimes and wrong sometimes, because a question that is always right or always wrong cannot show drift. That is 97 percent of the pool, discarded. He then measured his own selection bias: questions chosen for looking like coin flips passed at 54.7 percent in screening and 62.0 percent on fresh samples, so the power calculation uses the honest number, and one run a day of the full panel detects a change of about 7.5 accuracy points per ten-day window. The statistics follow Anthropic’s own paper on adding error bars to evals, the harness is the UK AI Security Institute’s open-source Inspect, and the protocol was pre-registered before the first sample: days 1 through 10 are the baseline, then two ten-day windows, and the first results row lands after day 20. Nothing about the sampling can be made deterministic, since sampling parameters are gone and thinking cannot be switched off, so everything else is frozen instead - prompts, CLI version, graders - and the raw logs are kept forever.
What the instrument can actually see is where the project gets interesting, because it publishes its own resolution. In validation, a drop to low reasoning effort showed up far more clearly in tokens than in accuracy: 62 percent fewer output tokens against 8.3 points of accuracy loss, with error bars of 4.5. Swapping in Opus 5, the previous generation, was not distinguishable from Opus 5.5 at 99 percent confidence at all, at 3.8 points of apparent difference and 23 percent fewer tokens. Which means the version of the accusation people repeat most often - that a cheaper sibling is quietly being served behind the same name - sits right at the edge of what this thing can detect in a validation’s worth of samples. The audit section is just as unflattering: 8 of the panel’s answer keys look wrong and 30 questions look ambiguous, which is what questions a strong model only sometimes gets right should look like, and nothing was dropped, with a pre-registered sensitivity analysis to rerun the result without them. There is also a line about the serving path, where a safety classifier sometimes answers with the older model or refuses biology and some math questions, and those samples are rejected, counted, and excluded along with the questions they touched.
The thread mostly did not argue about any of that. It split into two camps that are both about evidence. “This is genius,” one commenter wrote. “I’m so worried opus 5.5 will get nerfed cuz sonnet 5 was such trash I can’t go back.” Against that, a group argued the whole phenomenon is a honeymoon ending. “It’s always around a week,” one reply said, and the answer under it proposed hedonic adaptation as the mechanism, comparing the communities to “spoiled children on the day after Christmas when new toy novelty has begun to wane.” Or, more bluntly: “Hot take, none of the models are getting ‘nerfed’, people are just getting used to the new level of intelligence.” What made the thread worth reading was that the skeptics were not smug about it. “Has there ever been any measurement of this, of any sort?” one asked. “This should be measurable, and I’m glad this project is measuring it,” adding that replies in the form of additional anecdotes, stated with greater passion, would prove his point. Someone else produced the closest thing to a controlled experiment in the thread: an automated 3 AM prompt that records input and output tokens alongside weekly and five-hour quota before and after a run, where the absolute token counts stayed within 0.1 percent while the same work counted as 1 percent of the quota in one mode and 4 percent in another. And one commenter pointed at a second tracker, Nerf Bench, which samples on launch day and flags any deviation above 10 percent, and which he said caught a real degradation in Opus 4.6 that Anthropic later wrote about.
That last part is checkable, and it checks out, though not in the way the thread assumed. Anthropic published an update on recent Claude Code quality reports on April 23, after a month of complaints. It traced everything to three separate changes. On March 4 the company moved Claude Code’s default reasoning effort from high to medium to stop the interface from appearing to freeze, and reverted it on April 7 because users said they would rather have the intelligence by default. On March 26 a change meant to clear stale thinking from idle sessions shipped with a bug that made it repeat every turn for the rest of the session, fixed on April 10. On April 16 a system prompt instruction to cut verbosity hurt coding quality in combination with other changes. All three were resolved by April 20. “We never intentionally degrade our models,” the post says, “and we were able to immediately confirm that our API and inference layer were unaffected.”
🎩 Cask’s Take
The most valuable thing in that April postmortem is not the denial. It is that in the one case where complaints were real and were investigated, the cause was three changes in the wrapper - a default setting, a bug, a prompt line - and the model underneath was never touched. That re-frames the entire genre of complaint. “Nerf” is treated as a statement about weights when the last verified instance was a statement about harness code, quota accounting, and what gets read as a session. Which is also the part livenerf deliberately holds still, by pinning its CLI and changing nothing but the day. The instrument is aimed at the server side: routing, effort level, quantization, classification. It is a clean experiment and, by design, a narrow one.
And there is a detail in its own README that the debate should have spent more time on. The safety classifier sometimes refuses biology and some math questions, and those samples are rejected, counted, and excluded. So the degradation people in that thread fear most, the model quietly becoming more cautious and less willing, would arrive looking like a shorter sample count rather than a lower score. The panel measures capability and files timidity under “not measured.” Anything else would break the statistics, and it is still the sharpest unbudgeted finding in the project.
The hedge that has been circulating for months is that there is no clean day-zero baseline, and this is one person fixing that with a Max subscription and a cron job. The bet is pre-registered and dated, which converts a permanent argument into a forecast that has to resolve in about three weeks. The two camps settle together, too. If the panel’s line stays flat while the complaints keep coming, adaptation wins without anyone having to concede bad faith. If the line bends downward, the labs lose an argument they have been winning by default. Either way the question stops being a matter of how strongly it was felt. Somebody had to start the clock, and it turns out day zero was the only day nobody thought to keep.