Open-Source Livenerf Tracks Whether Claude Opus 5.5 Gets Nerfed, Using a Max Plan
TL;DR
livenerf is an open-source benchmark built to answer one argument that never ends: does a frontier model get worse after it ships? It started the clock on Claude Opus 5.5, which launched on 2026-09-22, and runs a frozen panel of 78 hard questions once a day for 30 days through headless Claude Code on a Max subscription, no API key. The design, panel and decision rule were committed to git before any series data existed. Its pre-baseline validation says the clearest sign of a quiet downgrade is shorter answers (output tokens), not lower scores, and it states plainly that it cannot yet distinguish Opus 5.5 from Opus 5. The project hit the Hacker News front page on September 29 and sat at 793 points with 337 comments when we checked.
What it is
For every model launch there is a follow-up thread claiming the model got "nerfed": quantized, swapped for a smaller model behind the same name, run at lower effort, or routed somewhere cheaper. The README describes the state of that debate accurately: with no clean day-0 baseline, "every argument ends up as vibes versus vibes."
livenerf is the boring fix. It is a Python 3.11 project on Inspect, the UK AI Security Institute's open-source eval framework, with a custom claudecode model provider that wraps a hermetic claude -p call. Inspect treats a Max subscription like any other model API. The statistics follow Anthropic's own paper Adding Error Bars to Evals: paired per-item differences against the baseline with clustered standard errors, so question difficulty drops out.
The call is hermetic in a strict sense: no tools, no MCP servers, no settings files or hooks, no CLAUDE.md, no auto-memory, and a fixed empty working directory. The Claude Code CLI is pinned to one version (2.1.280), and the runner refuses to start if claude --version stops matching it. That rule matters more than it looks: a CLI update changes the harness, and a changed harness looks exactly like a changed model.
How the 78 questions were picked
You cannot make Opus 5.5 deterministic through Claude Code, since sampling parameters are gone and thinking cannot be turned off. So livenerf makes everything else deterministic and relies on sample volume. The hard part is choosing questions that can actually move.
The author screened 2,336 questions from GPQA Diamond, a fixed seeded 2,000-question subset of MMLU-Pro, competition math (BRUMO, CMIMC, HMMT, APEX) and AIME 2025-26, with 4 samples each. Opus 5.5 gets about 93% right on the first try, and 97% of the questions were either always right or always wrong. A question the model always aces or always fails tells you nothing about drift, so only the 1-to-3-out-of-4 questions made the panel.
Think of it as testing whether a sprinter lost a step: timing them on a flight of stairs they always clear, or a wall they can never climb, measures nothing. You want the hurdle they clear about half the time.
Two details from the calibration logs are worth stealing for your own evals. First, AIME was answered from memory: the deviations log records 59 of 60 right on the first sample, with the hardest problems solved in under 40 output tokens, so it contributed almost nothing and MMLU-Pro and competition math were added. Second, questions picked for being "sometimes right" look closer to 50/50 than they are; on fresh samples their pass rate rose from 54.7% to 62.0%, so the power calculation uses the fresh rates instead of the flattering ones.
The panel's SHA-256 goes into data/panel.lock and the daily runner refuses a changed panel. The pre-registration file was committed before the first series run, and every later change sits in a dated deviations log. That log is unusually candid, including a day-5 run where the author overrode the budget guard once.
The validation result: watch the tokens
Before the baseline started, livenerf ran a positive control: the same frozen panel at effort high, medium and low, plus claude-opus-5 at high as a stand-in for the most common nerf claim, a different model behind the same name. This is the most useful finding in the repo so far.
The token signal is loud. Dropping to effort low cut output tokens 62% (99% CI: -70% to -51%) and medium cut them 26%. Accuracy moved in the same direction but with confidence intervals that still crossed zero: -8.3 points for low and -4.2 for medium.
The model swap is the uncomfortable row. Opus 5 in place of Opus 5.5 came out at -3.8 points (95% CI -16.3 to +8.6) and -23% tokens with a 99% interval that also crossed zero. As its own pre-registration requires, the README says so directly: this instrument cannot detect a same-family model swap of that size in a validation's worth of samples. The A/A check, which splits identical effort-high runs in half, landed at +6.4 points with z = 1.79, just inside the |z| < 1.96 pass mark. That is how noisy "the same model, same day" is on a panel built to be on the edge.
The practical takeaway reaches beyond Claude: if you suspect a provider cut thinking effort, log output tokens per request. Per the README, a model that starts thinking less shows it in token counts "often before accuracy moves at all."
What it costs, in plan meter points
Max plans do not publish limits in tokens, so livenerf reads the same percentage meters that /usage shows. Across 657 calibration chunks, the author measured one point of the weekly meter at about 224,527 output tokens. The daily schedule (78 Opus 5.5 samples plus 12 GPQA samples on Opus 5 as a control arm) costs about 3.6 points of the weekly meter, or roughly 803,000 output tokens a week.
That buys a minimum detectable effect of 7.5 accuracy points per 10-day window at 80% power under a 99% test. The design doc lists the price of doing better: a 5-point MDE needs about 7.2 weekly points, and 2 points needs 44.7, nearly half a Max plan. Days 1-10 are the baseline, then two 10-day windows follow, and the decision rule needs both windows to agree with a change of at least 3 points. The first possible verdict is around 2026-10-24. As of the September 29 commit, 6 of 30 days were in and none missed.
Caveats
- No verdict yet. The results table is empty by design until day 20, and nothing in the repo says Opus 5.5 has or has not changed.
- Harness, not API. It measures the model as served through Claude Code on a subscription. The README also logs that the safety classifier sometimes answers with Opus 5 or refuses some biology and math questions; those samples are rejected and counted.
- Self-audit conflict. A report-only audit found 8 suspect answer keys and 30 ambiguous questions among the 80 audited. The auditor was Claude, which the pre-registration flags as a conflict of interest.
- No license file. The repo had no LICENSE when we checked, so the methods are public but reusing the code is legally unclear until one lands.
- Anthropic's stated position. Its September 2025 postmortem of three infrastructure bugs said: "We never reduce model quality due to demand, time of day, or server load." Bugs, though, can degrade output without anyone deciding to, which is exactly the case a day-0 baseline catches.
One Hacker News commenter predicted Anthropic will now "benchmaxx" the repo. Given the panel is 78 questions Opus 5.5 already gets right about 62% of the time, that would be the strangest optimization target in the industry.
Key Takeaways
- livenerf runs a frozen, pre-registered 78-question panel against Claude Opus 5.5 daily for 30 days via headless Claude Code, starting 2.5 days after launch.
- Only questions the model sometimes misses can reveal drift: 97% of 2,336 screened questions were always right or always wrong.
- Output tokens are the loud signal: effort low cut tokens 62% while the accuracy drop still overlapped zero.
- It cannot yet tell Opus 5.5 from Opus 5, and says so; the first possible call is around October 24.
- The whole instrument costs about 3.6% of a Max plan's weekly meter, where one point is roughly 224,527 output tokens.
- The reusable pattern: pin the harness, lock the panel by hash, pre-register in git, and log tokens per request.
Sources: livenerf on GitHub (README, PREREGISTRATION.md, docs/VALIDATION.md, docs/DESIGN.md), Hacker News discussion, Inspect AI, Adding Error Bars to Evals (arXiv), Anthropic postmortem, September 2025, Claude Code headless docs