Has Opus 5.5 Been Quietly Nerfed? This Live Tracker Knows.
In this article
Bottom line: Livenerf is a public GitHub benchmark that started on Claude Opus 5.5's launch day, September 22, 2026.
It re-runs frozen prompts with a pinned CLI and exact graders, and it keeps the raw logs. It reports deltas against launch-week performance, so "it feels dumber" becomes a number you can check.
A separate tracker, NerfBench, reported Opus 5.5 at 99.2% of its starting power in its first retest. Eight days in, the public evidence doesn't show a nerf.
It does show that we finally have a way to catch one.
You have probably typed "it got dumber" into a chat window at least once. I have, more than once, and I had nothing to back it up.
No baseline, no logs, just a vibe and a screenshot I couldn't reproduce.
Opus 5.5 launched on September 22, and within days the complaints were rolling in: dumb answers, oddly cautious refusals, and "this feels like an older model."
Then a repo called Livenerf hit Hacker News, and it asked the question I'd been dodging: compared to what?
A note before we start. I didn't run my own eight-day experiment for this piece, and I'm not going to invent one. What follows is what the tracker does, what it can and can't prove, and how to read it.
The Problem With "It Feels Dumber"
Nerf complaints follow a pattern. A model launches, everyone's impressed, and a few weeks later the forums fill with people swearing it got worse. Sometimes they're right.
Often they aren't. Nobody can tell which, because the only evidence is anecdotes from people who aren't running the same prompts.
Here are the usual suspects when a model seems to decline:
- Real changes. A quantized model, a swapped checkpoint, a tweaked system prompt, or new safety tuning.
- Novelty decay. You were impressed at launch and now you notice the flaws.
- Harder prompts. You trust the model more, so you give it tougher work.
- Plain variance. Language models are stochastic, and one bad run feels like a trend.
- Tooling changes. A CLI update or a different context setup changes behavior without touching the model.
Only the first one is a nerf. The other four feel identical from the inside. That's why a baseline matters more than a hot take.
What Livenerf Actually Does
Livenerf started on launch day and treats the whole thing as a measurement problem. According to its GitHub repo, it makes the test deterministic wherever it can:
- Frozen prompts: the same inputs every run, no edits.
- Pinned CLI: the tooling can't quietly change underneath the model.
- Exact graders: answers are scored by code, not by another model's opinion.
- Raw logs, kept forever: anyone can audit what happened.
The scripts repeatedly query Opus 5.5 on fixed prompt sets and log things like output length, factual accuracy, and refusal rates.
The headline artifact is a running table that compares each day to the launch-week baseline. A negative delta means the model is doing worse than it did in week one.
That is a small idea, and I think it's the right one. The value isn't in any single score. It's that the yardstick doesn't move, so if the numbers drift, something real changed.
Why the Design Choices Matter
I've read a lot of "model got worse" posts, and the weakest ones share a flaw: the judge is as unstable as the thing being judged.
If you ask one LLM to grade another LLM's answers, you've added a second source of drift. Livenerf's exact graders sidestep that.
A regex or a unit test gives the same verdict on Tuesday as it did on Friday.
Pinning the CLI matters for a less obvious reason. Many people use these models through agentic tools, and the tool is part of the system.
If the harness changes how it packs context or which instructions it injects, the "model" seems to change even when the weights didn't.
Freezing the harness isolates the variable you actually care about.
Then there's the logging. Raw logs forever means a skeptic can open the files and check the grader's work. That is the difference between a tracker and a press release.
What the Early Signal Says
As of this writing, the picture is calm. NerfBench, a separate effort, reported in its first retest that Opus 5.5 scored 99.2% of its starting power, with no nerf detected.
That figure is NerfBench's, not Livenerf's, so don't mix them up. It is one independent data point pointing the same direction.
Meanwhile, the social side told a different story. Coverage of quick "nerf" complaints days after launch shows the usual launch-week cycle of hype, disappointment, and accusation.
Complaints arrived fast, before anyone had meaningful longitudinal data.
So far the score is vibes: loud, benchmarks: quiet. That doesn't mean the vibes are wrong. It means they haven't been confirmed.
What a Tracker Can't Tell You
I want to be straight about the limits, because trackers have a way of getting over-read.
It only measures what it measures
A fixed prompt set covers a fixed slice of behavior.
If a model gets worse at long, messy, multi-file refactors, and the benchmark is mostly short factual questions, the dashboard stays green while your experience sours.
A clean tracker means "no detectable change on these tasks," not "no change."
Variance hides small shifts
Even with frozen prompts, sampling noise exists. A 1% dip might be a real degradation or just a bad day. You need enough runs over enough days to separate a trend from a wobble, and eight days is early.
Provider-side changes can be conditional
Behavior can differ by plan, by region, by time of day, or by which infrastructure serves your request. A tracker running from one setup sees one slice of that.
If your experience differs, you may be seeing something the tracker structurally can't.
Launch-week is the baseline, for better or worse
Livenerf compares against launch-week performance.
That's a sensible anchor, but if launch-week behavior was itself unusual (extra compute, a special configuration), every later delta inherits that bias. I'm not saying that happened.
I'm saying it's a thing to keep in mind.
How to Read the Table Like a Skeptic
If you open the tracker this week, here's the checklist I'd use:
- Look at the trend, not the day. One dip means little. Three consecutive days of drift on the same metric means more.
- Check all the metrics together. Output length, accuracy, and refusal rate can move independently. A refusal spike with flat accuracy is a different story from an accuracy drop.
- Match it to your work. Do the prompts resemble what you do all day? If not, discount the result.
- Keep your own baseline. Save ten of your real prompts and outputs from this week. Re-run them in a month.
- Treat one source as a hypothesis. Two independent trackers agreeing beats one tracker being confident.
Why This Matters Beyond One Model
The bigger shift is cultural. For years, the relationship between AI companies and users has been asymmetric: the provider can change the model any time, and the user can only complain.
A public, reproducible tracker moves a little power back toward users.
It also raises the cost of quiet changes. If anyone can run a frozen benchmark and publish the deltas, silent degradation becomes harder to get away with.
And it cuts the other way too: it gives providers a way to disprove false accusations with data instead of PR.
I'm not naive about it. A benchmark can be gamed, and a model could in principle be tuned to ace known public prompts while slipping elsewhere. Keeping your own private prompts is the defense.
But the existence of these trackers changes the conversation from "trust me" to "show me."
For context on how fast these shifts in trust move, see our piece on the "You're Absolutely Right" phenomenon, which covers how model behavior quirks spread through the community before anyone measures them.
What I'd Do This Week
If you rely on Opus 5.5 for real work, don't panic and don't shrug. Do three cheap things:
- Save your baseline. Pick five to ten prompts you use regularly, run them now, and keep the outputs.
- Watch the tracker's trend. Check it weekly, not hourly. Daily checking turns noise into anxiety.
- Note your tooling. Write down your CLI or app version so you can rule it out later.
If your own re-runs start to diverge from the tracker, you've found something worth reporting. If they match, you can stop wondering.
The Part That Surprised Me
I expected to come away either vindicated or debunked. Instead I found myself more interested in the instrument than the verdict.
Whether Opus 5.5 gets nerfed is a question that more weeks of data will start to answer.
The bigger change is that a stranger with a GitHub repo can now hold a frontier model to its own launch-week performance, in public.
That changes how I'll write about these models from here on. I'm going to ask "compared to what?" a lot more often, and I'd like you to as well.
So here's my question: do you have a baseline for the AI tools you use every day, or are you running on vibes like I was? I'd love to hear what you'd put in your own frozen prompt set.
Sources: Livenerf on GitHub, NerfBench first retest coverage, launch-week nerf complaints.