Your coding assistant feels dumber this week, a Reddit thread with thousands of upvotes agrees with you, and a new YouTube tracker claims it can prove the drop. This guide gives you a calmer test: what the public nerf trackers have actually found so far, why a chat window cannot settle the question, and a 10-prompt check you can freeze today and rerun every week.
Has any tracker actually called a nerf yet?
No public tracker has called a nerf yet. The best-known effort, LiveNerf, is still collecting its baseline for Claude Opus 5.5, so its results table is empty by design. Anyone sharing a “LiveNerf already shows a drop” chart is misreading calibration data, not the live series.
LiveNerf started its clock just after Opus 5.5 shipped on 22 September 2026. As explainx.ai summarised on 30 September 2026, the series was on day 6 of 30, with 6 of 10 baseline days collected at 90 samples a day and no days missed. The first possible verdict needs two full 10-day windows after the baseline, which puts the earliest call around 24 October 2026. Until then, the honest headline is “no call”, in either direction.
That timing matters because the anxiety is real. The daily trend watch for 6 October 2026 logged a 2,691-upvote Reddit thread claiming Opus 5.5 had degraded a week after launch, plus daily-benchmark videos tracking the same fear. Social velocity proves people are worried. It does not prove the served model changed.
What does “nerfed” actually mean?
“Nerfed” is five different failures wearing one label. Each one feels identical in a chat window — shorter, lazier, slightly wrong answers — but each needs different evidence. Separating them is the whole job of a tracker.
| What people mean | What would have changed | What evidence fits |
|---|---|---|
| Quieter weights | Quantisation or a smaller model behind the same name | Accuracy drop on frozen items with tokens roughly stable |
| Thinking less | Lower default effort or reasoning budget | Output tokens fall first, accuracy follows |
| Different route | Routing, system prompt or safety layer changed | One product path moves, the raw API does not |
| New harness | CLI, SDK or tool scaffold updated | Drop starts exactly with a version change |
| Nothing changed | Harder tasks, tired reviewer, hedonic adaptation | Frozen items stay flat while new work feels worse |
LiveNerf is explicit that it measures change in served behaviour without assuming a mechanism. Its plan lists quantisation, a smaller model, lower effort, routing, system-prompt changes and infrastructure bugs as possibilities to narrow down later — not as conclusions. Keep that discipline in your own testing: log the symptom, then earn the explanation.
Why does your chat window feel like proof?
A single re-run proves almost nothing because frontier models are not deterministic. On Opus 5.5, the sampling knobs are gone and thinking cannot be switched off, so the same prompt can legitimately return different answers. LiveNerf’s plan is blunt on this point: deterministic inputs and deterministic grading, judged statistically over many samples, are the only workable definition of “same test”.
History also warns against jumping to “they swapped the weights”. Reporting on 1 October 2026 notes that Anthropic investigated earlier Claude Code quality complaints and traced the impact to an inference-intensity change made on 4 March plus separate software issues, while stating the API and underlying inference system were not affected. The explainx.ai walkthrough of the Hacker News discussion makes the same point with earlier incidents: confirmed quality drops were traced to bugs and a changed product default, not a deliberate weight replacement. Users were right that output got worse. The cause was still not a secret nerf.
There is a quieter confounder too. A launch-week model handles your easy prompts so well that you promote it to harder work. When it hits the ceiling it always had, the model feels nerfed even though nothing moved. A frozen panel dodges that trap because the tasks never get more ambitious.
How does LiveNerf freeze the instrument?
LiveNerf freezes everything except the model, then watches thousands of graded samples for drift. The harness is pinned, the prompts never change, grading is exact rather than judged by another model, and the decision rule was written down before the data arrived.
The v0 setup runs Opus 5.5 as served through Claude Code on a Max subscription using headless claude -p, with no API key. The CLI is pinned (version 2.1.280 in the September snapshot), auto-update is disabled, and each call uses a short frozen system prompt with no tools, no MCP servers and no project files. Effort is always set explicitly, never left at the product default — a direct lesson from past incidents where a default-effort change looked exactly like a model drop.
livenerf measures Opus 5.5 as served through Claude Code on a subscription.
— LiveNerf README, via explainx.ai, 30 September 2026
The question panel is biased on purpose. The project screened 2,336 questions from GPQA Diamond, MMLU-Pro, competition maths and AIME 2025–26, then kept only the 78 items the model sometimes got right. Items it always passes or always fails cannot show a drop, so they were removed. A control arm reruns the older Opus 5 on part of the panel each day: if both models move together, you are looking at a platform or harness event, not an Opus 5.5 nerf.
Validation shows both the power and the limit. In pre-registered tests, dropping effort from high to low cut output tokens by 62% and accuracy by 8.3 points, while swapping Opus 5 for 5.5 was not distinguishable at 99% confidence. In plain English: a “thinking less” change should show up in token counts before it clears the accuracy bar, but this instrument cannot prove a same-family weight swap. Its minimum detectable effect is about 7.5 accuracy points per 10-day window — a smoke alarm, not a microscope. Figures above are LiveNerf’s published validation and plan values (September–October 2026), not our own test.
If you are deciding what to pay for rather than chasing drift, start with our Claude Opus 5.5 vs Sonnet 5.5 routing guide and our Codex $100 Pro vs Claude Code comparison. A drift bench that has not finished its baseline cannot make that buying decision for you.
How can you build your own 10-prompt nerf check?
Freeze 10 prompts you can grade exactly, pin every setting you control, and rerun the identical sheet weekly. You are copying LiveNerf’s method at kitchen-table scale: same items, same harness, tokens logged, no verdict from a single bad afternoon. Our sheet below is the original element of this guide — desk-designed from the published method, not hands-on validated against a live model, so treat your first two runs as your baseline, not as proof.
- Pick the path you actually pay for. Test Claude Code, Codex or the API exactly as you use it. Never mix paths in one sheet — a CLI result says nothing about the API.
- Write 10 prompts you can mark without taste. Use the sheet below: exact sums, a string reversal, a constrained paragraph, a small function with hidden tests. If grading needs an opinion, replace the prompt.
- Pin the harness. Record the CLI or SDK version, set effort explicitly, use one fixed system prompt, one turn, no tools. Note the model ID exactly as returned.
- Run each prompt 4 times. Sampling varies, so score the average per prompt, not the best or worst single answer. Save every raw output.
- Log output tokens for every run. Tokens are your early-warning gauge. A sustained token fall with flat accuracy suggests less thinking; a token fall plus an accuracy fall deserves a second week, not a headline.
- Add one control. Rerun 3 prompts on a second model or path you trust. If everything drops together, suspect your harness, your network or your grading before you suspect the vendor.
- Apply a two-window rule. Only worry when the average score moves at least 3 points in the same direction two weeks running, with the control flat. Anything smaller is weather, not climate.
| # | Prompt type (write your own wording) | Exact grade | What it catches |
|---|---|---|---|
| 1–2 | Multi-step arithmetic chain | Final number matches | Precision, skipped steps |
| 3 | Reverse / transform a long random string | Character-perfect match | Token fidelity |
| 4 | Paragraph with hard constraints (bullet count, banned letter) | Parser checks each rule | Instruction following |
| 5–6 | Small coding task with 3 hidden tests | Tests passed out of 3 | General coding ability |
| 7 | Logic grid / scheduling puzzle | Verifier accepts or rejects | Multi-step reasoning |
| 8 | Fill a JSON schema from messy text | Schema valid + fields correct | Format regressions |
| 9 | Needle question in a long pasted document | Exact quoted answer | Long-context retrieval |
| 10 | Your real weekly task, frozen verbatim | Checklist of required facts | The work you actually feel |
Keep the prompts private and unchanged. The moment you edit a prompt to make it “fairer”, you have started a new baseline. Our how-we-test methodology uses the same rule for tool reviews: freeze first, argue later.
Can you rely on detection claims you see shared?
Only when the claim names its path, window and control. Most viral nerf posts fail at least one of those three. Use this verdict table before you switch tools, cancel a plan or forward the chart.
| Claim you will see | Verdict | Why |
|---|---|---|
| “It felt worse on Friday” | Ignore | One session, unpinned harness, harder tasks likely |
| “My 5 prompts dropped this week” | Watch | Right instinct; needs a second window and token logs |
| “Tracker X shows a 4-point drop” | Check the rule | Ask: frozen items? two windows? control arm flat? |
| “They swapped in the old model” | Unproven | LiveNerf validation could not distinguish that swap |
| “Tokens fell 60% at same accuracy” | Take seriously | Matches the known low-effort signature; verify effort pin |
The same scepticism applies to other trackers. The Hacker News discussion pointed to a second bench tracking Opus 5.5 and GPT-6-class models with a launch-day baseline and a 10% change threshold. Different registry, different threshold, different product under test. Two trackers agreeing on direction is interesting; copying a threshold from one onto the other is not.
What should you do when you think it got worse?
Check the boring causes first, in this order. Most “nerfs” die at step 1 or 2, which is exactly why the trackers pin those layers before measuring anything exotic.
- Version: did your CLI, app or SDK update this week? Compare the logged version with last week’s sheet.
- Effort: is effort still explicit, or did an update reset it to a default? Rerun one prompt at explicit high effort.
- Tokens: compare median output tokens on identical prompts. A cliff suggests thinking budget, not weights.
- Control: run your 3 control prompts elsewhere. A shared drop points at you, not the vendor.
- Patience: if the sheet still looks bad, wait for the second window. Then decide with data you would defend in public.
If your frozen sheet stays flat while daily work still feels worse, believe both results. The model probably did not change; your work outgrew week one. That is a routing and prompting problem — cheaper to fix than a phantom nerf.
Frequently asked questions
Has Claude Opus 5.5 been nerfed?
No tracker has called it. LiveNerf was still in baseline collection at the end of September 2026, its results table stays empty until day 20, and the earliest possible call under its pre-registered rule is around 24 October 2026.
Why do my identical prompts give different answers?
Because hosted frontier models sample their reasoning. With sampling controls removed and thinking always on, variation is expected. Judge the average over repeated runs on frozen prompts, never a single reply.
What is the fastest signal of a real downgrade?
Median output tokens on identical prompts. In LiveNerf’s validation, lowering effort cut tokens by 62% while accuracy moved 8.3 points. A sustained token cliff at explicit effort is worth investigating before any accuracy chart.
Should I switch models because of nerf posts?
Not on vibes. Switch when your own two-week frozen sheet, with a flat control, says the path you pay for moved — or when price and features favour the alternative on their own merits.
Does an API result settle a Claude Code complaint?
No. LiveNerf deliberately measures the subscription Claude Code path because that is what most complaints describe. API, chat and CLI are different products sharing a model name; test the one you use.
Sources and methodology
This guide rests on LiveNerf’s published plan and validation, independent reporting of that project, and the 6 October 2026 trend record. We did not run LiveNerf or test a live model for this article; the 10-prompt sheet is an original desk adaptation of the published method. Verification date: 6 October 2026.
- LiveNerf measurement plan (GitHub) — served-path definition, pinning rules, validation and statistics.
- explainx.ai: Livenerf — Has Opus 5.5 Been Nerfed Yet? — day-6 status, decision rule and discussion summary, 30 September 2026.
- Digital Today: Open-source benchmark Livenerf tracks model performance — project origin and Anthropic investigation summary, 1 October 2026.
Where sources describe community claims (Reddit, YouTube, Hacker News), we report them as claims, not findings. Correlations between effort, tokens and accuracy in validation are experimental within that rig; they do not prove any vendor changed a production model.
Arva Rangwala covers AI news, AI tools, guides and prompts for OpenAIMaster — what is new, what is worth using, and how to put AI to work.
Feel free to email us at contact@openaimaster.ai — we are happy to help!


