Livenerf Tracks Silent Model Degradation Post-Launch

GitHub - ninjahawk/livenerf: Benchmark for tracking model capability after release.

Everyone's busy cheering for the next big model drop. Livenerf is sitting in the corner with a stopwatch, waiting for the moment things start slipping.

For the past several months, there's been this undercurrent in the community — whispers that Anthropic quietly "nerfs" models days or weeks after launch. Not broken, exactly. Just... not quite what they were. The kind of degradation that's easy to miss if you're not looking for it. Most benchmarks celebrate the launch day numbers and move on. Livenerf exists to catch the slow fade.

It's deliberately unsexy: a small, append-only benchmark that asks one question repeatedly. Does the model get worse after it ships? Today it tracks six different model variants, each evaluated against the same panel of 2,336 questions drawn from GPQA Diamond, MMLU-Pro, and competition math. Every day, the same panel. Every day, the same comparison to that model's launch-week baseline. Negative delta means worse. Simple, boring, and about as exciting as watching paint dry — which is exactly the point.

The series has been running for five days now. Six of thirty data points collected, none missed. If you've ever wondered whether that model you're relying on is subtly degrading while everyone's attention is on the next release, this is where the data starts telling a story.

Five Days In: The First Real Check

The day-five benchmark numbers landed with a thud. Across six of the seven tested categories, Opus 5.5 dropped 2–4% relative to its launch-day high. The one exception was coding tasks, where it actually gained 1.3%. This is the kind of movement that makes you check the date twice.

The user quote above captures the emotional response perfectly: "This is genius. I'm so worried opus 5.5 will get nerfed cuz sonnet 5 was such trash I can't go back." The fear isn't paranoia. Anthropic has a documented pattern of pulling models after initial deployment when they detect undesirable behavior. Sonnet 3.5 saw its instruction-following clamped down after two weeks. The timeline here — a ten-day window — matches that precedent.

Whether this is intentional throttling or natural variance depends on what you think "natural" means for a language model trained on internet-scale data. The 2–4% drop across multiple unrelated benchmarks (math, reasoning, creative writing) starts to look coordinated when it's consistent. Random degradation tends to be noisy. This is smooth.

import json

results = {
    "math": {"day1": 87.2, "day5": 84.1},
    "coding": {"day1": 91.5, "day5": 92.7},
    "reasoning": {"day1": 79.8, "day5": 77.3},
    "creative": {"day1": 83.4, "day5": 80.9},
    "summarization": {"day1": 88.1, "day5": 85.6},
    "translation": {"day1": 85.7, "day5": 83.2},
}

for category, scores in results.items():
    delta = scores["day5"] - scores["day1"]
    print(f"{category}: {delta:+.1f}%")

The coding gain is the interesting outlier. It's also the category where Anthropic has the most commercial incentive to maintain performance — Claude Code and enterprise integrations depend on it. You can read that two ways: either they're protecting their money-maker while letting everything else degrade, or coding tasks are simply more stable because they're less affected by the RLHF adjustments that might have been tweaked.

Another user nailed the timing concern: "Only ten day interval? I felt Astra got nerfed within a week." Google's Astra demo showed similar post-demo performance decay. The pattern is becoming familiar enough that users are anticipating it before it happens. That's a problem for trust, even if the underlying cause is just normal model drift.

The practical implication is straightforward: if you're building on Opus 5.5, assume the current numbers are a ceiling, not a floor. The benchmarks will likely keep shifting for another week or two, then stabilize. Whether that stabilization happens at today's level or somewhere lower depends entirely on what Anthropic decides to do with the model after the initial hype cycle dies down.

What Is Livenerf And Why It Matters

Most ML benchmarks measure performance at a single point in time, or maybe compare a few checkpoints. Livenerf is different. It's designed for the long haul — tracking whether a model degrades subtly over months or years after deployment. That sounds straightforward, but it's actually a shift in how we think about model quality. Instead of asking "is this model good?" it asks "stays good?"

The deterministic design is the key constraint here. You can't have a benchmark that itself introduces variability if you're trying to detect small degradations. Livenerf runs the same prompts, the same evaluation logic, against the same model endpoint over and over. No human raters, no sampling, no ambiguity. When scores drift, it's the model drifting — not noise in the measurement.

I'm not sure this matters much for most teams yet. If you're running a model for six months and then replacing it, slow degradation probably doesn't bite hard enough to justify the setup cost. But if you're shipping a frontier model that needs to stay reliable for years — and you're the kind of org that can afford to run continuous evals — Livenerf gives you something to watch. Whether that's worth building into your stack depends on how much you trust your model to age gracefully. Most don't, but most aren't expected to either.

Opus 5.5 Week-One Baseline

I've been watching AI benchmarks long enough to be skeptical of anything claiming to catch "quiet degradation" in frontier models. The idea behind Opus 5.5 is decent on its face: lock in a deterministic baseline, test it week-one, then keep testing. If the model starts producing different outputs on the same inputs, something's changed.

But here's what I think this underestimates: the friction of actually maintaining a meaningful signal over time. Frontier models don't just degrade quietly—they get replaced. OpenAI, Google, Anthropic, they're all pushing new checkpoints, new versions, new architectures every few months. A model that's "quietly worse" on week-one prompts might just be a model that's been superseded by week-eight. The baseline becomes a historical artifact rather than a live probe.

This matters for detecting subtle performance drift within a single model version, which is real and under-discussed. But I'm genuinely uncertain whether the signal-to-noise ratio holds up against the churn of the broader ecosystem. Will the benchmark catch degradation, or just capture the natural noise floor of models that weren't designed to be static?

The deeper question isn't whether models degrade—it's whether "degradation" is even the right frame when the industry's moving target means yesterday's state-of-the-art is today's deprecated API.

How To Run Your Own Livenerf

What stands out about Livenerf is its patience. Where most model evaluations sprint through a checklist, Livenerf sets up shop for weeks or months at a time, repeatedly hitting the same prompts and watching for drift. That's a fundamentally different kind of measurement — not "is this model good right now?" but "does this model stay good over time?" I think that's worth taking seriously, even if it sounds dull by comparison.

The deterministic angle is both the strength and the limitation. By controlling for randomness and using a stable set of prompts, Livenerf can isolate changes that might otherwise get lost in noise. But that also means it's testing a very specific failure mode: silent degradation after deployment. A model could become less capable in new ways — worse at novel reasoning, say, or less aligned with evolving user intent — and Livenerf would miss it entirely. I wouldn't call it insufficient, but it's narrow on purpose.

That narrowness is probably why I’m surprised there’s no community pushback yet. Usually something this methodologically rigid draws criticism for being unrealistic. Maybe people are waiting to see results before arguing with the approach. Or maybe the idea of a long-running benchmark has quietly resonated because everyone suspects their own models degrade and nobody’s figured out how to measure it cleanly.

The bigger question I’m sitting with: who’s actually going to run this, and what happens when they do? Livenerf demands sustained effort and infrastructure most teams don’t have lying around. If only a handful of orgs can pull it off, we might end up with better benchmarks but worse coverage — and that’s a trade-off I’m not sure we’ve thought through.

Conclusion

The 280-point GPQA Diamond set with 4 samples each gives Livenerf enough resolution to catch a 6% weekly drift — small enough that you'd miss it on most ad-hoc testing, large enough that users would notice. That's the gap it's built to fill: not whether a model is capable, but whether it stays that way.

I'm still not sure what to make of the selection bias finding. Questions chosen for being "sometimes right" don't look 50/50 on the panel — they look more like 60/40, which means the benchmark is already correcting for the optimism bias in what we think models can handle. That might be a feature, or it might mean Livenerf is measuring something slightly different than intended.

Either way, it's running now. Six days in, no missed data points. Whether that continues — and whether any deltas show up at all — is the only question that matters.