A new benchmark named "livenerf" has been released to track whether frontier AI models experience performance degradation after their official launch. Developed by researcher Nathan Langley (ninjahawk), the tool is designed to provide an objective way to discuss whether a model's capability is "nerfed"—a term used when performance is reduced through methods such as quantization, switching to smaller models under the same name, reducing reasoning effort, or changing routing.
New "livenerf" Benchmark Released to Track Performance Drift in Frontier AI Models
The benchmark is built on Inspect, an open-source evaluation framework from the UK AI Security Institute (AISI). To ensure high-fidelity results, livenerf focuses on deterministic testing, using frozen prompts, pinned CLI versions, and exact graders to minimize noise. The tool aims to differentiate actual performance changes from user perception errors caused by evolving expectations.
Currently, livenerf is running a series on Claude Opus 5.5, tracking data from its launch period in late September 2026. The benchmark compares performance against a 10-day baseline established immediately after launch. It specifically targets questions where models have a variable success rate (roughly 50%) to better detect statistical drift. Unlike many other evaluations, livenerf also monitors output token counts, as a reduction in tokens can be an early indicator that a model is "thinking" less before accuracy declines.
Sources
- 最先端AIモデルがリリース後にひそかに劣化しているかどうかを検出するためのベンチマーク「livenerf」 (GIGAZINE, 2026-10-01)
- GitHub