conversational data analysis

When AI Models Seem to Lose Their Edge, Listen to the Users

Recent user complaints about Anthropic's Claude models, particularly Opus 4.6 and Claude Code, have sparked a heated debate within the AI community. Developers claim they are experiencing performance degradation,…

3 min readVentureBeat
When AI Models Seem to Lose Their Edge, Listen to the Users

The noise around Claude Opus 4.6 and Claude Code is not just another wave of internet outrage. It is a signal, and a valuable one at that. When developers who live inside your tool every day start posting reams of session logs and token-level analyses, they are not trying to be difficult. They are telling you exactly what they value: reliability, predictability, and the freedom to trust the system with complex work. The fact that a senior leader at AMD took the time to analyze nearly 235,000 tool calls speaks to a deeper truth. These users are not casual observers; they are partners in the workflow, and their perception of degradation is their reality, regardless of whether the underlying weights changed.

Anthropic's response has been measured and fact-based, which we respect. Clarifying that the shift to medium effort or the hiding of thinking summaries are product decisions, not secret downgrades, is an important distinction. But in the court of public opinion, intent matters less than experience. When a user sees a model abandon a task midway or burn through tokens without making progress, they do not care if the cause is a cache TTL change or a new default. They only know the tool feels less capable. That is the gap between a technical explanation and a human outcome. The company is explaining *why* the behavior changed, but the users are telling them *how* it feels. Both are valid, but only one of them builds trust.

What makes this moment more combustible is the backdrop of confirmed capacity management. Adjusting session limits during peak hours is a transparent move, but it primes the pump for suspicion. When users hit a quota wall and then see a benchmark score drop, even a flawed one, the narrative writes itself. We should not dismiss the benchmark rebuttals, they are thorough and necessary, but we also cannot ignore that the public is connecting dots that Anthropic has not fully closed. The company says it is not degrading models to manage demand, and we believe them. But the perception of a link between demand pressure and performance is a trust issue that no changelog entry can fully resolve.

The real lesson here is about communication and expectation setting. If you change the default effort level, the cache duration, or the visibility of reasoning, you are changing the product. Telling users to type `/effort high` to get the old experience is a workaround, not a solution. It asks the user to re-engineer their own environment to match a prior standard. That is friction, and friction erodes confidence. For developers, the tool is an extension of their own cognition. When that extension feels less sharp, it is not just an inconvenience; it is a reason to look elsewhere. The companies that will win this era are not the ones that never make changes, but the ones that treat their most demanding users as co-pilots in the evolution of the product, not as test subjects for it.

From VentureBeat

A growing number of developers and AI power users are taking to social media to accuse Anthropic of degrading the performance of Claude Opus 4.6 and Claude Code — intentionally or as an outcome of compute limits — arguing that the company’s flagship coding model feels less capable, less reliable and more wasteful with tokens than it did just weeks ago.

Read the original at VentureBeat