The researchers we spoke with are not asking for a free pass. They are the people who find unknown vulnerabilities and build the tools that exploit them, often so that the rest of us never have to worry about those holes being used against us. Their work is adversarial by design, and that is precisely why it is valuable. Yet when they sit down to use frontier models from OpenAI or Anthropic, they run into guardrails that treat their legitimate research as if it were a threat. The result is a quiet tax on the very people who make our digital world safer.
This is not a complaint about safety measures existing. It is a question of whether those measures are aimed at the right target. The guardrails are built to stop bad actors from using AI for harm, but they also stop good actors from using AI to find harm before it finds us. We have seen this tension before in how we talk about the technology itself. In Talking to My AI Clone Taught Me to Question the Tech, the author grapples with the unease of interacting with something that feels familiar but is fundamentally not human. That same unease is at play here, except now it is reversed. The researchers are not worried about the AI becoming too human; they are frustrated that the AI's constraints make their own expertise less useful.
The practical cost is real. Offensive security work relies on iteration, on testing a hypothesis and adjusting on the fly. When a model refuses a request that a developer can see is benign, or when it hedges on a technical detail because it might be used for something it was not explicitly allowed to do, the researcher loses momentum. They lose time. And in a field where the window between a vulnerability being discovered and exploited is measured in hours, that time has a real price. We are not suggesting that these companies should hand out exploit code like candy. But the current approach treats all requests for offensive capabilities as equally dangerous, which is a blunt instrument for a subtle problem.
The deeper issue is one of trust. The companies building these models are making a judgment call about who should be allowed to use them and for what purpose. That judgment is understandable, but it is also a product decision that has consequences for the broader security ecosystem. If the best tools are effectively off-limits to the people who need them most, we will end up with a system where the defenders are working with one hand tied behind their backs. We saw a similar shift in how job requirements have changed in the AI field, where the lines between software engineering and machine learning have blurred, as covered in Navigating AI/ML Job Requirements: A Shift in Expected Skills. The point is that the landscape is shifting under everyone's feet, and the rules we set now will shape who gets to play.
What would we tell a researcher who asked us about this? We would say that their frustration is justified, but that the answer is not to rip out the guardrails. It is to push for more nuance, for models that can distinguish between a security researcher doing their job and a malicious actor probing for weaknesses. That is a hard problem, but it is not an impossible one. The open question is whether the companies building these models are willing to invest in that nuance, or whether they will continue to err on the side of saying no. Watch for the first time a major model vendor publicly acknowledges that its safeguards have slowed down a legitimate security research effort. That moment will tell us whether the guardrails are here to protect us, or just to protect themselves.
