24 Jul 2026, Fri

AI’s Cybersecurity Paradox: How Guardrails Meant to Protect Are Stifling the Defenders

For months, leading artificial intelligence developers have meticulously crafted specialized, vetted programs and implemented stringent guardrails to prevent their powerful models from falling into the hands of malicious actors. However, this robust defense strategy is now inadvertently creating significant hurdles for legitimate network defenders and offensive cybersecurity researchers, hindering their critical work in identifying and mitigating threats before they can be exploited. The very safeguards designed to keep the digital world safe are, in effect, becoming obstacles for those tasked with its protection.

The recent U.S. government’s imposition of export control restrictions on Anthropic’s highly anticipated AI models, Mythos and Fable, in June, serves as a stark illustration of this emerging challenge. This regulatory action was reportedly spurred, at least in part, by a report detailing potential vulnerabilities that could allow users to circumvent the models’ built-in protections, thereby enabling the creation and execution of malicious cyberattacks. While the precise motivations behind this governmental intervention remain a subject of debate, with some suggesting it was more about controlling advanced AI capabilities than a genuine "jailbreak" fear, the practical outcome has been a tightening of access to these potent tools.

Anthropic itself has, in the past, heavily marketed Mythos as a "doomsday cybermachine," capable of both immense good and catastrophic harm, underscoring the need for extreme caution and restricted access. This perception, coupled with the government’s export controls, led to a situation where the most powerful versions of these models were only accessible to carefully vetted entities. Although the export controls on Fable 5 and Mythos 5 have since been partially lifted, with Fable 5 returning to general access on July 1 and Mythos 5 being reintroduced to vetted U.S. organizations under government review, the underlying issue of restricted access for cybersecurity professionals persists.

This pattern of gatekeeping is not unique to Anthropic’s Mythos. Both Anthropic, with its other AI offerings, and OpenAI have established dedicated programs designed to grant cybersecurity researchers access to their models with fewer restrictions. OpenAI’s "Trusted Access for Cyber" program and Anthropic’s "Cyber Verification Program" are examples of these initiatives, requiring applicants to undergo a vetting process to demonstrate their legitimate intent and expertise. The intention behind these programs is to foster collaboration and accelerate the discovery of vulnerabilities by allowing trusted researchers to probe AI capabilities in a controlled environment.

However, these very guardrails have become a significant point of contention and criticism, particularly among researchers whose primary role is to proactively identify unknown vulnerabilities in software and systems and develop methods to exploit them before malicious actors can. These individuals argue that overly cautious restrictions can cripple their ability to conduct thorough security assessments and develop effective defenses.

Mark Dowd, a renowned security researcher with decades of experience in discovering and selling "zero-day" vulnerabilities – previously unknown software flaws and the exploits that leverage them – to Western governments, has voiced his strong concerns. Speaking on a recent cybersecurity podcast, Dowd articulated his discomfort with the current state of affairs: "It’s not really comfortable to me that these random large companies are making arbitrary decisions about what is safe in security and what’s not." His work, which involves keeping vulnerabilities secret for intelligence purposes rather than reporting them for patching, highlights a different, albeit crucial, facet of cybersecurity research that is directly impacted by AI model restrictions. Governments often pay a premium for these undisclosed vulnerabilities precisely because their continued existence is valuable for intelligence gathering and operations.

Dowd acknowledges that his perspective may be influenced by his profession, but he is far from alone in his sentiments. Numerous individuals involved in offensive cybersecurity – the practice of proactively probing systems for weaknesses – have shared their experiences with TechCrunch, detailing how they utilize AI tools and the challenges posed by their inherent guardrails.

Chris Anley, Chief Scientist at the security consulting firm NCC Group, emphasizes that one of the critical steps in confirming a genuine vulnerability is to ask an AI model to attempt an exploitation. This process is vital for understanding the severity and impact of a flaw. However, when an AI’s guardrail prompts an outright refusal to engage with the request, it directly impedes the defensive process. Anley eloquently explains this dichotomy: "This is where the whole offensive versus defensive and guardrails part comes in, because ‘fix this code’ as a prompt is both an essential mechanism for defense but also a roadmap for finding critical vulnerabilities in the code base. So at the same time, the same tool is both an offensive tool and a defensive tool, and the two can’t really be unpicked." He likens the situation to a hammer: "You can’t build a house without a hammer. It’s definitely a tool but it’s also irreducibly a weapon as well."

When faced with these AI-induced roadblocks, Anley and his colleagues often resort to open-source AI models that come without any guardrails, allowing for more unfettered exploration and analysis.

Paolo Stagno, CTO of Crowdfense, a company specializing in the acquisition and sale of undisclosed vulnerabilities to government agencies, echoes Dowd’s sentiments. He criticizes AI companies for treating their customers like children needing constant supervision with their heavily restricted and vetted programs. Stagno notes that while his team does use frontier AI models, they are primarily for reverse engineering purposes. They deliberately avoid using AI to assist in vulnerability discovery or exploit development. The primary reason for this avoidance is the inherent risk of data leakage and the potential for sensitive vulnerability information to be absorbed into future AI training datasets when using cloud-based models. For these critical tasks, they exclusively rely on open-source models run locally, ensuring that data never leaves their secure environment.

Giuseppe Cali, another security researcher focused on discovering zero-days and developing exploits, presents a slightly different perspective. He states that guardrails do not impede his work because he chooses not to employ AI for offensive tasks. Instead, he leverages AI tools for initial reverse engineering, to gain a deeper understanding of the code he is analyzing, and to develop supporting utilities. For these purposes, AI can significantly accelerate the process, allowing him to dedicate more time and cognitive resources to the actual discovery of vulnerabilities. "I still want to own the actual bug discovery and weaponization myself and that wouldn’t change if all guardrails were lifted tomorrow," Cali asserts. "I am jealous of my bugs, and I like this game too much to let models play it for me." This highlights a segment of researchers who value the intellectual challenge and personal ownership of their discoveries.

However, for many, the restrictions remain a significant impediment. A researcher at a smartphone-component manufacturer, who requested anonymity due to authorization constraints, revealed that their employer is not part of Anthropic’s Cyber Verification Program. Consequently, the AI tools available to them are "barely useful for finding vulnerabilities because the guardrails are too strict." They lament, "If it catches wind we’re doing anything security related, it just stops and isn’t usable." This sentiment underscores the frustration experienced by legitimate security professionals whose day-to-day work is hampered by overly cautious AI implementations.

Chris Thompson, CEO of cybersecurity firm RemoteThreat and founder of Offensive AI Con, an event dedicated to offensive security and AI, shares his observations on the practical impact of these guardrails. He notes that they can be inconsistent and behave unpredictably, even within the more relaxed parameters of Anthropic’s and OpenAI’s vetted programs. "The practical impact is you spend a lot of time negotiating with the model instead of working on the core security program," Thompson explains. "Instead of analyzing a vulnerability and reasoning through the exploitability, you’re trying to find why you’re getting inconsistent results or why are models over-sanitizing the output." This constant battle to circumvent or work around AI limitations diverts valuable time and resources away from actual security tasks.

As a result of these restrictive measures, Thompson observes a concerning trend: researchers are increasingly turning to or being pushed towards open-source AI models originating from China, such as GLM. These models are freely downloadable and can be run locally, offering complete freedom from vetting processes and usage restrictions. "You have these responsible researchers that are being pushed away from U.S.-governed systems to foreign-owned systems," Thompson states. "I think it’s more harmful than good to have these guardrails in place." This shift raises geopolitical concerns and the potential for valuable cybersecurity research and development to migrate away from Western-aligned systems.

Thompson advocates for a different approach, urging AI frontier labs to open up their programs, provide responsible access, and focus on holding those who abuse their tools accountable, rather than implementing broad restrictions that stifle legitimate users. He warns of a looming crisis: "There’s this big storm coming. There’s this big wave of attacks that are going to happen at speed and scale like never before. But the same security consulting firms and legit researchers that are trying to make a difference are being stifled right now." The current AI development landscape, with its emphasis on stringent, often overly broad, guardrails, risks leaving the digital world more vulnerable in the face of escalating cyber threats. The paradox lies in the fact that the very tools designed to enhance security are, in their current implementation, hindering the proactive efforts of those best equipped to defend against the evolving threat landscape.

Leave a Reply

Your email address will not be published. Required fields are marked *