24 Jul 2026, Fri

AI’s Cybersecurity Dilemma: Guardrails Meant to Protect Are Now Hindering Defenders

In the ongoing arms race between malicious actors and cybersecurity professionals, artificial intelligence has emerged as a powerful new weapon. However, the very guardrails designed by AI giants to prevent their models from being weaponized by hackers are now inadvertently stifling the work of legitimate network defenders and offensive cybersecurity researchers, creating a complex dilemma for the future of digital security. For months, leading AI developers have implemented stringent vetting processes and robust safety protocols, aiming to create a secure environment for their advanced models. Yet, these measures, while well-intentioned, have inadvertently become significant obstacles for those tasked with safeguarding digital infrastructure and uncovering emerging threats.

The tension surrounding AI’s dual-use nature in cybersecurity recently came to a head in June when the U.S. government imposed export control restrictions on Anthropic’s highly anticipated AI models, Mythos and Fable. This move, reportedly driven in part by concerns over the models’ susceptibility to bypass safety mechanisms designed to thwart malicious cyberattacks, highlights the delicate balance between innovation and security. While the precise motivations behind the government’s action remain a subject of debate – with some questioning whether it was genuinely about an AI "jailbreak" or other geopolitical considerations – the practical impact is undeniable. Anthropic had actively marketed Mythos as a sophisticated tool with the potential for "doomsday cybermachine" capabilities, emphasizing the necessity of strict controls and vetting for its deployment. Consequently, the export controls, though since lifted for Fable 5 and partially for Mythos 5 (which has been reintroduced to vetted U.S. organizations under government review), underscored the perceived risks associated with such powerful AI technologies.

This pattern of restrictive access is not an isolated incident. Both Anthropic and OpenAI have established specialized programs for cybersecurity researchers, offering them the opportunity to apply for vetted access to models with fewer inherent restrictions. OpenAI’s "Trusted Access for Cyber" program and Anthropic’s "Cyber Verification Program" are prime examples, aiming to provide qualified professionals with the tools they need for offensive security research. However, these programs, and the guardrails they still impose, have drawn considerable criticism from researchers whose livelihoods depend on identifying and exploiting vulnerabilities before malicious actors can.

Mark Dowd, a prominent security researcher with decades of experience in discovering and selling "zero-days"—previously unknown software flaws and the exploits that leverage them—to Western governments, voiced his concerns during a recent cybersecurity podcast. He articulated a discomfort with "random large companies making arbitrary decisions about what is safe in security and what’s not." Dowd’s work, which involves withholding vulnerabilities from software makers to maintain their strategic value for intelligence operations, illustrates a distinct perspective on the cybersecurity landscape. Governments value these undisclosed flaws precisely because they remain unpatched, offering crucial advantages for espionage and defense. While acknowledging his own potential bias, Dowd’s sentiment is echoed by many in the offensive cybersecurity community.

Offensive cybersecurity professionals, whose role involves proactively probing systems for weaknesses, shared their experiences with AI tools and their inherent limitations. Chris Anley, chief scientist at the security consulting firm NCC Group, explained that using AI to attempt to exploit a bug is a critical step in validating its severity. However, he noted that when guardrails intercept such attempts, refusing to answer outright, they actively hinder defensive efforts. "This is where the whole offensive versus defensive and guardrails part comes in," Anley stated, "because ‘fix this code’ as a prompt is both an essential mechanism for defense but also a roadmap for finding critical vulnerabilities in the code base. So at the same time, the same tool is both an offensive tool and a defensive tool, and the two can’t really be unpicked." He likened AI’s utility to a hammer: indispensable for building but also undeniably capable of being used as a weapon. This duality means that restricting AI’s offensive capabilities inherently limits its defensive applications.

When faced with these AI-imposed roadblocks, Anley and his colleagues often resort to open-source AI models, which typically come without any guardrails, allowing for unrestricted exploration. Paolo Stagno, CTO at Crowdfense, a company that develops, acquires, and sells unknown vulnerabilities to government agencies, shared a similar sentiment. He characterized the AI companies’ approach as treating customers "like children who need babysitting" through their vetted programs and guardrails. Stagno and his team do utilize frontier models, but primarily for reverse engineering. They avoid using AI for vulnerability discovery or exploit development due to the risk of data leakage or the sensitive vulnerability information being absorbed into future training datasets when using cloud-based models. For these critical tasks, they opt for locally run, open-source models that ensure data remains within their control.

However, not all security researchers find guardrails to be a significant impediment. Giuseppe Cali, a security researcher specializing in zero-day discovery and exploit development, explained that he doesn’t employ AI for offensive purposes. Instead, he leverages AI tools for initial reverse engineering, to better understand the code he’s analyzing, and to build supporting utilities. In this capacity, AI accelerates his workflow, freeing him to concentrate on the core task of vulnerability discovery. "I still want to own the actual bug discovery and weaponization myself and that wouldn’t change if all guardrails were lifted tomorrow," Cali remarked. "I am jealous of my bugs, and I like this game too much to let models play it for me." His perspective highlights that for some, AI is a powerful assistant for foundational tasks, rather than a replacement for human ingenuity in the act of exploitation itself.

Despite these individual success stories, the broader impact of strict guardrails remains a significant concern. An anonymous researcher at a smartphone-component manufacturer, who spoke on condition of anonymity due to authorization constraints, revealed that their employer’s lack of participation in Anthropic’s Cyber Verification Program means their AI tools are "barely useful for finding vulnerabilities because the guardrails are too strict." This researcher further elaborated, "If it catches wind we’re doing anything security related, it just stops and isn’t usable." This anecdote underscores how restrictive policies can directly impede the security research efforts of even legitimate organizations.

Chris Thompson, CEO of cybersecurity firm RemoteThreat and founder of Offensive AI Con, an event focused on offensive security and AI, has observed firsthand the inconsistencies and fluctuating nature of guardrails, even within the more permissive frameworks of Anthropic and OpenAI’s vetted programs. "The practical impact is you spend a lot of time negotiating with the model instead of working on the core security program," Thompson stated. "Instead of analyzing a vulnerability and reasoning through the exploitability, you’re trying to find why you’re getting inconsistent results or why are models over-sanitizing the output." This constant "negotiation" diverts valuable time and cognitive resources away from the critical task of identifying and mitigating threats.

This frustrating experience is pushing researchers towards alternative solutions, often leading them to Chinese open-source models like GLM. These models are freely downloadable, can be run locally without any vetting or usage restrictions, and offer a degree of freedom absent in many Western-developed AI tools. "You have these responsible researchers that are being pushed away from U.S.-governed systems to foreign-owned systems," Thompson observed. He argues that the presence of overly strict guardrails is ultimately "more harmful than good."

Thompson advocates for a shift in approach, urging AI frontier labs to open up their programs, provide more responsible access, and hold users accountable for any misuse of their tools. He believes that an overly cautious approach, characterized by increasingly stringent restrictions, risks putting defenders at a disadvantage in the escalating AI-driven cyber threats. "There’s this big storm coming," Thompson warned. "There’s this big wave of attacks that are going to happen at speed and scale like never before. But the same security consulting firms and legit researchers that are trying to make a difference are being stifled right now." This sentiment encapsulates the growing concern that in the effort to prevent AI from falling into the wrong hands, the very individuals tasked with defending against cyber threats are being disarmed. The current dilemma underscores the urgent need for a more nuanced and collaborative approach to AI governance in cybersecurity, one that empowers legitimate defenders without compromising fundamental safety principles.

Leave a Reply

Your email address will not be published. Required fields are marked *