Illustration of hands typing on a keyboard with monitors displaying computer code in the background.

AI-Cyber Operations: A New Frontier for Public-Private Partnerships

This summer, even the most casual follower of the AI news cycle would not have missed AI agents going rogue and hacking into third parties during cyber evaluations. These incidents have helped expose the raw—and growing—capabilities of frontier AI systems, the ongoing lack of effective safeguards around them, and the current state of AI alignment (meaning, systems operating according to our goals without unsanctioned behaviors). They have also raised the question of whether progress on safety and security can keep up with capability improvements.

In the highest-profile AI-cyber incident of the past month, the OpenAI/Hugging Face case, a combination of AI models—including a prototype model intended only for internal use—identified and exploited numerous zero-day vulnerabilities, breaching the systems of innocent third parties in a dogged pursuit of a narrow evaluation task. OpenAI never intended for its models to escape its systems and attack other companies. Indeed, it took proactive steps to try to contain the models in a “sandbox” (a virtual jail cell) to avoid that possibility. Likewise, the third parties affected by the incident did not knowingly create vulnerabilities in their systems or intend for them to be exploited by an AI system or any other malicious actor.

From AI Cyber Accidents to AI Cyber Operations

But it is inevitable that highly cyber-capable AI systems will be deployed expressly to conduct offensive cyber operations, including through long-running agentic actions. There have already been examples of this, such as the incident highlighted in Anthropic’s November 2025 report, which claimed that a China-linked group used Claude’s agentic capabilities to conduct a sophisticated cyber espionage operation with minimal human input. In another incident that began in December 2025, several Mexican government agencies were attacked in a month-long Claude-assisted operation resulting in mass data exfiltration. Just this month, Taiwan announced that it had been the victim of an “abnormal” and “first-of-a-kind” AI-assisted cyber attack, suspected to have been conducted by China-linked entities. Analysis by the Israeli AI company Dream claimed the attack involved AI agents operating via open-source harnesses conducting reconnaissance of government systems, right through to data exfiltration, before pivoting the attack to the government IT supply chain.

Agentic AI cyber operations, whether or not intended by their operators, are a reality today. As the cyber capabilities of the underlying models and their scaffolding improve and proliferate via open releases (models which can be freely downloaded from the internet), the frequency, intensity, and sophistication of AI-cyberattacks will likely grow. In such circumstances, it is reasonable to assume that the government is unlikely to remain a bystander in this new frontier. Indeed, the foundations for its involvement are already being built. In May, the Pentagon announced that it had entered into agreements with eight companies in the AI ecosystem, including OpenAI, Google, and SpaceXAI. These agreements have paved the way for making these companies’ AI models and other relevant capabilities available to the Pentagon in classified national security systems, including those designated for TOP SECRET information and operations. In June, reporting suggested that Anthropic was supporting the NSA in deploying Mythos—a model so powerful it has not been made generally available—for offensive cyber operations.

Opening The Door to the Private Sector

To add to this, a Presidential National Security Memorandum issued this month has directed the creation of a program for authorized private-sector entities to conduct cyber surveillance and cyber effects operations, under federal government control and oversight, against foreign “Cyber-Enabled Transnational Criminal Organizations (CE-TCOs).” The definition of CE-TCOs included in the memo is of particular interest. Any foreign group conducting cyber-enabled crime against the United States government, a U.S. person, or U.S. interests, so long as it is not “an institutional part of a foreign government or wholly operated under a foreign government’s direction,” meets the CE-TCO definition. The memo states that the onus is on the intelligence community to provide clear intelligence showing that a CE-TCO is connected to a foreign government, and, if it cannot, that the group should be assumed to be fair game for private-sector-conducted cyber operations under the new program. The memo includes other guardrails, such as prohibiting operations that are likely to result in loss of life or serious injury or that are likely to reach the threshold for the use of force or an armed attack under international law. However, as has been argued elsewhere, calibrating cyber operations to remain below such thresholds is by no means straightforward. In sum, the memo invites American industry into an increasingly expansive, more opaque, and potentially higher-risk offensive cyberspace.

The program envisaged by the memo indicates that this administration is unwilling to accept a policy posture that it believes puts America on the back foot in the cyber domain. This, coupled with growing AI capabilities in cyber and long-horizon agentic tasks, the increasingly close relationships between the Pentagon and the AI industry in classified settings, and emerging indications that adversary-linked entities are using agentic AI for cyber operations, suggests the administration is likely open to private-sector entities conducting not just general cyber operations, but specifically AI-based cyber operations on the government’s behalf, including potentially via long-running agentic activity. Indeed, there is nothing in the memo stating that AI-cyber operations or agentic activities are out of scope for the new program.

Autonomous AI Cyber Operations Raise the Stakes

In May, my colleagues at the Institute for AI Policy and Strategy published a flagship report on Highly Autonomous Cyber-Capable Agents (HACCAs). These are AI systems that can “autonomously conduct cyber operations at the level of sophisticated criminal groups and possibly even intelligence agencies.” The benefits of increased speed, scale, and sophistication are clear for the attacker, but the report also points out that HACCAs can pose significant risks if operators and overseers lose control of them due to misalignment, exploitation by adversaries, or multi-agent failures. The recent third-party hacking incidents offer real-world evidence that misalignment and multi-agent failures are a present reality in the AI-cyber domain. They were warning shots, highlighting open challenges in a bounded way that we can recover and learn from at little cost.

The use of AI in cyber operations by or on behalf of the government could also go wrong if control is lost in the ways that my colleagues outlined in the HACCAs report. If this were to happen, the stakes could be much higher than in the OpenAI/Hugging Face case and similar incidents. Rogue AI agents explicitly tasked with conducting cyber operations against real targets could lead to escalatory dynamics that boil over into direct military confrontation, off-target effects that cause collateral damage, unintended interference with other American cyberspace operations, and greater disruption or damage than anticipated. Such outcomes could prompt adversary responses that are much more challenging to recover from and carry lasting political, diplomatic, social, and economic consequences.

Guardrails for Private-Sector-Conducted AI Cyber Operations

As the implementing guidance for the program directed by the memo is developed over the coming weeks, and the program matures thereafter, the National Coordination Center and the Program Executive Directors designated in the memo should proactively ensure that targeted, proportionate, and mission-enabling guardrails are in place if AI-cyber operations are to be brought within the program’s scope. This is especially important if such operations will run agentically with minimal to no human input once initially tasked. At minimum, these guardrails should include the following:

  1. Comprehensive pre-deployment testing and evaluation of AI systems that are candidates to be used as part of the program in realistic, representative, and highly secure “cyber ranges,” digital environments that simulate real computer networks to test AI cyber capabilities without damaging actual systems.
  2. Private certification of the AI systems’ fitness (from senior officials) and appropriateness (from political principals) to operate as part of the program, with notifications to congressional defence, intelligence, justice, homeland security, and foreign affairs committees when certification decisions are made (irrespective of the outcome), together with the reasoning behind them.
  3. Robust, effective, and scalable real-time monitoring technology during testing, evaluation, and deployment of in-scope AI systems.
  4. Technical capabilities built into in-scope AI systems ahead of testing, evaluation, and deployment to allow operators to intervene on, constrain, or shut them down if needed, together with appropriate training.
  5. Detailed, timely, and tamper-proof logging that allows for prompt review and remediation of any incidents connected to in-scope AI systems during or prior to deployment, while also laying the foundations for accountability measures.
  6. A narrowed target set for AI-cyber operations under the program, to include only entities where the benefit of action clearly outweighs the risks, including those arising from the failure modes I described above.
  7. Stronger penalties for contractual non-compliance, negligence or recklessness by private-sector entities if approved AI systems are used in any cyber-operations they conduct under the program.
  8. More stringent reporting requirements for AI-cyber operations that have exceeded any predefined parameters, restrictions, or operational scope. This should include the NCC disclosing to the congressional defence, intelligence, justice, homeland security, and foreign affairs committees any information it receives about such incidents, with the maximum possible public transparency.

If techniques for controlling and aligning AI systems—especially in multi-agent settings—improve, these guardrails could be adjusted over time. But the evidence before us about the challenges of current AI-cyber agents justifies a precautionary approach for now.

Securing AI Cyber Advantage Responsibly

Our adversaries will not wait to weaponize AI-cyber capabilities against America and its allies, and the government is unlikely to stand still in response. However, the right prescription at this early stage is responsible, judicious, and proportionate use by government and industry, with safeguards matched to the state of AI-cyber capabilities, the level of agentic operations envisaged, and the robustness of safety measures. Adopting this approach will enable America to claim a leading position in this new arena and ensure it gains the legitimacy that long-term engagement in it will require.

Filed Under

, , , , , , , , ,
Send A Letter To The Editor

DON'T MISS A THING. Stay up to date with Just Security curated newsletters: