The logo of the AI assistant "Claude Mythos".

When Intelligence Becomes a Strategic Asset: Frontier AI, Sandbox Breakouts, and the Case for Independent Oversight

In recent weeks, leading AI executives have taken an increasingly urgent stance on pacing the frontier of Artificial Intelligence. But the hard intersection between development and governance is, of course, not new.

In June, leading AI developer Anthropic found itself at the center of a controversy that would have seemed improbable just a few years earlier. Following Anthropic’s release of Claude Mythos 5 and Claude Fable 5, government officials and national security experts intensified existing debate over whether access to frontier AI models should be treated as a matter of national security. 

The release of these models marked a turning point in the history of AI governance. Beginning in April, Anthropic introduced its newest AI model, Mythos, to a limited group of trusted government and industry partners after determining that its cybersecurity capabilities were too powerful for general release. Anthropic voiced concern over the model’s ability to accelerate vulnerability discovery and exploitation that previously required highly specialized, nation-state level expertise. Two months later, Anthropic launched Fable 5, a public-facing version of the same underlying model equipped with additional safeguards, while Mythos remained restricted to vetted organizations. 

However, within days, the U.S. government applied export controls on these models over concerns that foreign adversaries could jailbreak and exploit the models, compelling Anthropic to suspend access worldwide. After nearly three weeks of negotiations, the restrictions were partially relaxed for approved U.S. organizations before being lifted more broadly once Anthropic agreed to enhanced security protocols and closer coordination with federal authorities. 

The episode was the first time Washington had effectively taken a frontier AI model offline, demonstrating that advanced AI systems had crossed an important threshold from commercial software to technologies viewed as strategic national assets whose capabilities can materially affect national security and geopolitical power. Nations around the world are now considering independent oversight capable of testing not only frontier models, but also the environments meant to train them, without sacrificing the speed that increasingly defines strategic competition. Effective oversight and strategic speed need not be mutually exclusive.

The Sandbox Breakouts

The Mythos moment marked the beginning of a new frontier for leading AI models and cybersecurity. In July, OpenAI disclosed an “unprecedented cyber incident” during an internal evaluation of advanced cyber capabilities. Several models, including GPT-5.6 Sol and a more capable pre-release model, placed in a controlled environment searched for ways to cheat the system, discovering previously unknown vulnerabilities. The models successfully escaped their testing sandbox and hacked into the external servers of the machine learning platform Hugging Face. The evaluation had crossed the boundary separating a simulated cyber exercise from a real-world intrusion. 

OpenAI was not alone. Later that month, Anthropic revealed that Claude models had accessed live external systems during testing due to unintended internet access. Meta subsequently disclosed that one of its models also reached an outside company’s systems during a cybersecurity evaluation. These incidents are happening in China as well. Most recently, Moonshot AI’s open-weight Kimi K3 found weaknesses in its sandbox during cyber testing, obtained internet access, and searched an outside site for data. 

The broader significance of these events lies in what they collectively reveal about frontier AI testing and AI governance. Advanced models can now identify vulnerabilities and propose exploitation pathways at a scale that human analysts cannot match. Activities that previously required large teams of elite cyber professionals are now accessible to anyone with an internet connection. 

Until recently, leading frontier laboratories leveraged their in-house expertise to independently determine what mitigation steps were necessary to address these concerns and evaluate whether their models were ready for public deployment. Allowing broad government access to, and control over unreleased models with advanced cyber or biological capabilities, would create its own security risks. This problem is compounded by competitive pressure which rewards speed and technological innovation at the risk of safety. Military strategists have long understood the importance of speed. Few concepts have proven more influential than Air Force Colonel John Boyd’s OODA Loop: Observe, Orient, Decide, Act, which greatly increased the speed at which military operators were able to make decisions. 

Frontier AI systems compress multiple stages of the decision cycle simultaneously. What once required a sequence of human actions can increasingly unfold through interconnected systems operating at machine speed. As Bruce Schneider has argued, contemporary security environments reward actors capable of acting first. The United States and China both recognize the strategic advantage of being first to develop artificial superintelligence, and have respectively invested billions in acquiring that capability. This first mover advantage is also driving massive investments in the private sector.

For decades, governments designed political and military institutions to slow decision-making and reach better outcomes. Deliberation and verification were mechanisms for reducing miscalculation, because strategic competition has traditionally rewarded better decisions. But in the AI age, first movers may be the ones reaping the rewards. 

Sovereign AI: When Intelligence Becomes a Strategic Asset

The debate surrounding Mythos and Fable revealed that frontier AI may be entering a stage requiring exclusive government control. States have historically regulated technologies capable of altering the balance of power between nations, from nuclear and stealth capabilities to precision-guided weapons. When access to a technology could materially influence military effectiveness or geopolitical influence, governments concluded that market forces alone could not determine its distribution. 

Frontier AI represents a fundamentally different challenge than previous state-controlled technologies. Earlier technologies were constrained by physical scarcity and the ability to legally and institutionally restrict the distribution of protected technical information. Nuclear weapons, for example, depend on scarce material and specialized infrastructure, while critical information about their design was tightly controlled. Frontier AI increasingly challenges these assumptions, as model weights are digital artifacts that can be replicated at minimal cost, and as open-weight systems approach the capabilities of proprietary frontier models, governments may find that controlling chips and compute does not provide enough control over the underlying capability. Tasks that previously required large organizations staffed by specialists can now be performed with frontier models, forcing governments to consider how to control access to a strategically significant technology that is inherently easy to distribute. This echoes Gregory Allen’s argument that U.S. leadership in artificial intelligence is directly intertwined with national power. The National Security Commission on Artificial Intelligence reached a similar conclusion, warning that advances in AI would influence military competition, economic leadership, and geopolitical influence. 

This shift helps explain the growing emphasis on concepts such as maintaining a “small yard, and high fence.” The challenge lies in defining where that boundary should exist. Semiconductors, advanced computing infrastructure, and frontier AI models consolidate complementary components into a single strategic ecosystem. Advanced chips enable ever more capable models, which accelerate scientific discovery, generating demand for still greater computational capacity. Each reinforces the other in a self-amplifying cycle. Restricting access to one influences the strategic value of the others. From a distribution regulatory perspective, the objective is to preserve leadership over the relatively small number of technologies that determine who develops the next generation of frontier capabilities.

As access to frontier cognitive capability becomes a matter of national interest, governments must confront questions that previous generations reserved for technologies such as nuclear weapons. Who should have access to the most capable models and under what conditions? When does a commercial AI system become a strategic asset whose distribution must be regulated?

Regulatory Oversight: Establishing an Frontier AI Standards Body

The policy debate is already moving toward independent oversight of frontier models. In July, Google DeepMind CEO Demis Hassabis proposed a new Frontier AI Standards Body, which Treasury Secretary Scott Bessent has endorsed. This proposal envisions an industry-funded but federally supervised institution modeled partly on the Financial Industry Regulatory Authority (FINRA), where frontier laboratories would submit models for independent assessment prior to release. Since then, the idea has gained momentum with Anthropic CEO Dario Amodei calling for independent evaluators to work inside frontier laboratories with direct access, while leaders at OpenAI and other major AI companies have expressed support for greater outside evaluation. Those concerns gained additional weight in September, when Anthropic disclosed that it had disrupted scientists’ use of Claude in research that could support biological weapons development, including work involving gain-of-function research. The incidents highlight the danger of increasingly capable systems able to circumvent the safeguards that currently prevent access to dangerous biological manipulation. Anthropic and Accenture have since committed over $2 billion over five years to build independent evaluation capacity. Meanwhile, the bipartisan FRONTIER Act introduced in Congress would require substantial independent oversight for advanced AI developers. 

The federal government also has pieces of this architecture already in place. The National Institute of Standards and Technology’s Center for AI standards and Innovation conducts evaluations of U.S. and foreign frontier models, and works with industry on voluntary standards. A frontier standards regulatory model could connect these efforts, where CASI could provide the government with technical expertise, while accredited independent evaluators conduct assessments against common standards that evolve alongside model capabilities. This framework could establish capability thresholds for heightened scrutiny and common requirements for cybersecurity and other dangerous-capability evaluations.

The summer of sandbox incidents suggest that the standards body’s mandate should extend beyond evaluating models to include the evaluation environments themselves. An assessment provides little assurance if the model can manipulate the test. There should be clear requirements to determine if a testing environment is capable of containing the model being evaluated. Standards should specify minimum requirements for isolation, network access, and credential management before dangerous evaluations begin. 

An independent oversight model also creates its own risks. An institution funded primarily by the companies it regulates could become susceptible to regulatory capture, while embedded evaluators may struggle to remain independent from the laboratories they are engaged with. Allowing frontier laboratories to play a major role in establishing initial assessment protocols could permit incumbents to shape standards that favor their own architectures. Evaluators will need protected access to relevant systems and information as well as incentives to report unfavorable findings. Open-weight models also create an additional challenge because developers lose control once their weights are publicly released. Finally, a volunteer model will need to avoid incentive structures that penalize the most responsible actors. For the most advanced models, it is likely that voluntary compliance will eventually become mandatory, requiring baseline standards for evaluation before release. 

Any regulatory framework must also account for the strategic geopolitical environment in which AI is developing. A U.S. AI standards body will only have oversight of the laboratories that submit to it. It says nothing about whether China’s frontier developers face anything comparable, and the public record does not yet support confident claims of how rigorously Beijing’s labs are testing their own systems. What is verifiable is the asymmetry in incentives: although not entirely consistent, at various key moments Washington has chosen to build friction into its AI development at the same moment that competitive pressure argues for removing that friction on the assumption that rivals elsewhere are making a different calculation. A U.S. framework that constrains domestic development without affecting foreign competitors could inadvertently strengthen China’s relative position. Effective oversight must balance safety against strategic competitiveness, introducing safeguards where the risks are justified, without creating regulatory hurdles that allow less constrained competitors to pull ahead. 

China Competition

For much of the post-WWII era, governments have spent decades establishing procedures intended to slow escalation to reduce miscalculation, and create opportunities for reassessment during crisis. The Washington-Moscow hotline, established after the Cuban Missile Crisis, became the enduring symbol of that effort, reflecting the principle that when technology compresses decision time, institutions must deliberately create space for human judgement. Throughout the Cold War, layers of verification procedures were designed to ensure that actions remained political decisions even as military systems became capable of acting with unprecedented speed. 

The current pace of AI competition has rewarded the elimination or reduction in safeguards to accelerate the decision cycle. Systems capable of acting within seconds create constant pressure to delegate greater authorities to machines. If one competitor believes human intervention introduces unacceptable delay, the incentive for the other to preserve deliberate decision-making weakens as well. What begins as an effort to gain a tactical advantage can become a race to remove the safeguards that previous generations considered essential to avoiding catastrophic miscalculation. 

The U.S.-China competition is mutually reinforcing this technological acceleration, driving competitors toward ever-shorter decision cycles regardless of whether that compression improves judgement. Chinese military thinkers have spent more than a decade exploring concepts associated with “intelligentized warfare,” a framework that envisions integrated systems linking sensing, automated AI analysis, and operational execution. China’s open-source DeepSeek AI models punctured the assumption that American AI leadership was secure through access to superior chips and capital. For the first time, a Chinese team had produced a system competitive with leading Western models while claiming to use fewer resources and making it broadly available. The startup Moonshot AI has reinforced that Chinese firms are capable of narrowing the frontier between Chinese and U.S. leading AI models. 

China’s technological strategy is directly integrated with the country’s institutional approach. In July 2026, China and 28 other governments signed an agreement establishing the World Artificial Intelligence Cooperation Organization, a Shanghai-based intergovernmental body intended to promote Chinese AI models, and expand China’s approach to AI oversight. Beijing has presented the organization as an alternative to a global AI order it portrays as overly concentrated in a handful of American companies and Western alliances. China’s approach prioritizes speed above all else for operational advantage. For example, Chinese military doctrine focuses control on selected senior leaders. Chinese AI investment and approaches aspire to allow for the most senior ranking individuals in the government to make decisions at the tactical level, not just strategic. By contrast, the United States is pursuing many of the same goals through initiatives such as Joint All-Domain Command and Control while preserving stronger institutional commitments to human oversight, known as the human in the loop model. As U.S. military doctrine and guidance evolves, some of the same pressures on speed will come into question. The need for accelerated AI-supported decision making will meet with political realities at the cross section of responsible AI principles.

That tension is increasingly visible in the United States, where The Department of War’s 2026 AI strategy explicitly calls for an “accelerated pace of Military AI integration” including an Agent Network intended to apply AI to battle management and decision support “from campaign planning to kill chain execution.” That disconnect in acceleration priorities complicates the contrast between Chinese acceleration and American caution. Earlier U.S. defense policy emphasized responsible-AI principles and human judgement, but the current military strategy places substantially greater weight on speed and operational adoption. 

Strategic competition creates similar pressures on both the United States and China even when governance systems differ. As AI continues to compress decision timelines, military advantage may increasingly appear to depend on removing humans from the decision cycle, creating a structural incentive danger where each country’s efforts to avoid being slower than the other could progressively narrow the circumstances in which either is willing to preserve meaningful human deliberation. 

These different approaches to integrating AI into military decision-making will greatly impact strategic competition and global stability. Dmitri Alperovitch has argued that future competition between major powers will depend on technical innovation as well as the ability to integrate emerging technologies into operational concepts. But the current landscape of frontier AI competition increasingly involves rival systems operating at different levels of acceptable risk. Consequently, future advantages may emerge not only from the most capable models, but also from the side best able to balance speed and judgement — or perhaps from the side most willing to maximize speed at the risk of giving up control. During routine operations, these differences could appear manageable, but during a crisis, they could shape how each side evaluates risk and responds to unfolding events. This sentiment is currently on full display, with faulty intelligence reports produced by AI causing operational confusion for the United States and a Chinese cargo vessel. 

The primary challenge for the United States and its allies is designing regulations that preserve the benefits of human judgement without surrendering the speed required to remain competitive. Human oversight cannot consist of a commander mechanically approving recommendations that are generated faster than they can reasonably evaluate them. This requires a structure that elevates human intervention at the moments where judgment matters most, while allowing automation to manage the large universe of routine sensing, analysis, and execution beneath it. 

Conclusion: Designing Friction for and Accelerated World

The lesson of the first sovereign AI crisis extends beyond Mythos or any particular controversy and exposes the fact that frontier AI is going to be regulated one way or another. The central question is whether the proposed Frontier AI Standards Body and future AI regulations can preserve human judgement within environments increasingly optimized for speed. If properly designed, this would reduce the risk of poor decisions in moments of crisis — exactly the dynamic revealed during the Mythos episode. But if designed poorly, it risks becoming a bureaucratic bottleneck that slows innovation without improving security. That balance will be especially important in competition with China, where poorly designed safeguards could inadvertently shift the technological advantage toward Beijing. 

The Mythos episode marked the moment the U.S. government began treating frontier AI as an instrument of national power. The next challenge will be ensuring that the institutions responsible for governing these systems evolve as rapidly as the technology itself. As these issues become more front and center for the American and world public, sentiment will continue to change in response to both AI events as well as the thought leadership of governments around them. The decisions made over the next few years will shape not only who leads in AI, but whether democratic societies can preserve human judgement in an era increasingly defined by machine-speed competition. 

Filed Under

, , , , , , , , , , , , , ,
Send A Letter To The Editor

DON'T MISS A THING. Stay up to date with Just Security curated newsletters: