My earlier piece on “AI and the Commercial Data Loophole” focused on the risks of operating LLMs on the government’s purchases of Americans’ data from commercial brokers. But the Pentagon’s commercial purchases are only one part of the picture. Foreign intelligence surveillance programs also sweep up vast volumes of Americans’ communications without a warrant. The government does not treat these programs as “mass domestic surveillance” because they are not targeted at U.S. persons. But these programs result in the accumulation of enormous quantities of Americans’ information, including the content of their communications. Large language models deployed against this data would expand what the government can do with it, erode the rules meant to constrain it, and undermine the tradecraft standards that are supposed to ensure it produces reliable intelligence.
Two foreign intelligence authorities do most of this work. The best known is Section 702 of the FISA Amendments Act, which authorizes the targeting of specific foreign persons and entities overseas but inevitably sweeps up the communications of Americans who are in contact with them. The Foreign Intelligence Surveillance Court (FISC) approves the procedures governing the program but plays no role in deciding whom the NSA actually targets. Section 702 lapsed in June 2026, due to some lawmakers’ concerns over an intelligence-community nomination and longstanding differences among lawmakers on the extent to which the law needed reforms. Efforts are underway to reauthorize the law; in the meantime, the program continues to operate under earlier approvals from the FISC. The second, Executive Order 12333, is less well-known, but is broader still. It permits bulk collection of communications and data overseas without identifying any specific target, and the collection it authorizes is not subject to any judicial supervision whatsoever.
LLMs amplify both programs in ways that pose grave risks to Americans’ privacy and civil liberties, while eroding the efficacy of rules that were supposed to constrain those risks.
FISA Section 702
The NSA labels the collection of Americans’ communications under Section 702 “incidental.” It has steadfastly refused to quantify the volume of American communications swept up, which may be so large as to strain the “incidental” characterization. 702 information does not stay in the foreign intelligence realm, but seeps into domestic law enforcement through the Federal Bureau of Investigation. The Bureau has queried databases that include 702 communications for information about U.S. persons millions of times, often ignoring rules meant to constrain the scope of such searches. The FISC found the FBI’s violations to be “persistent and widespread,” including searches for people arrested during protests and checks on political donors.
Congress has not addressed the integration of AI into 702 surveillance programs, and to the extent the FISC has done so, its analysis remains classified. But as with commercial data, LLMs could magnify the reach and intrusiveness of this surveillance program.
First, AI-assisted targeting could be used to greatly expand the NSA’s pool of foreign targets (which in 2025 approached 350,000 individuals) and thus the volume of incidentally collected Americans’ communications. The scale of the NSA’s collection of Americans’ information in the context of Section 215 (the bulk phone-records program) and in the context of “abouts” collection (a subset of Section 702 surveillance which captured communications that merely mentioned a target) contributed to the discontinuation of those programs. But because Congress has allowed the agency to refuse to quantify the extent of its U.S. person collection, the NSA’s collection of even exponentially more information about Americans would not necessarily come to light.
Second, AI adds to the already substantial civil liberties risks of the Section 702 program. As early as 2013, the NSA was using sophisticated tools to search and analyze the communications it intercepted. LLMs add the ability to synthesize thousands of documents and interpret colloquial or allusive language, bringing knowledge gained from their training data to bear on communications that prior tools could read only in isolation. A model’s interpretations could prove valuable to the agency’s foreign intelligence mission. The problem is that the NSA’s database also includes a vast volume of Americans’ emails and phone calls. And as explained in my earlier piece on “AI and the Commercial Data Loophole,” the foreign intelligence purpose requirement may not serve as an effective safeguard. Rules limiting the purposes for which this database may be queried are supposed to protect Americans but have been repeatedly violated by the FBI. And minimization procedures of the type that the FISC has approved to protect Americans’ privacy are, as with the rest of the procedures discussed in this piece, substantially undermined by LLMs.
Executive Order 12333
As with Section 702, the government places certain E.O. 12333 programs outside the envelope of “domestic” surveillance because they are not targeted at U.S. persons. Section 702 is visible to the public. It must be periodically reauthorized, generating Congressional debate and press coverage. Some additional information about the program comes from declassified FISC opinions and regularly published transparency reports. In contrast, the Defense Department’s collection of foreign intelligence under E.O. 12333 remains firmly in the grip of the executive branch. These programs are rarely debated in Congress and are not subject to any judicial supervision or transparency reporting. What is known comes largely from the Snowden documents and the occasional disclosure elicited from the government by persistent senators.
Despite this dearth of information, it is almost certain that the volume of information collected under E.O. 12333 dwarfs the scale of the Section 702 program. This is because, unlike Section 702, E.O. 12333 permits bulk collection. Section 702 surveillance starts with a specific foreign target, using a “selector” such as a phone number or email address. E.O. 12333 doesn’t require a specific target or selector. Agencies can collect without a target “when necessary due to technical or operational considerations.” And they have, vacuuming up entire datasets, often from internet and telecommunications infrastructure, such as all calls into and out of a country or all data transiting a foreign data center. Under a program codenamed SOMALGET, the NSA recorded and stored the content of virtually every mobile phone call made in the Bahamas for a rolling 30-day window. The Central Intelligence Agency has run bulk collection programs that swept up Americans’ financial records as well as information about Americans in contact with foreign nationals.
Bulk collection under E.O. 12333 has long been criticized by civil society groups (including the Brennan Center) and former government officials for lack of transparency, its impact on Americans’ privacy, and the lack of any judicial supervision. Because E.O. 12333 collection operates at such scale, applying LLMs to that data exposes more Americans to their analytical capabilities. And the rules meant to minimize the impact of this collection of information on Americans’ privacy are weakened when these models are deployed.
The minimization rules agencies have adopted for information collected under E.O. 12333 broadly parallel the rules approved by the FISC for Section 702: limits on how long unreviewed data can be retained, masking of U.S. person identities, access controls, and limits on the purposes for which searches can be carried out. But these protections are generally less robust. For instance, unlike in the Section 702 context, the FBI is not subject to specific querying rules when searching for Americans in E.O. 12333 data. It must only abide by a general admonition that its activities serve a “valid purpose” and comply with applicable law, including the Constitution. Moreover, E.O. 12333 rules are written and supervised entirely by the executive branch. They vary from agency to agency and are subject to no judicial oversight, a vulnerability that becomes especially clear when those agencies are headed by appointees inclined to use them aggressively. The concerns described below apply to minimization under both authorities, but they are most acute under E.O. 12333, where the rules are thinnest and the oversight is weakest.
Minimization rules—which were already quite weak and riddled with exceptions—are rendered even weaker by LLMs.
Retention
To start, limits on retention operate on the assumption that most information about Americans picked up in foreign intelligence surveillance will be deleted after five years with some exceptions. The information may be kept longer if a senior official determines in writing that it serves an authorized foreign intelligence requirement.
LLMs put pressure on this limit in several ways. They dramatically lower the practical processing ceiling that constrained retention. A model can process enormous quantities of collected communications and generate foreign intelligence justifications at a scale no human workforce (or even earlier generations of processing tools) could match. So long as the designated official signs off on the foreign intelligence justifications put forward by the model, the underlying information can be kept indefinitely. But vast numbers of foreign intelligence justifications may dilute the efficacy of the sign-off requirement as a check. Even if LLMs are not used to generate foreign intelligence justifications, analysts may be more inclined to mark material as having potential foreign intelligence value on the theory that the model may be able to extract connections that a human reviewer would not see. Together, these dynamics undercut the privacy protective function of retention limits. The five-year clock would still run. But far less data may ever reach it because it would be marked as reviewed and of potential foreign intelligence value.
Masking
Minimization rules assume that masking Americans’ identities in disseminated reports will protect their privacy. In intelligence reports, an American’s identity is typically replaced with a generic designator (e.g., USP1). An analyst who wants to learn the identity of such a person must formally request unmasking from the originating agency, which evaluates whether the requestor has a legitimate need to know the identity of the person referenced. The request is documented and the decision is logged, providing some procedural protection.
LLMs can circumvent this process by the same kind of re-identification demonstrated in the studies highlighted in my earlier piece on “AI and the Commercial Data Loophole”: A model asked to analyze a set of intelligence reports can recover a masked identity with no specific request and no obvious paper trail. Despite this lack of safeguards, the LLM’s inference may shape what happens next: which databases get queried, which leads get followed, which information finds its way into an FBI investigation or an immigration proceeding.
Querying
The possibility of contextual inferences and incomplete paper trails also puts pressure on querying-based restrictions. In the Section 702 context, for instance, restrictions on retrieving Americans’ information are triggered when analysts run U.S.-person search terms. Someone must type an identifier (e.g., a name or email) to reach a particular American and each term is recorded, creating an auditable trail. But an LLM matches on meaning, not exact words: asked who is organizing a protest, it can surface a person whose own messages never use the word “protest” and without anyone entering that person’s identifier. It can even generate and refine its own search terms, or run them through an agent. In sum, a model can effectively evade limitations on accessing specific Americans’ information by accessing all Americans’ information.
Reliability
The problems with the rules are compounded by a problem that runs across every use described in this piece: the tools themselves are unreliable. It is well-established that LLMs hallucinate, producing confident, plausible-sounding outputs that are factually wrong and even fabricated. It is equally well-established that LLM’s outputs reflect biases in their training data that cannot always be identified or corrected.
Risk and impact assessments of the kind contemplated by the rescinded Biden framework—and likely to feature in whatever replaces it—are meant to address these concerns. They may not, however, be sufficient for high-stakes predictive uses. Verifying a model by seeing how it operates in the field may work when there is an objective truth against which to measure its performance (e.g., the model correctly identifies a military facility for targeting during a training exercise). But if an LLM flags someone as a potential threat who never carries out an attack, we can’t know whether the model generated a false positive or intervening events prevented an attack. Conversely, if the model doesn’t flag someone who subsequently carries out an attack, we can’t know if it produced a false negative or whether the behavior was genuinely unforeseeable. There is no fully observed ground truth against which to validate predictions about human behavior.
Another way to evaluate reliability is by examining training data and model weights. When an LLM generates erroneous outputs, whether in testing or the real world, access to this information can help a customer agency understand what went wrong and how to fix the model or system. Ideally, the provenance and sourcing of the training data should also be available, so that any biases and gaps in the data could be examined. But commercial vendors generally treat this data as proprietary and may not be willing to share it—even with the government. Even where such data is available a more fundamental limit remains. As AI companies concede, there is currently no reliable way to look inside the model when it is performing a task and check that it is reasoning correctly. As a result, agencies may not be able to reliably detect where the model breaks down or predict where it might err in the future and generate inaccurate outputs.
The Intelligence Community has acknowledged that LLMs may strain the tradecraft standards meant to ensure the reliability of intelligence products, directing that these tools should be designed to let personnel adhere to standards including Intelligence Community Directive (ICD) 203 and ICD 206. This high-level admonition does little to address the actual problem. Those directives require analysts to document the provenance of each claim in finished intelligence reports, describe the quality and credibility of underlying sources, and distinguish what is known from what is inferred. ICD 206 goes even further. It requires that a finished intelligence report include sourcing information so that a reader can independently pull and check sources. It is hard to see how an LLM can meet these standards. Because an LLM generates outputs from statistical patterns across an enormous training corpus rather than drawing on discrete sources, a given statement cannot be reliably traced back to the material that produced it. LLM outputs are also unstable. Asked the same question more than once, a model may give materially different answers. There may be no fixed result to trace back to sources. Its statements therefore may not be tied to specific, credibility-rated sources in the way intelligence tradecraft standards require.
Even where an LLM is used only to generate inputs—drafts, leads, candidate hypotheses to be checked—analysts may over-rely on what it produces. Automation bias studies suggest that analysts may be inclined to place too much credence in LLM outputs, which are presented in a confident and reasoned register even when the underlying evidence is weak.
A number of technical measures might mitigate some of these concerns (retrieval-augmented generation, for example, lets a model draw on and cite a curated set of sources rather than generating from its training data alone) but their effectiveness in this setting remains unproven, and the government has not disclosed whether it has adopted any.
Conclusion
As with its purchases of commercial data, the government places its collection of Americans’ information and communications under foreign intelligence surveillance authorities outside the “mass domestic surveillance” envelope. But these authorities, too, allow the government to sweep up vast quantities of Americans’ information. And unlike the commercial purchases, which are largely metadata, they routinely capture the substance of communications. The introduction of LLMs heightens the risks of this accumulation of data in the hands of the Pentagon because it can be used to create detailed profiles of beliefs, associations, and behavior at scale, which could be turned on the administration’s perceived foes. LLMs further compound those risks by eroding the rules—on retention, masking, and querying—that are supposed to protect Americans.
The Government Surveillance Reform Act, which the Brennan Center supports, would remedy many of the serious and longstanding failures of the current framework for foreign intelligence. But lawmakers must also address the use of LLMs against Americans’ data. Executive branch policies are easily rescinded, as was the case with the Biden administration’s National Security AI framework. In any event, even that framework did not grapple with how rules meant to protect Americans are undermined by LLMs. At a minimum, Congress should require agencies to disclose how these models are trained, how they perform, and how they impact Americans and to demonstrate how retention limits, masking, and querying rules continue to function when these models are in use. Otherwise, the safeguards meant to protect Americans will exist only on paper, describing rules that no longer constrain the surveillance actually being conducted.






