On August 21, 2026, a typically sleepy government agency published guidance on a wonky topic, an announcement largely ignored by the press, D.C. think tanks, and scholars. The National Archives and Records Administration (NARA)’s Guidance on Applying the Federal Records Act (FRA) to Artificial Intelligence (AI) Materials seemingly only affects the work of federal recordkeepers. However, this guidance, and its shortcomings, have major implications for the public’s understanding of—and trust in—how their government functions, and for how future generations will understand how this moment of profound technological change impacted how the government works. In this article, we analyze NARA’s AI guidance – the good, the bad, and what’s missing. We close by describing actions various stakeholders might take next.
Like the rest of society, federal employees are increasingly adopting generative AI tools, sometimes of their own accord and sometimes at the encouragement of supervisors. Federal use of AI can be highly consequential. The State Department is using AI to expedite visa revocations, raising potential First Amendment concerns. The Department of Veterans Affairs (VA) inspector general found clinicians using Microsoft Copilot for patient care, even though it lacked the guardrails necessary for high-impact AI tools (e.g., adverse impact monitoring). Military personnel reportedly used Anthropic’s Claude to support an operation to capture Nicolás Maduro. Further, Department of Government Efficiency (DOGE) staff may have submitted nonpublic federal data in prompts with Grok models in ways that raised conflict-of-interest concerns; used an OpenAI model to produce error-ridden analysis to identify VA contracts for cancellation; and deployed Meta’s Llama models to review employees’ responses to the infamous “fork-in-the-road” email.
In 1950, Congress passed the Federal Records Act (FRA), building on the Records Disposal Act of 1943, and reconfirming the need to preserve information concerning the “transaction of public business.” The Archives administers the FRA and other records laws, providing guidance on how they intersect with new technologies. As agencies began using email, social media, and messaging apps, NARA established pathways for their inclusion in the records regime. The most instructive example occurred in 2015, when NARA clearly declared that electronic messages of all types created in the course of agency business are federal records, regardless of whether federal employees use a personal account, even if an agency had to build new systems to capture that information.
In these examples, Congress, researchers, the press, and the public might want to know exactly which AI model versions were used, the details of procurement contracts that might prioritize confidentiality over government data, testing conducted before deployment, whether employees entered sensitive patient or government data into prompts, and if AI systems produced recommendations that, if followed, would otherwise violate the law. If federal agencies fail to create records related to their AI interactions, they effectively insulate their decisionmaking from oversight. Because the FRA governs what records agencies must keep, that information may not be recoverable. Once the data is gone, investigators cannot determine if an AI-driven visa revocation was based on prejudice, evidence, or an AI fabrication; Congress cannot assess whether contract cancellations are based on merit or improper influence; and courts cannot perform the “arbitrary and capricious” review required under the Administrative Procedure Act (APA) because the rationale—the AI interaction itself—has been erased.
With generative AI spreading quickly, the Archives has the opportunity to build on its history of telling agencies when and how to preserve federal information associated with a nascent technology that advances the public’s interest in better understanding its government. Unlike the messaging guidance from a decade ago, however, the new AI guidance is too narrowly focused, raising more questions than it answers.
What the NARA Guidance Does
The memo begins with the statutory definition of a federal record, citing 44 U.S.C. § 3301(a), information in any format, that is:
“made or received by a Federal agency under Federal law or in connection with the transaction of public business and preserved or appropriate for preservation by that agency or its legitimate successor as evidence of the organization, functions, policies, decisions, procedures, operations, or other activities of the United States Government or because of the informational value of data in them.”
This definition creates four possible categories: (1) information made and that is preserved, (2) information received and that is preserved, (3) information made that is appropriate for preservation, and (4) information received that is appropriate for preservation. The memo then notes that the context will determine what is considered a record, its “creation, maintenance, and use.”
The intent of the law—to preserve how the government is conducting the public business—and this broad set of categories would suggest that agencies should assume that most uses of AI could be records, or at least appropriate for a records expert evaluation to determine if they are records. But this is where the rest of the NARA guidance seems to deviate from the statutory definition and instead focuses on only one of the four possible categories: information that is received and already preserved. This framing significantly narrows how agencies will retain AI materials, determine what is a record, how they will write contracts with AI companies, retain information, and establish related processes.
The substantive portions of NARA’s recent memorandum to federal agencies consists of two parts: one clarifying what constitutes a federal record under the FRA and another addressing retention requirements.
Part I of the agency’s memorandum establishes that certain types of information are likely federal records, including user prompts and AI outputs, meeting summaries, chat histories (referred to as “audit trails”), procurement documents, copies of records used to train AI systems, and software code created by an agency or its contractors. While this is intended to be a non-exhaustive list, the rest of the guidance seems to treat it as exhaustive rather than exemplary. Further, it notably does not include broad use cases from national security agencies that are known, do not fit neatly into these boxes, and arguably warrant a much more prescriptive stance from NARA given their function. Part II focuses on record disposal. Under the FRA, agencies are not permitted to destroy anything considered a record without a NARA-approved records management schedule. Some information that might otherwise be a record is not permanent and is categorized as transitory (needed for less than 180 days), which can be destroyed when it’s no longer required; others are intermediary (used to create a subsequent record) and may be destroyed once that record exists. Other records are retained based on what they inform; for instance, AI prompts related to drafting an email should be retained on the same disposal schedule as emails generally. Critically, this section only relates to those items that agencies have already determined to be records, so concerns with the narrow nature of that category are upstream of the disposition guidance.
The Good
While we have a number of substantive critiques of this guidance, it is good it exists, presenting an opportunity for refinement. Though not flashy, records management is an indispensable part of rule of law and democracy. The guidance also establishes a useful framework of what AI materials might be in the context of federal records and where lines have to be drawn.
The Bad
Although the statute says that agencies must consider all information “made” or “received” that is either “preserved” or “appropriate for preservation,” NARA’s guidance takes a narrower approach. In describing what would likely constitute a record, it includes prompts, queries, and outputs from generative AI platforms with the caveat: “if they are captured and saved in an agency system.”
In this way, NARA takes a narrower stance than in previous comparable guidance, like how agencies should approach retention of records on cloud services or social media. In both cases, NARA guidance recognizes that occasionally information considered records will be held on a third-party server and need to be retrieved. In this guidance, NARA instead seems to suggest that before an agency can determine if information’s circumstances make it a record, it must actually possess the information. According to the NARA memorandum, AI material in a third-party system is not “received” by the agency until that information is captured in existing digital systems, effectively neutering the new guidance.
AI chatbots, like ChatGPT, Claude, Grok, Gemini, or Copilot, generally capture and store conversation history—prompts, outputs, chain-of-thought steps in reasoning models, metadata, and more. That information stays in those systems unless a user decides otherwise. All tools enable copying generated text or other media for pasting into other locations (e.g., Word document, email, social media post). Some tools let users share a link to the conversation or download their entire account history.
Under the NARA guidance, the first step for this information to be possibly considered a record is the user moving it to an existing digital system. That’s when the information is “received” by the agency, possibly triggering FRA obligations. Prompts and chatbot outputs are not considered agency records if they reside solely within a third-party AI system. This is a policy floor–agencies could go further–but few have an incentive to do so.
The word “receive” matters because the definition of “federal record,” defined earlier, turns on it. Received, in turn, is defined by NARA regulations as “the acceptance or collection of documentary materials by or on behalf of an agency or agency personnel in the course of their official duties regardless of their origin…and regardless of how transmitted (in person or by messenger, mail, electronic means, or by any other method).” On its own, this definition doesn’t limit records based on where information is stored. NARA confronted this issue when agencies migrated their information to cloud computing environments and used third-party software-as-a-service systems. In those instances, NARA guidance from 2010 clarified that records might reside in third-party computing environments.
The recent AI guidance took the opposite approach, clarifying in the definition of “received” that “Information created by AI and passively retained by a third-party platform such as ChatGPT or Claude is not necessarily ‘received’ by an agency, unless it is downloaded or otherwise captured in an agency system and used for official purposes.”
What’s notable about this provision is if an agency employee downloads or otherwise captures AI-generated information in an official system, that would likely be a record before this guidance. Viewed in that way, the guidance seemingly imposes no new responsibilities on agencies regarding the common use of AI. While this makes the guidance easy to administer for records managers, that policy design choice comes at the cost of transparency. It reduces public access to how agencies use AI, obscures decisionmaking processes, and limits insight into the efficacy of government technology.
It’s hard to imagine federal employees regularly copying over ChatGPT conversations in full, appending metadata and attachments, and saving that information as a Word document on a shared drive. More likely, employees will use AI systems to generate information for a document, memo, or other record, without any trace of having used AI.
Beyond where records reside, the guidance addresses other issues of note. Full chatbot account histories are only federal records if used to conduct official business, such as an investigation. The fact that an IT manager could access an employee’s chat history does not make it a federal record until it is used for an official function, which is an extension of the same logic of narrowing the definition of “received.”
The Missing
There are also several components of guidance that are absent or in need of more detail.
First, perhaps because of the overly narrow focus on information that has been received by agencies rather than information made and that should then be captured given its content, use, or application, the guidance does not include procurement guidance. Past NARA guidance on both cloud and social media specifically include draft language for inclusion in terms of service agreements. For example, both the cloud and social media guidance includes the following clause: “Managing the records includes, but is not limited to, secure storage, retrievability, and proper disposition of all federal records including transfer of permanently valuable records to NARA in a format and manner acceptable to NARA at the time of transfer.” This language—and the broader recommended clause it’s contained in—presumes that some records will need to be transferred to the government, that they may not yet have been received, but are appropriate for preservation. Further, the social media guidance includes a section specifically offering insight into how an agency could go about capturing records, including technical recommendations like “[u]sing platform-specific application programming interfaces (APIs) to pull content.”
Second, regarding source code for an AI application, the guidance states that software is a federal record when an agency creates it, pays a contractor to create it, or “significantly modifies” an off-the-shelf product. While bespoke agency-developed software is clearly a record and ChatGPT is clearly not, the line for “significant modification” is blurry. This leaves common use cases unresolved. Does that category include Meta’s Llama models deployed, and presumably fine-tuned, for DOGE to review federal employee resignation emails, or SweetREX, the Gemini-based tool used to flag regulations for elimination? Further guidance should clarify and provide case studies that speak specifically to this continuum.
Third, the guidance is silent on chain-of-thought reasoning, which advanced models often contain and which can illuminate how an AI system develops ultimate outputs. It also fails to address attachments, metadata (such as model versions and timestamps), and other data fields available in enterprise accounts that are in the public interest. There are likely cases where this data would be important to understanding how the government is conducting the public business. The evolution of record keeping to address the rise of government use of email in the early 1990s is instructive. In Armstrong v. Executive Office of the President, the court sided with the plaintiffs and ruled that hard-copy printouts of emails were insufficient to serve as the record because they did not contain all relevant context. As archivist David Bearman wrote, “[T]he printing of electronic mail messages would have resulted in the loss of structural and contextual information required to understand their significance, including the names of recipients and senders, the date and time of receipt, the link to prior messages, and full distribution lists.” In the same way, it is appropriate to assume that the secondary footprint of AI use may be appropriate for preservation.
Fourth, as noted earlier, the example use cases provided are unfortunately narrow and lacking in real-world complexity. Both case studies are about audit trails (chat histories). In the first, an employee has chat history enabled for her own later reference, never captures the material in an agency system, and never circulates or relies on it; NARA calls it a personal file. In the second, an agency that already retains audit trails for all employees captures one and uses it in an internal investigation, which NARA says makes it a record, though mere retention within the AI chatbot does not. That distinction is fine and mostly understandable, though one can imagine a number of scenarios where a prompt never captured in agency systems and not sent to anyone else would plainly constitute a record if developed in other contexts. But rather the concern is that these are remarkably narrow cases for a type of technology that warrants much broader consideration, particularly for that information that is “made” and “appropriate for preservation,” which raises the question of default preservation by vendors and procurement terms. Additional cases to address these prevalent but less obvious delineations could include: a prompt and output used to prepare military lethal targeting recommendations; a prompt used to produce recommendations for changes or cuts to a certain federal program or office; or the use of an AI agent to edit a legislative proposal prior to transfer to the Office of Management and Budget or Congress.
Finally, previous guidance has usually included, in a helpful question and answer format, an outline of the specific responsibilities the guidance creates for agencies. This guidance does not include that and should. Both the narrow framework provided by NARA, or what we would argue the more appropriate broader method of consideration, warrants an implementation plan from agencies. The guidance notes that there is not a one-size-fits-all approach, but that agencies will have to use their discretion to determine how the FRA and guidance applies. NARA—and Congress—should make the statutory requirements and their expectations more clear.
Things Could be Better
Our critiques do not stem from an unworkable ideal. Consider NARA’s 2015 guidance on electronic messaging. It applied broadly to instant messages, texts, and social media DMs, stating clearly that “electronic messages created or received in the course of agency business are Federal records.” It directed agencies to capture relevant info, train employees, and ensure metadata and attachments were preserved. It created a duty to capture information rather than looking for excuses to exclude it. The current AI guidance inverts this logic by deeming material in ChatGPT or Claude “not received” rather than establishing a duty to capture it.
Ultimately, several stakeholders can act to improve how records law responds to AI. Agencies could voluntarily exceed the NARA guidance floor. Certain litigants could sue NARA or agencies, arguing that the guidance fails to satisfy FRA’s statutory requirements. The Department of Justice’s Office of Information Policy (OIP) could provide Freedom of Information Act (FOIA) guidance that supplements NARA guidance. Finally, Congress could pass a law, as it has done in the past, to update record-keeping laws for the digital age, or start asking pointed questions about implementation of this guidance.
NARA’s AI guidance is easy to administer because it asks so little. By deeming material in third-party tools “not received,” it allows much of how the government uses AI to remain beyond the inspection of affected parties, researchers, other policymakers, and members of the public. This is a deliberate policy choice and does not reflect the best reading of the statute. The result is an accountability gap where it matters: a fast-moving decisionmaking tool, operating in domains where public oversight is necessary. Whether society closes this gap is now on agencies, courts, and Congress being willing to say the guidance isn’t enough.






