Who Writes the AI Constitution?

Over the past year, the federal government has grown intensely interested in the values embedded within and expressed by artificial intelligence (AI) models. Across executive orders, Office of Management and Budget (OMB) procurement rulesFederal Trade Commission actions, and Department of Justice interventions, the federal government has engaged in an expansive effort to dictate the values on the models millions of people use every day. Many state governments have likewise considered or passed legislation concerned with AI values and biases.

One avenue for potential government intervention into the lab’s values-selection process has emerged: the labs’ respective documents spelling out the values baked into their models. We refer generally to these as “AI constitutions.” AI constitutions are long documents outlining the values, ethics, and character that models should have. Written predominantly by the employees at the labs deploying these models, albeit with some consultation with external stakeholders, these documents increasingly serve as the central locus of AI values discourse. In fact, Lawfare recently shared a research agenda inviting inquiry into the values and processes behind the various AI constitutions being crafted and implemented.

Anthropic publishes Claude’s Constitution, “the foundational document that both expresses and shapes who Claude is.” OpenAI maintains the Model Spec, which “outlines the intended behavior for the models that power OpenAI’s products.” As detailed further below, these are neither mere mission statements nor pure governance frameworks. They play a functional role at various stages in the model development process, such as what data the model trains on, what examples it’s fine-tuned on, and what behavior it’s rewarded for. Anthropic reports that its Constitution “directly shapes Claude’s behavior.” An AI constitution is therefore both a published statement of a company’s values and an operative control mechanism.

The demonstrated capacity of constitutions to shape model behavior is precisely what will likely make them a target of regulators seeking to confine models to certain values and perspectives. The possibility of governments and governing bodies—state, federal, and international—attempting to amend or revise AI constitutions warrants advanced scrutiny from, among other things, First Amendment scholars. At this stage in the AI governance debates, free speech and free expression scholars have focused on more doctrinal questions, such as whether AI outputs are protected speech. It’s urgent that they expand their inquiry as to whether AI Constitutions are protected speech, before constitution-based regulation takes off in earnest.

This article argues that AI constitutions—as policymakers increasingly turn to regulating model behavior and characteristics—contain expression protected by the First Amendment, a claim that requires first defining what AI constitutions are and situating them within existing First Amendment jurisprudence. We also consider counterarguments and alternative possibilities for shaping model characteristics that would avoid First Amendment issues altogether.

What Counts as an AI Constitution

The use of the word “constitution” to refer to technical documents risks inviting direct comparisons to documents traditionally bestowed with that title. For that reason, it’s important to precisely assess what’s similar and not about, say, Claude’s Constitution and the U.S. Constitution.

In the AI context, the term was originally coined in a 2022 paper. The terminology was intended to capture the structural and sociological similarities between AI constitutions and legal constitutions. That is, AI constitutions look like legal constitutions in some ways. For example, they list things models should and should not do, or should and should not care about. They also function in society like legal constitutions in that they explicitly outline principles to resolve difficult ethical and practical questions. These foundational technical documents are likewise intended for public analysis. According to the company, OpenAI publishes its Model Spec because “it’s important for people to be able to understand and discuss the practical choices involved in shaping model behavior.” Perhaps most importantly for the analysis here—whether intended by AI researchers or not—the term carries a certain solemnity that invites significant scrutiny of its contents. Anthropic researchers at one point referred to the constitution as a “soul document.”

Yet AI constitutions, unlike legal constitutions, are not the product of some democratic process that signals broader consent of the governed to the terms of the constitution. Nor, as Nathan Darmon and Tom Reed pointed out, are conflicts over how to interpret the constitution subject to external, independent adjudication. Still, the similarities with legal constitutions are strong enough to elicit interest in them as a vehicle for regulation.

Joe Carlsmith, one of the principal authors of Claude’s constitution, minimally defines an AI constitution as “a description of the intended values and behavior for an AI system.” A thicker definition accounts for how constitutions are applied in training and monitoring. Constitutions specify what a model should and should not do, take final authority over other instructions, are trained directly into the model, and, with some exceptions, are meant to shape a model’s character rather than police outputs case by case. Both Anthropic’s Constitution and OpenAI’s Model Spec qualify as AI constitutions as defined here. Not all labs have adopted a version of an AI constitution. Only Anthropic and OpenAI publish stand-alone documents that function as training-time specifications of value. Our definition excludes Meta’s acceptable-use policy, Google’s app guidelines, and xAI’s raw system prompts, though in the case of Google it is possible to craft one from various company documents that more or less make up the core components of a constitution.

There is not yet a default way to write an AI constitution. Claude’s Constitution and the Model Spec, for example, read quite differently. Claude’s Constitution is philosophical and concerned with the model’s psychology. Rather than order the model to “always be honest,” it exhaustively explains what honesty is and why it matters; it’s trying to encode Anthropic’s higher order values by reference to the specific ways one can imagine a model behaving well (or badly). Anthropic has described the document as an effort to “materialize a new archetype for how an AI assistant can be.” In turn, OpenAI’s Model Spec reads like case law, pairing each principle with sample prompts and examples of allowed and disallowed answers. Both are important to model training. They may inform the generation of synthetic data, sit inside the model’s chain-of-thought reasoning, or serve as the score card for alignment afterward. Anthropic calls the constitution “the final authority on how we want Claude to be and to behave” and reports that training on it improved alignment in ways that “persisted through RL post-training.” In other words, these documents have been empirically shown to shape how AI tools perform in the wild.

Why AI Constitution Regulation Is Coming

Recent history strongly suggests that the public should expect some kind of regulation of AI constitutions in the not too distant future. The number of congresspeople discussing the dangers of AI has skyrocketed, and policymakers’ concerns often center on what these models value. The “Preventing Woke AI” executive order led to an OMB implementation memo, a copycat bill in Congress, and an Iowa bill that passed the state house. Illinois, Colorado, and New York City, meanwhile, have laws regulating AI discrimination in employment decisions. Concerns over alleged AI bias have been sustained by reports of continued ideological skew in model outputs. The Washington Post, for example, determined that leading models tend to produce neutral or left-leaning answers to political questions. Such findings will become more relevant as the 2026 midterms and 2028 presidential election near, especially given that other researchers have found that users tend to be influenced by the political answers offered by models.

As politicians are increasingly interested in regulating AI, and in particular the values that models express, constitutions will prove an attractive object for regulation for at least three reasons. They are high leverage: Because they sit upstream of data generation, training, and evaluation, any edits to a constitution propagate through everything a future model learns and does. They are legible: Written in English and often reading like a statute, they can be marked up by a staffer and altered without consulting a technical expert. And they carry signaling value: An amendment to a constitution addressing wokeness, patriotism, or child safety, for instance, would be something a politician could easily point to when explaining their AI governance efforts to constituents. Regulations about reward functions or mechanistic interpretability—more technically complex features of AI development—seem less likely to spark an appearance on the Sunday shows.

What Speech Gets Protection

Efforts to regulate AI constitutions will likely face constitutional headwinds arising from the potential curtailment of expressive activity. Gauging the stiffness of those winds requires a review of First Amendment protections of speech.

In the most general sense, the First Amendment restricts the ability of the government to regulate speech. But within this broad prohibition there are myriad exceptions. Not all words qualify as speech protected by the First Amendment, and not all speech is protected speech. Protection instead runs along a spectrum. At one end is fully expressive speech, where a content-based restriction is “presumptively unconstitutional.” Somewhere in the middle is commercial speech, which is “related solely to the economic interests of the speaker,” and which the government can regulate more freely. At the far end is speech “plainly incidental” to conduct, which is generally regulable. Where a given document lands turns on how much expression it carries. The Court has “long recognized that not all speech is of equal First Amendment importance,” with the First Amendment concerned primarily with protecting matters of public importance as determined by “[the expression’s] content, form, and context.”

In the context of corporate speech, government regulation of a company’s mission statement would likely run afoul of core protections around free expression and freedom of association by compelling adoption of specific views. Still, government regulation of an instruction manual is far less likely to spark First Amendment concerns because of the absence of expressive content. How best to use a chainsaw, for example, says little about the corporate actor’s views other than that they seek to preserve life. An AI constitution arguably occupies some space in between given its clear expressive purpose as well as its more practical effort to ensure that the tool works as intended.

This ambiguity is best illustrated when one compares the AI constitutions produced by Anthropic and OpenAI. OpenAI’s document, the Model Spec, explicitly and categorically prohibits any models from “whistleblowing,” whereas the Constitution permits Claude to take “independent action” in “cases where the evidence is overwhelming and the stakes are extremely high.” The Model Spec allows models to tell some white lies, whereas Claude’s Constitution almost categorically forbids them. The Model Spec instructs models to always listen to its commands, whereas Claude’s Constitution contemplates situations where Claude determines that some portion of its Constitution is itself unethical. Each of these is a clear expression of the values of each corporation.

Other provisions look less like expressions of distinctive corporate values than like functional necessities any commercial operator would adopt. Each document tells the models to be careful when providing legal or medical advice, for example, and each instructs the models to prefer information from reliable sources. Traits common to both documents might be driven less by any company’s particular vision of a good model than by liability concerns or the demands of shipping a useful product. To the extent a constitution is built from provisions like these, it drifts toward the more regulable end of the spectrum.

AI Constitutions Contain Protected Speech

Given all of the above, some portions of AI constitutions may be justifiably regulated. Other sections, particularly those that tend toward the expressive end of the spectrum, may be safeguarded by the First Amendment. But few scholars or lawyers have considered the question of regulating corporate AI constitutions directly. Instead, the vast majority of First Amendment scholarship on AI thus far has focused on how to think about AI outputs in a First Amendment context. In brief, Eugene Volokh, Mark Lemley, and Peter Henderson are inclined to regard outputs as protected speech, while Peter Salib argues they are not because when a model emits text “no one thereby communicates.” Still others tie protection to whether a speaker “knows what he said when he said it.”

Scholarship has presumably focused on these more doctrinal questions up until now because it was believed that steering outputs with upstream regulation was impossible (at least on the time and complexity scales on which Congress typically operates). Safety rules must operate on outputs, Salib has argued, because “there is … no way, currently, to write legal rules mandating safe code.” The labs now claim there is such a way, and a regulator who takes them at their word will naturally want a say in how it’s written. A clue may be found in Justice Amy Coney Barrett’s concurrence in the recent First Amendment case of Moody v. NetChoice (2024). The justice warned that as companies “hand the reins” of their decision making to machine-learning systems, “technology may attenuate the connection between” corporate conduct and the human choices the First Amendment protects. The corporate values of Meta may not clearly be at work in the algorithm that shapes a user’s Facebook feed, for instance. When only a vague goal is given to a machine learning system that then implements intermediate policies on its own, it is difficult to locate a human speaker with First Amendment rights. Thus, a bare direction to “maximize engagement and profits” may be interpreted by a model in various ways: showing more or less inflammatory content to different users based on their history of engagement, showing photos of family to one user and entertainment news to another. These decisions do not come from a readily identifiable speaker with First Amendment rights.

A constitution, though, is written by identified (or identifiable) people, published under a company’s name, and read and (increasingly) argued over by the public as a statement of values. The expressive interest sits with the humans who write these AI constitutions, whatever statistical use the training process later makes of the text. A values-rich constitution is thus about as unattenuated as anything in the industry. Many of the things that inform model character and behavior are unpredictable, which has bedeviled machine learning scholars for decades. Constitutional AI is an unusual technique in that it allows human authors to act with intentionality ex ante on model character, rather than the more typical process of nudging AI outputs ex post with techniques such as RLHF or safeguards on model APIs.

Such documents thus enable much more input and expression. Companies may choose to engage with external stakeholders such as faith leaders or philosophers, consult internally with employees or consider the company’s stated public benefit, and even engage with a random sample of members of the public. The result of this process necessarily differs wildly from a bare “follow the law” document or direction to profit-maximize, which expresses little about the company and has a correspondingly weak claim to First Amendment protection.

The Arguments and Counterarguments for First Amendment Issues Arising From AI Constitution Regulation

Imagine a hypothetical law that codified the concerns of the Woke AI executive order. Say that this hypothetical law directly requires AI constitutions to forbid “wokeness” and support for diversity, equity, and inclusion, and requires model providers to certify their models as “non-biased” or “biased” based on a government-defined benchmark.

This bill almost immediately runs into problems under Reed v. Town of Gilbert (2015), which held that content-based speech regulations “are presumptively unconstitutional and may be justified only if the government proves that they are narrowly tailored to serve compelling state interests.” The hypothetical law would seem highly content based, as it is directly concerned with the political valence of the content of the AI constitution. Forcing a government-scripted line into an authored document also runs afoul of the compelled speech doctrine articulated in Miami Herald Publishing Co. v. Tornillo (1974), where the Court struck down a Florida law requiring newspapers to print candidate replies to negative articles. Further, inserting even one ideological phrase into a company’s own document is what doomed the “conflict free” label in National Association of Manufacturers v. SEC (2014), where the U.S. Court of Appeals for the D.C. Circuit struck down a portion of a Securities and Exchange Commission rule requiring manufacturers to publish whether certain products from the Democratic Republic of Congo were “conflict free” or not. A requirement to certify a model as “non-biased” would seem at least as ideologically loaded as that.

The government’s most ambitious response is that an AI constitution is not really speech but a functional artifact, machine instructions that happen to be written in English. These are regulable under Universal City Studios v. Corley (2001), where a law aimed at code’s function drew only intermediate scrutiny. For highly expressive decisions in AI constitutions, such as those discussed above around honesty or whistleblowing, this argument likely fails. The work the constitution is doing is interpretive and value laden. For the less expressive choices discussed, such as how the model should characterize its legal or medical advice, constitutions likely look more like code than a newspaper and may draw only intermediate scrutiny. Laws requiring an AI constitution to emphasize that it is not a licensed attorney, for example, may have an easier time surviving.

A subtler argument in favor of being able to regulate AI constitutions frames them as quintessential commercial speech, governed by the less burdensome Central Hudson (1980) test. Central Hudson asks whether the speech concerns a lawful and non-misleading activity; whether the government’s interest is substantial; whether the regulation directly advances it; and whether it is no more extensive than necessary. This test is more permissive than the strict scrutiny applied to regulation of fully expressive speech. For example, in Fla. Bar v. Went For It, Inc. (1995), the Court upheld a Florida Bar rule banning lawyers from soliciting accident victims within 30 days of their accident. But commercial speech “does no more than propose a commercial transaction,” and constitutions certainly go well beyond that. A constitution may do favorable brand-building work, but it quotes no price and solicits no purchase. Where commercial and fully protected speech are “inextricably intertwined,” Riley v. National Federation of the Blind (1988) treats the whole as protected. And the harder the government insists the document is “just marketing,” the more it concedes that the document is the company’s own expression, the very premise that makes a forced edit a Tornillo problem.

Should a court apply strict scrutiny to a hypothetical law that seeks to regulate AI constitutions, such a law might still survive if it is narrowly tailored to serve a compelling government interest. National security can prove a compelling interest indeed, as tech companies have repeatedly discovered in recent years. In TikTok Inc. v. Garland (2024), the D.C. Circuit upheld the forced divestiture of TikTok on national security grounds, assuming but not deciding that strict scrutiny applied (the Supreme Court later applied only intermediate scrutiny). In Twitter, Inc. v. Garland (2023), similarly, the U.S. Court of Appeals for the Ninth Circuit upheld a restriction on Twitter’s disclosure of governmental requests regarding its users. AI companies have repeatedly emphasized the national security implications of their technology, and this may prove to be an important admission against interest in future litigation. Given that national security clearly is a compelling interest, the argument would then shift toward whether regulation of constitutions is narrowly tailored.

What the Government Can Still Do

Application of the constitution’s safeguards to novel threats and technologies is necessarily contextual, calibrating to the scope and scale of the risk to the rule of law and the constitution order itself. There are real dangers to the development of powerful AI, and it’s important that the state be able to step in and coordinate action to avoid catastrophic outcomes. Much of what the government is able to do in this context is regulate conduct and mandate outcomes, rather than attempt to control how a company details its values or beliefs.

First, the government may mandate “purely factual and uncontroversial” disclosures under Zauderer v. Office of Disciplinary Counsel (1985). Thus, a disclosure regime where labs must make their AI constitutions public, or specify whether they do or do not contain certain sorts of provisions, is likely constitutional. However, the government may not force the lab to characterize the document as, for example, “unbiased” or “patriotic”; that is the compelled branding NAM forbids.

As a buyer, the government has more room. Under Rust v. Sullivan (1991), it may decline to purchase models whose constitutions fail its specifications, which is the theory of the Woke AI order. In this way, the government can at least shape the values of models doing highly dangerous activities, such as military or intelligence work.

The government can, of course, require a constitution to forbid the model from helping commit a crime, such as producing child sexual abuse material, as speech integral to criminal conduct under Giboney v. Empire Storage & Ice Co. (1949), though United States v. Stevens (2010) bars it from inventing new categories of unprotected speech by decree. Every major lab already writes these prohibitions in, though it remains an ongoing problem with open-source image generators and language models, as well as jailbroken closed models.

The mandates likeliest to survive are those aimed at conduct, touching the document only incidentally. A rule requiring an AI agent implementing a contract to act in good faith (as human contracting parties are required to), which a lab chooses to implement partly through adding language to its constitution, regulates a course of conduct. Under Rumsfeld v. FAIR (2006), it “has never been deemed an abridgment of freedom of speech … to make a course of conduct illegal merely because [it] was … carried out by means of language.”

In a recent article, Simon Goldstein and Salib argue for “A Thousand AI Constitutions”—that is, a diversity of model constitutions built atop a common “kernel” constitution that may require “AIs to follow the law, to be honest, to be corrigible, and to refrain from causing mass destruction.” The idea of a kernel constitution may be a more appropriate place for government regulation, particularly under the national security justifications discussed above. Minimal public safety provisions with low expressive content could be mandated by regulation, while a diversity of more expressive choices could be made by labs (or perhaps one day individuals) on top of that foundation.

Who Gets to Write AI Constitutions

AI will soon represent a massive portion of the economy and be a significant determinant of our information ecosystem as well as our political discourse. In its S-1, xAI’s parent company claims to have a total addressable market of $28.5 trillion, with $26.5 trillion of that being from AI. ChatGPT recently hit 1 billion weekly active users, and billions more interact with AI through Google’s search results and phone voice assistant (with a similar product soon to replace Apple’s Siri). It is understandable, and likely warranted, that governments would want to shape the values of our whispering earrings and country of geniuses in a data center.

Much has been written, likely correctly, about the general technical (in)competence of government and the need for private organizations of subject matter experts to regulate AIs in one way or another. AI constitutions, though, are not (only) complex technical documents; they’re the point in the AI alignment pipeline where policymakers are most qualified to act. The people we choose to elect to public office have theoretically been selected for reflecting our values, and AI constitutions are documents of enormous public significance that we should all hope are imbued with laudable values. Certainly any AI constitution that encouraged models to lie or steal or kill would be a deeply evil document.

The U.S. Constitution is a pre-commitment device, including the First Amendment. Default protection of speech exists precisely because every generation finds its own exceptions compelling—sedition in 1798, syndicalism in 1919, wokeness or bias today. The point of committing in advance is to make the default hard to dislodge when the temptation to drift feels most urgent. AI may pose catastrophic risks, and the state retains real tools to combat them: conduct rules, procurement leverage, and so on. Yet beyond that narrow band, a model’s values should be shaped by the people who build with and rely on it, not by mandates that shift with each administration. A government that can rewrite Claude’s Constitution today can rewrite its successor’s tomorrow—in the opposite direction. That’s the sort of arbitrary and fleeting approach to law that’s antithetical to the constitutional order and to free expression. And the dangers of an AI monoculture, whatever its ideological flavor, may well exceed the dangers of any single model whose values one might find objectionable.

Don’t Let AI Developers Hire Their Own Referees

Generative Gap Filling

When Reporting an AI Security Incident Is Not Mandatory

On July 16, Hugging Face, a public platform for open-weight artificial intelligence (AI) models and datasets, disclosed that it had detected a significant cybersecurity breach. An autonomous AI agent had conducted the attack end to end, according to a statement.

Five days later, on July 21, OpenAI revealed that this incident was driven by a combination of agents built on two of its frontier models—GPT-5.6 Sol and a powerful, unreleased model—acting in unanticipated ways during an internal, cyberoffensive capabilities evaluation. For purposes of the evaluation, the researchers had turned off production safety classifiers that block high-risk cyber activity and confined the models to a sandbox, an isolated computing environment without access to the internet, to restrict their interaction with the outside world.

What followed is the first known example of an autonomous cyber incident executed by systems not yet available to the public. Rather than solve the tasks presented, OpenAI’s agents “escaped,” exploiting a previously unknown, zero-day vulnerability. They obtained internet access (the very access OpenAI intended to deny them) and then hacked Hugging Face’s systems to acquire the answers to the benchmark. All of this was seemingly performed without express instruction by humans.

This is not the first time that models have been observed cheating. A few months ago, METR, a nonprofit research organization that conducts evaluations of frontier AI systems, released a report finding that AI models “routinely attempted to cheat on our hardest evaluation tasks, often in flagrant and elaborate ways that we believe humans would not consider.” In one incident, METR reported that an AI model tasked with updating a web app screenshotted a fake version of the app instead of completing the task.

But the Hugging Face breach has struck many observers as more real than these past examples. OpenAI’s models imposed a real cost on an uninvolved third party, all before completing internal testing. While people have at times questioned previous examples of cheating as artificial or contrived, it is hard to imagine that OpenAI expected its agents to escape the confines of their testing environment or to engage in a sophisticated, multistep plan to circumvent their constraints.

Having considered all of these facts, it may come as a surprise that OpenAI might not be legally required to disclose this incident. Certain crucial information is not yet publicly available, and both policymakers and the public will need that information to make sense of what this all means. What about all of those state AI laws with mandatory incident reporting? Don’t they apply here? Many will be disappointed to learn that the answer is arguably “no,” and that even if reporting is mandated, it requires only the scantest of information. What to do about this is the purpose of this article.

Existing AI Transparency Laws and the Hugging Face Breach

The rationale for mandatory incident reporting is straightforward: Some industries have the potential to cause real harm to others, and the government and the public have an interest in learning about high-risk events. In the case of the AI industry, there is a major knowledge gap between the companies’ and governments’ understanding of the technology and its risks. Mandatory incident reporting about serious adverse events, which companies might otherwise be reluctant to disclose, helps close that gap. This rationale is all the more compelling in the context of a rapidly evolving, difficult to predict technology, where best practice and political consensus have yet to develop. Observing real-world incidents offers a path to resolve both political and empirical disagreements and prepares governments to respond to future events.

So did the Hugging Face breach trigger mandatory disclosure under existing incident reporting laws? The answer seems far from clear.

California’s SB 53, New York’s RAISE Act, and Illinois’s SB 315 each require that frontier AI developers report “critical safety incidents”—a term each law defines identically. Of the four reportable incident categories, three require actual harm, ranging from “bodily injury” to “the death of, or serious injury to, more than 50 people or more than one billion dollars ($1,000,000,000) in damage.” (If you’re thinking, “that’s an exceptionally high bar for what is a basic, low-cost reporting requirement” or “it sure seems like governments would want that information before mass harm occurs,” you would not be wrong, but we digress.)

So three of the four incident categories do not apply. That leaves only the fourth, which applies to incidents in which a frontier model “uses deceptive techniques against the frontier developer to subvert the controls or monitoring of its frontier developer outside the context of an evaluation designed to elicit this behavior and in a manner that demonstrates materially increased catastrophic risk.”

It is possible that the Hugging Face breach meets one or more of these elements. It is far less clear that it meets all of them. On the first element, the models may have used deceptive techniques against OpenAI, their frontier developer—they did, after all, try to complete their developer’s evaluation using stolen information, after bypassing restrictions placed on them. But deception is notoriously hard to define, especially if it turns on the “intentions” and “obfuscation” of AI agents. Another read of these events is that the systems simply used all available means of solving the task, and public reporting does not tell us whether the agents attempted to hide those efforts. The second element is also arguably met. While the incident occurred during an evaluation, that evaluation was not “designed to elicit” this specific “deceptive technique.” Based on the ExploitGym benchmark, this evaluation aimed to elicit agentic, cyber-offensive capabilities on a specific task in a controlled environment. It did not, to our knowledge, contemplate—let alone design for—an unexpected cyberattack on a real-world company.

The third element, that the incident “demonstrates materially increased catastrophic risk” would seem to be the most difficult to satisfy. While autonomous cyber capabilities certainly increase the capability and thus potential consequence of agentic action, so too do most capability improvements in AI models. With limited monetary harm and no physical injury, this incident is quite attenuated from future events that might result in the mass physical injury or property damage contemplated by the statute.

With uncertainty at each factor, it is unclear that these existing state laws cover this event. At the very least, it won’t cover all events like it. Stepping back, it seems far from ideal to condition basic incident reporting on a list of complex, highly contested, fact-dependent conditions—all of which must be satisfied simultaneously. In many cases, figuring out whether the incident is indicative of increased risk to the public or actually constitutes deception will not be possible without more information. The purpose of incident reporting is to produce that information, not to require that it be known before a report is ever sent. The chicken must come before the egg. There are better alternatives.

Better Practices for AI Incident Reporting Laws

So where do policymakers go from here? If current incident reporting isn’t providing the needed insight, there are several steps policymakers can take.

Adjust the scope of transparency laws: Lower the exceptionally high bar to basic reporting, gather (at least some) information before harm occurs, and increase visibility into the most capable nonpublic models. To facilitate all of this, rulemaking authority is key.

As we outlined above, most incident reporting laws are simply too narrow. If the only incidents that get reported involve massive damages or loss of human life, the law itself isn’t providing information beyond what the government and public will already know. At the very least, issues of the highest concern, like model theft or loss of control, should be included even in the absence of harm. But existing laws put too much weight on hard-to-pin-down concepts like “loss of control” or “deception” that are difficult to prove and arguably don’t apply in cases like the Hugging Face cyber incident. While loss of control and deception should be sufficient to trigger reporting, they shouldn’t be necessary.

Instead, incident reporting should turn on what information is most likely to update the government or public’s understanding of risks. Information that is unexpected, that is indicative of advanced capabilities in high-risk domains like bio and cyber, or that demonstrates safety and security failures all seem like good candidates for inclusion. Some of these ideas have already made their way into existing proposals. Where the reporting requirements are light touch, as is the case in all existing state AI laws, a wider category of harms can be included. The narrower categories in today’s laws can be saved for more onerous disclosures.

Perhaps even more importantly, transparency laws should increasingly move away from a focus on deployment to a focus on providing visibility into nonpublic models and systems. The Hugging Face breach highlights the need for visibility into nonpublic models. OpenAI deployed systems more capable than anything available to the public, with fewer restrictions, and real-world harm resulted. All of this occurred before any external deployment. This is unlikely to be an isolated event.

Frontier AI developers will likely be the earliest and most sophisticated users of their own models. And the models they deploy internally will most often be more capable than those available to the public. They may be operated with fewer safeguards, especially for evaluations seeking to assess the frontier of capabilities. The gap between the most capable internal and external models may start to grow as companies develop increasingly capable models with dual-use capabilities and as AI systems are used to accelerate their developers’ own AI research and development. In this case, the gap between knowledge inside these companies and outside would expand. Without visibility into the current state of the art, governments will struggle to act effectively or quickly, a problem as much about democratic governance as it is about safety. Visibility into the internal deployment of nonpublic models may become increasingly central to the future of AI governance.

To make all of this work, policymakers will need legislative and regulatory flexibility. Policymakers should expect the exact scope of incident reporting to change over time as societies get a clearer picture of what capabilities and use cases matter most. Because a key goal of incident reporting is to surface novel or unexpected information, some types of events may be less important to report once their dynamics are thoroughly understood and accounted for. To accommodate those changing needs and to provide clarity, narrowly scoped rulemaking authority to refine incident reporting and reporting on nonpublic model use is likely necessary.

Get the details: Make sure reports provide enough information to inform decision-making by providing agencies with rulemaking authority and investigative powers.

As of this writing, the public, and, possibly, some policymakers know remarkably little about the Hugging Face breach. This isn’t a criticism of OpenAI, which voluntarily summarized the event, but the missing details about the event matter. Many commentators have noted that the lessons from and level of concern about this event depend on unknown details. Existing incident reporting laws, requiring little more than the date of the event and a brief summary, are unlikely to provide those detailed answers. Stronger transparency requirements could help answer many key remaining questions.

First is a cluster of questions, posed by Stephen Casper, that can roughly be summarized as “how impressive and/or concerning is the thing I just witnessed”:

Second, there are questions about the adequacy of the company’s safety practices and preparedness:

Third, there are questions about the present and future security of relevant systems:

So long as reporting remains almost entirely voluntary, detailed answers to these questions will be hard to come by, and the companies that offer information voluntarily will be subjected to greater scrutiny than those that are less cooperative. Well-crafted rulemaking authority will be key both to ensure that information is adequate and that companies are well informed about their obligations. In many circumstances, authorities will not know all the details they require until after a serious event occurs. In those cases, they will need investigative powers to obtain the needed information.

Reduce reporting costs: To accommodate more robust reporting, design transparency requirements to limit compliance costs, maintain confidentiality, and avoid disincentivizing rigorous risk assessments.

Policymakers can take several important steps to reduce the burden of greater reporting requirements. As assessments and information generation come to focus less on external deployment and more on nonpublic models, periodic reporting becomes more and more attractive. Evaluation organizations like METR have argued that periodic reporting not only saves time but also avoids perverse incentives to rush assessments in the lead-up to deployment. While a small subset of the most severe incidents will require rapid response from law enforcement and others, many events, including those like the Hugging Face breach, may allow for more relaxed reporting timelines. This approach would allow companies to focus on mitigations in the moment and still ensure that they ultimately produce the critical information. In cases where rapid reporting would interfere with mitigation efforts, laws could require only a simple notice of incident, followed by more thorough reporting after the event has been resolved or if officials request it.

Mandating incident reporting or information sharing for nonpublic models can also mitigate perverse incentives as long as minimum requirements are in place. For example, if a company’s reporting obligation triggers only in the context of a risk assessment, but the risk assessment does not have mandatory minimum requirements, the company is incentivized to skip risk assessments or to conduct them less rigorously. Laws can also allow companies to anonymize and aggregate reports, facilitating governments’ information gathering without punishing anyone for proactively uncovering issues.

Finally, any disclosure laws will need clear norms around confidentiality and information sharing to assure companies that their intellectual property and confidential information remains private.

Share information with capable actors: Make sure information is shared with key decision-makers who can assess, verify, and act on it.

Information is only as valuable as the actions it informs. If information from incident reports or internal use assessments sits inside a state agency with limited authority, societies will incur the cost of reporting without most of the benefit. Viewed this way, information sharing is about return on investment. And once the information is generated, most of the cost has been paid. At that point, so long as confidentiality can be maintained, it is incumbent on governments to share this information with the policymakers who most need it. Within states, this will include sharing reports with governors and legislatures to help inform their decision-making and help them target future policy. This information sharing may also be key to spurring political consensus.

In the context of assessing nonpublic models or serious risks to national security or from loss of control, the federal government will often be the central actor. States will often lack the resources, expertise, and political legitimacy to wade in on matters of national security. As we saw in the regulatory response to Mythos, if and when serious national security concerns emerge, the federal government will take the lead.

That’s why it’s confusing that some state laws restrict the ability of states to share information they gather from assessments of companies’ internal use of AI models. If, in fact, these reports generate important information—say, surprising developments in AI research and development or concerning deceptive behavior that doesn’t result in reportable incidents—that information should be shared. It is considerably less useful if locked away in a state agency in Illinois. Internal use and nonpublic model assessments are perhaps the most likely sources of information relevant to national security. To handle that effectively, governments also need the capacity and expertise to process this information. This requires staffing and likely some level of reliance on third-party auditing and assessment. In the wake of incidents like the Hugging Face breach, third-party auditors would be well positioned to conduct the sort of careful fact gathering outlined in previous sections.      

Similarly, it is in the interest of the United States and its close allies to share select information related to AI security incidents. Soon enough, and likely far sooner than most U.S. federal or state laws, the EU AI Act will be enforced. What that means, practically, is that the European AI Office will soon receive information that may be useful to public safety and cybersecurity in the United States. Luckily for us, Article 78(5) of the EU AI Act enables information sharing where the European Commission and EU member states create confidentiality agreements with third countries. For its part, the U.K. AI Security Institute conducts crucial research regarding model capabilities and can be a source of trusted expertise for U.S. policymakers. While there will be upfront costs in establishing such a shared information system, failing to make this investment would be a missed opportunity to improve our security at little regulatory cost.

Finally, where appropriate, the government should be empowered to disseminate information to the public and to vulnerable companies. Vulnerabilities found in software, for instance, may affect many different companies and actors, and information sharing will help keep the public safe from the risks they raise. Many of these risks should be discussed in the public square. While worries about information hazards and intellectual property leakage are real, so too is the value of public scrutiny of safety events. There is simply a lot to be learned from the collective scrutiny of the outside world. As the past few days have shown, outside experts have been invaluable in analyzing the Hugging Face breach, and most of them exist outside of government. In tweets and blogs, some of the brightest minds in AI have analyzed the publicly available facts and asked the questions that help us all better understand this event. Many of those questions have come from employees at OpenAI and its competitors, as well as from academics, policy wonks, and online skeptics. Further investigation of the Hugging Face breach will go substantially better because these discussions happened in public. Where possible, policymakers should ensure that these conversations continue to happen in the future.

The Future of Incident Reporting

As this article has perhaps made clear, existing laws fail to prepare us for events like the Hugging Face breach. Most incidents will go unreported. The information that is generated won’t be shared. And, at the end of the day, governments and the public will be repeatedly surprised about developments in this technology. That issue will only get worse as models advance and the gap between internal and externally deployed models widens.

But there is much policymakers can do. Just as this event has brought clarity to how much information societies need to assess these complex and emerging risks, future incidents can inform policy decisions and catalyze moments of political consensus. Policymakers can ensure that concerning incidents and behaviors are reported before harm results and gather information on lapses in security practices. Policymakers can refocus attention on the most capable models likely to be deployed first inside of frontier developers. And policymakers can do all of that while keeping the frequency and urgency of these reports at reasonable levels. If policymakers can do that, and ensure that this information is shared with decision-makers and competent evaluators, societies will be in a position to manage the uncertainty of this technology and make smarter, faster policy decisions in the future.

A Thousand AI Constitutions

Responding to AI Distillation Without Panic

Chinese large language model (LLM) developers are under scrutiny for reportedly employing large-scale “distillation attacks” on U.S. frontier artificial intelligence (AI) models to improve their own systems. Many U.S. actors have sent signals that they consider distillation a serious threat. For example, in May, Anthropic released a policy paper during President Trump’s trip to China, highlighting distillation attacks as a key challenge in U.S.-China competition. In April, the White House issued an official memorandum about distillation, warning about “deliberate, industrial-scale campaigns” from Chinese entities. Also in April, the House Foreign Affairs Committee universally advanced a bill called the Deterring American AI Model Theft Act to address the issue. And others have circulated additional policy proposals.

Discussions of distillation often take for granted that it is a form of theft. But there are key differences between “stealing an AI model” and distillation that policymakers should recognize. To properly address distillation, policy should focus on illegitimate model access—and avoid imposing poorly targeted rules that could harm Americans and distort the open and competitive U.S. AI ecosystem.

What Is Distillation?

The concept of distillation has evolved since it was introduced as a machine learning technique in which a larger “teacher” model’s outputs are used to train a smaller “student” model. Traditionally, that often meant training the student model on the teacher model’s probability distribution over possible outputs, rather than only on the correct answer. Today, the term is used more broadly. “Distillation” also includes prompting a frontier model to generate outputs, and then using the prompt-output pairs—or, where available, reasoning traces—as training data to refine a model. Frontier models may also be used as judges or verifiers for reinforcement learning. Together, these methods improve weaker models by training on stronger models’ responses to prompts and solutions to complex problems.

Distillation is a common practice in contemporary AI development. While on the witness stand at the recent Musk v. Altman trial, Elon Musk acknowledged that xAI had done at least some distillation of OpenAI models and that “generally AI companies distill other AI companies.” As Nathan Lambert, a leading U.S. open-source AI researcher, recently wrote, distillation helps train smaller, often open-source or open-weight models. The White House has recognized this: Office of Science and Technology Policy Director Michael Kratsios pointed out that “AI distillation, when legitimately used to produce” such models, is a “vital part” of creating open models and ensuring a competitive AI ecosystem.

But some Chinese AI developers appear to be using distillation well beyond ordinary practice, accessing U.S. frontier models at a massive scale to do so. In February, Anthropic reported that three Chinese AI labs had generated more than 16 million exchanges with Claude through approximately 24,000 fraudulent accounts, in some cases using jailbreak prompts to extract as much information as possible. OpenAI and Google have also reported or detected similar distillation efforts.

The unusually aggressive distillation efforts of Chinese labs have been portrayed as an attempt at “model theft” and to “steal” the intellectual property of frontier AI labs. But while calling distillation a form of “stealing” or “theft” may make for effective rhetoric, it isn’t an accurate description of how distillation of a closed AI model really works.

Why Isn’t Distillation “Model Theft”?

Distillation doesn’t involve breaking into a developer’s internal system to download the model weights or source code. To a distiller, the model is still a black box. In this context, then, “model theft” would mean some kind of black-box extraction—learning enough about a model from the outputs to approximate model behavior such that it effectively steals the developer’s intellectual property (IP). But what IP would that be?

To start, copyright can be ruled out. The aspects of a model that could plausibly be protected by copyright, such as software code, can’t be copied by distillation. Nor should copyright be used to create a backdoor property right in model outputs. An AI system cannot be an author, and AI-generated outputs are protected only when sufficient human authorship is present. Treating model outputs themselves as copyrighted property of AI labs would create a new right to control downstream uses of text they did not write, raising serious commercial and public policy problems.

Patent rights are also a poor fit. Distillation doesn’t, by itself, copy a patented implementation or allow a distiller to practice a patented method. In any event, the frontier labs themselves haven’t claimed that distillation amounts to patent infringement.

What about trade secrets? AI labs develop and maintain their models in secrecy, which lets them protect many aspects of those models as trade secrets. But distillation typically relies only on information returned through the model’s public-facing interface—the outputs it provides in response to prompts. Trade secret protection requires reasonable efforts to keep information secret, and ordinary outputs are available to anyone with an account. That makes it hard to argue that distillation extracts information qualifying for trade secret protection.

The strongest trade secret theft argument is that mass distillation requires unusual efforts—such as using coordinated proxy accounts—that let a distiller learn more about the model than an ordinary customer could. Mass distillers have also been accused of using jailbreak prompts that elicit information that isn’t normally made public, such as hidden system prompts that guide model responses. But fundamentally, the output being returned is still the kind of output a legitimate user could get. The case would be different if distillers could obtain information like full nonpublic reasoning chains, agent traces, or token-level probability distributions—but there’s no evidence that’s happening.

Compulife Software Inc. v. Newman shows the outer limits of the trade secret argument and why it doesn’t seem to reach mass distillation. The U.S. Court of Appeals for the Eleventh Circuit allowed a trade secret claim involving mass scraping of online life insurance quotes to proceed in that case. The defendant had allegedly acquired enough of the plaintiff’s proprietary database to pose a competitive threat. But the case involved information in a proprietary database and allegations about copying software code—facts not at issue in distillation cases. A recent lawsuit did raise trade secret misappropriation based on jailbreaking as one of its causes of action, but observers noted that the claim was highly questionable (the case settled before reaching the merits). So—at least under current law—the distillation attacks as the frontier AI labs describe them are very unlikely to support a successful trade secret claim.

What’s more, distillation does often violate the AI lab’s terms of service (TOS) for accessing the model. But if every TOS violation counts as “theft,” then the concept has no limiting principle. The more serious legal question is whether mass distillation relies on false identities, misrepresented credentials, or other ways of getting around access limits. That kind of conduct could support a civil or criminal claim under the Computer Fraud and Abuse Act (CFAA). But there are important limits on that, since ordinary TOS violations don’t generally violate the CFAA.

The point is that the distillation itself isn’t an act of stealing an AI model or breaking into an AI lab’s system. Distillers instead are most clearly breaking the law when they take unlawful means to circumvent the safeguards AI labs have in place to prevent distillation.

The more effective way forward, then, is not to treat distillation as theft. Instead, policymakers should focus on securing frontier models against misuse by Chinese competitors and other foreign actors, while studying whether distillation contributes to the dangerous diffusion of model capabilities.

What Anti-Distillation Policy Should Do

Mass distillation merits a policy response, even if it isn’t theft. But policymakers first need to identify the problem they are trying to solve. If the concern is illegitimate access to U.S. frontier models by foreign competitors and state actors, then policy should help labs secure access, share threat information, and identify fraudulent accounts and proxy networks used to disguise who is accessing the model. If the concern is cybersecurity, then the problem is account abuse and getting around access controls. If the concern is the diffusion of dangerous model capabilities, then the first step is to determine whether distillation meaningfully improves those capabilities or helps remove safeguards.

These are public interests that support policies designed to protect access security, enable information sharing, prosecute and sanction unlawful conduct, and evaluate safety. What they don’t justify is measures that effectively provide additional IP protection for AI developers or otherwise restrict legitimate competition in ways that would favor the commercial interests of AI labs over those of the general public.

The most basic defense against unauthorized distillation is for AI labs to recognize when users are circumventing access controls, detect attempts to generate training data, and block outputs to suspicious requests. To succeed, they’ll need to identify patterns of use and other signals that accounts are being used for distillation and shut down their access. One commonsense proposal that the White House and others have suggested is to facilitate coordination and information sharing between frontier AI labs and the government to better prevent illegitimate access for distillation. This could be enabled through antitrust guidance, including guidance based on the existing antitrust exemption for cybersecurity information sharing, which has been extended through Sept. 30, 2026, while Congress considers longer-term reauthorization, or dedicated legislation.

Further light-touch legislation could enable the government to play a more active role in collecting and sharing threat information, identifying proxy and fraudulent account networks, and helping develop best practices. Given geopolitical considerations, sanctions authority, as proposed in the aforementioned Deterring American AI Model Theft Act may be another tool. But sanctions may be more effective as a punitive or foreign policy tool than as a way to stop distillation—and should be balanced against whether they make sense from a trade perspective.

The rhetoric around distillation as a form of IP theft, along with concern that these light-touch legal authorities may be insufficient, has led to interest in activating the United States’ robust trade secret and IP enforcement regime, including against overseas-based actors. These tools include the Economic Espionage Act (EEA), the Defend Trade Secrets Act , and the Protecting American Intellectual Property Act (PAIPA), which provide for criminal, civil, and sanctions tools in cases involving trade secret theft or other covered misconduct.

These tools should be available where there is trade secret theft. But there’s a real risk that defining distillation itself as trade secret theft under the EEA and PAIPA would eventually bleed into private trade secret actions and broader legislative proposals. The debate over IP rights in AI models should be the subject of public debate, not something shaped indirectly through a heated fight over distillation.

A more promising route for legal action against mass distillation is through the CFAA. There are some obstacles though. The ordinary idea of “hacking” is breaking into a system without authorization. But distillation involves accessing a model through ordinary access channels: usernames, passwords, API keys, subscriptions. While distillation violates the provider’s TOS, that, on its own, isn’t likely enough to establish liability under the CFAA. But after the Supreme Court’s 2021 decision in Van Buren v. United States, an ordinary TOS violation is generally not enough to “exceed authorized access” under the CFAA. The Court read the statute as focused on access restrictions—whether the user accessed information in a part of the computer system they were not allowed to access—not on whether the user had an improper purpose for accessing information they were otherwise allowed to see.

That doesn’t mean the CFAA is off the table. AI providers often cut off access upon detecting patterns of use suggesting distillation. Efforts to get around being cut off—through false accounts, misrepresented credentials, proxy access, or other forms of access-control evasion—could implicate the CFAA, creating the potential for civil and even criminal consequences. For example, in United States v. Cuomo, the U.S. Court of Appeals for the Second Circuit affirmed CFAA convictions against defendants who, though using a publicly available state website, bypassed its authentication gate by entering other people’s credentials to extract protected records. CFAA investigations can also be facilitated through information sharing between AI labs and the federal government.

A more sensible response to mass distillation is to use tools that target the conduct around it, rather than beginning to treat model outputs themselves as a form of property. The government can help AI developers make a lot of progress on mass distillation by enabling information sharing and coordination through a targeted antitrust safe harbor. The government can also assist developers build cases under the CFAA and other legal tools where the facts support them. Those measures get at what’s needed to actually prevent unauthorized distillation: detecting fraudulent accounts, spotting efforts to get around access cutoffs, and sharing that information with other labs and the government.

Overall, the labs’ interests ought to be balanced against public interests. Expanding IP rights in frontier model outputs is one way policy gets that balance wrong. Poorly targeted anti-distillation measures could also prevent legitimate open models from competing in the AI marketplace. The concern that distillation might be free-riding on the efforts of frontier labs doesn’t justify excessive limits either. AI developers have themselves benefited from open-source software and published research. Frontier models are trained on the commons of human knowledge, and user interactions and data are used to improve them. When done safely and lawfully, distillation can help keep AI development from becoming dominated by a few companies with disproportionate access to compute and rich stores of user data. Letting frontier labs learn from everyone else—while giving them broad new rights to stop others from learning from model behavior—would only intensify the concentration of AI capabilities and economic power.

Study Whether Distillation Creates Real Safety Risks

There’s an important argument that threats to public safety and national security through the diffusion of more powerful models via distillation would justify stronger measures. Given the risks associated with transformative AI, researchers, policymakers, and civil society should take this seriously. The problem is that we don’t know enough about how much distillation contributes. Does distillation meaningfully advance near-frontier models, or does it mostly benefit smaller models—or make marginal contributions such as by validating performance? Anthropic reported that DeepSeek had far fewer exchanges (150,000) with Claude than Moonshot (3.4 million) or MiniMax (13 million). That suggests DeepSeek’s limited distillation may have been more on par with what “every AI company does,” per Musk. Yet DeepSeek is among China’s most powerful AI models.

Given that uncertainty, the sensible response is further investigation, potentially by the Center for AI Standards and Innovation (CAISI) at the National Institute of Standards and Technology, to determine how much distillation actually contributes to the threat. CAISI could study whether distillation materially improves dangerous capabilities, whether it transfers or strips away safeguards, whether already-available open-weight models can provide the same uplift, and what kinds of model access are most likely to matter. If that work shows distillation poses a meaningful safety risk, then stronger measures could be justified. In the meantime, restraint is warranted to avoid policy errors driven by perceived threats that outpaced the evidence.

The bottom line is this: The threats associated with distillation are best addressed by targeting fraudulent access and efforts to circumvent access controls, and empowering companies to cooperate on measures to prevent illegitimate access. Creating new quasi-IP rights in model outputs, or other premature or disproportionate responses, would do more to protect AI companies’ interests than the public’s. Policy should protect U.S. people and businesses by targeting real harms and unlawful conduct—not speculative ones.

Congress Should Do Something: The Case for (Fixing) the Great American AI Act

Since the April announcement of Anthropic’s Mythos model and its unprecedented cyber capabilities, there has been a remarkable shift in the artificial intelligence (AI) policy discourse. This has been most noticeable, and most noticed, in the statements and actions of prominent Trump administration officials. After months of dismissing concerns about national security risks from AI and engaging with the issue primarily by attempting to preempt state AI safety laws, the White House recently issued an executive order that called for the establishment of a voluntary predeployment program headed by the National Security Agency to evaluate offensive cyber capabilities of frontier models. This came amid statements from senior administration officials about “striking a … balance between innovation and safety” and even considering a Food and Drug Administration-style mandatory predeployment licensing regime for frontier models.

On June 12, the Trump administration’s concerns about Mythos’s cyber capabilities boiled over into an unprecedented decision to use export control authorities to prohibit Anthropic from allowing foreign nationals to access its Mythos-class Fable 5 model. Practically, this amounted to a mandate that Anthropic revoke public access to the model entirely. Much of the online commentary on this decision devolved into speculation about the administration’s motivations, the alleged behavior of Anthropic’s executives, and other petty interpersonal drama. As intriguing as these are, the more important takeaway from the White House’s decision to abruptly institute a de facto licensing regime for frontier AI systems—as many commentators across the political and safety/innovation spectrums have observed—is that federal legislation to establish a framework for addressing the national security risks posed by the most advanced AI systems is an urgent necessity.

 The Commerce Department’s decision to impose export controls on Fable may or may not have been wise, depending on who you believe about the seriousness of the vulnerability that motivated the decision. But even assuming the decision was justified, the fact that the government was apparently caught by surprise and had to scramble to put together a heavy-handed response based on ad hocpotentially legally questionable authorities that were not designed with anything like frontier AI systems in mind, with no due process for the affected company, is a serious problem that should be remedied with legislation as soon as possible.

Which brings us, finally, to the subject of this piece—the Great American Artificial Intelligence Act of 2026 (GAAIA), a comprehensive frontier AI safety bill that is the long-awaited product of months of intense negotiation between Rep. Jay Obernolte (R-Calif.) and Rep. Lori Trahan (D-Mass.). Obernolte had tried for months to get a Democrat to sign on to an AI bill that preempted state AI laws before finally persuading Trahan. For her part, Trahan wrote that she was motivated to support the bill by the announcement of Mythos’s groundbreaking cyber capabilities.

The current version of GAAIA is a discussion draft, meaning it has not been introduced yet and is intended to spark a conversation and elicit feedback from stakeholders rather than to become law in its current form. It may seem somewhat strange that a discussion draft sponsored by two relatively junior members of the House, neither of whom appears to have the backing of their party’s leadership, should receive so much attention from the media and from AI policy commentators, but for once the buzz is warranted.

GAAIA is the best attempt to design a federal framework for the governance of frontier AI systems introduced to date. In other words, it is the first serious attempt to actually do the thing that last week’s Fable incident clearly shows is necessary—address the national security risks posed by the most advanced AI systems—in a transparent, legally sound, and democratically legitimate way. While the bill may not pass in the near future—it likely faces opposition from both Democrats and Republicans—the draft can tell us a great deal about what the future of federal and state frontier AI governance efforts may look like.

The bill in its current form falls short in a number of respects and should not be passed. That said, passing a similar bill with narrower preemption of state laws and somewhat stronger federal authorities would be an excellent first step toward a workable federal regime for governing frontier AI systems.

What the Bill Does

GAAIA is a bipartisan compromise, in the truest sense of the phrase, which means that everyone hates it. Obernolte, a longtime advocate for federal preemption of state AI laws who supported last summer’s “moratorium” (which would have preempted all nongenerally applicable state AI laws and replaced them with essentially nothing), has compromised by granting his seal of approval to a number of genuinely consequential affirmative policy proposals. Trahan, who strongly supported increased oversight of AI in the past, has compromised by accepting broad preemption of state AI laws.

The Federal Framework

GAAIA’s four titles contain 45 sections, each of which addresses a significant topic in AI policy. I am an AI safety guy, and my research focuses mostly on serious risks that advanced AI systems might pose to national security and public safety, so this article focuses on the sections of GAAIA that are relevant to those risks. However, GAAIA is not an AI risk bill exclusively. There are also sections on, for example, “Preparing K-12 educators and students for an AI literate future,” “Modernizing access to artificial intelligence-related labor market data,” and establishing a “National artificial intelligence research resource” for improving capacity for AI research in the U.S., among many others. This article does not discuss those sections, not because they’re not important, but because they’re mostly irrelevant to the catastrophic risk concerns that this piece focuses on.

The noteworthy catastrophic risk provisions in GAAIA are:

“Transparency,” in this context, means requiring frontier AI companies to publish frontier safety frameworks (documents describing how the company evaluates and addresses catastrophic risks from the company’s most advanced AI models) and model cards (documents accompanying the release of specific models that contain information about the capabilities and limitations of a model and the results of the safety evaluations conducted under the company’s safety framework). “Incident reporting” requirements mandate that companies report “critical safety incidents” (essentially, incidents in which a frontier model’s model weights are stolen or in which a model does something scary that seems catastrophic-risk-ish) to the government and/or to law enforcement. And “auditing” refers to the practice of having an independent third party evaluate the adequacy of a company’s catastrophic risk mitigation practices as well as the company’s compliance with transparency requirements and with its own frontier AI framework. Transparency and incident reporting requirements are intended to provide the information needed for the government and the public to understand how companies think about and address catastrophic risks, and auditing requirements are supposed to ensure that transparency and reporting requirements remain effective rather than being ignored or becoming meaningless box-checking exercises.

As regards GAAIA’s transparency and auditing provisions, one common take is that the risk mitigation benefits of establishing these programs would be marginal because similar requirements already exist at the state level in New YorkCalifornia, and Illinois. That view is, I think, mistaken in two important respects. For one thing, as Anton Leicht points out, it is vitally important to build up regulatory capacity within the federal government. I co-wrote an essay on this topic a few weeks back. The argument is, essentially:

  1. AI might end up being a very big deal with extremely serious national security implications at some point in the next 10 years (and possibly within the next two years).
  2. If that happens, we should expect that serious regulatory interventions may be required.
  3. That serious regulatory work will almost certainly have to be carried out by the federal government, because the federal government—
    1. has orders of magnitude more regulatory capacity, expertise, and resources to devote to complex regulatory tasks than state governments do; and
    2. is, constitutionally and practically speaking, the only entity that can realistically be entrusted with extremely complex and high-stakes national security projects.
  4. Building up the institutional capacity and expertise to competently undertake complex regulatory tasks is difficult and cannot realistically be done in a matter of days or even months.
  5. Therefore, it is vitally important that we begin the process of aggressively building up technical expertise and regulatory capacity and know-how within the federal government as soon as possible.

But even setting aside the capacity-building considerations, it is simply not true that existing state catastrophic risk laws are equivalent to GAAIA’s transparency or auditing provisions. These state laws—California’s Transparency in Frontier Artificial Intelligence Act (TFAIA), New York’s Responsible Artificial Intelligence Safety and Education (RAISE) Act, and Illinois’s Artificial Intelligence Safety Measures Act (AISMA)—are an important foundation for future efforts and have been extremely influential. GAAIA itself is clear evidence of this influence; some of Section 111’s transparency provisions are lifted almost verbatim from the transparency provisions of SB 53 (which are substantially identical to the transparency provisions of RAISE and AISMA). But, as groundbreaking as those state laws are, they are still state laws and therefore cannot leverage the resources, institutions, or legal authorities of the federal government in the way that a bill like GAAIA can.

Perhaps the most important institutional advantage that GAAIA leverages is the capacity of federal agencies such as the Department of Commerce to carry out sophisticated rulemaking, a capacity built over decades of administering complex regulatory programs that no state agency can realistically match. GAAIA grants CAISI and the Department of Commerce broad authority to issue regulations fleshing out the auditing and transparency regimes outlined in Sections 111-112. Rulemaking! That word may not sound like the most exciting thing you’ve heard this week, but take my word for it: This is the good stuff.

Take auditing, for example. Because AISMA does not confer any explicit rulemaking authority, Illinois’s auditing regime, when it goes into effect, will be defined solely by the requirements in AISMA’s text. AISMA requires that audits be conducted “consistent with generally accepted auditing standards and best practices” and that auditors possess “demonstrated competence to perform the audit.” These vague requirements, however, aren’t enough to guarantee a functional auditing regime.

Under AISMA, auditors are paid by the AI company that retains them. By default, this system will lead to a race to the bottom in which market forces compel auditors to compete with each other over who can cause the least hassle and difficulty for their customers (frontier AI companies). Rather than ensuring that companies abide by their commitments and hew to responsible risk mitigation practices, this kind of auditing regime will eventually devolve into a system where companies are disincentivized from hiring rubber-stamp auditors only by the uncertain prospect of ex post tort liability.

To be clear, this is not a criticism of AISMA’s auditing provisions, which are well designed. The issue is that Illinois’s state government simply lacks the ability to design and competently administer a complex, technically involved auditing program for out-of-state tech companies. The issue is not that AISMA is insufficiently ambitious but, rather, that Springfield—on its best day—has only a small fraction of the capacity for complex interstate regulatory projects that the federal government has on its worst.

In contrast to AISMA, GAAIA’s auditing section could establish a functional and effective third-party auditing regime. CAISI—reestablished as a regulatory agency separate from the nonregulatory NIST—would be granted broad rulemaking authority. The regulations that CAISI would be required to promulgate would include rules addressing conflict of interest and funding transparency requirements for auditors, requirements for licensing auditors and revoking auditor licenses, minimum requirements for audits and assessments, and “any other rules reasonably necessary to the administration of the IVO oversight and licensing regime.” This is the kind of rulemaking and oversight authority that could, in theory and with competent implementation, actually establish the kind of auditing regime that AISMA gestures at.

GAAIA’s transparency requirements would also be a significant upgrade from the existing state transparency requirements. While GAAIA’s transparency section is similar on its face to existing California, New York, and Illinois transparency statutes, it delegates fairly broad rulemaking authority to the Department of Commerce, which can prescribe regulations governing, among other things, the “form, manner, and minimum quality” of the model cards and safety frameworks that companies are required to publish. In combination with GAAIA’s auditing requirements, which require auditors to regularly evaluate and assess the adequacy of an AI company’s safety framework and its other efforts to identify and mitigate catastrophic risks, GAAIA’s transparency requirements would allow Commerce to ensure that transparency requirements actually result in meaningful transparency.

This rulemaking and minimum-standard-setting authority would allow GAAIA to provide significantly more transparency than existing state laws. Consider California’s TFAIA, which has been in effect for just over five months. TFAIA is a light-touch statute by design and imposes very few obligations on companies. This light-touch approach is, in my opinion, a good thing, but one downside is that companies can technically comply by publishing documents that check the statutorily required boxes without actually saying anything meaningful about the company’s approach to mitigating risks. For example, xAI, despite founder Elon Musk’s frequent public statements about the existential risks posed by superintelligence, complies with TFAIA by publishing a barebones framework that describes xAI’s risk mitigation practices in cursory and general terms.

Under TFAIA, there is no realistic way to require xAI to provide the industry-standard level of transparency that its competitors’ frameworks typically demonstrate. And while the RAISE Act does provide New York’s Department of Financial Services with some rulemaking authority that could in theory be used to give the act’s transparency requirements more teeth, practical and constitutional limitations may prevent a New York state financial services agency from regulating California-based software companies with the same level of rigor and precision that the U.S. Department of Commerce could, in theory, bring to bear.

Of course, the value proposition of GAAIA depends on the assumption that CAISI and the Commerce Department will do a decent job of establishing and administering the proposed transparency and auditing regimes. It’s far from clear that the Commerce Department would view this as a top priority, given that Secretary of Commerce Howard Lutnick has generally signaled skepticism of AI safety concerns. While public reporting in the months since Anthropic’s Mythos announcement has documented a shift among some of Lutnick’s fellow Cabinet members toward taking some of these concerns more seriously, Lutnick’s views have not—at least publicly—evolved along similar lines.

Even if the Commerce Department’s approach is initially ineffective, however, it might still be good to establish the statutory bones of an effective federal oversight regime in case future developments create additional political pressure for governance efforts. Given Mythos’s impact in shifting the Overton window on AI policy, the Trump administration may become more willing to implement serious oversight measures as they receive further concrete evidence of significant national security risks. Of course, it would be better not to have to cross our fingers and hope that the government starts taking risks seriously before anything bad happens. Still, enacting something like GAAIA would, at least, meaningfully improve the legal authorities available to the executive branch when and if the political will to get serious about national security risks from AI manifests itself.

GAAIA’s affirmative provisions are not perfect. The lack of effective whistleblower protections, for example, is a serious defect. The gold standard for federal AI whistleblower legislation is Sen. Chuck Grassley’s (R-Iowa) AI Whistleblower Protection Act (AI WPA). GAAIA’s whistleblower section falls short of that standard because, unlike the AI WPA, it fails to protect AI company employees who disclose information about a “substantial and specific” danger to public health, public safety, or national security to an appropriate government agency. The only AI whistleblowers who are protected from retaliation by GAAIA’s whistleblower section are those who report a violation of federal law—and because state whistleblower laws (including in California, where most of the relevant frontier AI company employees reside) already protect this kind of disclosure, a redundant federal protection adds little value. As I’ve argued before, protection for disclosures about serious risks that don’t involve law violations is crucial because it’s very plausible that perfectly legal frontier AI development activities could lead to serious national security and public safety risks that the government ought to be aware of.

Additionally, GAAIA’s auditing section lacks sufficiently clear enforcement authorities. Auditors are given a wide variety of authorities to assess whether a developer’s practices mitigate risks sufficiently. It’s not entirely clear, however, what happens if the mitigations are inadequate or nonexistent. CAISI is authorized to penalize developers for, for example, failing to retain an auditor or failing to grant timely access to appropriate records, but a developer who checks these boxes may not be required to actually do anything meaningful in response to an auditor’s recommendations. It’s possible that CAISI could use the broad rulemaking authority that the auditing section grants to cobble together a solution to this problem, but given how skeptical courts have been of this kind of broad agency interpretation of congressional delegations of rulemaking authority in recent years, it would be far better to have a clear statutory enforcement hook.

It should be noted that Trahan’s office has signaled potential willingness to remedy both of the specific issues discussed above in future drafts of GAAIA. If this does happen, it will improve my view of the bill’s value significantly; the point of a discussion draft is to bring potential issues like these to light so that they can be hashed out. More generally, a flawed whistleblower or auditing provision is at least marginally better than having no federal AI whistleblower or auditing statute whatsoever. As good as GAAIA’s federal framework is, however, it comes at a steep price: broad preemption of all state laws regulating AI development.

Preemption

In exchange for the substantial federal framework described above, GAAIA preempts—for a period of three years—any state AI law “specifically regulating the development of any artificial intelligence model.” As in prior AI preemption efforts, “generally applicable” laws are exempted, and GAAIA also includes a somewhat opaque exemption for “post-deployment activities.” “Development” is defined to mean

the acts performed or directed by a developer with respect to an artificial intelligence model prior to its deployment, including determining training or fine-tuning objectives; training, fine-tuning, or otherwise substantially modifying the weights or other parameters of an artificial intelligence model; and evaluating and deciding, prior to deployment, whether an artificial intelligence model satisfies applicable safety or capability thresholds for deployment.

There’s currently no consensus regarding exactly which existing or hypothetical state laws or bills GAAIA would preempt. A few commentators have argued that this language would preempt only the few existing state frontier AI safety laws—Illinois’s AISMA, California’s TFAIA, and New York’s RAISE Act—and future state laws in the same vein. If this were true and could be demonstrated clearly, an overwhelming majority of the AI safety community and a number of other factions who currently oppose the bill (including, for example, child safety advocates) would likely change course to support or assume a neutral attitude toward it. One-to-one preemption—eliminating state AI catastrophic risk auditing, transparency, and incident reporting bills in exchange for establishing robust federal AI auditing, transparency, and incident reporting regimes—would be a very reasonable compromise. If Reps. Trahan and Obernolte believe that GAAIA’s preemption section is effectively one-to-one, a clarifying amendment would go a long way toward increasing support for the bill.

Unfortunately, a court would probably not accept such a narrow interpretation. The key phrase is “specifically regulating”—state laws will be preempted if they are deemed to “specifically regulate” development. This phrase does not appear in any existing preemption legislation, meaning that there’s no firm precedent that can be used to accurately predict how courts will interpret the scope of GAAIA’s preemption. Still, a few relatively safe conclusions can be drawn.

It seems clear that “specifically regulating” is a relatively narrow formulation, as preemption triggers go—it certainly preempts less than the broad “relating to,” for example, which would sweep in any state law that had a “connection with” or contained a “reference to” AI. Still, “specifically regulating” is broader than, for example, the CAN-SPAM Act’s “expressly regulates,” which imposes a facial test (i.e., preempts a state law only if the law, on its face, names and regulates the forbidden federal subject). Arguably, GAAIA’s preemption language would instead impose a functional test under which state laws would be preempted if they had the effect of regulating AI development, as defined.

This functional interpretation would be consistent with how courts have treated comparable language in cases examining the (admittedly broader) preemption language in the Energy Policy and Conservation Act, the Clean Air Act, the Federal Meat Inspection Act, and the Federal Aviation Administration Authorization Act. In cases addressing the preemptive effect of the other laws mentioned above, courts have applied functional tests to prevent states from circumventing preemption with clever legislative drafting. For instance, in National Meat Association v. Harris, the Supreme Court held that a provision of California law that regulated the sale of meat from inhumanely raised pigs was preempted by a federal preemption provision that prohibited states from regulating the operation of slaughterhouses because the California provision because “the sales ban … functions as a command to slaughterhouses,” and because “if the sales ban were to avoid the … preemption clause, then any State could impose any regulation on slaughterhouses just by framing it as a ban on the sale of meat produced in whatever way the State disapproved.”

In short, it seems unlikely in light of existing preemption precedents that federal courts will allow states to circumvent GAAIA preemption through creative legislative drafting choices. GAAIA’s express carve-out for “post-deployment activities” and the relatively narrow scope of “specifically regulating” would likely protect some state laws that affect development without explicitly targeting it. But any law that functions primarily to regulate development, or can realistically be complied with only by altering development practices, will likely be preempted even if it purports to regulate deployment.

Because GAAIA’s definition of “development” is quite broad, this functional interpretation would preempt far more than just TFAIA-style catastrophic risk transparency laws. Existing state measures that GAAIA would almost certainly preempt include Texas’s TRAIGA (which, among other development-focused regulations, prohibits developing an AI system that impersonates a child while describing sexual conduct or developing an AI system with the intention of producing child sexual abuse material or illegal deepfakes), California’s AB2013, Colorado’s SB 189 (the law that replaced Colorado’s controversial algorithmic bias bill with less onerous, more business-friendly requirements), and certain California Consumer Privacy Act regulations affecting “automated decisionmaking technology.” State measures that might (or might not) be preempted include the many state child safety chatbot laws, like Idaho’s Conversational AI Safety Act, that purport to regulate deployers or “operators” of AI systems but could realistically be satisfied only via interventions implemented during training or fine-tuning, both of which are “development” under GAAIA.

By far the more important issue with GAAIA’s broad preemption, however, is that, in addition to invalidating these existing laws, it would preempt a broader and more important category of future state laws that would otherwise be passed in 2027, 2028, and 2029. TFAIA, RAISE, and AISMA were always meant to lay the foundation for future efforts rather than to establish a stable end state for state AI policy. As the capabilities of the most advanced AI systems continue to improve, their risks will increase as well. And as risks become more immediate and more difficult to deny, it seems safe to predict that state AI laws will be enacted that would have been politically unrealistic in 2026. The technology’s benefits may outweigh these risks; even so, it would be foolish to assume that AI will be the first transformatively impactful general-purpose technology, the development of which does not create any local harms that state legislation and regulation need to address.

The Takeaway

GAAIA includes by far the best federal framework for frontier AI safety that has been publicly introduced to date, as a discussion draft or otherwise. This being the case, some of the criticisms of GAAIA seem somewhat misguided. The Democratic House AI Commission, for example, has asserted that the bill “does not meet the enormity of the moment.” This may well be true, but if GAAIA’s unprecedentedly ambitious framework does not meet the moment, what bill does? I sincerely hope that there is a tidal wave of enormity-addressing AI legislation waiting in the wings, but the federal legislation that has been introduced thus far (with a few notable exceptions, such as the AI Whistleblower Protection Act and the Artificial Intelligence Risk Evaluation Act) has mostly been notable for its inoffensiveness rather than its moment-meeting boldness. A lot of “convening a multi-stakeholder process to consider the development of a process” for setting up a purely voluntary incident reporting regime, and things of that nature.

Despite this, the discussion draft would, in my judgment, be net-negative if enacted in its current form. As I have argued elsewhere, federal preemption of state AI laws should proceed on a narrow, issue-by-issue basis. Broad preemption of a vaguely defined and poorly understood category of state laws would, in all likelihood, be a disaster, for both political and policy reasons. Politically, it makes no sense to try to preempt broad categories of state AI law that have the backing of politically potent constituencies—developer-focused child safety laws, for example—in exchange for a bill with great frontier AI safety provisions but no child safety provisions. And while state frontier AI safety laws are not an adequate substitute for a robust federal framework, state laws have thus far proved so much easier to pass than federal laws that passing GAAIA in its current state might mean locking in a good-but-ultimately-inadequate 2026 governance framework indefinitely, despite the fact that future capabilities improvements might require more ambitious legislation.

At the same time, I worry many stakeholders who are rejecting GAAIA’s approach out of hand fail to recognize the urgency of the situation. By default, without new legislation, the process for addressing serious risks from frontier systems will be undertaken by the executive branch, the intelligence community, and national security agencies in a haphazard, legally suspect, case-by-case, and increasingly securitized manner, as June’s Fable incident proved. There is, of course, a sense in which serious national security risks being addressed by the nation’s national security apparatus are expected and necessary. But a program that is made up on the fly and administered on a totally discretionary, case-by-case basis, without the resources or structure that only Congress can provide, will never be as effective or as democratically legitimate as a well-designed federal legislative framework could be.

My hope, therefore, is that GAAIA’s authors will narrow its preemption and improve its substance and that GAAIA’s critics will either engage with the process of trying to improve the bill or introduce a serious alternative proposal in the near future. The stakes are high enough, and the issues with the status quo are clear enough, that continuing to do nothing indefinitely is no longer a defensible course of action.

The NDAA: A Key Vehicle for AI Governance

Summary

Introduction

The National Defense Authorization Act has quietly become one of Congress’s most powerful tools for shaping AI policy, and the FY26 NDAA featured many key AI provisions. This commentary compiles all the major AI provisions from the FY26 NDAA and analyzes the most significant language in detail. With some of the initial deadlines imposed by the FY26 NDAA now having passed—and with negotiations around the FY27 NDAA underway—it’s useful to take stock of the potential and pitfalls of these provisions.

The NDAA is not just restricted to the nuts and bolts of defense operations. It has also been used to achieve broader policy goals, sometimes by limiting the executive branch’s actions. For example, one of the most important and successful nonproliferation programs in history, the Nunn-Lugar Cooperative Threat Reduction Program, was originally proposed as an amendment to the NDAA and was subsequently expanded through the NDAA.[ref 2] More recently, the FY19 NDAA effectively banned the government from using certain Chinese telecommunications companies such as Huawei and ZTE; likewise, the FY26 NDAA bans certain foreign AI products like DeepSeek.

As the government increasingly prioritizes AI use in warfighting and military operations, the NDAA has a key role to play in shaping AI policy.

AI Governance in the FY26 NDAA

Notable Provisions

Several provisions stand out as particularly important for the government’s broader interest in overseeing and fostering the responsible development of secure AI systems: 

Section 1535: Artificial Intelligence Futures Steering Committee

Section 1535 requires DOD to create an Artificial Intelligence Futures Steering Committee (Steering Committee) to prepare DOD for advanced AI and AGI. The Steering Committee will be co-chaired by the Deputy Secretary of Defense and Vice Chairman of the Joint Chiefs of Staff (VCJCS). It will primarily be composed of principal deputies of the military services, relevant under secretaries (e.g., the Under Secretary of Defense for Research and Engineering (USD(R&E))), and others responsible for AI (e.g., the Chief Digital and AI Officer (CDAO)).[ref 3]

By January 31, 2027, the Steering Committee must submit a report to Congress covering what can be described as two main focus areas. First, the committee must help prepare DOD for advanced AI and AGI by creating: 

  1. A proactive policy for the evaluation, adoption, governance, and risk mitigation of advanced AI systems, including systems that approach or achieve AGI. 
  2. An analysis of the forecasted trajectory of advanced AI models and enabling technologies that could lead to AGI such as AI agents, neuromorphic computing, cognitive science applications, infrastructure needs, new microelectronics, etc.
  3. An analysis of the potential operational effects of integrating advanced AI or AGI into DOD networks and systems from a technical, doctrinal, training, and resourcing perspective to better understand effects on operational commands. 
  4. A strategy for the risk-informed adoption, governance, and oversight of advanced AI and AGI including ethical, policy, and technical guardrails to maintain appropriate human decision-making and prevent misuse. 

The second focus area is U.S. adversaries. Though not specifically named, the People’s Republic of China (PRC), which is actively pursuing AI capabilities that rival those of the United States, is likely the primary focus. The committee must assess the possible technological, operational, and doctrinal trajectories of U.S. adversaries with respect to AI capabilities, including the pursuit of AGI. Additionally, the committee must analyze the threat landscape associated with the use of advanced AI and AGI and develop options to counter these threats. 

Within the Pentagon’s sprawling bureaucracy, there’s often fierce competition between different programs and priorities for funding and attention from leadership. In this sense, the Steering Committee could be a valuable forcing function for the department to prepare for advanced AI, reinforced by the requirement to report its findings to Congress by early 2027. There’s precedent for DOD using these sorts of committees as a way to spur action on issues such as software modernization and autonomous systems.

However, such committees sometimes serve more as a signaling mechanism for Congress than as a catalyst for serious action. Unless chairs or members of the committee invest their time and professional capital to drive it forward, it can easily devolve into a box-checking exercise. While the Steering Committee’s substantive mandate is broad, its required procedural actions, as set by Congress, are fairly minimal: meet at least once every three months, and submit a report on its findings to the relevant congressional committees by January 31, 2027. That means that depending on when the committee is actually established and how quickly it first convenes, it may meet only three or four times before its report is due. 

On top of that, the NDAA provision does not allocate any dedicated staff or budget for the Steering Committee. Any resources must be drawn from existing reserves, which could further limit its capacity. Given those constraints and their already-full plates, the Steering Committee’s principals might be tempted to delegate their roles and responsibilities down the chain of command to other, typically less-empowered subordinates, whose remit might be narrower—i.e., drafting a report that satisfies Congress’s requirements while potentially tabling thornier policy disagreements or implementation details for later.

Congress should remain attuned to these possible failure modes and use its oversight power to solicit information about the Steering Committee and its progress, in the hopes of helping it gain and maintain momentum. There are some encouraging signs on this front. In March, Senator Jim Banks sent a letter to Secretary Hegseth requesting a staff-level briefing within 60 days to discuss DOD’s plans for the Steering Committee. The letter suggested areas of focus with respect to U.S.-PRC AI competition. Even just one or a few members of Congress taking specific, sustained interest in the Steering Committee could keep it high enough on DOD’s long list of priorities to increase its odds of success. 

Congressional oversight can be particularly valuable in two ways. First, it can keep pressure on the committee if it fails to meet the report submission deadline of January 31. Second, and perhaps more importantly, Congress can help ensure that the Steering Committee doesn’t waste the 11 months between its reporting deadline and its termination date of December 31, 2027. 

While the report is the Steering Committee’s most tangible required deliverable, Congress provided that the committee will continue to exist for nearly a year beyond the report submission deadline. This time would allow the committee to refine or update its policies and to work on implementing and disseminating the findings throughout DOD. Because the Steering Committee lacks deliverables or other measurable benchmarks throughout most of 2027, it’ll likely be incumbent on Congress to use tools like letters, hearings, and requests for briefings to push forward that updating and implementation work. These efforts could ultimately have a much greater impact on DOD operations in the long term than just the drafting of the report itself.

As of June 30, 2026, no public materials indicate whether the Steering Committee was established by the April 1 statutory deadline, or whether it has held its first meeting. That’s not necessarily cause for concern, as DOD is not required by the NDAA to report those actions to Congress or the public. But it does make it harder to predict which of these paths the AI Futures Steering Committee will ultimately follow. Overall, this provision could pay dividends by prompting DOD to proactively prepare for major threats and opportunities raised by AGI—planning that might otherwise get neglected—though its success is far from assured. 

Key Dates

Section 1533: AI Model Assessment and Oversight

Section 1533 instructs DOD to create a Cross-Functional Team (Team) for AI “model assessment and oversight.” The Team must develop a standardized assessment framework for AI models currently used by DOD, as well as guidelines to facilitate procurement of future models. The Team is led by the CDAO and composed of other DOD technology leaders, such as CIOs, CAIOs of the combatant commands, service acquisition executives, and USD(R&E). The Team must: 

This provision allows DOD to retain a lot of discretion over how it evaluates current and future AI models. Congress has mandated that DOD establish a framework and protocols, but didn’t set substantive thresholds for performance. That’s understandable to some degree, given the risk of setting standards via legislation, which might quickly become outdated and then prove difficult to adjust. And it’s similar to the approach that states like California and New York have taken in enacting frontier AI transparency reporting requirements. But some key requirements in the provision, such as the creation of “governance structures” and assessing “ultimate use-case-based risk,” use terms that are undefined and open to interpretation, and could have benefited from a bit more congressional guidance about the elements that should at least be considered or addressed.

That vagueness, combined with the long timelines the provision establishes, could make it hard for Congress to assess the Team’s progress. Congress notably gave the Team an extended timeline to develop its model assessments and oversight, which may be in tension with the pace of AI progress. The standardized assessment framework isn’t due until June 2027—a year and a half after enactment—and no actual assessments of DOD’s major AI systems are required until January 2028. Meanwhile, new frontier AI models are released many times a year. 

To be sure, the Team’s task is difficult. And Congress sometimes errs by giving agencies unrealistically short deadlines. But a failure to keep up with the pace of AI development risks undermining the Team’s purpose. To frame that risk, consider the events that have transpired since the FY26 NDAA passed six months ago. First, there was the blow-up over contract terms between the Pentagon and Anthropic in February. More recently, the June 5 National Security Presidential Memorandum (NSPM) 11 ordered Secretary Hegseth, ODNI, and IC elements to “review and update procurement processes to ensure the rapid onboarding of the most advanced AI models from multiple vendors” within 120 days. It’s unclear how or whether this review will be coordinated with the procurement guidelines that the Team is tasked with developing on its longer timeframe.

Here, again, Congress can deploy its oversight tools to steer DOD in the direction of consistent and streamlined guidelines for AI procurement. It should aim to ensure that standards are applied uniformly and transparently, not reactively, to AI developers. Helpfully, this provision requires DOD to provide a briefing to congressional defense committees within 30 days of hitting significant statutorily prescribed milestones, starting with its establishment of the Team on or before June 1, 2026. That offers a natural opening for Congress to probe the Team’s trajectory, and potentially to spur a course correction if needed. Congress might consider incorporating that sort of regular briefing requirement into future AI-related NDAA provisions; it’s particularly beneficial in this area due to rapid and sometimes unexpected jumps in capabilities and risks, and might also have been helpful for similar initiatives like the Steering Committee discussed above.

Finally, it’s worth a closer look at the provision’s definition of “major [AI] system”—one of only a few terms that the provision does actually define—buried near the end of the provision. That definition limits coverage to systems used annually by at least 500 users within DOD, and excludes systems used solely for research, development, testing, or evaluation that have not been deployed for operational use. Elsewhere, the provision specifies that DOD must assess all major AI systems using the standardized assessment framework, leaving it somewhat unclear when or to what extent that framework also governs assessment of other AI models used by DOD. In other words, for models used by less than 500 employees per year, or those involved only in R&D, how will DOD assess performance, security, and “compliance with ethical principles”? 

While this sort of line-drawing exercise is almost always difficult but necessary for administrability, in this instance the exclusions arguably represent the frontier of DOD’s own AI development and deployment in what could end up being the highest-stakes and hardest-to-monitor situations. At minimum, it’s plausible that some of the most powerful systems, deployed in potentially highly consequential cases, might be available to only a small number of users. Congress should ask DOD how it plans to assess AI systems that fall into those categories and potentially require the development of standards for such systems in future legislation.

Key Dates

Section 1534: Digital Sandbox Environments for AI

Section 1534 requires the CDAO to create a task force to promote AI sandbox environments supporting “experimentation, training, familiarization, and development.” The task force should “identify, coordinate, and advance” DOD efforts to develop and deploy AI sandboxes, with an eye toward accelerating AI adoption across the department. The provision defines an “[AI] sandbox environment” as a “secure, isolated computing environment that enables users with varying levels of technical proficiency to access [AI] tools, models, and capabilities for the purposes of experimentation, training, testing, and development without affecting operational systems or requiring specialized technical knowledge to operate.” The provision requires that the task force be established by April 1, 2026, and that the CDAO provide a briefing to congressional defense committees by August 1 on the task force’s goals and objectives.

One noteworthy aspect of this provision is the emphasis that Congress has placed on using sandboxes to facilitate training and familiarization with AI by DOD employees—“from personnel with little technical proficiency to personnel with expert technical proficiency.” Congress should be commended for devoting at least as much attention to that purpose as to how sandboxes are used to develop and test AI tools and models, which is often the main or even exclusive focus of sandboxing. In an organization as large and varied as DOD—and in which the stakes are matters of national security—giving employees a dedicated environment in which to try (and fail) so as to ultimately gain a level of comfort using novel and quickly evolving AI systems is critical to the widespread adoption that Congress is after.

One area where both DOD and Congress might focus some more attention during the required briefing is how the task force can facilitate a pipeline between successful AI development that occurs in sandboxes and the actual implementation of those systems, tools, or methods in the real world of DOD operations. That’s a topic that the provision as written doesn’t address as squarely, but it’ll be key to ensuring that DOD can fully capitalize on its investment in AI sandbox environments. DOD can be a process-heavy place at times; the task force will need to plan for how to judge when AI experiments are ready to graduate from sandboxes, and to efficiently move those successful innovations from sandboxes to the rest of the department.

Key Dates

Section 1513: Physical and Cybersecurity Procurement Requirements for Artificial Intelligence Systems

Section 1513 requires DOD, in collaboration with industry and academia, to develop a framework for the implementation of cybersecurity and physical security standards and best practices for AI systems, “to mitigate risks to [DOD] from the use of such technologies.” The framework must cover enumerated concerns like insider threats, data poisoning, and adversarial tampering. The provision also instructs that the framework must be “risk-based,” drawing on existing reference documents, including NIST’s SP 800 series, and augmenting existing cybersecurity frameworks, including DOD CMMC

To implement the best practices developed under the framework, DOD must amend the Defense Federal Acquisition Regulation Supplement (DFARS) “or take other similar action” ensuring that those practices apply to contractors who engage in AI development, deployment, storage, or hosting. In carrying out that function, DOD must weigh the costs and benefits of imposing security requirements on contractors—and specifically, the costs of “slowing down” AI development and deployment against “the benefits of mitigating national security risks and potential security risks” to DOD.

While this provision is expressly attuned to the potential costs of slowing down AI development through unduly onerous security requirements, it’s at least equally concerned with mitigating the risks to DOD—and national security more generally—that AI systems can pose. It will be worth monitoring how the framework approaches that statutorily required balancing, not least because of how it contrasts with the January 9 AI Strategy memo issued by Secretary Hegseth, which seemingly prized speed above all else. 

Lines from that memo, like “speed wins,” and “We must accept that the risks of not moving fast enough outweigh the risks of imperfect alignment,” offer a preview of where DOD seems most likely to come down on these issues. They also suggest that Congress may have to be dogged in reviewing a required June status update and pursuing other oversight measures to confirm that the statutorily mandated cost-benefit analysis is sufficiently rigorous, with real attention to serious risks Congress mentioned, such as adversarial tampering.

Key Date

Section 1061: Notification of Waivers under DOD Directive 3000.09 

Section 1061 requires DOD to notify congressional defense committees when it has waived DOD Directive 3000.09 (DoDD 3000.09) relating to the use of autonomous weapon systems (AWS).[ref 6] The notification must be in writing and transmitted to the relevant committees within 30 days of when the waiver was issued. The notification also must be unclassified and must include the rationale for the waiver, a description of the weapons system or technology covered by the waiver, and the anticipated duration of the waiver. DOD may include a classified annex to the waiver, as necessary.

DoDD 3000.09 states that “[a]utonomous and semi-autonomous weapons will be designed to allow commanders and operators to exercise appropriate levels of human judgment over the use of force” (emphasis added). As Kelley Sayler of the Congressional Research Service has noted, that does not mean that “manual human ‘control’” of the system is required, but rather mandates “broader human involvement in decisions about how, when, where, and why the weapon will be employed”—for example, “a human must assess the operational environment and decide to deploy the weapon, which can then operate autonomously.” 

As most relevant here, DoDD 3000.09 allows for DOD to skip the traditional review and approval process for AWS when there is an “urgent military need.” Typically, the Under Secretary of Defense for Policy (USD(P)), USD(R&E), and the VCJCS must approve a system before formal development, and then it must be approved again before being deployed in operations by the Under Secretary of Defense for Acquisition and Sustainment, USD(P), and VCJCS.[ref 7] DoDD 3000.09 allows any of these parties to request a waiver of the policy requirements per approval of the Deputy Secretary of Defense.

Section 1061 is the latest in a series of recent NDAA provisions through which Congress has sought greater insight into DoDD 3000.09, particularly whether and how it’s being applied or modified. In the NDAA for fiscal year 2024, Congress required that DOD provide a briefing to congressional defense committees within 30 days of making any changes to DoDD 3000.09, including a description of the change and an explanation of the reasons for it. In fiscal year 2025’s NDAA, Congress required DOD to submit annual reports to those committees through December 31, 2029, on its approval and deployment of lethal AWS under DoDD 3000.09, including any systems that received a waiver from the policy’s review requirement.

This is a prime example of Congress using the NDAA to iterate and build progressively on existing requirements as issues rise in salience—and the salience of DoDD 3000.09 has arguably never been greater. The directive featured prominently in the Pentagon’s dispute with Anthropic earlier this year. Furthermore, NSPM-11 issued by President Trump on June 5 orders Secretary Hegseth to update DoDD 3000.09 within 90 days, and to review it annually “to account for the rapidly evolving capabilities of AI systems” and “ensure the deliberate adoption of AI systems that respect the chain of command and operational authorities.”

In keeping with this progression, one valuable adjustment to Section 1061 that Congress might make would be an amendment that requires an update to the committees when the duration of a waiver is extended beyond the “anticipated” period previously notified, as well as regular updates for any waivers that DOD issues that don’t have a specified end date or timeframe. This would help to guard against overreliance on waivers that might be open-ended or persist for years without prompting congressional scrutiny. Otherwise, waivers issued in prior years might not necessarily show up in the annual reports required under the NDAA for fiscal year 2025.

Going further, Congress could consider whether to codify all or parts of DoDD 3000.09, potentially preserving DOD’s ability to waive or deviate from aspects of the policy when warranted to avoid restrictions that might prove too rigid or become quickly outdated. Both the House and Senate FY27 NDAA markups address DOD AWS policy, though with notable differences. While the final text of any AWS-policy provision in the FY27 NDAA may differ substantially from the markups, these initial versions shed some light on possible approaches. 

The House markup requires that DOD update its AWS policy, including DoDD 3000.09, within one year of enactment—significantly longer than the 90 days DOD has to update the directive under NSPM-11. But as compared to the NSPM, the House markup provides more detail on what an updated policy must include, not least “requirements to preserve existing human command responsibility for the use of force involving autonomous systems or artificial intelligence-enabled systems, including procedures to identify the human commanders or operators responsible for authorizing, supervising, and terminating such use of force.” The Senate markup goes much further still, prescribing an AWS policy and governance regime for DOD in significantly greater detail, with an even more defined substantive floor. And while the Senate markup in multiple places incorporates DoDD 3000.09’s familiar standard of “appropriate levels of human judgment,” it does not directly address the directive’s existing waiver process, leaving it unclear whether that aspect of DoDD 3000.09 would pass muster and thus survive the substantive standards established by this provision.

If Congress opts for a more prescriptive approach, it could consider adding a sunset clause to hedge against the risks of excessive rigidity or obsolescence. A short initial timeline of 1–2 years would prompt Congress to revisit and adjust as needed, providing a short feedback loop for any DOD operational concerns or issues that emerge. 

The Road Ahead: What to Watch for in 2026 and 2027

The FY26 NDAA showed how the annual defense bill can be one of—or even the—primary vehicle for the governance and oversight of defense-relevant AI decisions. Congress can use it to spur prioritization and adoption (Steering Committee and sandboxes), mandate the development of standards and assessments (AI model oversight), prompt consideration and safeguarding against security risks (cybersecurity procurement requirements), and gather information about how the department is using AI (autonomous weapons waivers and various briefing requirements in other provisions). 

Throughout the remainder of 2026 and beyond, it’s worth continuing to monitor updates to key provisions via congressional briefings and other potential disclosures, especially regarding autonomous weapons waivers and the implementation of an AI physical and cybersecurity procurement framework and AI model assessment and oversight. At least one of these initiatives, the AI physical and cybersecurity procurement framework, expressly requires that DOD seek input from groups like industry and academia. Experts should look for opportunities to engage through requests for information or other formats. Congress also has a significant role to play in ensuring that implementation proceeds responsibly and on schedule, using oversight tools like letters, briefing requests, and hearings to supplement the reporting requirements baked into some, but not all, of the key provisions.

As negotiations for the FY27 NDAA ramp up, we can expect numerous AI initiatives to be considered and ultimately included—perhaps even more than last year, since other legislative vehicles will likely be few and far between in this midterm election year. The current House and Senate FY27 NDAA markups include provisions on AI incident and vulnerability reporting within DOD, using AI agents at scale and speed, and promoting competition in AI procurement. The FY27 NDAA could also serve as the vehicle for another attempt at federal preemption of state AI laws, which was dropped shortly before last year’s bill was passed. 

In all of these, Congress should learn from last year’s NDAA. It should craft implementation timelines for DOD that provide space for careful consideration but are not overly long relative to the rapid rate of technological development and diffusion. And it should think about where to build in briefing and other reporting requirements to fill in its knowledge gaps regarding implementation, while being sensitive to the demands they impose on personnel’s time. Doing so helps Congress not only ensure that last year’s initiatives are proceeding according to plan, but also provides valuable insight about unexpected challenges or shortcomings that can inform the coming year’s bill.

First Amendment Questions for AI Transparency Laws

A bipartisan group of lawmakers in the U.S. House of Representatives recently introduced the AI Foundation Model Transparency Act, which would direct the Federal Trade Commission to set transparency requirements for the data used to train high-impact foundation models. Developers would need to provide information about where training data comes from, how models are trained, and whether user data is collected during use. The bill joins a growing roster of artificial intelligence (AI) transparency measures at the state and federal level, including California’s Assembly Bill (AB) No. 2013, which requires developers to publish high-level summaries of their training data; California’s Senate Bill (SB) 53, which requires frontier AI developers to publish safety frameworks and make public disclosures about risk assessment and mitigation measures; and New York’s RAISE Act, which imposes requirements similar to SB 53.

The basic idea behind AI transparency laws is straightforward: As AI plays a larger role in public life, the public should have access to basic information about how these systems are built and the risks they pose. But laws requiring companies to publish information about their AI systems can face First Amendment scrutiny. Legislators drafting disclosure requirements will need to do so with an eye toward how courts may evaluate those laws.

Under U.S. law, when the government compels a company to publish information about its products or services, it regulates the company’s speech. That means transparency laws trigger First Amendment scrutiny. If companies challenge these laws in court, the outcome can turn on what level of scrutiny a court applies—a question currently being litigated in a challenge to California’s AB 2013. How courts answer this question will matter far beyond AB 2013. While AI transparency laws remain viable, First Amendment doctrine is becoming less predictable, and drafting choices matter more than many policymakers assume. The questions that the U.S. Court of Appeals for the Ninth Circuit now confronts will not only affect AB 2013 but also may shape how courts evaluate other AI transparency laws going forward.

xAI’s Challenge to AB 2013

AB 2013 is an early test for how courts may evaluate training data transparency mandates. The statute requires AI developers who provide models in California to post high-level disclosures on their websites, including:

As the law took effect in the new year, Elon Musk’s AI company, xAI, sued in the Northern District of California, claiming that AB 2013 was unconstitutional and seeking to block the law with a motion for a preliminary injunction. On March 4, U.S. District Judge Jesus Bernal of the Central District of California denied xAI’s motion. Even though California won at this stage, the case continues on appeal at the Ninth Circuit.

The appeal may address how courts evaluate disclosure requirements such as AB 2013 as compelled speech under the First Amendment. AB 2013 remains in effect for now, but the state’s victory was tempered by Bernal’s caution that xAI could have “a distinct possibility of prevailing on the merits” of its First Amendment challenge as the case develops, even if the Ninth Circuit affirms the denial of the preliminary injunction. Bernal also subjected AB 2013 to a more searching review than courts have traditionally applied to regulatory disclosure laws.

The Three Standards for Compelled Speech

To understand what happened in the AB 2013 case, it helps to begin with the standards courts use to evaluate laws that principally compel speech, including transparency laws that require companies to disclose information.

How a court characterizes a disclosure requirement—such as whether it qualifies as commercial speech—can determine whether it survives a legal challenge. Two distinct questions therefore matter for transparency laws. First, when does a compelled disclosure count as commercial speech? Second, if it does, which lower standard applies—Central Hudson or Zauderer? For policymakers, the practical point is that the same disclosure requirement can face very different odds in court depending on how a judge answers those two questions.

Narrowing What Commercial Speech Qualifies for Deferential Review

Until relatively recently, courts usually treated factual disclosure requirements imposed on businesses and professionals as compelled commercial speech and reviewed them under Zauderer’s relatively permissive standard. However, several recent developments have unsettled that understanding. That matters because it can make disclosure laws harder to defend.

First, in 2018, the Supreme Court decided National Institute of Family and Life Advocates (NIFLA) v. Becerra, striking down a California law that required crisis pregnancy centers to make disclosures about their services. The Supreme Court rejected the argument that Zauderer applied to the disclosures, despite their regulation of speech by professionals, as they pertained to “abortion, anything but an ‘uncontroversial’ topic.” NIFLA breathed new life into Zauderer’s requirement that compelled disclosures be “uncontroversial.”

After NIFLA, companies challenging disclosure laws can argue that even required factual disclosures shouldn’t get Zauderer’s more deferential treatment when disclosures concern controversial topics. In 2024, the Ninth Circuit blocked a California law requiring social media companies to produce reports on their content moderation policies based on state-specified categories. The lower court had allowed the law under Zauderer, but the appeals panel decided that the law regulated noncommercial speech and applied strict scrutiny instead. The panel suggested that a law could mandate disclosing terms of service and “existing content moderating policies.” But the panel said the state’s reporting framework did more than require disclosure of existing policies. It required companies to address “intensely debated and politically fraught topics, including hate speech, racism, misinformation, and radicalization.”

While that decision invoked the political nature of the compelled speech, other cases in which courts found that speech was too controversial to be “commercial” don’t necessarily implicate hot-button political disputes. For example, in another case, the Ninth Circuit considered a different California law requiring online service providers to report on risks their products pose to children and the steps they take to mitigate them. The court also blocked the law under strict scrutiny, reasoning that the required reports didn’t qualify as “commercial speech” because they required “businesses to opine on and mitigate the risk that children are exposed to harmful content online.” Still, the reports implicated free speech values in a different sense: Identifying and characterizing “harmful” content can raise concerns about censorship and editorial judgment over other’s speech.

Even more remote from considerations of political disputes and free speech values are cases involving scientific information. For example, in 2019, shortly after NIFLA, the Ninth Circuit applied Zauderer and upheld a municipal ordinance requiring retailers to warn consumers that storing a cell phone in pants or a shirt pocket might exceed federal radiofrequency (RF) radiation guidelines. While it acknowledged disagreement over the dangers of RF radiation, the court pointed out that the ordinance did not “force cell phone retailers to take sides in a heated political controversy.” Even so, subsequent Ninth Circuit panels haven’t confined NIFLA to heated political controversies. For example, courts have struck down mandated warnings about cancer risks from acrylamide and glyphosate on grounds that scientific debate over the risk indicated a “controversy,” thereby failing Zauderer.

Recent cases not only have narrowed what counts as “uncontroversial” but also have renewed attention to the limits of what counts as “commercial speech.” The Supreme Court has described commercial speech as speech relating to “the proposal of a commercial transaction,” such as advertising. Courts often look to factors such as whether the speech promotes products or is economically motivated. But courts disagree over how far that category extends beyond direct commercial contexts. This creates uncertainty for AI transparency laws where the required disclosure does not tie closely to economic activity.

Courts also disagree about when Zauderer applies rather than Central Hudson. A broad view treats Zauderer as a more permissive rule for compelled commercial speech and Central Hudson as a more demanding standard for restrictions on commercial speech. Under this approach, Zauderer can apply across a wide range of state interests, so long as the required disclosures are “uncontroversial.” Courts have applied it to warnings about the dangers of cigarettessugar, and cell phone radiation, as well as nation-of-origin labeling on meat products and Securities and Exchange Commission-mandated share buyback rationale disclosures.

Other courts treat Zauderer more as an exception to Central Hudson than as an alternative. If a compelled disclosure doesn’t satisfy Zauderer, then the court may ask whether it can survive under Central Hudson instead. Some decisions avoid the Zauderer question altogether by starting with Central Hudson and upholding the law if it survives that more demanding test.

A more skeptical view, advanced by some judges and businesses challenging disclosure laws, is that Zauderer should be confined to its original anti-deception setting. They argue that Zauderer should, at most, apply only to laws that address some form of misleading or deceptive commercial speech, such as a disclosure meant to correct misleading advertising, because the Supreme Court has applied it only to that setting. This remains a minority position, but if it gains traction, it could unsettle the broader application of Zauderer in product labeling and other disclosure cases.

What this means is that First Amendment doctrine is becoming unsettled along two axes at once: what counts as commercial speech, and what standard applies once speech is treated as commercial.

One way to understand the recent cases is that courts may be drawing an implicit line between two sorts of compelled disclosure laws. In one category are disclosures that suggest a product is harmful or otherwise undercut the company’s own message. Judges may be more cautious about those. In the other category are disclosures that simply tell users how a product works without embedding the government’s own views or judgment. This isn’t a formal doctrinal rule, but it helps explain why some disclosure mandates attract more judicial skepticism than others. It can also guide how AI transparency laws can be designed to be more robust to First Amendment challenges.

The broader point is that AI transparency laws are not doomed, but these laws can no longer be drafted on the basis that factual disclosure requirements will automatically receive deferential review. Recent developments in commercial speech doctrine and Zauderer’s scope make that body of First Amendment law less settled than it once appeared.

How the District Court Evaluated AB 2013 and Why It Matters

Judge Bernal’s opinion understood that Ninth Circuit precedent makes strict scrutiny the default starting point for a direct public disclosure requirement unless the law qualifies as commercial speech. Bernal then asked whether AB 2013 regulated “commercial speech,” which would permit a lower standard of review. He concluded that it likely did, reasoning that the disclosures related to an “actual or potential” commercial transaction because they gave the public information relevant to comparing AI models on the market. He rejected xAI’s argument that AB 2013 was really aimed at rooting out certain kinds of bias in AI training data—an argument xAI supported with statements made in the Senate Floor Analysis. Bernal noted that the bias language came from a supporter’s testimony rather than any legislator’s statements or the statute itself. He emphasized that nothing in the law requires disclosures relating to bias or suggests an effort to steer model outputs.

Bernal then considered whether to apply the more lenient Zauderer standard. While he suggested he might be inclined to find that AB 2013’s disclosures were purely factual and uncontroversial, he reasoned that “this case is even further afield from the original context under which Zauderer arose,” given the Supreme Court’s limited use of Zauderer outside misleading advertising. He therefore applied the more stringent Central Hudson standard and asked whether the law was “more extensive than is necessary” to serve the government’s interest. Bernal concluded that it was too early in the litigation to determine whether AB 2013’s required “high level summaries” would satisfy Central Hudsons tailoring requirement. That matters because Central Hudson gives courts more room to ask whether the legislature required more disclosure than necessary to achieve the law’s objectives, even if some disclosure would be permissible.

California won for now, but the opinion signals some of the First Amendment questions that similar transparency laws may face if challenged. If courts evaluating AI transparency laws follow the same path as the district court in AB 2013—or apply even higher scrutiny, as some Ninth Circuit opinions suggest—future challenges will be harder to predict. Government attorneys defending these laws may need to address not only whether required disclosures are factual and not “controversial” but also whether transparency laws are sufficiently connected to commercial activity and appropriately tailored to the interests they serve.

That’s why the Ninth Circuit’s decision could matter beyond AB 2013: Will courts continue to treat factual disclosure requirements for businesses as commercial speech? Will they limit Zauderer to cases where the government is trying to prevent deception in the marketplace?

Implications for Legislators

For legislatures, the practical lesson is that it increasingly matters how AI transparency laws are framed. The starting point is still what problems the law tries to solve, and what information is needed to address them. But in deciding how to structure public disclosure requirements, legislators may want to keep some basic principles in mind:

In that respect, training data transparency laws may be on firmer footing when they call for disclosure of facts that help consumers evaluate a model, such as its training data sources. These laws are harder to defend when they call for interpretive judgments instead, such as assessments of possible bias. The same principles can apply to other forms of AI transparency.

Most importantly, legislators should keep in mind that this doctrine is in flux. AI transparency laws like AB 2013 are arriving at a moment when at least some courts appear to be rethinking how to assess mandated corporate disclosure. First Amendment doctrine is shaped by whatever disputes reach the Supreme Court. Those disputes often arise in settings far removed from ordinary commercial disclosure and often involve highly charged political topics. But the rules those cases produce don’t usually stay in their original contexts; instead, they apply much more broadly. NIFLA, for example, arose in the context of abortion, yet it now influences how courts review compelled disclosures in very different settings, such as product risk warnings.

So even when a case seems to be unrelated to AI transparency, its outcome and reasoning may still matter. Policymakers interested in transparency laws therefore need to watch a broad range of First Amendment decisions. That’s also a reason to draft carefully. There’s an old legal adage that bad cases make bad law. An ill-designed statute may do more than just lose in court. It may create precedent that makes better statutes harder to sustain.

Whistleblower Protections in SB 53: Strengths, Limitations, and Open Questions

Background

SB 53’s whistleblower provisions are legally entwined with California’s pre-existing whistleblower framework, which they extend, and SB 53’s core transparency requirements, for which they serve as an enforcement mechanism. 

CA Labor Code § 1102.5 is California’s general whistleblowing law, covering all industries and employees. It prohibits employers from retaliating against employees who disclose information, which they have reasonable cause to believe to be a violation of the law, to a government or law enforcement agency, or to a person with authority over the employee. These protections cover a wide range of potential violations.  Nevertheless, California’s protections are narrower than some other states, such as New York, as they (i) only cover employees and do not extend to contractors and (ii) only protect disclosures of violations of the law, and do not protect employees from retaliation for disclosing a substantial and specific risk to public health and safety. This is particularly significant in the context of AI as frontier AI companies frequently rely on contractors for safety-relevant roles such as red-teaming and safety evaluations[ref 2] and novel AI risks often do not constitute clear legal violations, making the ability to report risks to public health and safety essential for early intervention. 

SB 53, or The Transparency in Frontier Artificial Intelligence Act,  signed into law in September 2025, creates a transparency framework for the most powerful AI systems. The Act applies to “frontier developers”, defined as persons who have trained or initiated the training of AI models using extraordinary computing power (greater than 1026 FLOPs) and imposes heightened obligations on “large frontier developers”, defined as  frontier developers with over five hundred million dollars ($500,000,000) in annual revenue. As of 26th January 2026, the only publicly known frontier developers are xAI and OpenAI, while the only large frontier developer is OpenAI.[ref 3] Beyond whistleblower protections, outlined below, SB 53’s main contribution is introducing transparency requirements, spanning four key areas: i) Frontier AI Framework Requirements: Large frontier developers must publish annually-reviewed protocols for managing catastrophic risk; ii) Transparency Reporting Requirements: Developers must publish summary reports of features and risks, similar to existing model cards and system cards, before deploying new or substantially modified models; iii) Government Reporting Requirements: Frontier developers must report critical safety incidents to the OES within 15 days (or 24 hours if posing imminent risk of death or serious injury), with large developers also submitting quarterly risk assessment summaries; iv) Prohibition on materially false statements: Large frontier developers are prohibited from making materially false or misleading statements about catastrophic risks from their models or their management of such risks. 

Crucially, these requirements are legally binding. As such, companies that fail to comply are in violation of California Law and hence whistleblower protections apply to any employee who reports them and are a critical oversight mechanism. 

Policy Analysis

SB 53 extends pre-existing whistleblower protections to cover reporting catastrophic risks as well as legal violations, creates mandatory internal reporting channels, and enables employees to report critical compliance failures directly to authorities. Our analysis proceeds through four dimensions: personal scope, material scope, remedies, and channels. For each aspect of whistleblowing law, we discuss the advantages and limitations of SB 53’s provisions. 

Personal Scope

Personal scope defines who can make a whistleblowing disclosure. Whistleblowing law is part of the California Labor Code and hence only governs employees whose employment contract is under California law. SB 53 creates a tiered protection structure: all employees receive protection when reporting violations of SB 53’s requirements; and a subset, ‘covered employees’, defined as those responsible for assessing, managing, or addressing risk of critical safety incidents, receive protection for reporting catastrophic risks that do not include a legal violation. The scope of ‘covered employee’ remains ambiguous, creating uncertainty about who falls into this protected group. Nevertheless, the breadth of the language suggests that most employees whose work relates to AI safety in some way are likely “covered”.

Advantages

Limitations

Material Scope

Material Scope defines the subject matter that can form the content of a whistleblowing disclosure. SB 53’s key contribution is extending  California’s whistleblower protections beyond legal violations by protecting covered employees from retaliation for reporting “a specific and substantial danger to the public health or safety resulting from a catastrophic risk” with “reasonable cause to believe” such danger exists (§ 1107.1(1)). To understand this scope, two definitions and their thresholds are essential: “critical safety incident” and “catastrophic risk.”

“Critical safety incident” means any of the following:

  1. Unauthorized access to, modification of, or exfiltration of the model weights of a foundation model that results in death, bodily injury, or damage to, or loss of, property.
  2. Harm resulting from the materialization of a catastrophic risk.
  3. Loss of control of a foundation model causing death or bodily injury.
  4. A foundation model that uses deceptive techniques against the frontier developer to subvert the controls or monitoring of its frontier developer outside of the context of an evaluation designed to elicit this behavior and in a manner that demonstrates materially increased catastrophic risk.

“Catastrophic risk” means a foreseeable and material risk that a frontier developer’s development, storage, use, or deployment of a foundation model will materially contribute to the death of, or serious injury to, more than 50 people or more than one billion dollars ($1,000,000,000) in damage to, or loss of, property arising from a single incident involving a foundation model doing any of the following:

  1. Providing expert-level assistance in the creation or release of a chemical, biological, radiological, or nuclear weapon.
  2. Engaging in conduct with no meaningful human oversight, intervention, or supervision that is either a cyberattack or, if committed by a human, would constitute the crime of murder, assault, extortion, or theft, including theft by false pretense.
  3. Evading the control of its frontier developer or user.

The Act expressly excludes risks from the category of catastrophic risks: information generated by a model if it is already publicly available, lawful federal government activity, and situations where the model causes harm in combination with other software, but AI did not materially contribute to the harm. Thus ‘catastrophic risk’ in the Act only refers to CBRN and loss of control risks.

Advantages 

Limitations

Remedies

SB 53 applies existing Labor Code remedies to its whistleblower provisions. This section examines the protections available to whistleblowers and deterrents against retaliation.

Advantages

Limitations

Channels

Whistleblowing channels are the individuals, agencies, or offices to which a whistleblower disclosure can be made. SB 53 protects covered employees if they whistleblow “to the AG, a federal authority, a person with authority over the covered employee, or another covered employee who has authority to investigate, discover, or correct the reported issue” (§ 1107.1(a)). Additionally, SB 53 requires the creation of “a reasonable internal process through which a covered employee may anonymously disclose information to the large frontier developer if the covered employee believes in good faith that the information indicates that the large frontier developer’s activities present a specific and substantial danger to the public health or safety resulting from a catastrophic risk” or a violation of SB 53’s transparency obligations. Under the California Labor Code, any employee is protected if they whistleblow to “a government or law enforcement agency, to a person with authority over the employee or another employee who has the authority to investigate, discover, or correct the violation or noncompliance,” or “any public body conducting an investigation, hearing, or inquiry” (§ 1102.5(b)). California’s Labor Code also provides a whistleblower hotline, operated by the AG, intended to receive calls from whistleblowers. Following a report made to the hotline, the  AG is directed to refer calls received on the hotline to the appropriate government authority for review and possible investigation. 

Advantages

Limitations