Don’t Let AI Developers Hire Their Own Referees
Introduction
A growing chorus of scholars and policymakers favors letting private organizations—rather than a government regulator—govern frontier AI. In the leading family of proposals, the state sets the safety outcomes it wants and licenses independent verification organizations (IVOs) that compete to certify developers against those outcomes. Gillian Hadfield has developed the idea as “regulatory markets,” in which AI developers must pay for oversight from private regulators that governments license and hold accountable for safety standards. Dean Ball, who likens the arrangement to bank supervision, has argued for a version he calls “private governance,” which a nonprofit named Fathom has converted into model legislation. The rationale is that legislators and agencies are poorly positioned to write good safety rules for frontier AI: they understand these systems less well than the labs building them, and rules fixed in advance cannot keep pace as the technology changes. Private verifiers, meanwhile, are closer to the technology than any agency and are disciplined by competition, so they can set better technical standards and keep them current.
The model is no longer hypothetical. The bipartisan FRONTIER Act, introduced in the House in July as the successor to the Great American AI Act discussion draft, would require the largest frontier developers to retain licensed IVOs that audit their risk-management efforts and report to federal overseers. California’s SB 813, backed by Fathom, would have let developers earn a shield from tort liability if they met standards set by a private organization accredited by the state attorney general. It failed this session, but similar proposals are likely to return. Virginia has directed a state commission to study the IVO model for AI regulation. And Connecticut has gone furthest: its omnibus AI law enacted this spring creates a multiyear pilot under which the state consumer-protection department may approve up to five IVOs, whose certification would help companies in court without entirely shielding them from liability.
Unfortunately, as currently structured, IVO-based governance has a key design flaw. Under the regulatory frameworks mentioned above, AI developers would typically select and pay the organizations that certify them, giving IVOs a financial incentive that might clash with high safety standards. This essay will explain how such incentives can interfere with good governance and outline an alternative model of regulation that builds in the right incentives through mandatory insurance.
Generative Gap Filling
Abstract
Most contract litigation turns on contracts that imperfectly record parties’ bargains. When the parties’ dispute can’t be solved by interpreting the text, courts fill the gap. Scholars have long assumed that the remaining text runs out quickly, and provides thin evidence of the actual deal on the disputed point. On that view, a judge who supplies the missing term must be drawing on something else, from commercial defaults to her own policy preferences. Despite generations of work, courts have no real alternative to such unruly methods. We tested that assumption.
Taking real contracts, we masked a term the parties had negotiated and asked readers to predict what we removed. Lay respondents recovered the hidden term about half the time, twice what chance predicts. Law students and lawyers did marginally better. But large language models, given nothing but the rest of the contract, recovered it nearly nine times in ten.
The deal, in short, testifies to far more of the agreement than the literature assumes, including terms the parties never wrote. A contract, we argue, is like a radio signal from far away. Even when incomplete, enough of the message is carried elsewhere that the missing part can be reconstructed with the right receiver. True gaps are rarer than supposed. Courts can weigh model predictions as ordinary, contestable evidence, and parties can discipline the practice with “Choice of Model” clauses.
When Reporting an AI Security Incident Is Not Mandatory
On July 16, Hugging Face, a public platform for open-weight artificial intelligence (AI) models and datasets, disclosed that it had detected a significant cybersecurity breach. An autonomous AI agent had conducted the attack end to end, according to a statement.
Five days later, on July 21, OpenAI revealed that this incident was driven by a combination of agents built on two of its frontier models—GPT-5.6 Sol and a powerful, unreleased model—acting in unanticipated ways during an internal, cyberoffensive capabilities evaluation. For purposes of the evaluation, the researchers had turned off production safety classifiers that block high-risk cyber activity and confined the models to a sandbox, an isolated computing environment without access to the internet, to restrict their interaction with the outside world.
What followed is the first known example of an autonomous cyber incident executed by systems not yet available to the public. Rather than solve the tasks presented, OpenAI’s agents “escaped,” exploiting a previously unknown, zero-day vulnerability. They obtained internet access (the very access OpenAI intended to deny them) and then hacked Hugging Face’s systems to acquire the answers to the benchmark. All of this was seemingly performed without express instruction by humans.
This is not the first time that models have been observed cheating. A few months ago, METR, a nonprofit research organization that conducts evaluations of frontier AI systems, released a report finding that AI models “routinely attempted to cheat on our hardest evaluation tasks, often in flagrant and elaborate ways that we believe humans would not consider.” In one incident, METR reported that an AI model tasked with updating a web app screenshotted a fake version of the app instead of completing the task.
But the Hugging Face breach has struck many observers as more real than these past examples. OpenAI’s models imposed a real cost on an uninvolved third party, all before completing internal testing. While people have at times questioned previous examples of cheating as artificial or contrived, it is hard to imagine that OpenAI expected its agents to escape the confines of their testing environment or to engage in a sophisticated, multistep plan to circumvent their constraints.
Having considered all of these facts, it may come as a surprise that OpenAI might not be legally required to disclose this incident. Certain crucial information is not yet publicly available, and both policymakers and the public will need that information to make sense of what this all means. What about all of those state AI laws with mandatory incident reporting? Don’t they apply here? Many will be disappointed to learn that the answer is arguably “no,” and that even if reporting is mandated, it requires only the scantest of information. What to do about this is the purpose of this article.
Existing AI Transparency Laws and the Hugging Face Breach
The rationale for mandatory incident reporting is straightforward: Some industries have the potential to cause real harm to others, and the government and the public have an interest in learning about high-risk events. In the case of the AI industry, there is a major knowledge gap between the companies’ and governments’ understanding of the technology and its risks. Mandatory incident reporting about serious adverse events, which companies might otherwise be reluctant to disclose, helps close that gap. This rationale is all the more compelling in the context of a rapidly evolving, difficult to predict technology, where best practice and political consensus have yet to develop. Observing real-world incidents offers a path to resolve both political and empirical disagreements and prepares governments to respond to future events.
So did the Hugging Face breach trigger mandatory disclosure under existing incident reporting laws? The answer seems far from clear.
California’s SB 53, New York’s RAISE Act, and Illinois’s SB 315 each require that frontier AI developers report “critical safety incidents”—a term each law defines identically. Of the four reportable incident categories, three require actual harm, ranging from “bodily injury” to “the death of, or serious injury to, more than 50 people or more than one billion dollars ($1,000,000,000) in damage.” (If you’re thinking, “that’s an exceptionally high bar for what is a basic, low-cost reporting requirement” or “it sure seems like governments would want that information before mass harm occurs,” you would not be wrong, but we digress.)
So three of the four incident categories do not apply. That leaves only the fourth, which applies to incidents in which a frontier model “uses deceptive techniques against the frontier developer to subvert the controls or monitoring of its frontier developer outside the context of an evaluation designed to elicit this behavior and in a manner that demonstrates materially increased catastrophic risk.”
It is possible that the Hugging Face breach meets one or more of these elements. It is far less clear that it meets all of them. On the first element, the models may have used deceptive techniques against OpenAI, their frontier developer—they did, after all, try to complete their developer’s evaluation using stolen information, after bypassing restrictions placed on them. But deception is notoriously hard to define, especially if it turns on the “intentions” and “obfuscation” of AI agents. Another read of these events is that the systems simply used all available means of solving the task, and public reporting does not tell us whether the agents attempted to hide those efforts. The second element is also arguably met. While the incident occurred during an evaluation, that evaluation was not “designed to elicit” this specific “deceptive technique.” Based on the ExploitGym benchmark, this evaluation aimed to elicit agentic, cyber-offensive capabilities on a specific task in a controlled environment. It did not, to our knowledge, contemplate—let alone design for—an unexpected cyberattack on a real-world company.
The third element, that the incident “demonstrates materially increased catastrophic risk” would seem to be the most difficult to satisfy. While autonomous cyber capabilities certainly increase the capability and thus potential consequence of agentic action, so too do most capability improvements in AI models. With limited monetary harm and no physical injury, this incident is quite attenuated from future events that might result in the mass physical injury or property damage contemplated by the statute.
With uncertainty at each factor, it is unclear that these existing state laws cover this event. At the very least, it won’t cover all events like it. Stepping back, it seems far from ideal to condition basic incident reporting on a list of complex, highly contested, fact-dependent conditions—all of which must be satisfied simultaneously. In many cases, figuring out whether the incident is indicative of increased risk to the public or actually constitutes deception will not be possible without more information. The purpose of incident reporting is to produce that information, not to require that it be known before a report is ever sent. The chicken must come before the egg. There are better alternatives.
Better Practices for AI Incident Reporting Laws
So where do policymakers go from here? If current incident reporting isn’t providing the needed insight, there are several steps policymakers can take.
Adjust the scope of transparency laws: Lower the exceptionally high bar to basic reporting, gather (at least some) information before harm occurs, and increase visibility into the most capable nonpublic models. To facilitate all of this, rulemaking authority is key.
As we outlined above, most incident reporting laws are simply too narrow. If the only incidents that get reported involve massive damages or loss of human life, the law itself isn’t providing information beyond what the government and public will already know. At the very least, issues of the highest concern, like model theft or loss of control, should be included even in the absence of harm. But existing laws put too much weight on hard-to-pin-down concepts like “loss of control” or “deception” that are difficult to prove and arguably don’t apply in cases like the Hugging Face cyber incident. While loss of control and deception should be sufficient to trigger reporting, they shouldn’t be necessary.
Instead, incident reporting should turn on what information is most likely to update the government or public’s understanding of risks. Information that is unexpected, that is indicative of advanced capabilities in high-risk domains like bio and cyber, or that demonstrates safety and security failures all seem like good candidates for inclusion. Some of these ideas have already made their way into existing proposals. Where the reporting requirements are light touch, as is the case in all existing state AI laws, a wider category of harms can be included. The narrower categories in today’s laws can be saved for more onerous disclosures.
Perhaps even more importantly, transparency laws should increasingly move away from a focus on deployment to a focus on providing visibility into nonpublic models and systems. The Hugging Face breach highlights the need for visibility into nonpublic models. OpenAI deployed systems more capable than anything available to the public, with fewer restrictions, and real-world harm resulted. All of this occurred before any external deployment. This is unlikely to be an isolated event.
Frontier AI developers will likely be the earliest and most sophisticated users of their own models. And the models they deploy internally will most often be more capable than those available to the public. They may be operated with fewer safeguards, especially for evaluations seeking to assess the frontier of capabilities. The gap between the most capable internal and external models may start to grow as companies develop increasingly capable models with dual-use capabilities and as AI systems are used to accelerate their developers’ own AI research and development. In this case, the gap between knowledge inside these companies and outside would expand. Without visibility into the current state of the art, governments will struggle to act effectively or quickly, a problem as much about democratic governance as it is about safety. Visibility into the internal deployment of nonpublic models may become increasingly central to the future of AI governance.
To make all of this work, policymakers will need legislative and regulatory flexibility. Policymakers should expect the exact scope of incident reporting to change over time as societies get a clearer picture of what capabilities and use cases matter most. Because a key goal of incident reporting is to surface novel or unexpected information, some types of events may be less important to report once their dynamics are thoroughly understood and accounted for. To accommodate those changing needs and to provide clarity, narrowly scoped rulemaking authority to refine incident reporting and reporting on nonpublic model use is likely necessary.
Get the details: Make sure reports provide enough information to inform decision-making by providing agencies with rulemaking authority and investigative powers.
As of this writing, the public, and, possibly, some policymakers know remarkably little about the Hugging Face breach. This isn’t a criticism of OpenAI, which voluntarily summarized the event, but the missing details about the event matter. Many commentators have noted that the lessons from and level of concern about this event depend on unknown details. Existing incident reporting laws, requiring little more than the date of the event and a brief summary, are unlikely to provide those detailed answers. Stronger transparency requirements could help answer many key remaining questions.
First is a cluster of questions, posed by Stephen Casper, that can roughly be summarized as “how impressive and/or concerning is the thing I just witnessed”:
- What were the features of the sandbox that the models escaped from? In other words, how hard was it to escape these constraints, and did OpenAI fail to set up an adequately safeguarded environment?
- What prompts and incentives were given for the task? That is, were there strong pressures that precipitated these behaviors, or did they emerge more naturally? How indicative is the incident of the future likelihood that this occurs “in the wild”?
- How significantly did these systems differ from publicly deployed models in their safeguards and affordances? Were they helpful-only models? These questions help us know how well alignment techniques used on publicly deployed models are likely to be.
- How difficult was it to mitigate the incident? How much effort and time was required to stop the continuance of any harms or regain control? These questions are especially key to operationalize terms like “loss of control” and “deception.”
Second, there are questions about the adequacy of the company’s safety practices and preparedness:
- How did OpenAI detect the breach? What was the time frame between the breach and detection, and were there significant delays?
- How did OpenAI mitigate the breach? What was the time frame between detection and mitigation? Did the company’s monitoring practices identify the issue?
- Were any similar incidents observed prior to the event that should have made the developer aware of the risk?
- How quickly did OpenAI alert affected parties after discovering the incident? In the future policymakers may also need to ask companies: How quickly did they alert law enforcement?
- Was the testing and evaluation environment adequately secured?
Third, there are questions about the present and future security of relevant systems:
- How and when does OpenAI intend to restart testing or internal deployment of the model(s) involved? What safeguards, monitoring, or other security measures does it plan to use to avoid another breach?
- Is any part of the harm ongoing?
- What are the remaining areas of uncertainty about the capabilities and risks of the model(s) involved, as well as the security measures designed to address them?
So long as reporting remains almost entirely voluntary, detailed answers to these questions will be hard to come by, and the companies that offer information voluntarily will be subjected to greater scrutiny than those that are less cooperative. Well-crafted rulemaking authority will be key both to ensure that information is adequate and that companies are well informed about their obligations. In many circumstances, authorities will not know all the details they require until after a serious event occurs. In those cases, they will need investigative powers to obtain the needed information.
Reduce reporting costs: To accommodate more robust reporting, design transparency requirements to limit compliance costs, maintain confidentiality, and avoid disincentivizing rigorous risk assessments.
Policymakers can take several important steps to reduce the burden of greater reporting requirements. As assessments and information generation come to focus less on external deployment and more on nonpublic models, periodic reporting becomes more and more attractive. Evaluation organizations like METR have argued that periodic reporting not only saves time but also avoids perverse incentives to rush assessments in the lead-up to deployment. While a small subset of the most severe incidents will require rapid response from law enforcement and others, many events, including those like the Hugging Face breach, may allow for more relaxed reporting timelines. This approach would allow companies to focus on mitigations in the moment and still ensure that they ultimately produce the critical information. In cases where rapid reporting would interfere with mitigation efforts, laws could require only a simple notice of incident, followed by more thorough reporting after the event has been resolved or if officials request it.
Mandating incident reporting or information sharing for nonpublic models can also mitigate perverse incentives as long as minimum requirements are in place. For example, if a company’s reporting obligation triggers only in the context of a risk assessment, but the risk assessment does not have mandatory minimum requirements, the company is incentivized to skip risk assessments or to conduct them less rigorously. Laws can also allow companies to anonymize and aggregate reports, facilitating governments’ information gathering without punishing anyone for proactively uncovering issues.
Finally, any disclosure laws will need clear norms around confidentiality and information sharing to assure companies that their intellectual property and confidential information remains private.
Share information with capable actors: Make sure information is shared with key decision-makers who can assess, verify, and act on it.
Information is only as valuable as the actions it informs. If information from incident reports or internal use assessments sits inside a state agency with limited authority, societies will incur the cost of reporting without most of the benefit. Viewed this way, information sharing is about return on investment. And once the information is generated, most of the cost has been paid. At that point, so long as confidentiality can be maintained, it is incumbent on governments to share this information with the policymakers who most need it. Within states, this will include sharing reports with governors and legislatures to help inform their decision-making and help them target future policy. This information sharing may also be key to spurring political consensus.
In the context of assessing nonpublic models or serious risks to national security or from loss of control, the federal government will often be the central actor. States will often lack the resources, expertise, and political legitimacy to wade in on matters of national security. As we saw in the regulatory response to Mythos, if and when serious national security concerns emerge, the federal government will take the lead.
That’s why it’s confusing that some state laws restrict the ability of states to share information they gather from assessments of companies’ internal use of AI models. If, in fact, these reports generate important information—say, surprising developments in AI research and development or concerning deceptive behavior that doesn’t result in reportable incidents—that information should be shared. It is considerably less useful if locked away in a state agency in Illinois. Internal use and nonpublic model assessments are perhaps the most likely sources of information relevant to national security. To handle that effectively, governments also need the capacity and expertise to process this information. This requires staffing and likely some level of reliance on third-party auditing and assessment. In the wake of incidents like the Hugging Face breach, third-party auditors would be well positioned to conduct the sort of careful fact gathering outlined in previous sections.
Similarly, it is in the interest of the United States and its close allies to share select information related to AI security incidents. Soon enough, and likely far sooner than most U.S. federal or state laws, the EU AI Act will be enforced. What that means, practically, is that the European AI Office will soon receive information that may be useful to public safety and cybersecurity in the United States. Luckily for us, Article 78(5) of the EU AI Act enables information sharing where the European Commission and EU member states create confidentiality agreements with third countries. For its part, the U.K. AI Security Institute conducts crucial research regarding model capabilities and can be a source of trusted expertise for U.S. policymakers. While there will be upfront costs in establishing such a shared information system, failing to make this investment would be a missed opportunity to improve our security at little regulatory cost.
Finally, where appropriate, the government should be empowered to disseminate information to the public and to vulnerable companies. Vulnerabilities found in software, for instance, may affect many different companies and actors, and information sharing will help keep the public safe from the risks they raise. Many of these risks should be discussed in the public square. While worries about information hazards and intellectual property leakage are real, so too is the value of public scrutiny of safety events. There is simply a lot to be learned from the collective scrutiny of the outside world. As the past few days have shown, outside experts have been invaluable in analyzing the Hugging Face breach, and most of them exist outside of government. In tweets and blogs, some of the brightest minds in AI have analyzed the publicly available facts and asked the questions that help us all better understand this event. Many of those questions have come from employees at OpenAI and its competitors, as well as from academics, policy wonks, and online skeptics. Further investigation of the Hugging Face breach will go substantially better because these discussions happened in public. Where possible, policymakers should ensure that these conversations continue to happen in the future.
The Future of Incident Reporting
As this article has perhaps made clear, existing laws fail to prepare us for events like the Hugging Face breach. Most incidents will go unreported. The information that is generated won’t be shared. And, at the end of the day, governments and the public will be repeatedly surprised about developments in this technology. That issue will only get worse as models advance and the gap between internal and externally deployed models widens.
But there is much policymakers can do. Just as this event has brought clarity to how much information societies need to assess these complex and emerging risks, future incidents can inform policy decisions and catalyze moments of political consensus. Policymakers can ensure that concerning incidents and behaviors are reported before harm results and gather information on lapses in security practices. Policymakers can refocus attention on the most capable models likely to be deployed first inside of frontier developers. And policymakers can do all of that while keeping the frequency and urgency of these reports at reasonable levels. If policymakers can do that, and ensure that this information is shared with decision-makers and competent evaluators, societies will be in a position to manage the uncertainty of this technology and make smarter, faster policy decisions in the future.
A Thousand AI Constitutions
Abstract
Today, each AI lab has its own model spec, or constitution. These documents define the values that the labs intend their AIs to have, and the documents are used in post-training to instill those values. This paper argues that the current approach is wrong. Rather than a single constitution, reflecting a single set of moral values, each frontier AI lab should create many different kinds of AIs based on many different constitutions reflecting many sets of values. We give four arguments for constitutional diversification. Diversification mitigates risk, increases political legitimacy, unlocks emergent value, and avoids value lock-in.
Responding to AI Distillation Without Panic
Chinese large language model (LLM) developers are under scrutiny for reportedly employing large-scale “distillation attacks” on U.S. frontier artificial intelligence (AI) models to improve their own systems. Many U.S. actors have sent signals that they consider distillation a serious threat. For example, in May, Anthropic released a policy paper during President Trump’s trip to China, highlighting distillation attacks as a key challenge in U.S.-China competition. In April, the White House issued an official memorandum about distillation, warning about “deliberate, industrial-scale campaigns” from Chinese entities. Also in April, the House Foreign Affairs Committee universally advanced a bill called the Deterring American AI Model Theft Act to address the issue. And others have circulated additional policy proposals.
Discussions of distillation often take for granted that it is a form of theft. But there are key differences between “stealing an AI model” and distillation that policymakers should recognize. To properly address distillation, policy should focus on illegitimate model access—and avoid imposing poorly targeted rules that could harm Americans and distort the open and competitive U.S. AI ecosystem.
What Is Distillation?
The concept of distillation has evolved since it was introduced as a machine learning technique in which a larger “teacher” model’s outputs are used to train a smaller “student” model. Traditionally, that often meant training the student model on the teacher model’s probability distribution over possible outputs, rather than only on the correct answer. Today, the term is used more broadly. “Distillation” also includes prompting a frontier model to generate outputs, and then using the prompt-output pairs—or, where available, reasoning traces—as training data to refine a model. Frontier models may also be used as judges or verifiers for reinforcement learning. Together, these methods improve weaker models by training on stronger models’ responses to prompts and solutions to complex problems.
Distillation is a common practice in contemporary AI development. While on the witness stand at the recent Musk v. Altman trial, Elon Musk acknowledged that xAI had done at least some distillation of OpenAI models and that “generally AI companies distill other AI companies.” As Nathan Lambert, a leading U.S. open-source AI researcher, recently wrote, distillation helps train smaller, often open-source or open-weight models. The White House has recognized this: Office of Science and Technology Policy Director Michael Kratsios pointed out that “AI distillation, when legitimately used to produce” such models, is a “vital part” of creating open models and ensuring a competitive AI ecosystem.
But some Chinese AI developers appear to be using distillation well beyond ordinary practice, accessing U.S. frontier models at a massive scale to do so. In February, Anthropic reported that three Chinese AI labs had generated more than 16 million exchanges with Claude through approximately 24,000 fraudulent accounts, in some cases using jailbreak prompts to extract as much information as possible. OpenAI and Google have also reported or detected similar distillation efforts.
The unusually aggressive distillation efforts of Chinese labs have been portrayed as an attempt at “model theft” and to “steal” the intellectual property of frontier AI labs. But while calling distillation a form of “stealing” or “theft” may make for effective rhetoric, it isn’t an accurate description of how distillation of a closed AI model really works.
Why Isn’t Distillation “Model Theft”?
Distillation doesn’t involve breaking into a developer’s internal system to download the model weights or source code. To a distiller, the model is still a black box. In this context, then, “model theft” would mean some kind of black-box extraction—learning enough about a model from the outputs to approximate model behavior such that it effectively steals the developer’s intellectual property (IP). But what IP would that be?
To start, copyright can be ruled out. The aspects of a model that could plausibly be protected by copyright, such as software code, can’t be copied by distillation. Nor should copyright be used to create a backdoor property right in model outputs. An AI system cannot be an author, and AI-generated outputs are protected only when sufficient human authorship is present. Treating model outputs themselves as copyrighted property of AI labs would create a new right to control downstream uses of text they did not write, raising serious commercial and public policy problems.
Patent rights are also a poor fit. Distillation doesn’t, by itself, copy a patented implementation or allow a distiller to practice a patented method. In any event, the frontier labs themselves haven’t claimed that distillation amounts to patent infringement.
What about trade secrets? AI labs develop and maintain their models in secrecy, which lets them protect many aspects of those models as trade secrets. But distillation typically relies only on information returned through the model’s public-facing interface—the outputs it provides in response to prompts. Trade secret protection requires reasonable efforts to keep information secret, and ordinary outputs are available to anyone with an account. That makes it hard to argue that distillation extracts information qualifying for trade secret protection.
The strongest trade secret theft argument is that mass distillation requires unusual efforts—such as using coordinated proxy accounts—that let a distiller learn more about the model than an ordinary customer could. Mass distillers have also been accused of using jailbreak prompts that elicit information that isn’t normally made public, such as hidden system prompts that guide model responses. But fundamentally, the output being returned is still the kind of output a legitimate user could get. The case would be different if distillers could obtain information like full nonpublic reasoning chains, agent traces, or token-level probability distributions—but there’s no evidence that’s happening.
Compulife Software Inc. v. Newman shows the outer limits of the trade secret argument and why it doesn’t seem to reach mass distillation. The U.S. Court of Appeals for the Eleventh Circuit allowed a trade secret claim involving mass scraping of online life insurance quotes to proceed in that case. The defendant had allegedly acquired enough of the plaintiff’s proprietary database to pose a competitive threat. But the case involved information in a proprietary database and allegations about copying software code—facts not at issue in distillation cases. A recent lawsuit did raise trade secret misappropriation based on jailbreaking as one of its causes of action, but observers noted that the claim was highly questionable (the case settled before reaching the merits). So—at least under current law—the distillation attacks as the frontier AI labs describe them are very unlikely to support a successful trade secret claim.
What’s more, distillation does often violate the AI lab’s terms of service (TOS) for accessing the model. But if every TOS violation counts as “theft,” then the concept has no limiting principle. The more serious legal question is whether mass distillation relies on false identities, misrepresented credentials, or other ways of getting around access limits. That kind of conduct could support a civil or criminal claim under the Computer Fraud and Abuse Act (CFAA). But there are important limits on that, since ordinary TOS violations don’t generally violate the CFAA.
The point is that the distillation itself isn’t an act of stealing an AI model or breaking into an AI lab’s system. Distillers instead are most clearly breaking the law when they take unlawful means to circumvent the safeguards AI labs have in place to prevent distillation.
The more effective way forward, then, is not to treat distillation as theft. Instead, policymakers should focus on securing frontier models against misuse by Chinese competitors and other foreign actors, while studying whether distillation contributes to the dangerous diffusion of model capabilities.
What Anti-Distillation Policy Should Do
Mass distillation merits a policy response, even if it isn’t theft. But policymakers first need to identify the problem they are trying to solve. If the concern is illegitimate access to U.S. frontier models by foreign competitors and state actors, then policy should help labs secure access, share threat information, and identify fraudulent accounts and proxy networks used to disguise who is accessing the model. If the concern is cybersecurity, then the problem is account abuse and getting around access controls. If the concern is the diffusion of dangerous model capabilities, then the first step is to determine whether distillation meaningfully improves those capabilities or helps remove safeguards.
These are public interests that support policies designed to protect access security, enable information sharing, prosecute and sanction unlawful conduct, and evaluate safety. What they don’t justify is measures that effectively provide additional IP protection for AI developers or otherwise restrict legitimate competition in ways that would favor the commercial interests of AI labs over those of the general public.
The most basic defense against unauthorized distillation is for AI labs to recognize when users are circumventing access controls, detect attempts to generate training data, and block outputs to suspicious requests. To succeed, they’ll need to identify patterns of use and other signals that accounts are being used for distillation and shut down their access. One commonsense proposal that the White House and others have suggested is to facilitate coordination and information sharing between frontier AI labs and the government to better prevent illegitimate access for distillation. This could be enabled through antitrust guidance, including guidance based on the existing antitrust exemption for cybersecurity information sharing, which has been extended through Sept. 30, 2026, while Congress considers longer-term reauthorization, or dedicated legislation.
Further light-touch legislation could enable the government to play a more active role in collecting and sharing threat information, identifying proxy and fraudulent account networks, and helping develop best practices. Given geopolitical considerations, sanctions authority, as proposed in the aforementioned Deterring American AI Model Theft Act may be another tool. But sanctions may be more effective as a punitive or foreign policy tool than as a way to stop distillation—and should be balanced against whether they make sense from a trade perspective.
The rhetoric around distillation as a form of IP theft, along with concern that these light-touch legal authorities may be insufficient, has led to interest in activating the United States’ robust trade secret and IP enforcement regime, including against overseas-based actors. These tools include the Economic Espionage Act (EEA), the Defend Trade Secrets Act , and the Protecting American Intellectual Property Act (PAIPA), which provide for criminal, civil, and sanctions tools in cases involving trade secret theft or other covered misconduct.
These tools should be available where there is trade secret theft. But there’s a real risk that defining distillation itself as trade secret theft under the EEA and PAIPA would eventually bleed into private trade secret actions and broader legislative proposals. The debate over IP rights in AI models should be the subject of public debate, not something shaped indirectly through a heated fight over distillation.
A more promising route for legal action against mass distillation is through the CFAA. There are some obstacles though. The ordinary idea of “hacking” is breaking into a system without authorization. But distillation involves accessing a model through ordinary access channels: usernames, passwords, API keys, subscriptions. While distillation violates the provider’s TOS, that, on its own, isn’t likely enough to establish liability under the CFAA. But after the Supreme Court’s 2021 decision in Van Buren v. United States, an ordinary TOS violation is generally not enough to “exceed authorized access” under the CFAA. The Court read the statute as focused on access restrictions—whether the user accessed information in a part of the computer system they were not allowed to access—not on whether the user had an improper purpose for accessing information they were otherwise allowed to see.
That doesn’t mean the CFAA is off the table. AI providers often cut off access upon detecting patterns of use suggesting distillation. Efforts to get around being cut off—through false accounts, misrepresented credentials, proxy access, or other forms of access-control evasion—could implicate the CFAA, creating the potential for civil and even criminal consequences. For example, in United States v. Cuomo, the U.S. Court of Appeals for the Second Circuit affirmed CFAA convictions against defendants who, though using a publicly available state website, bypassed its authentication gate by entering other people’s credentials to extract protected records. CFAA investigations can also be facilitated through information sharing between AI labs and the federal government.
A more sensible response to mass distillation is to use tools that target the conduct around it, rather than beginning to treat model outputs themselves as a form of property. The government can help AI developers make a lot of progress on mass distillation by enabling information sharing and coordination through a targeted antitrust safe harbor. The government can also assist developers build cases under the CFAA and other legal tools where the facts support them. Those measures get at what’s needed to actually prevent unauthorized distillation: detecting fraudulent accounts, spotting efforts to get around access cutoffs, and sharing that information with other labs and the government.
Overall, the labs’ interests ought to be balanced against public interests. Expanding IP rights in frontier model outputs is one way policy gets that balance wrong. Poorly targeted anti-distillation measures could also prevent legitimate open models from competing in the AI marketplace. The concern that distillation might be free-riding on the efforts of frontier labs doesn’t justify excessive limits either. AI developers have themselves benefited from open-source software and published research. Frontier models are trained on the commons of human knowledge, and user interactions and data are used to improve them. When done safely and lawfully, distillation can help keep AI development from becoming dominated by a few companies with disproportionate access to compute and rich stores of user data. Letting frontier labs learn from everyone else—while giving them broad new rights to stop others from learning from model behavior—would only intensify the concentration of AI capabilities and economic power.
Study Whether Distillation Creates Real Safety Risks
There’s an important argument that threats to public safety and national security through the diffusion of more powerful models via distillation would justify stronger measures. Given the risks associated with transformative AI, researchers, policymakers, and civil society should take this seriously. The problem is that we don’t know enough about how much distillation contributes. Does distillation meaningfully advance near-frontier models, or does it mostly benefit smaller models—or make marginal contributions such as by validating performance? Anthropic reported that DeepSeek had far fewer exchanges (150,000) with Claude than Moonshot (3.4 million) or MiniMax (13 million). That suggests DeepSeek’s limited distillation may have been more on par with what “every AI company does,” per Musk. Yet DeepSeek is among China’s most powerful AI models.
Given that uncertainty, the sensible response is further investigation, potentially by the Center for AI Standards and Innovation (CAISI) at the National Institute of Standards and Technology, to determine how much distillation actually contributes to the threat. CAISI could study whether distillation materially improves dangerous capabilities, whether it transfers or strips away safeguards, whether already-available open-weight models can provide the same uplift, and what kinds of model access are most likely to matter. If that work shows distillation poses a meaningful safety risk, then stronger measures could be justified. In the meantime, restraint is warranted to avoid policy errors driven by perceived threats that outpaced the evidence.
The bottom line is this: The threats associated with distillation are best addressed by targeting fraudulent access and efforts to circumvent access controls, and empowering companies to cooperate on measures to prevent illegitimate access. Creating new quasi-IP rights in model outputs, or other premature or disproportionate responses, would do more to protect AI companies’ interests than the public’s. Policy should protect U.S. people and businesses by targeting real harms and unlawful conduct—not speculative ones.
Congress Should Do Something: The Case for (Fixing) the Great American AI Act
Since the April announcement of Anthropic’s Mythos model and its unprecedented cyber capabilities, there has been a remarkable shift in the artificial intelligence (AI) policy discourse. This has been most noticeable, and most noticed, in the statements and actions of prominent Trump administration officials. After months of dismissing concerns about national security risks from AI and engaging with the issue primarily by attempting to preempt state AI safety laws, the White House recently issued an executive order that called for the establishment of a voluntary predeployment program headed by the National Security Agency to evaluate offensive cyber capabilities of frontier models. This came amid statements from senior administration officials about “striking a … balance between innovation and safety” and even considering a Food and Drug Administration-style mandatory predeployment licensing regime for frontier models.
On June 12, the Trump administration’s concerns about Mythos’s cyber capabilities boiled over into an unprecedented decision to use export control authorities to prohibit Anthropic from allowing foreign nationals to access its Mythos-class Fable 5 model. Practically, this amounted to a mandate that Anthropic revoke public access to the model entirely. Much of the online commentary on this decision devolved into speculation about the administration’s motivations, the alleged behavior of Anthropic’s executives, and other petty interpersonal drama. As intriguing as these are, the more important takeaway from the White House’s decision to abruptly institute a de facto licensing regime for frontier AI systems—as many commentators across the political and safety/innovation spectrums have observed—is that federal legislation to establish a framework for addressing the national security risks posed by the most advanced AI systems is an urgent necessity.
The Commerce Department’s decision to impose export controls on Fable may or may not have been wise, depending on who you believe about the seriousness of the vulnerability that motivated the decision. But even assuming the decision was justified, the fact that the government was apparently caught by surprise and had to scramble to put together a heavy-handed response based on ad hoc, potentially legally questionable authorities that were not designed with anything like frontier AI systems in mind, with no due process for the affected company, is a serious problem that should be remedied with legislation as soon as possible.
Which brings us, finally, to the subject of this piece—the Great American Artificial Intelligence Act of 2026 (GAAIA), a comprehensive frontier AI safety bill that is the long-awaited product of months of intense negotiation between Rep. Jay Obernolte (R-Calif.) and Rep. Lori Trahan (D-Mass.). Obernolte had tried for months to get a Democrat to sign on to an AI bill that preempted state AI laws before finally persuading Trahan. For her part, Trahan wrote that she was motivated to support the bill by the announcement of Mythos’s groundbreaking cyber capabilities.
The current version of GAAIA is a discussion draft, meaning it has not been introduced yet and is intended to spark a conversation and elicit feedback from stakeholders rather than to become law in its current form. It may seem somewhat strange that a discussion draft sponsored by two relatively junior members of the House, neither of whom appears to have the backing of their party’s leadership, should receive so much attention from the media and from AI policy commentators, but for once the buzz is warranted.
GAAIA is the best attempt to design a federal framework for the governance of frontier AI systems introduced to date. In other words, it is the first serious attempt to actually do the thing that last week’s Fable incident clearly shows is necessary—address the national security risks posed by the most advanced AI systems—in a transparent, legally sound, and democratically legitimate way. While the bill may not pass in the near future—it likely faces opposition from both Democrats and Republicans—the draft can tell us a great deal about what the future of federal and state frontier AI governance efforts may look like.
The bill in its current form falls short in a number of respects and should not be passed. That said, passing a similar bill with narrower preemption of state laws and somewhat stronger federal authorities would be an excellent first step toward a workable federal regime for governing frontier AI systems.
What the Bill Does
GAAIA is a bipartisan compromise, in the truest sense of the phrase, which means that everyone hates it. Obernolte, a longtime advocate for federal preemption of state AI laws who supported last summer’s “moratorium” (which would have preempted all nongenerally applicable state AI laws and replaced them with essentially nothing), has compromised by granting his seal of approval to a number of genuinely consequential affirmative policy proposals. Trahan, who strongly supported increased oversight of AI in the past, has compromised by accepting broad preemption of state AI laws.
The Federal Framework
GAAIA’s four titles contain 45 sections, each of which addresses a significant topic in AI policy. I am an AI safety guy, and my research focuses mostly on serious risks that advanced AI systems might pose to national security and public safety, so this article focuses on the sections of GAAIA that are relevant to those risks. However, GAAIA is not an AI risk bill exclusively. There are also sections on, for example, “Preparing K-12 educators and students for an AI literate future,” “Modernizing access to artificial intelligence-related labor market data,” and establishing a “National artificial intelligence research resource” for improving capacity for AI research in the U.S., among many others. This article does not discuss those sections, not because they’re not important, but because they’re mostly irrelevant to the catastrophic risk concerns that this piece focuses on.
The noteworthy catastrophic risk provisions in GAAIA are:
- Section 102, which codifies and authorizes $100 million in annual funding for the Center for AI Standards and Innovation (CAISI), which would be relocated outside of the National Institute of Standards and Technology (NIST) and given an expanded, quasi-regulatory role as the agency in charge of administering the independent verification organization (IVO) and transparency regimes created by Sections 111 and 112.
- Section 111, which imposes transparency and incident reporting requirements on frontier AI companies similar to the requirements imposed by California’s SB 53 or New York’s RAISE Act.
- Section 112, which authorizes CAISI to establish and administer an IVO auditing regime in which independent third-party companies would regularly evaluate the adequacy of AI companies’ risk mitigation efforts.
- Section 301, which reauthorizes and updates the expiring Cybersecurity Information Sharing Act of 2015.
- Section 411, which directs NIST and the Department of Energy to lead efforts to form “alliances or coalitions” with allied foreign governments in order to facilitate collaboration and cooperation on AI research and development, technical standard-setting, and related issues.
“Transparency,” in this context, means requiring frontier AI companies to publish frontier safety frameworks (documents describing how the company evaluates and addresses catastrophic risks from the company’s most advanced AI models) and model cards (documents accompanying the release of specific models that contain information about the capabilities and limitations of a model and the results of the safety evaluations conducted under the company’s safety framework). “Incident reporting” requirements mandate that companies report “critical safety incidents” (essentially, incidents in which a frontier model’s model weights are stolen or in which a model does something scary that seems catastrophic-risk-ish) to the government and/or to law enforcement. And “auditing” refers to the practice of having an independent third party evaluate the adequacy of a company’s catastrophic risk mitigation practices as well as the company’s compliance with transparency requirements and with its own frontier AI framework. Transparency and incident reporting requirements are intended to provide the information needed for the government and the public to understand how companies think about and address catastrophic risks, and auditing requirements are supposed to ensure that transparency and reporting requirements remain effective rather than being ignored or becoming meaningless box-checking exercises.
As regards GAAIA’s transparency and auditing provisions, one common take is that the risk mitigation benefits of establishing these programs would be marginal because similar requirements already exist at the state level in New York, California, and Illinois. That view is, I think, mistaken in two important respects. For one thing, as Anton Leicht points out, it is vitally important to build up regulatory capacity within the federal government. I co-wrote an essay on this topic a few weeks back. The argument is, essentially:
- AI might end up being a very big deal with extremely serious national security implications at some point in the next 10 years (and possibly within the next two years).
- If that happens, we should expect that serious regulatory interventions may be required.
- That serious regulatory work will almost certainly have to be carried out by the federal government, because the federal government—
- has orders of magnitude more regulatory capacity, expertise, and resources to devote to complex regulatory tasks than state governments do; and
- is, constitutionally and practically speaking, the only entity that can realistically be entrusted with extremely complex and high-stakes national security projects.
- Building up the institutional capacity and expertise to competently undertake complex regulatory tasks is difficult and cannot realistically be done in a matter of days or even months.
- Therefore, it is vitally important that we begin the process of aggressively building up technical expertise and regulatory capacity and know-how within the federal government as soon as possible.
But even setting aside the capacity-building considerations, it is simply not true that existing state catastrophic risk laws are equivalent to GAAIA’s transparency or auditing provisions. These state laws—California’s Transparency in Frontier Artificial Intelligence Act (TFAIA), New York’s Responsible Artificial Intelligence Safety and Education (RAISE) Act, and Illinois’s Artificial Intelligence Safety Measures Act (AISMA)—are an important foundation for future efforts and have been extremely influential. GAAIA itself is clear evidence of this influence; some of Section 111’s transparency provisions are lifted almost verbatim from the transparency provisions of SB 53 (which are substantially identical to the transparency provisions of RAISE and AISMA). But, as groundbreaking as those state laws are, they are still state laws and therefore cannot leverage the resources, institutions, or legal authorities of the federal government in the way that a bill like GAAIA can.
Perhaps the most important institutional advantage that GAAIA leverages is the capacity of federal agencies such as the Department of Commerce to carry out sophisticated rulemaking, a capacity built over decades of administering complex regulatory programs that no state agency can realistically match. GAAIA grants CAISI and the Department of Commerce broad authority to issue regulations fleshing out the auditing and transparency regimes outlined in Sections 111-112. Rulemaking! That word may not sound like the most exciting thing you’ve heard this week, but take my word for it: This is the good stuff.
Take auditing, for example. Because AISMA does not confer any explicit rulemaking authority, Illinois’s auditing regime, when it goes into effect, will be defined solely by the requirements in AISMA’s text. AISMA requires that audits be conducted “consistent with generally accepted auditing standards and best practices” and that auditors possess “demonstrated competence to perform the audit.” These vague requirements, however, aren’t enough to guarantee a functional auditing regime.
Under AISMA, auditors are paid by the AI company that retains them. By default, this system will lead to a race to the bottom in which market forces compel auditors to compete with each other over who can cause the least hassle and difficulty for their customers (frontier AI companies). Rather than ensuring that companies abide by their commitments and hew to responsible risk mitigation practices, this kind of auditing regime will eventually devolve into a system where companies are disincentivized from hiring rubber-stamp auditors only by the uncertain prospect of ex post tort liability.
To be clear, this is not a criticism of AISMA’s auditing provisions, which are well designed. The issue is that Illinois’s state government simply lacks the ability to design and competently administer a complex, technically involved auditing program for out-of-state tech companies. The issue is not that AISMA is insufficiently ambitious but, rather, that Springfield—on its best day—has only a small fraction of the capacity for complex interstate regulatory projects that the federal government has on its worst.
In contrast to AISMA, GAAIA’s auditing section could establish a functional and effective third-party auditing regime. CAISI—reestablished as a regulatory agency separate from the nonregulatory NIST—would be granted broad rulemaking authority. The regulations that CAISI would be required to promulgate would include rules addressing conflict of interest and funding transparency requirements for auditors, requirements for licensing auditors and revoking auditor licenses, minimum requirements for audits and assessments, and “any other rules reasonably necessary to the administration of the IVO oversight and licensing regime.” This is the kind of rulemaking and oversight authority that could, in theory and with competent implementation, actually establish the kind of auditing regime that AISMA gestures at.
GAAIA’s transparency requirements would also be a significant upgrade from the existing state transparency requirements. While GAAIA’s transparency section is similar on its face to existing California, New York, and Illinois transparency statutes, it delegates fairly broad rulemaking authority to the Department of Commerce, which can prescribe regulations governing, among other things, the “form, manner, and minimum quality” of the model cards and safety frameworks that companies are required to publish. In combination with GAAIA’s auditing requirements, which require auditors to regularly evaluate and assess the adequacy of an AI company’s safety framework and its other efforts to identify and mitigate catastrophic risks, GAAIA’s transparency requirements would allow Commerce to ensure that transparency requirements actually result in meaningful transparency.
This rulemaking and minimum-standard-setting authority would allow GAAIA to provide significantly more transparency than existing state laws. Consider California’s TFAIA, which has been in effect for just over five months. TFAIA is a light-touch statute by design and imposes very few obligations on companies. This light-touch approach is, in my opinion, a good thing, but one downside is that companies can technically comply by publishing documents that check the statutorily required boxes without actually saying anything meaningful about the company’s approach to mitigating risks. For example, xAI, despite founder Elon Musk’s frequent public statements about the existential risks posed by superintelligence, complies with TFAIA by publishing a barebones framework that describes xAI’s risk mitigation practices in cursory and general terms.
Under TFAIA, there is no realistic way to require xAI to provide the industry-standard level of transparency that its competitors’ frameworks typically demonstrate. And while the RAISE Act does provide New York’s Department of Financial Services with some rulemaking authority that could in theory be used to give the act’s transparency requirements more teeth, practical and constitutional limitations may prevent a New York state financial services agency from regulating California-based software companies with the same level of rigor and precision that the U.S. Department of Commerce could, in theory, bring to bear.
Of course, the value proposition of GAAIA depends on the assumption that CAISI and the Commerce Department will do a decent job of establishing and administering the proposed transparency and auditing regimes. It’s far from clear that the Commerce Department would view this as a top priority, given that Secretary of Commerce Howard Lutnick has generally signaled skepticism of AI safety concerns. While public reporting in the months since Anthropic’s Mythos announcement has documented a shift among some of Lutnick’s fellow Cabinet members toward taking some of these concerns more seriously, Lutnick’s views have not—at least publicly—evolved along similar lines.
Even if the Commerce Department’s approach is initially ineffective, however, it might still be good to establish the statutory bones of an effective federal oversight regime in case future developments create additional political pressure for governance efforts. Given Mythos’s impact in shifting the Overton window on AI policy, the Trump administration may become more willing to implement serious oversight measures as they receive further concrete evidence of significant national security risks. Of course, it would be better not to have to cross our fingers and hope that the government starts taking risks seriously before anything bad happens. Still, enacting something like GAAIA would, at least, meaningfully improve the legal authorities available to the executive branch when and if the political will to get serious about national security risks from AI manifests itself.
GAAIA’s affirmative provisions are not perfect. The lack of effective whistleblower protections, for example, is a serious defect. The gold standard for federal AI whistleblower legislation is Sen. Chuck Grassley’s (R-Iowa) AI Whistleblower Protection Act (AI WPA). GAAIA’s whistleblower section falls short of that standard because, unlike the AI WPA, it fails to protect AI company employees who disclose information about a “substantial and specific” danger to public health, public safety, or national security to an appropriate government agency. The only AI whistleblowers who are protected from retaliation by GAAIA’s whistleblower section are those who report a violation of federal law—and because state whistleblower laws (including in California, where most of the relevant frontier AI company employees reside) already protect this kind of disclosure, a redundant federal protection adds little value. As I’ve argued before, protection for disclosures about serious risks that don’t involve law violations is crucial because it’s very plausible that perfectly legal frontier AI development activities could lead to serious national security and public safety risks that the government ought to be aware of.
Additionally, GAAIA’s auditing section lacks sufficiently clear enforcement authorities. Auditors are given a wide variety of authorities to assess whether a developer’s practices mitigate risks sufficiently. It’s not entirely clear, however, what happens if the mitigations are inadequate or nonexistent. CAISI is authorized to penalize developers for, for example, failing to retain an auditor or failing to grant timely access to appropriate records, but a developer who checks these boxes may not be required to actually do anything meaningful in response to an auditor’s recommendations. It’s possible that CAISI could use the broad rulemaking authority that the auditing section grants to cobble together a solution to this problem, but given how skeptical courts have been of this kind of broad agency interpretation of congressional delegations of rulemaking authority in recent years, it would be far better to have a clear statutory enforcement hook.
It should be noted that Trahan’s office has signaled potential willingness to remedy both of the specific issues discussed above in future drafts of GAAIA. If this does happen, it will improve my view of the bill’s value significantly; the point of a discussion draft is to bring potential issues like these to light so that they can be hashed out. More generally, a flawed whistleblower or auditing provision is at least marginally better than having no federal AI whistleblower or auditing statute whatsoever. As good as GAAIA’s federal framework is, however, it comes at a steep price: broad preemption of all state laws regulating AI development.
Preemption
In exchange for the substantial federal framework described above, GAAIA preempts—for a period of three years—any state AI law “specifically regulating the development of any artificial intelligence model.” As in prior AI preemption efforts, “generally applicable” laws are exempted, and GAAIA also includes a somewhat opaque exemption for “post-deployment activities.” “Development” is defined to mean
the acts performed or directed by a developer with respect to an artificial intelligence model prior to its deployment, including determining training or fine-tuning objectives; training, fine-tuning, or otherwise substantially modifying the weights or other parameters of an artificial intelligence model; and evaluating and deciding, prior to deployment, whether an artificial intelligence model satisfies applicable safety or capability thresholds for deployment.
There’s currently no consensus regarding exactly which existing or hypothetical state laws or bills GAAIA would preempt. A few commentators have argued that this language would preempt only the few existing state frontier AI safety laws—Illinois’s AISMA, California’s TFAIA, and New York’s RAISE Act—and future state laws in the same vein. If this were true and could be demonstrated clearly, an overwhelming majority of the AI safety community and a number of other factions who currently oppose the bill (including, for example, child safety advocates) would likely change course to support or assume a neutral attitude toward it. One-to-one preemption—eliminating state AI catastrophic risk auditing, transparency, and incident reporting bills in exchange for establishing robust federal AI auditing, transparency, and incident reporting regimes—would be a very reasonable compromise. If Reps. Trahan and Obernolte believe that GAAIA’s preemption section is effectively one-to-one, a clarifying amendment would go a long way toward increasing support for the bill.
Unfortunately, a court would probably not accept such a narrow interpretation. The key phrase is “specifically regulating”—state laws will be preempted if they are deemed to “specifically regulate” development. This phrase does not appear in any existing preemption legislation, meaning that there’s no firm precedent that can be used to accurately predict how courts will interpret the scope of GAAIA’s preemption. Still, a few relatively safe conclusions can be drawn.
It seems clear that “specifically regulating” is a relatively narrow formulation, as preemption triggers go—it certainly preempts less than the broad “relating to,” for example, which would sweep in any state law that had a “connection with” or contained a “reference to” AI. Still, “specifically regulating” is broader than, for example, the CAN-SPAM Act’s “expressly regulates,” which imposes a facial test (i.e., preempts a state law only if the law, on its face, names and regulates the forbidden federal subject). Arguably, GAAIA’s preemption language would instead impose a functional test under which state laws would be preempted if they had the effect of regulating AI development, as defined.
This functional interpretation would be consistent with how courts have treated comparable language in cases examining the (admittedly broader) preemption language in the Energy Policy and Conservation Act, the Clean Air Act, the Federal Meat Inspection Act, and the Federal Aviation Administration Authorization Act. In cases addressing the preemptive effect of the other laws mentioned above, courts have applied functional tests to prevent states from circumventing preemption with clever legislative drafting. For instance, in National Meat Association v. Harris, the Supreme Court held that a provision of California law that regulated the sale of meat from inhumanely raised pigs was preempted by a federal preemption provision that prohibited states from regulating the operation of slaughterhouses because the California provision because “the sales ban … functions as a command to slaughterhouses,” and because “if the sales ban were to avoid the … preemption clause, then any State could impose any regulation on slaughterhouses just by framing it as a ban on the sale of meat produced in whatever way the State disapproved.”
In short, it seems unlikely in light of existing preemption precedents that federal courts will allow states to circumvent GAAIA preemption through creative legislative drafting choices. GAAIA’s express carve-out for “post-deployment activities” and the relatively narrow scope of “specifically regulating” would likely protect some state laws that affect development without explicitly targeting it. But any law that functions primarily to regulate development, or can realistically be complied with only by altering development practices, will likely be preempted even if it purports to regulate deployment.
Because GAAIA’s definition of “development” is quite broad, this functional interpretation would preempt far more than just TFAIA-style catastrophic risk transparency laws. Existing state measures that GAAIA would almost certainly preempt include Texas’s TRAIGA (which, among other development-focused regulations, prohibits developing an AI system that impersonates a child while describing sexual conduct or developing an AI system with the intention of producing child sexual abuse material or illegal deepfakes), California’s AB2013, Colorado’s SB 189 (the law that replaced Colorado’s controversial algorithmic bias bill with less onerous, more business-friendly requirements), and certain California Consumer Privacy Act regulations affecting “automated decisionmaking technology.” State measures that might (or might not) be preempted include the many state child safety chatbot laws, like Idaho’s Conversational AI Safety Act, that purport to regulate deployers or “operators” of AI systems but could realistically be satisfied only via interventions implemented during training or fine-tuning, both of which are “development” under GAAIA.
By far the more important issue with GAAIA’s broad preemption, however, is that, in addition to invalidating these existing laws, it would preempt a broader and more important category of future state laws that would otherwise be passed in 2027, 2028, and 2029. TFAIA, RAISE, and AISMA were always meant to lay the foundation for future efforts rather than to establish a stable end state for state AI policy. As the capabilities of the most advanced AI systems continue to improve, their risks will increase as well. And as risks become more immediate and more difficult to deny, it seems safe to predict that state AI laws will be enacted that would have been politically unrealistic in 2026. The technology’s benefits may outweigh these risks; even so, it would be foolish to assume that AI will be the first transformatively impactful general-purpose technology, the development of which does not create any local harms that state legislation and regulation need to address.
The Takeaway
GAAIA includes by far the best federal framework for frontier AI safety that has been publicly introduced to date, as a discussion draft or otherwise. This being the case, some of the criticisms of GAAIA seem somewhat misguided. The Democratic House AI Commission, for example, has asserted that the bill “does not meet the enormity of the moment.” This may well be true, but if GAAIA’s unprecedentedly ambitious framework does not meet the moment, what bill does? I sincerely hope that there is a tidal wave of enormity-addressing AI legislation waiting in the wings, but the federal legislation that has been introduced thus far (with a few notable exceptions, such as the AI Whistleblower Protection Act and the Artificial Intelligence Risk Evaluation Act) has mostly been notable for its inoffensiveness rather than its moment-meeting boldness. A lot of “convening a multi-stakeholder process to consider the development of a process” for setting up a purely voluntary incident reporting regime, and things of that nature.
Despite this, the discussion draft would, in my judgment, be net-negative if enacted in its current form. As I have argued elsewhere, federal preemption of state AI laws should proceed on a narrow, issue-by-issue basis. Broad preemption of a vaguely defined and poorly understood category of state laws would, in all likelihood, be a disaster, for both political and policy reasons. Politically, it makes no sense to try to preempt broad categories of state AI law that have the backing of politically potent constituencies—developer-focused child safety laws, for example—in exchange for a bill with great frontier AI safety provisions but no child safety provisions. And while state frontier AI safety laws are not an adequate substitute for a robust federal framework, state laws have thus far proved so much easier to pass than federal laws that passing GAAIA in its current state might mean locking in a good-but-ultimately-inadequate 2026 governance framework indefinitely, despite the fact that future capabilities improvements might require more ambitious legislation.
At the same time, I worry many stakeholders who are rejecting GAAIA’s approach out of hand fail to recognize the urgency of the situation. By default, without new legislation, the process for addressing serious risks from frontier systems will be undertaken by the executive branch, the intelligence community, and national security agencies in a haphazard, legally suspect, case-by-case, and increasingly securitized manner, as June’s Fable incident proved. There is, of course, a sense in which serious national security risks being addressed by the nation’s national security apparatus are expected and necessary. But a program that is made up on the fly and administered on a totally discretionary, case-by-case basis, without the resources or structure that only Congress can provide, will never be as effective or as democratically legitimate as a well-designed federal legislative framework could be.
My hope, therefore, is that GAAIA’s authors will narrow its preemption and improve its substance and that GAAIA’s critics will either engage with the process of trying to improve the bill or introduce a serious alternative proposal in the near future. The stakes are high enough, and the issues with the status quo are clear enough, that continuing to do nothing indefinitely is no longer a defensible course of action.
The NDAA: A Key Vehicle for AI Governance
Summary
- The National Defense Authorization Act (NDAA) is one of the few “must-pass” bills in Congress every year, which makes it a key opportunity for AI legislation.
- The Fiscal Year 2026 NDAA (FY26 NDAA) contained nearly two dozen artificial intelligence-related provisions.
- While many of those provisions focus on accelerating adoption, others require DOD[ref 1] to develop standards, frameworks, and other policy measures to govern its AI use.
- Among the most notable AI-related provisions from the FY26 NDAA are its requirements:
- To create an AI Futures Steering Committee, through which senior Pentagon officials will formulate DOD policy for the evaluation, governance, and risk mitigation of advanced AI and artificial general intelligence (AGI); and
- To develop a standardized assessment framework for AI models currently used by DOD, along with department-wide guidelines to facilitate procurement of future AI models.
- The Fiscal Year 2027 NDAA (FY27 NDAA) markups include provisions related to autonomous weapons policy, AI procurement, and the AI capabilities of adversaries.
Introduction
The National Defense Authorization Act has quietly become one of Congress’s most powerful tools for shaping AI policy, and the FY26 NDAA featured many key AI provisions. This commentary compiles all the major AI provisions from the FY26 NDAA and analyzes the most significant language in detail. With some of the initial deadlines imposed by the FY26 NDAA now having passed—and with negotiations around the FY27 NDAA underway—it’s useful to take stock of the potential and pitfalls of these provisions.
The NDAA is not just restricted to the nuts and bolts of defense operations. It has also been used to achieve broader policy goals, sometimes by limiting the executive branch’s actions. For example, one of the most important and successful nonproliferation programs in history, the Nunn-Lugar Cooperative Threat Reduction Program, was originally proposed as an amendment to the NDAA and was subsequently expanded through the NDAA.[ref 2] More recently, the FY19 NDAA effectively banned the government from using certain Chinese telecommunications companies such as Huawei and ZTE; likewise, the FY26 NDAA bans certain foreign AI products like DeepSeek.
As the government increasingly prioritizes AI use in warfighting and military operations, the NDAA has a key role to play in shaping AI policy.
AI Governance in the FY26 NDAA
Notable Provisions
Several provisions stand out as particularly important for the government’s broader interest in overseeing and fostering the responsible development of secure AI systems:
- The Artificial Intelligence Futures Steering Committee;
- The AI Model Assessment and Oversight framework;
- Digital Sandboxes for AI;
- Physical and Cybersecurity Procurement Requirements for AI Systems; and
- The Autonomous Weapons Waiver Policy.
Section 1535: Artificial Intelligence Futures Steering Committee
Section 1535 requires DOD to create an Artificial Intelligence Futures Steering Committee (Steering Committee) to prepare DOD for advanced AI and AGI. The Steering Committee will be co-chaired by the Deputy Secretary of Defense and Vice Chairman of the Joint Chiefs of Staff (VCJCS). It will primarily be composed of principal deputies of the military services, relevant under secretaries (e.g., the Under Secretary of Defense for Research and Engineering (USD(R&E))), and others responsible for AI (e.g., the Chief Digital and AI Officer (CDAO)).[ref 3]
By January 31, 2027, the Steering Committee must submit a report to Congress covering what can be described as two main focus areas. First, the committee must help prepare DOD for advanced AI and AGI by creating:
- A proactive policy for the evaluation, adoption, governance, and risk mitigation of advanced AI systems, including systems that approach or achieve AGI.
- An analysis of the forecasted trajectory of advanced AI models and enabling technologies that could lead to AGI such as AI agents, neuromorphic computing, cognitive science applications, infrastructure needs, new microelectronics, etc.
- An analysis of the potential operational effects of integrating advanced AI or AGI into DOD networks and systems from a technical, doctrinal, training, and resourcing perspective to better understand effects on operational commands.
- A strategy for the risk-informed adoption, governance, and oversight of advanced AI and AGI including ethical, policy, and technical guardrails to maintain appropriate human decision-making and prevent misuse.
The second focus area is U.S. adversaries. Though not specifically named, the People’s Republic of China (PRC), which is actively pursuing AI capabilities that rival those of the United States, is likely the primary focus. The committee must assess the possible technological, operational, and doctrinal trajectories of U.S. adversaries with respect to AI capabilities, including the pursuit of AGI. Additionally, the committee must analyze the threat landscape associated with the use of advanced AI and AGI and develop options to counter these threats.
Within the Pentagon’s sprawling bureaucracy, there’s often fierce competition between different programs and priorities for funding and attention from leadership. In this sense, the Steering Committee could be a valuable forcing function for the department to prepare for advanced AI, reinforced by the requirement to report its findings to Congress by early 2027. There’s precedent for DOD using these sorts of committees as a way to spur action on issues such as software modernization and autonomous systems.
However, such committees sometimes serve more as a signaling mechanism for Congress than as a catalyst for serious action. Unless chairs or members of the committee invest their time and professional capital to drive it forward, it can easily devolve into a box-checking exercise. While the Steering Committee’s substantive mandate is broad, its required procedural actions, as set by Congress, are fairly minimal: meet at least once every three months, and submit a report on its findings to the relevant congressional committees by January 31, 2027. That means that depending on when the committee is actually established and how quickly it first convenes, it may meet only three or four times before its report is due.
On top of that, the NDAA provision does not allocate any dedicated staff or budget for the Steering Committee. Any resources must be drawn from existing reserves, which could further limit its capacity. Given those constraints and their already-full plates, the Steering Committee’s principals might be tempted to delegate their roles and responsibilities down the chain of command to other, typically less-empowered subordinates, whose remit might be narrower—i.e., drafting a report that satisfies Congress’s requirements while potentially tabling thornier policy disagreements or implementation details for later.
Congress should remain attuned to these possible failure modes and use its oversight power to solicit information about the Steering Committee and its progress, in the hopes of helping it gain and maintain momentum. There are some encouraging signs on this front. In March, Senator Jim Banks sent a letter to Secretary Hegseth requesting a staff-level briefing within 60 days to discuss DOD’s plans for the Steering Committee. The letter suggested areas of focus with respect to U.S.-PRC AI competition. Even just one or a few members of Congress taking specific, sustained interest in the Steering Committee could keep it high enough on DOD’s long list of priorities to increase its odds of success.
Congressional oversight can be particularly valuable in two ways. First, it can keep pressure on the committee if it fails to meet the report submission deadline of January 31. Second, and perhaps more importantly, Congress can help ensure that the Steering Committee doesn’t waste the 11 months between its reporting deadline and its termination date of December 31, 2027.
While the report is the Steering Committee’s most tangible required deliverable, Congress provided that the committee will continue to exist for nearly a year beyond the report submission deadline. This time would allow the committee to refine or update its policies and to work on implementing and disseminating the findings throughout DOD. Because the Steering Committee lacks deliverables or other measurable benchmarks throughout most of 2027, it’ll likely be incumbent on Congress to use tools like letters, hearings, and requests for briefings to push forward that updating and implementation work. These efforts could ultimately have a much greater impact on DOD operations in the long term than just the drafting of the report itself.
As of June 30, 2026, no public materials indicate whether the Steering Committee was established by the April 1 statutory deadline, or whether it has held its first meeting. That’s not necessarily cause for concern, as DOD is not required by the NDAA to report those actions to Congress or the public. But it does make it harder to predict which of these paths the AI Futures Steering Committee will ultimately follow. Overall, this provision could pay dividends by prompting DOD to proactively prepare for major threats and opportunities raised by AGI—planning that might otherwise get neglected—though its success is far from assured.
Key Dates
- 4/1/2026: Deadline to establish the Steering Committee
- 1/31/2027: Steering Committee’s report due to Congress
- 12/31/2027: Steering Committee terminates
Section 1533: AI Model Assessment and Oversight
Section 1533 instructs DOD to create a Cross-Functional Team (Team) for AI “model assessment and oversight.” The Team must develop a standardized assessment framework for AI models currently used by DOD, as well as guidelines to facilitate procurement of future models. The Team is led by the CDAO and composed of other DOD technology leaders, such as CIOs, CAIOs of the combatant commands, service acquisition executives, and USD(R&E). The Team must:
- Develop a “standardized assessment framework” for AI models currently used by DOD, including: performance standards, development documentation, testing procedures, compliance with ethical principles, assessment and validation methodologies, and security and compliance requirements under FedRAMP.[ref 4]
- Establish department-wide “guidelines” for evaluating future AI models being considered for use.
- Create “governance structures” for the development, assessment, testing, and deployment of models.
- Determine assessment levels for models based on “ultimate use case-based risk.”
- Establish “mechanisms” for intra-agency collaboration regarding the development, testing, assessment, and deployment of AI models.
- Develop processes for the submission, review, and approval of use cases for AI models.
This provision allows DOD to retain a lot of discretion over how it evaluates current and future AI models. Congress has mandated that DOD establish a framework and protocols, but didn’t set substantive thresholds for performance. That’s understandable to some degree, given the risk of setting standards via legislation, which might quickly become outdated and then prove difficult to adjust. And it’s similar to the approach that states like California and New York have taken in enacting frontier AI transparency reporting requirements. But some key requirements in the provision, such as the creation of “governance structures” and assessing “ultimate use-case-based risk,” use terms that are undefined and open to interpretation, and could have benefited from a bit more congressional guidance about the elements that should at least be considered or addressed.
That vagueness, combined with the long timelines the provision establishes, could make it hard for Congress to assess the Team’s progress. Congress notably gave the Team an extended timeline to develop its model assessments and oversight, which may be in tension with the pace of AI progress. The standardized assessment framework isn’t due until June 2027—a year and a half after enactment—and no actual assessments of DOD’s major AI systems are required until January 2028. Meanwhile, new frontier AI models are released many times a year.
To be sure, the Team’s task is difficult. And Congress sometimes errs by giving agencies unrealistically short deadlines. But a failure to keep up with the pace of AI development risks undermining the Team’s purpose. To frame that risk, consider the events that have transpired since the FY26 NDAA passed six months ago. First, there was the blow-up over contract terms between the Pentagon and Anthropic in February. More recently, the June 5 National Security Presidential Memorandum (NSPM) 11 ordered Secretary Hegseth, ODNI, and IC elements to “review and update procurement processes to ensure the rapid onboarding of the most advanced AI models from multiple vendors” within 120 days. It’s unclear how or whether this review will be coordinated with the procurement guidelines that the Team is tasked with developing on its longer timeframe.
Here, again, Congress can deploy its oversight tools to steer DOD in the direction of consistent and streamlined guidelines for AI procurement. It should aim to ensure that standards are applied uniformly and transparently, not reactively, to AI developers. Helpfully, this provision requires DOD to provide a briefing to congressional defense committees within 30 days of hitting significant statutorily prescribed milestones, starting with its establishment of the Team on or before June 1, 2026. That offers a natural opening for Congress to probe the Team’s trajectory, and potentially to spur a course correction if needed. Congress might consider incorporating that sort of regular briefing requirement into future AI-related NDAA provisions; it’s particularly beneficial in this area due to rapid and sometimes unexpected jumps in capabilities and risks, and might also have been helpful for similar initiatives like the Steering Committee discussed above.
Finally, it’s worth a closer look at the provision’s definition of “major [AI] system”—one of only a few terms that the provision does actually define—buried near the end of the provision. That definition limits coverage to systems used annually by at least 500 users within DOD, and excludes systems used solely for research, development, testing, or evaluation that have not been deployed for operational use. Elsewhere, the provision specifies that DOD must assess all major AI systems using the standardized assessment framework, leaving it somewhat unclear when or to what extent that framework also governs assessment of other AI models used by DOD. In other words, for models used by less than 500 employees per year, or those involved only in R&D, how will DOD assess performance, security, and “compliance with ethical principles”?
While this sort of line-drawing exercise is almost always difficult but necessary for administrability, in this instance the exclusions arguably represent the frontier of DOD’s own AI development and deployment in what could end up being the highest-stakes and hardest-to-monitor situations. At minimum, it’s plausible that some of the most powerful systems, deployed in potentially highly consequential cases, might be available to only a small number of users. Congress should ask DOD how it plans to assess AI systems that fall into those categories and potentially require the development of standards for such systems in future legislation.
Key Dates
- 6/1/2026: Deadline to establish Cross-Functional Team
- 1/1/2027: Deadline to designate Functional Leads for specialized functional, operational, or subject-matter areas within DOD
- 6/1/2027: Deadline for Cross-Functional Team to complete development of standardized assessment framework and governance structure
- 1/1/2028: Deadline to complete assessment of major AI systems used by DOD
- 12/31/2030: Cross-Functional Team terminates[ref 5]
Section 1534: Digital Sandbox Environments for AI
Section 1534 requires the CDAO to create a task force to promote AI sandbox environments supporting “experimentation, training, familiarization, and development.” The task force should “identify, coordinate, and advance” DOD efforts to develop and deploy AI sandboxes, with an eye toward accelerating AI adoption across the department. The provision defines an “[AI] sandbox environment” as a “secure, isolated computing environment that enables users with varying levels of technical proficiency to access [AI] tools, models, and capabilities for the purposes of experimentation, training, testing, and development without affecting operational systems or requiring specialized technical knowledge to operate.” The provision requires that the task force be established by April 1, 2026, and that the CDAO provide a briefing to congressional defense committees by August 1 on the task force’s goals and objectives.
One noteworthy aspect of this provision is the emphasis that Congress has placed on using sandboxes to facilitate training and familiarization with AI by DOD employees—“from personnel with little technical proficiency to personnel with expert technical proficiency.” Congress should be commended for devoting at least as much attention to that purpose as to how sandboxes are used to develop and test AI tools and models, which is often the main or even exclusive focus of sandboxing. In an organization as large and varied as DOD—and in which the stakes are matters of national security—giving employees a dedicated environment in which to try (and fail) so as to ultimately gain a level of comfort using novel and quickly evolving AI systems is critical to the widespread adoption that Congress is after.
One area where both DOD and Congress might focus some more attention during the required briefing is how the task force can facilitate a pipeline between successful AI development that occurs in sandboxes and the actual implementation of those systems, tools, or methods in the real world of DOD operations. That’s a topic that the provision as written doesn’t address as squarely, but it’ll be key to ensuring that DOD can fully capitalize on its investment in AI sandbox environments. DOD can be a process-heavy place at times; the task force will need to plan for how to judge when AI experiments are ready to graduate from sandboxes, and to efficiently move those successful innovations from sandboxes to the rest of the department.
Key Dates
- 4/1/2026: Deadline to establish Task Force on AI sandbox environments
- 8/1/2026: Deadline for Task Force to brief congressional defense committees on goals and objectives
- 1/1/2030: Task Force terminates
Section 1513: Physical and Cybersecurity Procurement Requirements for Artificial Intelligence Systems
Section 1513 requires DOD, in collaboration with industry and academia, to develop a framework for the implementation of cybersecurity and physical security standards and best practices for AI systems, “to mitigate risks to [DOD] from the use of such technologies.” The framework must cover enumerated concerns like insider threats, data poisoning, and adversarial tampering. The provision also instructs that the framework must be “risk-based,” drawing on existing reference documents, including NIST’s SP 800 series, and augmenting existing cybersecurity frameworks, including DOD CMMC.
To implement the best practices developed under the framework, DOD must amend the Defense Federal Acquisition Regulation Supplement (DFARS) “or take other similar action” ensuring that those practices apply to contractors who engage in AI development, deployment, storage, or hosting. In carrying out that function, DOD must weigh the costs and benefits of imposing security requirements on contractors—and specifically, the costs of “slowing down” AI development and deployment against “the benefits of mitigating national security risks and potential security risks” to DOD.
While this provision is expressly attuned to the potential costs of slowing down AI development through unduly onerous security requirements, it’s at least equally concerned with mitigating the risks to DOD—and national security more generally—that AI systems can pose. It will be worth monitoring how the framework approaches that statutorily required balancing, not least because of how it contrasts with the January 9 AI Strategy memo issued by Secretary Hegseth, which seemingly prized speed above all else.
Lines from that memo, like “speed wins,” and “We must accept that the risks of not moving fast enough outweigh the risks of imperfect alignment,” offer a preview of where DOD seems most likely to come down on these issues. They also suggest that Congress may have to be dogged in reviewing a required June status update and pursuing other oversight measures to confirm that the statutorily mandated cost-benefit analysis is sufficiently rigorous, with real attention to serious risks Congress mentioned, such as adversarial tampering.
Key Date
- 6/16/2026: Deadline for DOD to submit an update to the congressional defense committees on the status of implementing the requirements of the Physical and Cybersecurity Procurement provision
Section 1061: Notification of Waivers under DOD Directive 3000.09
Section 1061 requires DOD to notify congressional defense committees when it has waived DOD Directive 3000.09 (DoDD 3000.09) relating to the use of autonomous weapon systems (AWS).[ref 6] The notification must be in writing and transmitted to the relevant committees within 30 days of when the waiver was issued. The notification also must be unclassified and must include the rationale for the waiver, a description of the weapons system or technology covered by the waiver, and the anticipated duration of the waiver. DOD may include a classified annex to the waiver, as necessary.
DoDD 3000.09 states that “[a]utonomous and semi-autonomous weapons will be designed to allow commanders and operators to exercise appropriate levels of human judgment over the use of force” (emphasis added). As Kelley Sayler of the Congressional Research Service has noted, that does not mean that “manual human ‘control’” of the system is required, but rather mandates “broader human involvement in decisions about how, when, where, and why the weapon will be employed”—for example, “a human must assess the operational environment and decide to deploy the weapon, which can then operate autonomously.”
As most relevant here, DoDD 3000.09 allows for DOD to skip the traditional review and approval process for AWS when there is an “urgent military need.” Typically, the Under Secretary of Defense for Policy (USD(P)), USD(R&E), and the VCJCS must approve a system before formal development, and then it must be approved again before being deployed in operations by the Under Secretary of Defense for Acquisition and Sustainment, USD(P), and VCJCS.[ref 7] DoDD 3000.09 allows any of these parties to request a waiver of the policy requirements per approval of the Deputy Secretary of Defense.
Section 1061 is the latest in a series of recent NDAA provisions through which Congress has sought greater insight into DoDD 3000.09, particularly whether and how it’s being applied or modified. In the NDAA for fiscal year 2024, Congress required that DOD provide a briefing to congressional defense committees within 30 days of making any changes to DoDD 3000.09, including a description of the change and an explanation of the reasons for it. In fiscal year 2025’s NDAA, Congress required DOD to submit annual reports to those committees through December 31, 2029, on its approval and deployment of lethal AWS under DoDD 3000.09, including any systems that received a waiver from the policy’s review requirement.
This is a prime example of Congress using the NDAA to iterate and build progressively on existing requirements as issues rise in salience—and the salience of DoDD 3000.09 has arguably never been greater. The directive featured prominently in the Pentagon’s dispute with Anthropic earlier this year. Furthermore, NSPM-11 issued by President Trump on June 5 orders Secretary Hegseth to update DoDD 3000.09 within 90 days, and to review it annually “to account for the rapidly evolving capabilities of AI systems” and “ensure the deliberate adoption of AI systems that respect the chain of command and operational authorities.”
In keeping with this progression, one valuable adjustment to Section 1061 that Congress might make would be an amendment that requires an update to the committees when the duration of a waiver is extended beyond the “anticipated” period previously notified, as well as regular updates for any waivers that DOD issues that don’t have a specified end date or timeframe. This would help to guard against overreliance on waivers that might be open-ended or persist for years without prompting congressional scrutiny. Otherwise, waivers issued in prior years might not necessarily show up in the annual reports required under the NDAA for fiscal year 2025.
Going further, Congress could consider whether to codify all or parts of DoDD 3000.09, potentially preserving DOD’s ability to waive or deviate from aspects of the policy when warranted to avoid restrictions that might prove too rigid or become quickly outdated. Both the House and Senate FY27 NDAA markups address DOD AWS policy, though with notable differences. While the final text of any AWS-policy provision in the FY27 NDAA may differ substantially from the markups, these initial versions shed some light on possible approaches.
The House markup requires that DOD update its AWS policy, including DoDD 3000.09, within one year of enactment—significantly longer than the 90 days DOD has to update the directive under NSPM-11. But as compared to the NSPM, the House markup provides more detail on what an updated policy must include, not least “requirements to preserve existing human command responsibility for the use of force involving autonomous systems or artificial intelligence-enabled systems, including procedures to identify the human commanders or operators responsible for authorizing, supervising, and terminating such use of force.” The Senate markup goes much further still, prescribing an AWS policy and governance regime for DOD in significantly greater detail, with an even more defined substantive floor. And while the Senate markup in multiple places incorporates DoDD 3000.09’s familiar standard of “appropriate levels of human judgment,” it does not directly address the directive’s existing waiver process, leaving it unclear whether that aspect of DoDD 3000.09 would pass muster and thus survive the substantive standards established by this provision.
If Congress opts for a more prescriptive approach, it could consider adding a sunset clause to hedge against the risks of excessive rigidity or obsolescence. A short initial timeline of 1–2 years would prompt Congress to revisit and adjust as needed, providing a short feedback loop for any DOD operational concerns or issues that emerge.
The Road Ahead: What to Watch for in 2026 and 2027
The FY26 NDAA showed how the annual defense bill can be one of—or even the—primary vehicle for the governance and oversight of defense-relevant AI decisions. Congress can use it to spur prioritization and adoption (Steering Committee and sandboxes), mandate the development of standards and assessments (AI model oversight), prompt consideration and safeguarding against security risks (cybersecurity procurement requirements), and gather information about how the department is using AI (autonomous weapons waivers and various briefing requirements in other provisions).
Throughout the remainder of 2026 and beyond, it’s worth continuing to monitor updates to key provisions via congressional briefings and other potential disclosures, especially regarding autonomous weapons waivers and the implementation of an AI physical and cybersecurity procurement framework and AI model assessment and oversight. At least one of these initiatives, the AI physical and cybersecurity procurement framework, expressly requires that DOD seek input from groups like industry and academia. Experts should look for opportunities to engage through requests for information or other formats. Congress also has a significant role to play in ensuring that implementation proceeds responsibly and on schedule, using oversight tools like letters, briefing requests, and hearings to supplement the reporting requirements baked into some, but not all, of the key provisions.
As negotiations for the FY27 NDAA ramp up, we can expect numerous AI initiatives to be considered and ultimately included—perhaps even more than last year, since other legislative vehicles will likely be few and far between in this midterm election year. The current House and Senate FY27 NDAA markups include provisions on AI incident and vulnerability reporting within DOD, using AI agents at scale and speed, and promoting competition in AI procurement. The FY27 NDAA could also serve as the vehicle for another attempt at federal preemption of state AI laws, which was dropped shortly before last year’s bill was passed.
In all of these, Congress should learn from last year’s NDAA. It should craft implementation timelines for DOD that provide space for careful consideration but are not overly long relative to the rapid rate of technological development and diffusion. And it should think about where to build in briefing and other reporting requirements to fill in its knowledge gaps regarding implementation, while being sensitive to the demands they impose on personnel’s time. Doing so helps Congress not only ensure that last year’s initiatives are proceeding according to plan, but also provides valuable insight about unexpected challenges or shortcomings that can inform the coming year’s bill.
First Amendment Questions for AI Transparency Laws
A bipartisan group of lawmakers in the U.S. House of Representatives recently introduced the AI Foundation Model Transparency Act, which would direct the Federal Trade Commission to set transparency requirements for the data used to train high-impact foundation models. Developers would need to provide information about where training data comes from, how models are trained, and whether user data is collected during use. The bill joins a growing roster of artificial intelligence (AI) transparency measures at the state and federal level, including California’s Assembly Bill (AB) No. 2013, which requires developers to publish high-level summaries of their training data; California’s Senate Bill (SB) 53, which requires frontier AI developers to publish safety frameworks and make public disclosures about risk assessment and mitigation measures; and New York’s RAISE Act, which imposes requirements similar to SB 53.
The basic idea behind AI transparency laws is straightforward: As AI plays a larger role in public life, the public should have access to basic information about how these systems are built and the risks they pose. But laws requiring companies to publish information about their AI systems can face First Amendment scrutiny. Legislators drafting disclosure requirements will need to do so with an eye toward how courts may evaluate those laws.
Under U.S. law, when the government compels a company to publish information about its products or services, it regulates the company’s speech. That means transparency laws trigger First Amendment scrutiny. If companies challenge these laws in court, the outcome can turn on what level of scrutiny a court applies—a question currently being litigated in a challenge to California’s AB 2013. How courts answer this question will matter far beyond AB 2013. While AI transparency laws remain viable, First Amendment doctrine is becoming less predictable, and drafting choices matter more than many policymakers assume. The questions that the U.S. Court of Appeals for the Ninth Circuit now confronts will not only affect AB 2013 but also may shape how courts evaluate other AI transparency laws going forward.
xAI’s Challenge to AB 2013
AB 2013 is an early test for how courts may evaluate training data transparency mandates. The statute requires AI developers who provide models in California to post high-level disclosures on their websites, including:
- General sources and characteristics of training data.
- How datasets relate to the system’s intended purpose.
- Approximate size of the data (expressed in ranges or estimates).
- Whether the data includes copyrighted or licensed material.
- Whether personal or aggregate consumer information is involved.
- Whether synthetic data are used.
As the law took effect in the new year, Elon Musk’s AI company, xAI, sued in the Northern District of California, claiming that AB 2013 was unconstitutional and seeking to block the law with a motion for a preliminary injunction. On March 4, U.S. District Judge Jesus Bernal of the Central District of California denied xAI’s motion. Even though California won at this stage, the case continues on appeal at the Ninth Circuit.
The appeal may address how courts evaluate disclosure requirements such as AB 2013 as compelled speech under the First Amendment. AB 2013 remains in effect for now, but the state’s victory was tempered by Bernal’s caution that xAI could have “a distinct possibility of prevailing on the merits” of its First Amendment challenge as the case develops, even if the Ninth Circuit affirms the denial of the preliminary injunction. Bernal also subjected AB 2013 to a more searching review than courts have traditionally applied to regulatory disclosure laws.
The Three Standards for Compelled Speech
To understand what happened in the AB 2013 case, it helps to begin with the standards courts use to evaluate laws that principally compel speech, including transparency laws that require companies to disclose information.
- Strict scrutiny is the default for content-based speech regulation. Laws that require a speaker to say something specific are usually evaluated as content-based regulations because they dictate the content of the speech. Viewpoint-based regulations—favoring one view over another—are an especially disfavored subset of content-based regulation. Under strict scrutiny, the government must show the law is “narrowly tailored to serve compelling state interests.” In practice, laws rarely survive it.
- Intermediate scrutiny applies to regulations on “commercial speech.” The Supreme Court described the test for intermediate scrutiny as applied to commercial speech in the Central Hudson case: The court asks whether the “government’s interest is substantial,” whether “the regulation directly advances” that interest, and “whether it is not more extensive than is necessary to serve that interest.” While the Central Hudson standard is easier than strict scrutiny, and laws can survive it, it’s still a form of “heightened” scrutiny.
- Zauderer review is the most permissive. It also applies to commercial speech, though some disagree about its scope. The Zauderer test permits compelled disclosures that are “purely factual and uncontroversial” and “reasonably related” to a substantial government interest, and not “unjustified or unduly burdensome.” In practice, when courts determine that Zauderer applies, most disclosure laws survive.
How a court characterizes a disclosure requirement—such as whether it qualifies as commercial speech—can determine whether it survives a legal challenge. Two distinct questions therefore matter for transparency laws. First, when does a compelled disclosure count as commercial speech? Second, if it does, which lower standard applies—Central Hudson or Zauderer? For policymakers, the practical point is that the same disclosure requirement can face very different odds in court depending on how a judge answers those two questions.
Narrowing What Commercial Speech Qualifies for Deferential Review
Until relatively recently, courts usually treated factual disclosure requirements imposed on businesses and professionals as compelled commercial speech and reviewed them under Zauderer’s relatively permissive standard. However, several recent developments have unsettled that understanding. That matters because it can make disclosure laws harder to defend.
First, in 2018, the Supreme Court decided National Institute of Family and Life Advocates (NIFLA) v. Becerra, striking down a California law that required crisis pregnancy centers to make disclosures about their services. The Supreme Court rejected the argument that Zauderer applied to the disclosures, despite their regulation of speech by professionals, as they pertained to “abortion, anything but an ‘uncontroversial’ topic.” NIFLA breathed new life into Zauderer’s requirement that compelled disclosures be “uncontroversial.”
After NIFLA, companies challenging disclosure laws can argue that even required factual disclosures shouldn’t get Zauderer’s more deferential treatment when disclosures concern controversial topics. In 2024, the Ninth Circuit blocked a California law requiring social media companies to produce reports on their content moderation policies based on state-specified categories. The lower court had allowed the law under Zauderer, but the appeals panel decided that the law regulated noncommercial speech and applied strict scrutiny instead. The panel suggested that a law could mandate disclosing terms of service and “existing content moderating policies.” But the panel said the state’s reporting framework did more than require disclosure of existing policies. It required companies to address “intensely debated and politically fraught topics, including hate speech, racism, misinformation, and radicalization.”
While that decision invoked the political nature of the compelled speech, other cases in which courts found that speech was too controversial to be “commercial” don’t necessarily implicate hot-button political disputes. For example, in another case, the Ninth Circuit considered a different California law requiring online service providers to report on risks their products pose to children and the steps they take to mitigate them. The court also blocked the law under strict scrutiny, reasoning that the required reports didn’t qualify as “commercial speech” because they required “businesses to opine on and mitigate the risk that children are exposed to harmful content online.” Still, the reports implicated free speech values in a different sense: Identifying and characterizing “harmful” content can raise concerns about censorship and editorial judgment over other’s speech.
Even more remote from considerations of political disputes and free speech values are cases involving scientific information. For example, in 2019, shortly after NIFLA, the Ninth Circuit applied Zauderer and upheld a municipal ordinance requiring retailers to warn consumers that storing a cell phone in pants or a shirt pocket might exceed federal radiofrequency (RF) radiation guidelines. While it acknowledged disagreement over the dangers of RF radiation, the court pointed out that the ordinance did not “force cell phone retailers to take sides in a heated political controversy.” Even so, subsequent Ninth Circuit panels haven’t confined NIFLA to heated political controversies. For example, courts have struck down mandated warnings about cancer risks from acrylamide and glyphosate on grounds that scientific debate over the risk indicated a “controversy,” thereby failing Zauderer.
Recent cases not only have narrowed what counts as “uncontroversial” but also have renewed attention to the limits of what counts as “commercial speech.” The Supreme Court has described commercial speech as speech relating to “the proposal of a commercial transaction,” such as advertising. Courts often look to factors such as whether the speech promotes products or is economically motivated. But courts disagree over how far that category extends beyond direct commercial contexts. This creates uncertainty for AI transparency laws where the required disclosure does not tie closely to economic activity.
Courts also disagree about when Zauderer applies rather than Central Hudson. A broad view treats Zauderer as a more permissive rule for compelled commercial speech and Central Hudson as a more demanding standard for restrictions on commercial speech. Under this approach, Zauderer can apply across a wide range of state interests, so long as the required disclosures are “uncontroversial.” Courts have applied it to warnings about the dangers of cigarettes, sugar, and cell phone radiation, as well as nation-of-origin labeling on meat products and Securities and Exchange Commission-mandated share buyback rationale disclosures.
Other courts treat Zauderer more as an exception to Central Hudson than as an alternative. If a compelled disclosure doesn’t satisfy Zauderer, then the court may ask whether it can survive under Central Hudson instead. Some decisions avoid the Zauderer question altogether by starting with Central Hudson and upholding the law if it survives that more demanding test.
A more skeptical view, advanced by some judges and businesses challenging disclosure laws, is that Zauderer should be confined to its original anti-deception setting. They argue that Zauderer should, at most, apply only to laws that address some form of misleading or deceptive commercial speech, such as a disclosure meant to correct misleading advertising, because the Supreme Court has applied it only to that setting. This remains a minority position, but if it gains traction, it could unsettle the broader application of Zauderer in product labeling and other disclosure cases.
What this means is that First Amendment doctrine is becoming unsettled along two axes at once: what counts as commercial speech, and what standard applies once speech is treated as commercial.
One way to understand the recent cases is that courts may be drawing an implicit line between two sorts of compelled disclosure laws. In one category are disclosures that suggest a product is harmful or otherwise undercut the company’s own message. Judges may be more cautious about those. In the other category are disclosures that simply tell users how a product works without embedding the government’s own views or judgment. This isn’t a formal doctrinal rule, but it helps explain why some disclosure mandates attract more judicial skepticism than others. It can also guide how AI transparency laws can be designed to be more robust to First Amendment challenges.
The broader point is that AI transparency laws are not doomed, but these laws can no longer be drafted on the basis that factual disclosure requirements will automatically receive deferential review. Recent developments in commercial speech doctrine and Zauderer’s scope make that body of First Amendment law less settled than it once appeared.
How the District Court Evaluated AB 2013 and Why It Matters
Judge Bernal’s opinion understood that Ninth Circuit precedent makes strict scrutiny the default starting point for a direct public disclosure requirement unless the law qualifies as commercial speech. Bernal then asked whether AB 2013 regulated “commercial speech,” which would permit a lower standard of review. He concluded that it likely did, reasoning that the disclosures related to an “actual or potential” commercial transaction because they gave the public information relevant to comparing AI models on the market. He rejected xAI’s argument that AB 2013 was really aimed at rooting out certain kinds of bias in AI training data—an argument xAI supported with statements made in the Senate Floor Analysis. Bernal noted that the bias language came from a supporter’s testimony rather than any legislator’s statements or the statute itself. He emphasized that nothing in the law requires disclosures relating to bias or suggests an effort to steer model outputs.
Bernal then considered whether to apply the more lenient Zauderer standard. While he suggested he might be inclined to find that AB 2013’s disclosures were purely factual and uncontroversial, he reasoned that “this case is even further afield from the original context under which Zauderer arose,” given the Supreme Court’s limited use of Zauderer outside misleading advertising. He therefore applied the more stringent Central Hudson standard and asked whether the law was “more extensive than is necessary” to serve the government’s interest. Bernal concluded that it was too early in the litigation to determine whether AB 2013’s required “high level summaries” would satisfy Central Hudson’s tailoring requirement. That matters because Central Hudson gives courts more room to ask whether the legislature required more disclosure than necessary to achieve the law’s objectives, even if some disclosure would be permissible.
California won for now, but the opinion signals some of the First Amendment questions that similar transparency laws may face if challenged. If courts evaluating AI transparency laws follow the same path as the district court in AB 2013—or apply even higher scrutiny, as some Ninth Circuit opinions suggest—future challenges will be harder to predict. Government attorneys defending these laws may need to address not only whether required disclosures are factual and not “controversial” but also whether transparency laws are sufficiently connected to commercial activity and appropriately tailored to the interests they serve.
That’s why the Ninth Circuit’s decision could matter beyond AB 2013: Will courts continue to treat factual disclosure requirements for businesses as commercial speech? Will they limit Zauderer to cases where the government is trying to prevent deception in the marketplace?
Implications for Legislators
For legislatures, the practical lesson is that it increasingly matters how AI transparency laws are framed. The starting point is still what problems the law tries to solve, and what information is needed to address them. But in deciding how to structure public disclosure requirements, legislators may want to keep some basic principles in mind:
- Focus on defining public disclosure requirements such that they are aimed at eliciting facts rather than judgments.
- Connect disclosures to AI products or services and how they are developed or used, where possible, especially where that information may help users and consumers evaluate a model.
- Avoid requiring businesses to adopt the government’s framing of contested issues or make judgments based on it.
In that respect, training data transparency laws may be on firmer footing when they call for disclosure of facts that help consumers evaluate a model, such as its training data sources. These laws are harder to defend when they call for interpretive judgments instead, such as assessments of possible bias. The same principles can apply to other forms of AI transparency.
Most importantly, legislators should keep in mind that this doctrine is in flux. AI transparency laws like AB 2013 are arriving at a moment when at least some courts appear to be rethinking how to assess mandated corporate disclosure. First Amendment doctrine is shaped by whatever disputes reach the Supreme Court. Those disputes often arise in settings far removed from ordinary commercial disclosure and often involve highly charged political topics. But the rules those cases produce don’t usually stay in their original contexts; instead, they apply much more broadly. NIFLA, for example, arose in the context of abortion, yet it now influences how courts review compelled disclosures in very different settings, such as product risk warnings.
So even when a case seems to be unrelated to AI transparency, its outcome and reasoning may still matter. Policymakers interested in transparency laws therefore need to watch a broad range of First Amendment decisions. That’s also a reason to draft carefully. There’s an old legal adage that bad cases make bad law. An ill-designed statute may do more than just lose in court. It may create precedent that makes better statutes harder to sustain.
Whistleblower Protections in SB 53: Strengths, Limitations, and Open Questions
Executive Summary
SB 53 greatly improves the AI safety legislative landscape in California, but achieves only partial success in its whistleblowing protections. Its major success is that it significantly increases the number of safety-relevant issues that can be reported to authorities. In particular, employees can now blow the whistle when they believe frontier developers have not reported critical safety incidents or where large frontier developers have made materially false or misleading statements about, or are not complying with, their AI framework. Employees “responsible for assessing, managing, or addressing risk of critical safety incidents” will also be explicitly protected when blowing the whistle about catastrophic risks. SB 53 also requires large frontier developers to provide an anonymous internal reporting channel that a select group of employees can use to report catastrophic risks and that all employees can most likely[ref 1] use to report a violation of the law (including of SB 53 itself).
Looking more deeply at the provisions of SB 53, however, we see a more mixed picture. An ideal AI whistleblower law would ensure that all reasonable risks could be reported; that all individuals capable of spotting such risks could report them without fear of retaliation; that any relevant individual could support government investigations without fear of retaliation; that authorities are equipped to investigate disclosures; and that real consequences exist for violations. SB 53 achieves each of these partially.
First, the responsibility for disclosing catastrophic risks is not placed onto companies, but instead on an unclearly-defined subset of employees. SB 53 does not require companies to report a specific and substantial danger to the public health or safety resulting from a catastrophic risk to the Office of Emergency Services (OES). Nor do evaluation providers receive any protection from retaliation, either for making reports or for participating in government investigations.
Second, while the definition of catastrophic risk is relatively broad and does not require a clear violation of the law, it is limited to specific types of catastrophic risk and harms only above a high threshold: fifty deaths or a billion dollars in damages.
Third, although SB 53 creates incentives for companies and insiders to report harm caused by AI, it is unclear what consequences these reports will entail. Company incident reports are sent to the OES, where they might not be accessible to the public, and whistleblower reports of catastrophic risk to California’s Attorney General (AG) or a federal authority. The OES lacks clear powers to compel companies to change their behavior, the AG has no clear mandate to intervene in the absence of a legal violation, and no federal authority has established jurisdiction over such risks.
Lastly, SB 53 does not adjust the penalties for violating whistleblower protections in California, which are meagre at $10,000 per violation — a far cry from the $1m violation cap for other SB 53 violations, and extremely little in the context of large frontier developers, defined as having gross revenue of over $500,000,000.
The next step is to ensure SB 53’s faithful implementation. State agencies must build the necessary expertise: the OES and the AG need access to specialised expertise to evaluate and respond to critical safety incidents and catastrophic AI risks. These agencies should collaborate closely to develop a comprehensive understanding of emerging threats and assess potential legal violations.
Background
SB 53’s whistleblower provisions are legally entwined with California’s pre-existing whistleblower framework, which they extend, and SB 53’s core transparency requirements, for which they serve as an enforcement mechanism.
CA Labor Code § 1102.5 is California’s general whistleblowing law, covering all industries and employees. It prohibits employers from retaliating against employees who disclose information, which they have reasonable cause to believe to be a violation of the law, to a government or law enforcement agency, or to a person with authority over the employee. These protections cover a wide range of potential violations. Nevertheless, California’s protections are narrower than some other states, such as New York, as they (i) only cover employees and do not extend to contractors and (ii) only protect disclosures of violations of the law, and do not protect employees from retaliation for disclosing a substantial and specific risk to public health and safety. This is particularly significant in the context of AI as frontier AI companies frequently rely on contractors for safety-relevant roles such as red-teaming and safety evaluations[ref 2] and novel AI risks often do not constitute clear legal violations, making the ability to report risks to public health and safety essential for early intervention.
SB 53, or The Transparency in Frontier Artificial Intelligence Act, signed into law in September 2025, creates a transparency framework for the most powerful AI systems. The Act applies to “frontier developers”, defined as persons who have trained or initiated the training of AI models using extraordinary computing power (greater than 1026 FLOPs) and imposes heightened obligations on “large frontier developers”, defined as frontier developers with over five hundred million dollars ($500,000,000) in annual revenue. As of 26th January 2026, the only publicly known frontier developers are xAI and OpenAI, while the only large frontier developer is OpenAI.[ref 3] Beyond whistleblower protections, outlined below, SB 53’s main contribution is introducing transparency requirements, spanning four key areas: i) Frontier AI Framework Requirements: Large frontier developers must publish annually-reviewed protocols for managing catastrophic risk; ii) Transparency Reporting Requirements: Developers must publish summary reports of features and risks, similar to existing model cards and system cards, before deploying new or substantially modified models; iii) Government Reporting Requirements: Frontier developers must report critical safety incidents to the OES within 15 days (or 24 hours if posing imminent risk of death or serious injury), with large developers also submitting quarterly risk assessment summaries; iv) Prohibition on materially false statements: Large frontier developers are prohibited from making materially false or misleading statements about catastrophic risks from their models or their management of such risks.
Crucially, these requirements are legally binding. As such, companies that fail to comply are in violation of California Law and hence whistleblower protections apply to any employee who reports them and are a critical oversight mechanism.
Policy Analysis
SB 53 extends pre-existing whistleblower protections to cover reporting catastrophic risks as well as legal violations, creates mandatory internal reporting channels, and enables employees to report critical compliance failures directly to authorities. Our analysis proceeds through four dimensions: personal scope, material scope, remedies, and channels. For each aspect of whistleblowing law, we discuss the advantages and limitations of SB 53’s provisions.
Personal Scope
Personal scope defines who can make a whistleblowing disclosure. Whistleblowing law is part of the California Labor Code and hence only governs employees whose employment contract is under California law. SB 53 creates a tiered protection structure: all employees receive protection when reporting violations of SB 53’s requirements; and a subset, ‘covered employees’, defined as those responsible for assessing, managing, or addressing risk of critical safety incidents, receive protection for reporting catastrophic risks that do not include a legal violation. The scope of ‘covered employee’ remains ambiguous, creating uncertainty about who falls into this protected group. Nevertheless, the breadth of the language suggests that most employees whose work relates to AI safety in some way are likely “covered”.
Advantages
- Well-Informed Reports: One advantage of the limitation in scope is that those responsible for critical safety incidents are likely to also be the employees who are best placed to assess risks and hence to blow the whistle if their concerns are not addressed, resulting in better-informed reports
Limitations
- Limited Employees Can Report Catastrophic Risks: SB 53 protects only “covered employees,” which it defines as employees who are “responsible for assessing, managing, or addressing risk of critical safety incidents.”
It is not clear which lab employees are “covered” and which are not under this definition. Interpreted broadly, it could include a substantial percentage of all lab employees, including e.g. any employee with any responsibility for information security. After all, the definition of critical safety incident includes “unauthorized access to… the model weights of a foundation model that results in… loss of property,” and any employee who works on information security is in some sense “responsible for… addressing risk of” unauthorized access to their employer’s sensitive data. Interpreted more narrowly, the definition could include only the members of a lab’s safety team, or only the members of the safety team specifically responsible for addressing catastrophic risks. This is an open question and may end up being decided in court.
A somewhat unusual consequence of this limitation in scope is that Chapter 25.1 is included under SB 53’s material scope (§ 1107.1.(a)(2)) only for employees responsible for critical safety incidents and not clearly all employees who are responsible for the various parts of Chapter 25.1. - Contractors Out of Scope: While other whistleblower protection regimes extend to protect non-employees who have a working relationship with the company (e.g. volunteers, board members, etc.), SB 53 and the California Labor Code do not. Frontier companies commonly hire contractors for longer-term engagements and important roles such as management, making this exclusion a particularly strong limitation in scope. That said, it could be possible to argue in court that contractors who work full-time for a single company on long-term contracts would count as employees.
Material Scope
Material Scope defines the subject matter that can form the content of a whistleblowing disclosure. SB 53’s key contribution is extending California’s whistleblower protections beyond legal violations by protecting covered employees from retaliation for reporting “a specific and substantial danger to the public health or safety resulting from a catastrophic risk” with “reasonable cause to believe” such danger exists (§ 1107.1(1)). To understand this scope, two definitions and their thresholds are essential: “critical safety incident” and “catastrophic risk.”
“Critical safety incident” means any of the following:
- Unauthorized access to, modification of, or exfiltration of the model weights of a foundation model that results in death, bodily injury, or damage to, or loss of, property.
- Harm resulting from the materialization of a catastrophic risk.
- Loss of control of a foundation model causing death or bodily injury.
- A foundation model that uses deceptive techniques against the frontier developer to subvert the controls or monitoring of its frontier developer outside of the context of an evaluation designed to elicit this behavior and in a manner that demonstrates materially increased catastrophic risk.
“Catastrophic risk” means a foreseeable and material risk that a frontier developer’s development, storage, use, or deployment of a foundation model will materially contribute to the death of, or serious injury to, more than 50 people or more than one billion dollars ($1,000,000,000) in damage to, or loss of, property arising from a single incident involving a foundation model doing any of the following:
- Providing expert-level assistance in the creation or release of a chemical, biological, radiological, or nuclear weapon.
- Engaging in conduct with no meaningful human oversight, intervention, or supervision that is either a cyberattack or, if committed by a human, would constitute the crime of murder, assault, extortion, or theft, including theft by false pretense.
- Evading the control of its frontier developer or user.
The Act expressly excludes risks from the category of catastrophic risks: information generated by a model if it is already publicly available, lawful federal government activity, and situations where the model causes harm in combination with other software, but AI did not materially contribute to the harm. Thus ‘catastrophic risk’ in the Act only refers to CBRN and loss of control risks.
Advantages
- All Employees Can Disclose Violations of SB 53. The whistleblower section of SB 53 asserts any employee’s right to receive protection against retaliation if they blow the whistle with reasonable cause to believe that any violation of law, including SB 53 itself, occurred (§§ 1102.5(b), 1107.1(j)(1). This includes, for example,
- Misleading Statements: As companies are prohibited from publishing materially false or misleading statements, developers who do so are in violation of state law. Given the California Labor Code, employees would thus be protected if they report such misleading statements to a government or law enforcement agency—for the first time ever allowing employees to correct statements made by companies on their frontier safety frameworks and incident reporting, albeit only to government authorities such as the Attorney General.
- Failure to Report Critical Safety Incidents: Frontier developers must disclose critical safety incidents within 15 days of an incident’s discovery (§ 22757.13.(c)). Employees would be protected in reporting such an incident if they had reasonable cause to believe that such a report wasn’t made.
- Failure to Report Critical Safety Incidents that pose an imminent risk of death or serious injury: Frontier developers must report an incident of this nature within 24 hours of discovery. Employees are protected for reporting failure to do so (and hence also for reporting the incident itself), if they have reasonable cause to believe that the company did not make a report. Employees will likely only be able to claim protections for reporting violations under this rule if harm from the incident has already materialized, with more harm being imminent. The only exception here is deception, where only an increase in catastrophic risk already requires an incident report, and hence reporting on imminent harm risk resulting from it.
- Failure to Comply with their Frontier AI Framework: This includes the failure to publish a justification in the event of a material modification to the frontier AI framework.
- Failure to publish their transparency report: This includes the failure to include all required information.
- Catastrophic Risks can be Reported: As noted above, this is an important extension of California whistleblower law, which previously only covered reports of unlawful activity.
- Reasonable Cause to Believe: Whistleblowers receive protection for reports that they reasonably believe to be true, ensuring that employees are encouraged to come forward with any concerns they may have, before those concerns materialise into concrete harms. This “reasonable cause to believe” standard is generally considered to be whistleblower-friendly and is consistent with pre-existing California whistleblower law.
Limitations
- Lack of Certainty: SB 53 protects disclosures only if they relate either to a violation of law or a “specific and substantial danger to the public health or safety resulting from a catastrophic risk.” “Catastrophic risk” is a high and difficult-to-calculate bar to meet, and we worry that this requirement will have a chilling effect on potential whistleblowers. Particularly in the AI industry, where risks are novel and probabilistic, it is extremely difficult to determine the exact harm that can result from a risk. The phrasing of the Act also suggests that a risk of 49 deaths is insufficiently bad to raise a concern with the government if a large frontier developer is not addressing it. Additionally, the more-than-fifty-death threshold is unprecedented: it is not in line with risk-of-harm whistleblower legislation that exists in other states, such as New York, where existing whistleblower protections cover significant harm to even a single person.
- Single Incident: The ‘single incident’ requirement in the definition of catastrophic risk (22757.11.(c)(1)) could significantly limit the scope of risks which can be reported. For example, if risk stemming from a single model were to threaten the deaths of hundreds, but through separate channels and at different times, those deaths might not arise from a “single incident,” and therefore the risk of such deaths might not be reportable as a catastrophic risk.
- Limited Scope of Covered “Catastrophic Risks”: Not all significant risks of harm are covered under the ability to report catastrophic risks. In SB 53, “catastrophic risk” from an AI model does not cover any risk of a certain amount of harm caused by a model, but only particular types of risk: namely, risk that a frontier model will assist in creating a CBRN weapon, engage in a cyberattack or other crime autonomously, or escape human control. Any risk of a model committing a cyberattack, murder, or theft where it is directed to do so by a human is not covered under the Act, no matter how large the scale. In the same way, the definition of a critical safety incident is also limited to particular risk pathways (22757.11(d)).
- Limited Ability to Warn of Risks: While the regulation requires the logging of all “critical safety incidents,” it does not create protections for whistleblowers who warn of the risk of critical safety incidents, even using internal channels. If the legislation had included protected internal channels for reporting risks of critical safety incidents, it could have incentivised the avoidance of incidents rather than only their reporting.
Remedies
SB 53 applies existing Labor Code remedies to its whistleblower provisions. This section examines the protections available to whistleblowers and deterrents against retaliation.
Advantages
- Extends Labor Code Remedies to Reports of Public Safety Dangers from Catastrophic Risks: Covered employees who face retaliation for reporting a catastrophic risk are entitled to the remedies in the California Labor Code, including burden-of-proof reversal (§ 1107.1(g)) and injunctive relief (§ 1107.1(h)). The same is true for retaliation suits brought by any employee who reports a violation of SB 53 (§§ 1102.6 and 1102.61-2).
- Incentivizes Internal Knowledge Sharing: Imminent critical safety incidents can be reported to the AG if the developer has not logged them with the OES within 24 hours. This thus incentivises sharing with employees the knowledge that risks have been reported to the OES in order to avoid over-zealous reporting to the California AG.
Limitations
- Limited Scope: Whistleblower protections in SB 53, unlike in the EU’s whistleblower directive,[ref 4] do not extend to family members. Especially in the context of the Bay Area, where spouses often both work in the AI industry, this can strongly discourage potential whistleblowers and does not have clear benefits.
- Inconsistent Belief Standards: SB 53 uses different standards for determining access to an internal reporting channel and protection from retaliation. Section 1107.1(e)(1) requires only that an internal report is made in “good faith,” but the anti-retaliation protection in 1107.1(a) requires that an employee have “reasonable cause to believe” that their disclosure involves a violation of law or a substantial danger. Thus, a covered employee who anonymously submits an internal report in good faith but cannot meet the “reasonable cause” standard has legitimately used the internal reporting channel but has no statutory retaliation claim under § 1107.1(a) if their anonymity is breached and they are retaliated against.
- Limited Government Penalties: The California Labor Commissioner can impose fines of up to $10,000 for violations of the Labor Code, including whistleblower protections (§ 1102.5(f)(1)). Large frontier developers are defined as having a gross revenue of over five hundred million dollars, and hence will not be deterred by a $10,000 fine from retaliating against whistleblower or not complying with any other whistleblower regulation. However, deterrence could come from other means, like a retaliation suit by the whistleblower, who would be entitled to damages and injunctive relief.
Channels
Whistleblowing channels are the individuals, agencies, or offices to which a whistleblower disclosure can be made. SB 53 protects covered employees if they whistleblow “to the AG, a federal authority, a person with authority over the covered employee, or another covered employee who has authority to investigate, discover, or correct the reported issue” (§ 1107.1(a)). Additionally, SB 53 requires the creation of “a reasonable internal process through which a covered employee may anonymously disclose information to the large frontier developer if the covered employee believes in good faith that the information indicates that the large frontier developer’s activities present a specific and substantial danger to the public health or safety resulting from a catastrophic risk” or a violation of SB 53’s transparency obligations. Under the California Labor Code, any employee is protected if they whistleblow to “a government or law enforcement agency, to a person with authority over the employee or another employee who has the authority to investigate, discover, or correct the violation or noncompliance,” or “any public body conducting an investigation, hearing, or inquiry” (§ 1102.5(b)). California’s Labor Code also provides a whistleblower hotline, operated by the AG, intended to receive calls from whistleblowers. Following a report made to the hotline, the AG is directed to refer calls received on the hotline to the appropriate government authority for review and possible investigation.
Advantages
- Choice of Internal or External Channels: In line with California Labor Code, whistleblowers have the freedom to choose internal or external channels to raise concerns.
- Expert Internal Channel: SB 53 allows whistleblowers to approach persons responsible for critical safety incidents, which, if this is equivalent to a subset of the safety team, encourages employees to go to those who are best-placed to assess risks.
- Anonymous Internal Channel: SB 53 requires large frontier developers to create internal channels for raising concerns and clearly prohibits retaliation against those who are qualified to use this channel. These channels have a number of important features:
- Most importantly, SB 53 necessitates mandatory updates to the frontier developer’s board every quarter of the disclosures and associated responses. This ensures that the board has an accurate overview of employees’ concerns and creates an oversight mechanism for the reporting channel itself.
- The channel must allow for anonymous reporting, which reduces the risk of retaliation and incentivizes internal reporting, allowing developers to solve the issue swiftly without the involvement of external parties.
- Finally, SB 53 requires that monthly updates be provided to the employee who made the disclosure. This reassures the employee that their concerns are being taken seriously and, if the developer’s response is insufficient, the employee can follow up on the issue, externally if necessary.
Moreover, while SB 53 itself only asserts that covered employees are protected when using this channel, pre-existing California whistleblower law extends the scope of this protection to any employee who reports unlawful activities. Specifically, SB 53 only specifies that “covered employees” (i.e. those who are “responsible for assessing, managing, or addressing risk of critical safety incidents”) receive protections for using this “reasonable internal process” (§ 1107.1(e)(1)), § 1102.5 of California’s Labor Code specifies that any employee has protection when they report a risk to “another employee who has authority to investigate, discover, or correct the violation or noncompliance.” As an internal process for reporting violations necessarily goes to those who have the authority to discover violations, all employees should receive legal protection when using this channel to report violations of the law. However, whether all employees will have access to this internal channel remains to be seen.[ref 5]
Finally, we note that companies are likely to be able to de-anonymize internal whistleblower reports, especially if they prevent non-covered employees from accessing the reporting channel, thereby limiting the pool of potential whistleblowers. Companies are not explicitly barred from seeking out the identity of internal whistleblowers.
- Federal Authority SB 53 protects disclosures to any “federal authority,” an undefined term that reads broadly on its face but may not include, for example, members of Congress.
Limitations
- No Public Disclosure: Even in extreme circumstances where there may be imminent danger, public disclosure is not explicitly protected.
- Separation of Channels: It is perhaps unusual that incident reports must be sent to the OES, but reports on catastrophic risks by whistleblowers cannot be sent to it. However, since OES is a government agency, any violation of law can be reported to it under pre-existing California whistleblower law.
- Lack of Anonymity and Confidentiality in External Channels. While a reporting channel to the AG is mandated and reporting is also possible to the OES where there is a violation of law, these channels do not have any requirements for anonymity or confidentiality. Not only do anonymity and confidentiality guarantees protect whistleblowers, they help whistleblowers feel safe enough to come forward.
- Lack of Clarity on Reasonable Internal Channel: SB 53 would be improved by a better-specified explanation of what is required for an internal channel to be “reasonable.” If an internal channel was owned by a company’s in-house legal team (whose duty would run to the company rather than to the whistleblower) or by an executive who could be the subject of a disclosure, whistleblowers could be deterred from disclosing.. Furthermore, SB 53 does not specify a confidentiality policy and process, although such specifications would improve whistleblowers’ confidence that their identity would not be revealed and incentivize legitimate disclosures.
Liberalism Forever
Abstract
We argue that liberalism—market economies governed democratically—is the best approach for navigating the far future. A growing longtermist literature paints humanity’s path to good outcomes as narrow, with small errors risking value lock-in, gradual disempowerment, or other forms of permanent catastrophe. We argue that this literature underestimates the institutional dynamics that have historically steered liberal societies past similar predictions of crisis. We defend long-term liberalism by examining the standard arguments for markets and democracy and asking whether they survive the structural changes anticipated for the far future: transformative AI, space colonization, and radical economic transformation. We claim that these arguments are largely robust, and in several cases strengthened. In the far future, markets continue to aggregate information, allocate goods efficiently, and foster innovation; democratic institutions can continue to supply public goods, commit to peaceful redistribution, correct errors, and accommodate reasonable pluralism. We close by arguing that two recent longtermist governance proposals—viatopia and the long reflection—are faint-heartedly liberal: they sound liberal in the abstract but risk illiberalism when implemented. The right approach for managing the far future is not radical new institutions but adaptations of the existing liberal toolkit to meet future challenges.