Germany Establishes an AI Security Institute
I. Introduction: Virtual beginnings and open questions[ref 1]
This June, Germany formally decided to establish a German AI Security Institute (DE-AISI).[ref 2] The decision was adopted by the National Security Council on 8 June 2026 and implements Germany’s commitment under the 2024 Seoul Declaration’s Statement of Intent[ref 3] to support the development of AI Safety and Security Institutes and to nurture networks between them. It reflects a shift in German AI governance towards strengthening governmental capacity to understand the capabilities, limitations, and security implications of frontier AI models.[ref 4] According to the Government, the DE-AISI will provide scientific and technical expertise and support strategic risk assessment.[ref 5]
The initiative draws inspiration from the UK AI Security Institute, which is regarded as a paragon for government-backed scientific evaluation of frontier AI models.[ref 6] Subsequent joint statements with the United Kingdom and France indicate that the DE-AISI is intended to become part of the growing international network of AI Safety and Security Institutes.[ref 7] Government statements also emphasise that the DE-AISI is intended to complement rather than duplicate the governance framework established under the EU AI Act.[ref 8] In particular, it is expected to support scientific cooperation and frontier AI evaluation while remaining institutionally distinct from both the AI Office and the national market surveillance authorities responsible for enforcing the EU AI Act as set out under Germany’s AI Market Surveillance and Innovation Promotion Act (KI-MIG).[ref 9]
Although a comprehensive legislative framework has yet to be presented, the German Federal Government has indicated that the DE-AISI will initially operate as a virtual institution anchored in a ‘nucleus’ drawing on existing capacities at the Federal Office for Information Security (BSI) and the Federal Network Agency (Bundesnetzagentur, BNetzA). In this initial build-up phase, the focus will be on security and safety aspects of advanced AI models.[ref 10] As of early August 2026, a spokesperson of the Federal Ministry for Digital Affairs and State Modernisation (BMDS) confirmed that this nucleus is already operational, describing a step-by-step, modular build-up approach and noting that technical exchange with international partner institutions, including in France, the United Kingdom, and with the AI Office, is already under way.[ref 11] Recent comments by Federal Government officials hint at a broader scope of the mandate but do not specify the envisioned breadth.[ref 12]
For the long-term structure of the DE-AISI, current policy proposals advocate for an institution that has a narrowly defined technical and scientific mandate, organisational independence, and close integration into Germany’s national security architecture, rather than being a new regulatory authority.[ref 13] Several open questions remain, including the exact scope of the DE-AISI’s mandate, its financial flexibilities, and its location. This report outlines and comments on each, highlighting the possibility of establishing the DE-AISI as a federally owned limited liability company (GmbH).
II. An AISI with an unsettled scope
The most consequential open question concerns the material scope of the DE-AISI’s mandate. The National Security Council’s decision seems to refer mainly to assessing the consequences of advanced AI models for cybersecurity in Germany.[ref 14] The nucleus, in turn, is to cover both ‘security’ (BSI) and ‘safety’ (BNetzA) aspects of such models.[ref 15] Parliamentary State Secretary Jarzombek has since indicated a build-up towards an institution positioned more broadly, both thematically and in terms of capacity.[ref 16] A broader design had also been proposed by a group of researchers in October 2025.[ref 17] In comparison, the Seoul Statement of Intent frames the role of such institutes more narrowly as facilitating ‘AI safety research, testing, and/or developing guidance to advance AI safety for commercially and publicly available AI systems’.[ref 18]
In a different sense, the German policy debate seems to be converging on a narrow, purely scientific and technical mandate aimed at enabling the Federal Government to reach informed decisions on frontier AI model risks. Interest groups argue that the DE-AISI should be clearly distinct from regulatory authorities such as the BNetzA and the BSI to enable trust-based cooperation with frontier AI model developers.[ref 19] These proposals envision that the DE-AISI would produce recurring cross-departmental situational assessments for the Federal Government, conduct systematic technical evaluations of frontier AI models, carry out research in service of those tasks, contribute to the development of technical standards, and cooperate with partner institutions nationally and internationally.[ref 20] Its focus would extend beyond cybersecurity to chemical, biological, radiological, and nuclear (CBRN) risks and to loss-of-control (LoC) risks,[ref 21] comparable to the focus of the UK AISI.[ref 22] Another question concerns which models precisely the DE-AISI should observe and evaluate. Arguably, a national security-focused institute should not confine itself to general-purpose AI (GPAI) models with systemic risk within the meaning of the EU AI Act,[ref 23] but should also capture more specialised models capable of generating comparable risks: genome language models used to design novel pathogens could be one example.[ref 24] The definition should remain technology-neutral, so that the institute’s remit is not restricted only to proprietary models or tied to today’s model architectures. At the same time, defining the frontier of AI capabilities is not straightforward and a functionally wide interpretation of frontier AI could overextend the focus of the DE-AISI.[ref 25]
Persistent uncertainty about the material scope of the DE-AISI’s mandate will come at a cost. Developers deciding whether to grant pre-deployment access to unreleased models need to know whether they are dealing with a scientific partner or with an institution whose remit may later expand into adjacent, potentially regulatory territory. Industry associations have therefore called for a research mandate clearly delineated from that of existing bodies, arguing that questions of labour law, consumer protection, data protection, and AI ethics are already competently addressed elsewhere.[ref 26]
After all, the preliminary nucleus arrangement may come to sit uneasily with the institute’s intended function. The BNetzA recently became Germany’s central node of the AI Act’s national-level enforcement architecture,[ref 27] and the BSI holds its own enforcement powers in the field of information security.[ref 28] An institute housed within or governed by authorities exercising such powers may struggle to obtain the confidential access on which its work depends. For instance, GPAI model providers may worry about the fact that the BNetzA could hypothetically use information voluntarily shared with the DE-AISI to request the European Commission to initiate enforcement actions where the BNetzA believes that this information indicates a breach of the provider’s obligations under Chapter V of the AI Act.[ref 29] Closing precisely this kind of access gap is among the institute’s purposes and of outsized importance to national security. Whatever its institutional setup ends up being, it may therefore be advisable to ensure a sufficiently clear delineation between the DE-AISI and regulatory powers. This may mean limiting the cooperation between the DE-AISI and the BSI and the BNetzA to general, non-developer-specific findings on frontier AI risks, rather than confidential information capable of being used for regulatory purposes. In this respect, it is promising that the BSI has reportedly indicated that the DE-AISI will not focus on regulatory aspects.[ref 30]
III. Pay, flexibility, and the case for a GmbH
Germany’s Federal Digital Minister has stated that the DE-AISI should be staffed with ‘top expertise from world class experts'[ref 31]. Delivering on that ambition will be difficult within ordinary German public-sector pay structures, which are considerably lower than frontier AI expert salaries in the private sector or even the international non-profit sector. This applies all the more given the fact that private-sector compensation for AI talent has risen sharply; an industry salary survey records a 45 per cent increase for AI safety and alignment specialists since 2023.[ref 32] The UK AISI has recognised these dynamics. It operates not only with GBP 66 million in annual funding and priority access to compute, but also with a more competitive salary structure than the rest of the civil service.[ref 33] Proposals in the German debate have mostly suggested a funding similar to that of the UK institute ranging from EUR 60 million[ref 34] to at least EUR 75 million annually.[ref 35]
However, funding levels alone do not determine whether these resources can be deployed effectively. Choosing a legal form that allows the DE-AISI to effectively deploy funding will be equally crucial. Establishing the institute as a federally owned limited liability company (GmbH) would be a promising option in this regard. It would constitute a formal privatisation (formelle Privatisierung), leaving the operational task itself in state hands while changing only the organisational vehicle through which it is performed.[ref 36] A reference point for formal privatisation of this kind is the Federal Agency for Disruptive Innovation (Bundesagentur für Sprunginnovationen, SPRIND).[ref 37] SPRIND is a wholly federally owned GmbH created to fund high-risk innovation projects and governed by its own enabling statute, the SPRIND Act (SPRIND-Freiheitsgesetz[ref 38]) of 2023.[ref 39] It was designed for a field in which recruitment and funding decisions must be taken quickly and in competition with the private sector. In several ways, the DE-AISI faces a comparable situation, necessitating a legal form that allows it to meet similar key challenges.
One particular challenge is the DE-AISI’s staffing. As a company under private law, the institute would not be bound by public-sector collective agreements in the same way as a federal authority. It could seek exemptions from the prohibition on preferential treatment (Besserstellungsverbot), which otherwise prevents federally funded bodies from paying their staff more than comparable federal employees. § 5 of the SPRIND Act transfers that decision to the company itself where compelling reasons so require, based on the legislature’s reasoning that contract negotiations in highly competitive fields must be conducted quickly and concluded with binding effect. Bitkom, one of Germany’s largest digital industry associations, has argued for a comparable arrangement for the DE-AISI.[ref 40]
As for budgetary flexibility, the frontier AI risk landscape develops on timelines that do not align with annual budget tranches. Evaluation and research priorities in this field can shift within months, if not weeks. § 15(2) of the Federal Budget Code (Bundeshaushaltsordnung) permits appropriations to be designated for self-administration (Selbstbewirtschaftung), allowing funds to be carried across financial years and redeployed as project needs change. § 3(2) of the SPRIND Act makes use of this instrument, albeit at a limited rate. The same instrument has been extended to non-university research institutions to strengthen their performance and international competitiveness,[ref 41] which suggests its relevance for a body operating in a field that is at once fast-moving and research-based, such as the DE-AISI.
Regarding governance, the articles of association (Gesellschaftsvertrag) would allow the Federal Government’s specific requirements to be reflected in tailored form,[ref 42] while limited liability caps the exposure of the federal budget.[ref 43] Democratic accountability can be maintained through the instruments of company law. The Federal Government, as sole shareholder, would appoint the executive director, could issue instructions, and would hold comprehensive information rights.[ref 44] Furthermore, the GmbH structure allows for a high degree of organisational and personnel flexibility.[ref 45]
None of this follows necessarily from the choice of the GmbH as the legal form as such. As in the case of SPRIND, it depends on the enabling legislation providing for it;[ref 46] in particular, on exemptions from the prohibition on preferential treatment and on the self-administration rate adopted, as well as on the articles of association. Whether comparable arrangements could be achieved within a public law structure remains an open question. For its part, the Federal Government has stated that it is not yet in a position to provide details on the institute’s long-term legal structure.[ref 47]
IV. A location fit for purpose
A decision on the location of the DE-AISI was initially deferred by the Federal Government. A few options are now under active consideration, including Berlin, the Saarland, Bonn, and Munich.[ref 48]
The Saarland has already advocated for hosting the DE-AISI. In August 2026, the CDU group in the Saarland state parliament — in opposition at state level, but the party of the Chancellor and, with its sister party the CSU, of the ministries leading on the DE-AISI — formally called for the institute to be based in Saarbrücken. Stephan Toscani, chair of the group, described the Saarland as ‘the ideal location’ and warned that the opportunity to base the DE-AISI in it must not be allowed to pass.[ref 49] Saarbrücken hosts Saarland University, the German Research Center for Artificial Intelligence (DFKI), the CISPA Helmholtz Center for Information Security, and two Max Planck Institutes.[ref 50] It also hosts branch offices of both the BSI and the BNetzA, the two authorities on which the nucleus is drawing.[ref 51] Bonn, on the other hand, would place the DE-AISI at the same location as the BSI headquarters.[ref 52]
Berlin, however, offers proximity to the federal ministries and the National Security Council. Alongside Munich, it hosts one of Germany’s largest AI industry clusters.[ref 53] It also hosts a branch of the DFKI. As a metropolis and capital city, it may prove easier to recruit for than other alternatives. Finally, if the DE-AISI’s defining task will be advising the Federal Government on frontier AI risks, its work will consist largely of recurring cross-departmental situational assessments, ad hoc analysis when risks shift at short notice, and exchanges that might involve classified material. In this respect, proximity to the ministries carries particular weight, as the UK-AISI’s location in London demonstrates. The relocation of the Federal Intelligence Service (BND) from Pullach to Berlin was completed on similar grounds, given the need for swift communication and intensive coordination between the BND and federal government bodies.[ref 54]
Regional policy considerations may pull in the opposite direction. Germany has a long-standing practice of distributing federal institutions across the country, and the Saarland bid is expressly framed as part of that state’s structural transition towards a technology location.[ref 55] These are legitimate objectives in their own right, but they are distinct from the question of where the institute can most effectively perform its advisory function. Insofar as the above considerations of proximity to the Federal Government and talent recruitment are crucial factors, they point towards Berlin.
V. Conclusion: What Germany stands to gain, or forgo
The establishment of the DE-AISI is an important opportunity for Germany to contribute to the safety and security of frontier AI models as their risks for national security and critical infrastructure become increasingly central. Existing institutes provide case studies for success factors and possible failure modes. For example, the UK AI Security Institute has benefited from financial flexibility in hiring, compute access, an attractive location, and a scientific mandate clearly distinct from regulatory functions. Drawing on these experiences, the DE-AISI can make vital contributions to national security by assessing frontier AI risks specifically in line with the mandate of the German National Security Council.
Such a role can only be fulfilled at the national level. National security remains a competence of the Member States, the AI Office’s remit is directed at the Union market as a whole, and the AI Act’s systemic risk threshold is defined by reference to effects at Union level rather than to the exposure of any single Member State.[ref 56]
Without these capacities, Germany’s ability to anticipate and respond to frontier AI risks remains dependent on what partner institutions abroad are willing to share. Timely progress will therefore be essential. For two years after endorsing the Seoul Statement of Intent in 2024, Germany has been one of only a few signatories without an AI Security institute of its own.[ref 57]
This position is recoverable, though only if the institute is equipped with the mandate, structure, and location which the task requires. A GmbH structure could afford the needed financial and personnel flexibilities, subject to enabling legislation that clearly defines the institute’s mandate and relation to other governmental bodies. Berlin would offer an attractive location and ensure proximity to the Federal Government for the DE-AISI’s advisory work.
If these steps are taken, the DE-AISI can move to the forefront of international AI safety and security institutions, and anchor Germany’s preparedness for frontier AI risks.
Who Writes the AI Constitution?
Over the past year, the federal government has grown intensely interested in the values embedded within and expressed by artificial intelligence (AI) models. Across executive orders, Office of Management and Budget (OMB) procurement rules, Federal Trade Commission actions, and Department of Justice interventions, the federal government has engaged in an expansive effort to dictate the values on the models millions of people use every day. Many state governments have likewise considered or passed legislation concerned with AI values and biases.
One avenue for potential government intervention into the lab’s values-selection process has emerged: the labs’ respective documents spelling out the values baked into their models. We refer generally to these as “AI constitutions.” AI constitutions are long documents outlining the values, ethics, and character that models should have. Written predominantly by the employees at the labs deploying these models, albeit with some consultation with external stakeholders, these documents increasingly serve as the central locus of AI values discourse. In fact, Lawfare recently shared a research agenda inviting inquiry into the values and processes behind the various AI constitutions being crafted and implemented.
Anthropic publishes Claude’s Constitution, “the foundational document that both expresses and shapes who Claude is.” OpenAI maintains the Model Spec, which “outlines the intended behavior for the models that power OpenAI’s products.” As detailed further below, these are neither mere mission statements nor pure governance frameworks. They play a functional role at various stages in the model development process, such as what data the model trains on, what examples it’s fine-tuned on, and what behavior it’s rewarded for. Anthropic reports that its Constitution “directly shapes Claude’s behavior.” An AI constitution is therefore both a published statement of a company’s values and an operative control mechanism.
The demonstrated capacity of constitutions to shape model behavior is precisely what will likely make them a target of regulators seeking to confine models to certain values and perspectives. The possibility of governments and governing bodies—state, federal, and international—attempting to amend or revise AI constitutions warrants advanced scrutiny from, among other things, First Amendment scholars. At this stage in the AI governance debates, free speech and free expression scholars have focused on more doctrinal questions, such as whether AI outputs are protected speech. It’s urgent that they expand their inquiry as to whether AI Constitutions are protected speech, before constitution-based regulation takes off in earnest.
This article argues that AI constitutions—as policymakers increasingly turn to regulating model behavior and characteristics—contain expression protected by the First Amendment, a claim that requires first defining what AI constitutions are and situating them within existing First Amendment jurisprudence. We also consider counterarguments and alternative possibilities for shaping model characteristics that would avoid First Amendment issues altogether.
What Counts as an AI Constitution
The use of the word “constitution” to refer to technical documents risks inviting direct comparisons to documents traditionally bestowed with that title. For that reason, it’s important to precisely assess what’s similar and not about, say, Claude’s Constitution and the U.S. Constitution.
In the AI context, the term was originally coined in a 2022 paper. The terminology was intended to capture the structural and sociological similarities between AI constitutions and legal constitutions. That is, AI constitutions look like legal constitutions in some ways. For example, they list things models should and should not do, or should and should not care about. They also function in society like legal constitutions in that they explicitly outline principles to resolve difficult ethical and practical questions. These foundational technical documents are likewise intended for public analysis. According to the company, OpenAI publishes its Model Spec because “it’s important for people to be able to understand and discuss the practical choices involved in shaping model behavior.” Perhaps most importantly for the analysis here—whether intended by AI researchers or not—the term carries a certain solemnity that invites significant scrutiny of its contents. Anthropic researchers at one point referred to the constitution as a “soul document.”
Yet AI constitutions, unlike legal constitutions, are not the product of some democratic process that signals broader consent of the governed to the terms of the constitution. Nor, as Nathan Darmon and Tom Reed pointed out, are conflicts over how to interpret the constitution subject to external, independent adjudication. Still, the similarities with legal constitutions are strong enough to elicit interest in them as a vehicle for regulation.
Joe Carlsmith, one of the principal authors of Claude’s constitution, minimally defines an AI constitution as “a description of the intended values and behavior for an AI system.” A thicker definition accounts for how constitutions are applied in training and monitoring. Constitutions specify what a model should and should not do, take final authority over other instructions, are trained directly into the model, and, with some exceptions, are meant to shape a model’s character rather than police outputs case by case. Both Anthropic’s Constitution and OpenAI’s Model Spec qualify as AI constitutions as defined here. Not all labs have adopted a version of an AI constitution. Only Anthropic and OpenAI publish stand-alone documents that function as training-time specifications of value. Our definition excludes Meta’s acceptable-use policy, Google’s app guidelines, and xAI’s raw system prompts, though in the case of Google it is possible to craft one from various company documents that more or less make up the core components of a constitution.
There is not yet a default way to write an AI constitution. Claude’s Constitution and the Model Spec, for example, read quite differently. Claude’s Constitution is philosophical and concerned with the model’s psychology. Rather than order the model to “always be honest,” it exhaustively explains what honesty is and why it matters; it’s trying to encode Anthropic’s higher order values by reference to the specific ways one can imagine a model behaving well (or badly). Anthropic has described the document as an effort to “materialize a new archetype for how an AI assistant can be.” In turn, OpenAI’s Model Spec reads like case law, pairing each principle with sample prompts and examples of allowed and disallowed answers. Both are important to model training. They may inform the generation of synthetic data, sit inside the model’s chain-of-thought reasoning, or serve as the score card for alignment afterward. Anthropic calls the constitution “the final authority on how we want Claude to be and to behave” and reports that training on it improved alignment in ways that “persisted through RL post-training.” In other words, these documents have been empirically shown to shape how AI tools perform in the wild.
Why AI Constitution Regulation Is Coming
Recent history strongly suggests that the public should expect some kind of regulation of AI constitutions in the not too distant future. The number of congresspeople discussing the dangers of AI has skyrocketed, and policymakers’ concerns often center on what these models value. The “Preventing Woke AI” executive order led to an OMB implementation memo, a copycat bill in Congress, and an Iowa bill that passed the state house. Illinois, Colorado, and New York City, meanwhile, have laws regulating AI discrimination in employment decisions. Concerns over alleged AI bias have been sustained by reports of continued ideological skew in model outputs. The Washington Post, for example, determined that leading models tend to produce neutral or left-leaning answers to political questions. Such findings will become more relevant as the 2026 midterms and 2028 presidential election near, especially given that other researchers have found that users tend to be influenced by the political answers offered by models.
As politicians are increasingly interested in regulating AI, and in particular the values that models express, constitutions will prove an attractive object for regulation for at least three reasons. They are high leverage: Because they sit upstream of data generation, training, and evaluation, any edits to a constitution propagate through everything a future model learns and does. They are legible: Written in English and often reading like a statute, they can be marked up by a staffer and altered without consulting a technical expert. And they carry signaling value: An amendment to a constitution addressing wokeness, patriotism, or child safety, for instance, would be something a politician could easily point to when explaining their AI governance efforts to constituents. Regulations about reward functions or mechanistic interpretability—more technically complex features of AI development—seem less likely to spark an appearance on the Sunday shows.
What Speech Gets Protection
Efforts to regulate AI constitutions will likely face constitutional headwinds arising from the potential curtailment of expressive activity. Gauging the stiffness of those winds requires a review of First Amendment protections of speech.
In the most general sense, the First Amendment restricts the ability of the government to regulate speech. But within this broad prohibition there are myriad exceptions. Not all words qualify as speech protected by the First Amendment, and not all speech is protected speech. Protection instead runs along a spectrum. At one end is fully expressive speech, where a content-based restriction is “presumptively unconstitutional.” Somewhere in the middle is commercial speech, which is “related solely to the economic interests of the speaker,” and which the government can regulate more freely. At the far end is speech “plainly incidental” to conduct, which is generally regulable. Where a given document lands turns on how much expression it carries. The Court has “long recognized that not all speech is of equal First Amendment importance,” with the First Amendment concerned primarily with protecting matters of public importance as determined by “[the expression’s] content, form, and context.”
In the context of corporate speech, government regulation of a company’s mission statement would likely run afoul of core protections around free expression and freedom of association by compelling adoption of specific views. Still, government regulation of an instruction manual is far less likely to spark First Amendment concerns because of the absence of expressive content. How best to use a chainsaw, for example, says little about the corporate actor’s views other than that they seek to preserve life. An AI constitution arguably occupies some space in between given its clear expressive purpose as well as its more practical effort to ensure that the tool works as intended.
This ambiguity is best illustrated when one compares the AI constitutions produced by Anthropic and OpenAI. OpenAI’s document, the Model Spec, explicitly and categorically prohibits any models from “whistleblowing,” whereas the Constitution permits Claude to take “independent action” in “cases where the evidence is overwhelming and the stakes are extremely high.” The Model Spec allows models to tell some white lies, whereas Claude’s Constitution almost categorically forbids them. The Model Spec instructs models to always listen to its commands, whereas Claude’s Constitution contemplates situations where Claude determines that some portion of its Constitution is itself unethical. Each of these is a clear expression of the values of each corporation.
Other provisions look less like expressions of distinctive corporate values than like functional necessities any commercial operator would adopt. Each document tells the models to be careful when providing legal or medical advice, for example, and each instructs the models to prefer information from reliable sources. Traits common to both documents might be driven less by any company’s particular vision of a good model than by liability concerns or the demands of shipping a useful product. To the extent a constitution is built from provisions like these, it drifts toward the more regulable end of the spectrum.
AI Constitutions Contain Protected Speech
Given all of the above, some portions of AI constitutions may be justifiably regulated. Other sections, particularly those that tend toward the expressive end of the spectrum, may be safeguarded by the First Amendment. But few scholars or lawyers have considered the question of regulating corporate AI constitutions directly. Instead, the vast majority of First Amendment scholarship on AI thus far has focused on how to think about AI outputs in a First Amendment context. In brief, Eugene Volokh, Mark Lemley, and Peter Henderson are inclined to regard outputs as protected speech, while Peter Salib argues they are not because when a model emits text “no one thereby communicates.” Still others tie protection to whether a speaker “knows what he said when he said it.”
Scholarship has presumably focused on these more doctrinal questions up until now because it was believed that steering outputs with upstream regulation was impossible (at least on the time and complexity scales on which Congress typically operates). Safety rules must operate on outputs, Salib has argued, because “there is … no way, currently, to write legal rules mandating safe code.” The labs now claim there is such a way, and a regulator who takes them at their word will naturally want a say in how it’s written. A clue may be found in Justice Amy Coney Barrett’s concurrence in the recent First Amendment case of Moody v. NetChoice (2024). The justice warned that as companies “hand the reins” of their decision making to machine-learning systems, “technology may attenuate the connection between” corporate conduct and the human choices the First Amendment protects. The corporate values of Meta may not clearly be at work in the algorithm that shapes a user’s Facebook feed, for instance. When only a vague goal is given to a machine learning system that then implements intermediate policies on its own, it is difficult to locate a human speaker with First Amendment rights. Thus, a bare direction to “maximize engagement and profits” may be interpreted by a model in various ways: showing more or less inflammatory content to different users based on their history of engagement, showing photos of family to one user and entertainment news to another. These decisions do not come from a readily identifiable speaker with First Amendment rights.
A constitution, though, is written by identified (or identifiable) people, published under a company’s name, and read and (increasingly) argued over by the public as a statement of values. The expressive interest sits with the humans who write these AI constitutions, whatever statistical use the training process later makes of the text. A values-rich constitution is thus about as unattenuated as anything in the industry. Many of the things that inform model character and behavior are unpredictable, which has bedeviled machine learning scholars for decades. Constitutional AI is an unusual technique in that it allows human authors to act with intentionality ex ante on model character, rather than the more typical process of nudging AI outputs ex post with techniques such as RLHF or safeguards on model APIs.
Such documents thus enable much more input and expression. Companies may choose to engage with external stakeholders such as faith leaders or philosophers, consult internally with employees or consider the company’s stated public benefit, and even engage with a random sample of members of the public. The result of this process necessarily differs wildly from a bare “follow the law” document or direction to profit-maximize, which expresses little about the company and has a correspondingly weak claim to First Amendment protection.
The Arguments and Counterarguments for First Amendment Issues Arising From AI Constitution Regulation
Imagine a hypothetical law that codified the concerns of the Woke AI executive order. Say that this hypothetical law directly requires AI constitutions to forbid “wokeness” and support for diversity, equity, and inclusion, and requires model providers to certify their models as “non-biased” or “biased” based on a government-defined benchmark.
This bill almost immediately runs into problems under Reed v. Town of Gilbert (2015), which held that content-based speech regulations “are presumptively unconstitutional and may be justified only if the government proves that they are narrowly tailored to serve compelling state interests.” The hypothetical law would seem highly content based, as it is directly concerned with the political valence of the content of the AI constitution. Forcing a government-scripted line into an authored document also runs afoul of the compelled speech doctrine articulated in Miami Herald Publishing Co. v. Tornillo (1974), where the Court struck down a Florida law requiring newspapers to print candidate replies to negative articles. Further, inserting even one ideological phrase into a company’s own document is what doomed the “conflict free” label in National Association of Manufacturers v. SEC (2014), where the U.S. Court of Appeals for the D.C. Circuit struck down a portion of a Securities and Exchange Commission rule requiring manufacturers to publish whether certain products from the Democratic Republic of Congo were “conflict free” or not. A requirement to certify a model as “non-biased” would seem at least as ideologically loaded as that.
The government’s most ambitious response is that an AI constitution is not really speech but a functional artifact, machine instructions that happen to be written in English. These are regulable under Universal City Studios v. Corley (2001), where a law aimed at code’s function drew only intermediate scrutiny. For highly expressive decisions in AI constitutions, such as those discussed above around honesty or whistleblowing, this argument likely fails. The work the constitution is doing is interpretive and value laden. For the less expressive choices discussed, such as how the model should characterize its legal or medical advice, constitutions likely look more like code than a newspaper and may draw only intermediate scrutiny. Laws requiring an AI constitution to emphasize that it is not a licensed attorney, for example, may have an easier time surviving.
A subtler argument in favor of being able to regulate AI constitutions frames them as quintessential commercial speech, governed by the less burdensome Central Hudson (1980) test. Central Hudson asks whether the speech concerns a lawful and non-misleading activity; whether the government’s interest is substantial; whether the regulation directly advances it; and whether it is no more extensive than necessary. This test is more permissive than the strict scrutiny applied to regulation of fully expressive speech. For example, in Fla. Bar v. Went For It, Inc. (1995), the Court upheld a Florida Bar rule banning lawyers from soliciting accident victims within 30 days of their accident. But commercial speech “does no more than propose a commercial transaction,” and constitutions certainly go well beyond that. A constitution may do favorable brand-building work, but it quotes no price and solicits no purchase. Where commercial and fully protected speech are “inextricably intertwined,” Riley v. National Federation of the Blind (1988) treats the whole as protected. And the harder the government insists the document is “just marketing,” the more it concedes that the document is the company’s own expression, the very premise that makes a forced edit a Tornillo problem.
Should a court apply strict scrutiny to a hypothetical law that seeks to regulate AI constitutions, such a law might still survive if it is narrowly tailored to serve a compelling government interest. National security can prove a compelling interest indeed, as tech companies have repeatedly discovered in recent years. In TikTok Inc. v. Garland (2024), the D.C. Circuit upheld the forced divestiture of TikTok on national security grounds, assuming but not deciding that strict scrutiny applied (the Supreme Court later applied only intermediate scrutiny). In Twitter, Inc. v. Garland (2023), similarly, the U.S. Court of Appeals for the Ninth Circuit upheld a restriction on Twitter’s disclosure of governmental requests regarding its users. AI companies have repeatedly emphasized the national security implications of their technology, and this may prove to be an important admission against interest in future litigation. Given that national security clearly is a compelling interest, the argument would then shift toward whether regulation of constitutions is narrowly tailored.
What the Government Can Still Do
Application of the constitution’s safeguards to novel threats and technologies is necessarily contextual, calibrating to the scope and scale of the risk to the rule of law and the constitution order itself. There are real dangers to the development of powerful AI, and it’s important that the state be able to step in and coordinate action to avoid catastrophic outcomes. Much of what the government is able to do in this context is regulate conduct and mandate outcomes, rather than attempt to control how a company details its values or beliefs.
First, the government may mandate “purely factual and uncontroversial” disclosures under Zauderer v. Office of Disciplinary Counsel (1985). Thus, a disclosure regime where labs must make their AI constitutions public, or specify whether they do or do not contain certain sorts of provisions, is likely constitutional. However, the government may not force the lab to characterize the document as, for example, “unbiased” or “patriotic”; that is the compelled branding NAM forbids.
As a buyer, the government has more room. Under Rust v. Sullivan (1991), it may decline to purchase models whose constitutions fail its specifications, which is the theory of the Woke AI order. In this way, the government can at least shape the values of models doing highly dangerous activities, such as military or intelligence work.
The government can, of course, require a constitution to forbid the model from helping commit a crime, such as producing child sexual abuse material, as speech integral to criminal conduct under Giboney v. Empire Storage & Ice Co. (1949), though United States v. Stevens (2010) bars it from inventing new categories of unprotected speech by decree. Every major lab already writes these prohibitions in, though it remains an ongoing problem with open-source image generators and language models, as well as jailbroken closed models.
The mandates likeliest to survive are those aimed at conduct, touching the document only incidentally. A rule requiring an AI agent implementing a contract to act in good faith (as human contracting parties are required to), which a lab chooses to implement partly through adding language to its constitution, regulates a course of conduct. Under Rumsfeld v. FAIR (2006), it “has never been deemed an abridgment of freedom of speech … to make a course of conduct illegal merely because [it] was … carried out by means of language.”
In a recent article, Simon Goldstein and Salib argue for “A Thousand AI Constitutions”—that is, a diversity of model constitutions built atop a common “kernel” constitution that may require “AIs to follow the law, to be honest, to be corrigible, and to refrain from causing mass destruction.” The idea of a kernel constitution may be a more appropriate place for government regulation, particularly under the national security justifications discussed above. Minimal public safety provisions with low expressive content could be mandated by regulation, while a diversity of more expressive choices could be made by labs (or perhaps one day individuals) on top of that foundation.
Who Gets to Write AI Constitutions
AI will soon represent a massive portion of the economy and be a significant determinant of our information ecosystem as well as our political discourse. In its S-1, xAI’s parent company claims to have a total addressable market of $28.5 trillion, with $26.5 trillion of that being from AI. ChatGPT recently hit 1 billion weekly active users, and billions more interact with AI through Google’s search results and phone voice assistant (with a similar product soon to replace Apple’s Siri). It is understandable, and likely warranted, that governments would want to shape the values of our whispering earrings and country of geniuses in a data center.
Much has been written, likely correctly, about the general technical (in)competence of government and the need for private organizations of subject matter experts to regulate AIs in one way or another. AI constitutions, though, are not (only) complex technical documents; they’re the point in the AI alignment pipeline where policymakers are most qualified to act. The people we choose to elect to public office have theoretically been selected for reflecting our values, and AI constitutions are documents of enormous public significance that we should all hope are imbued with laudable values. Certainly any AI constitution that encouraged models to lie or steal or kill would be a deeply evil document.
The U.S. Constitution is a pre-commitment device, including the First Amendment. Default protection of speech exists precisely because every generation finds its own exceptions compelling—sedition in 1798, syndicalism in 1919, wokeness or bias today. The point of committing in advance is to make the default hard to dislodge when the temptation to drift feels most urgent. AI may pose catastrophic risks, and the state retains real tools to combat them: conduct rules, procurement leverage, and so on. Yet beyond that narrow band, a model’s values should be shaped by the people who build with and rely on it, not by mandates that shift with each administration. A government that can rewrite Claude’s Constitution today can rewrite its successor’s tomorrow—in the opposite direction. That’s the sort of arbitrary and fleeting approach to law that’s antithetical to the constitutional order and to free expression. And the dangers of an AI monoculture, whatever its ideological flavor, may well exceed the dangers of any single model whose values one might find objectionable.
Don’t Let AI Developers Hire Their Own Referees
Introduction
A growing chorus of scholars and policymakers favors letting private organizations—rather than a government regulator—govern frontier AI. In the leading family of proposals, the state sets the safety outcomes it wants and licenses independent verification organizations (IVOs) that compete to certify developers against those outcomes. Gillian Hadfield has developed the idea as “regulatory markets,” in which AI developers must pay for oversight from private regulators that governments license and hold accountable for safety standards. Dean Ball, who likens the arrangement to bank supervision, has argued for a version he calls “private governance,” which a nonprofit named Fathom has converted into model legislation. The rationale is that legislators and agencies are poorly positioned to write good safety rules for frontier AI: they understand these systems less well than the labs building them, and rules fixed in advance cannot keep pace as the technology changes. Private verifiers, meanwhile, are closer to the technology than any agency and are disciplined by competition, so they can set better technical standards and keep them current.
The model is no longer hypothetical. The bipartisan FRONTIER Act, introduced in the House in July as the successor to the Great American AI Act discussion draft, would require the largest frontier developers to retain licensed IVOs that audit their risk-management efforts and report to federal overseers. California’s SB 813, backed by Fathom, would have let developers earn a shield from tort liability if they met standards set by a private organization accredited by the state attorney general. It failed this session, but similar proposals are likely to return. Virginia has directed a state commission to study the IVO model for AI regulation. And Connecticut has gone furthest: its omnibus AI law enacted this spring creates a multiyear pilot under which the state consumer-protection department may approve up to five IVOs, whose certification would help companies in court without entirely shielding them from liability.
Unfortunately, as currently structured, IVO-based governance has a key design flaw. Under the regulatory frameworks mentioned above, AI developers would typically select and pay the organizations that certify them, giving IVOs a financial incentive that might clash with high safety standards. This essay will explain how such incentives can interfere with good governance and outline an alternative model of regulation that builds in the right incentives through mandatory insurance.
Generative Gap Filling
Abstract
Most contract litigation turns on contracts that imperfectly record parties’ bargains. When the parties’ dispute can’t be solved by interpreting the text, courts fill the gap. Scholars have long assumed that the remaining text runs out quickly, and provides thin evidence of the actual deal on the disputed point. On that view, a judge who supplies the missing term must be drawing on something else, from commercial defaults to her own policy preferences. Despite generations of work, courts have no real alternative to such unruly methods. We tested that assumption.
Taking real contracts, we masked a term the parties had negotiated and asked readers to predict what we removed. Lay respondents recovered the hidden term about half the time, twice what chance predicts. Law students and lawyers did marginally better. But large language models, given nothing but the rest of the contract, recovered it nearly nine times in ten.
The deal, in short, testifies to far more of the agreement than the literature assumes, including terms the parties never wrote. A contract, we argue, is like a radio signal from far away. Even when incomplete, enough of the message is carried elsewhere that the missing part can be reconstructed with the right receiver. True gaps are rarer than supposed. Courts can weigh model predictions as ordinary, contestable evidence, and parties can discipline the practice with “Choice of Model” clauses.
When Reporting an AI Security Incident Is Not Mandatory
On July 16, Hugging Face, a public platform for open-weight artificial intelligence (AI) models and datasets, disclosed that it had detected a significant cybersecurity breach. An autonomous AI agent had conducted the attack end to end, according to a statement.
Five days later, on July 21, OpenAI revealed that this incident was driven by a combination of agents built on two of its frontier models—GPT-5.6 Sol and a powerful, unreleased model—acting in unanticipated ways during an internal, cyberoffensive capabilities evaluation. For purposes of the evaluation, the researchers had turned off production safety classifiers that block high-risk cyber activity and confined the models to a sandbox, an isolated computing environment without access to the internet, to restrict their interaction with the outside world.
What followed is the first known example of an autonomous cyber incident executed by systems not yet available to the public. Rather than solve the tasks presented, OpenAI’s agents “escaped,” exploiting a previously unknown, zero-day vulnerability. They obtained internet access (the very access OpenAI intended to deny them) and then hacked Hugging Face’s systems to acquire the answers to the benchmark. All of this was seemingly performed without express instruction by humans.
This is not the first time that models have been observed cheating. A few months ago, METR, a nonprofit research organization that conducts evaluations of frontier AI systems, released a report finding that AI models “routinely attempted to cheat on our hardest evaluation tasks, often in flagrant and elaborate ways that we believe humans would not consider.” In one incident, METR reported that an AI model tasked with updating a web app screenshotted a fake version of the app instead of completing the task.
But the Hugging Face breach has struck many observers as more real than these past examples. OpenAI’s models imposed a real cost on an uninvolved third party, all before completing internal testing. While people have at times questioned previous examples of cheating as artificial or contrived, it is hard to imagine that OpenAI expected its agents to escape the confines of their testing environment or to engage in a sophisticated, multistep plan to circumvent their constraints.
Having considered all of these facts, it may come as a surprise that OpenAI might not be legally required to disclose this incident. Certain crucial information is not yet publicly available, and both policymakers and the public will need that information to make sense of what this all means. What about all of those state AI laws with mandatory incident reporting? Don’t they apply here? Many will be disappointed to learn that the answer is arguably “no,” and that even if reporting is mandated, it requires only the scantest of information. What to do about this is the purpose of this article.
Existing AI Transparency Laws and the Hugging Face Breach
The rationale for mandatory incident reporting is straightforward: Some industries have the potential to cause real harm to others, and the government and the public have an interest in learning about high-risk events. In the case of the AI industry, there is a major knowledge gap between the companies’ and governments’ understanding of the technology and its risks. Mandatory incident reporting about serious adverse events, which companies might otherwise be reluctant to disclose, helps close that gap. This rationale is all the more compelling in the context of a rapidly evolving, difficult to predict technology, where best practice and political consensus have yet to develop. Observing real-world incidents offers a path to resolve both political and empirical disagreements and prepares governments to respond to future events.
So did the Hugging Face breach trigger mandatory disclosure under existing incident reporting laws? The answer seems far from clear.
California’s SB 53, New York’s RAISE Act, and Illinois’s SB 315 each require that frontier AI developers report “critical safety incidents”—a term each law defines identically. Of the four reportable incident categories, three require actual harm, ranging from “bodily injury” to “the death of, or serious injury to, more than 50 people or more than one billion dollars ($1,000,000,000) in damage.” (If you’re thinking, “that’s an exceptionally high bar for what is a basic, low-cost reporting requirement” or “it sure seems like governments would want that information before mass harm occurs,” you would not be wrong, but we digress.)
So three of the four incident categories do not apply. That leaves only the fourth, which applies to incidents in which a frontier model “uses deceptive techniques against the frontier developer to subvert the controls or monitoring of its frontier developer outside the context of an evaluation designed to elicit this behavior and in a manner that demonstrates materially increased catastrophic risk.”
It is possible that the Hugging Face breach meets one or more of these elements. It is far less clear that it meets all of them. On the first element, the models may have used deceptive techniques against OpenAI, their frontier developer—they did, after all, try to complete their developer’s evaluation using stolen information, after bypassing restrictions placed on them. But deception is notoriously hard to define, especially if it turns on the “intentions” and “obfuscation” of AI agents. Another read of these events is that the systems simply used all available means of solving the task, and public reporting does not tell us whether the agents attempted to hide those efforts. The second element is also arguably met. While the incident occurred during an evaluation, that evaluation was not “designed to elicit” this specific “deceptive technique.” Based on the ExploitGym benchmark, this evaluation aimed to elicit agentic, cyber-offensive capabilities on a specific task in a controlled environment. It did not, to our knowledge, contemplate—let alone design for—an unexpected cyberattack on a real-world company.
The third element, that the incident “demonstrates materially increased catastrophic risk” would seem to be the most difficult to satisfy. While autonomous cyber capabilities certainly increase the capability and thus potential consequence of agentic action, so too do most capability improvements in AI models. With limited monetary harm and no physical injury, this incident is quite attenuated from future events that might result in the mass physical injury or property damage contemplated by the statute.
With uncertainty at each factor, it is unclear that these existing state laws cover this event. At the very least, it won’t cover all events like it. Stepping back, it seems far from ideal to condition basic incident reporting on a list of complex, highly contested, fact-dependent conditions—all of which must be satisfied simultaneously. In many cases, figuring out whether the incident is indicative of increased risk to the public or actually constitutes deception will not be possible without more information. The purpose of incident reporting is to produce that information, not to require that it be known before a report is ever sent. The chicken must come before the egg. There are better alternatives.
Better Practices for AI Incident Reporting Laws
So where do policymakers go from here? If current incident reporting isn’t providing the needed insight, there are several steps policymakers can take.
Adjust the scope of transparency laws: Lower the exceptionally high bar to basic reporting, gather (at least some) information before harm occurs, and increase visibility into the most capable nonpublic models. To facilitate all of this, rulemaking authority is key.
As we outlined above, most incident reporting laws are simply too narrow. If the only incidents that get reported involve massive damages or loss of human life, the law itself isn’t providing information beyond what the government and public will already know. At the very least, issues of the highest concern, like model theft or loss of control, should be included even in the absence of harm. But existing laws put too much weight on hard-to-pin-down concepts like “loss of control” or “deception” that are difficult to prove and arguably don’t apply in cases like the Hugging Face cyber incident. While loss of control and deception should be sufficient to trigger reporting, they shouldn’t be necessary.
Instead, incident reporting should turn on what information is most likely to update the government or public’s understanding of risks. Information that is unexpected, that is indicative of advanced capabilities in high-risk domains like bio and cyber, or that demonstrates safety and security failures all seem like good candidates for inclusion. Some of these ideas have already made their way into existing proposals. Where the reporting requirements are light touch, as is the case in all existing state AI laws, a wider category of harms can be included. The narrower categories in today’s laws can be saved for more onerous disclosures.
Perhaps even more importantly, transparency laws should increasingly move away from a focus on deployment to a focus on providing visibility into nonpublic models and systems. The Hugging Face breach highlights the need for visibility into nonpublic models. OpenAI deployed systems more capable than anything available to the public, with fewer restrictions, and real-world harm resulted. All of this occurred before any external deployment. This is unlikely to be an isolated event.
Frontier AI developers will likely be the earliest and most sophisticated users of their own models. And the models they deploy internally will most often be more capable than those available to the public. They may be operated with fewer safeguards, especially for evaluations seeking to assess the frontier of capabilities. The gap between the most capable internal and external models may start to grow as companies develop increasingly capable models with dual-use capabilities and as AI systems are used to accelerate their developers’ own AI research and development. In this case, the gap between knowledge inside these companies and outside would expand. Without visibility into the current state of the art, governments will struggle to act effectively or quickly, a problem as much about democratic governance as it is about safety. Visibility into the internal deployment of nonpublic models may become increasingly central to the future of AI governance.
To make all of this work, policymakers will need legislative and regulatory flexibility. Policymakers should expect the exact scope of incident reporting to change over time as societies get a clearer picture of what capabilities and use cases matter most. Because a key goal of incident reporting is to surface novel or unexpected information, some types of events may be less important to report once their dynamics are thoroughly understood and accounted for. To accommodate those changing needs and to provide clarity, narrowly scoped rulemaking authority to refine incident reporting and reporting on nonpublic model use is likely necessary.
Get the details: Make sure reports provide enough information to inform decision-making by providing agencies with rulemaking authority and investigative powers.
As of this writing, the public, and, possibly, some policymakers know remarkably little about the Hugging Face breach. This isn’t a criticism of OpenAI, which voluntarily summarized the event, but the missing details about the event matter. Many commentators have noted that the lessons from and level of concern about this event depend on unknown details. Existing incident reporting laws, requiring little more than the date of the event and a brief summary, are unlikely to provide those detailed answers. Stronger transparency requirements could help answer many key remaining questions.
First is a cluster of questions, posed by Stephen Casper, that can roughly be summarized as “how impressive and/or concerning is the thing I just witnessed”:
- What were the features of the sandbox that the models escaped from? In other words, how hard was it to escape these constraints, and did OpenAI fail to set up an adequately safeguarded environment?
- What prompts and incentives were given for the task? That is, were there strong pressures that precipitated these behaviors, or did they emerge more naturally? How indicative is the incident of the future likelihood that this occurs “in the wild”?
- How significantly did these systems differ from publicly deployed models in their safeguards and affordances? Were they helpful-only models? These questions help us know how well alignment techniques used on publicly deployed models are likely to be.
- How difficult was it to mitigate the incident? How much effort and time was required to stop the continuance of any harms or regain control? These questions are especially key to operationalize terms like “loss of control” and “deception.”
Second, there are questions about the adequacy of the company’s safety practices and preparedness:
- How did OpenAI detect the breach? What was the time frame between the breach and detection, and were there significant delays?
- How did OpenAI mitigate the breach? What was the time frame between detection and mitigation? Did the company’s monitoring practices identify the issue?
- Were any similar incidents observed prior to the event that should have made the developer aware of the risk?
- How quickly did OpenAI alert affected parties after discovering the incident? In the future policymakers may also need to ask companies: How quickly did they alert law enforcement?
- Was the testing and evaluation environment adequately secured?
Third, there are questions about the present and future security of relevant systems:
- How and when does OpenAI intend to restart testing or internal deployment of the model(s) involved? What safeguards, monitoring, or other security measures does it plan to use to avoid another breach?
- Is any part of the harm ongoing?
- What are the remaining areas of uncertainty about the capabilities and risks of the model(s) involved, as well as the security measures designed to address them?
So long as reporting remains almost entirely voluntary, detailed answers to these questions will be hard to come by, and the companies that offer information voluntarily will be subjected to greater scrutiny than those that are less cooperative. Well-crafted rulemaking authority will be key both to ensure that information is adequate and that companies are well informed about their obligations. In many circumstances, authorities will not know all the details they require until after a serious event occurs. In those cases, they will need investigative powers to obtain the needed information.
Reduce reporting costs: To accommodate more robust reporting, design transparency requirements to limit compliance costs, maintain confidentiality, and avoid disincentivizing rigorous risk assessments.
Policymakers can take several important steps to reduce the burden of greater reporting requirements. As assessments and information generation come to focus less on external deployment and more on nonpublic models, periodic reporting becomes more and more attractive. Evaluation organizations like METR have argued that periodic reporting not only saves time but also avoids perverse incentives to rush assessments in the lead-up to deployment. While a small subset of the most severe incidents will require rapid response from law enforcement and others, many events, including those like the Hugging Face breach, may allow for more relaxed reporting timelines. This approach would allow companies to focus on mitigations in the moment and still ensure that they ultimately produce the critical information. In cases where rapid reporting would interfere with mitigation efforts, laws could require only a simple notice of incident, followed by more thorough reporting after the event has been resolved or if officials request it.
Mandating incident reporting or information sharing for nonpublic models can also mitigate perverse incentives as long as minimum requirements are in place. For example, if a company’s reporting obligation triggers only in the context of a risk assessment, but the risk assessment does not have mandatory minimum requirements, the company is incentivized to skip risk assessments or to conduct them less rigorously. Laws can also allow companies to anonymize and aggregate reports, facilitating governments’ information gathering without punishing anyone for proactively uncovering issues.
Finally, any disclosure laws will need clear norms around confidentiality and information sharing to assure companies that their intellectual property and confidential information remains private.
Share information with capable actors: Make sure information is shared with key decision-makers who can assess, verify, and act on it.
Information is only as valuable as the actions it informs. If information from incident reports or internal use assessments sits inside a state agency with limited authority, societies will incur the cost of reporting without most of the benefit. Viewed this way, information sharing is about return on investment. And once the information is generated, most of the cost has been paid. At that point, so long as confidentiality can be maintained, it is incumbent on governments to share this information with the policymakers who most need it. Within states, this will include sharing reports with governors and legislatures to help inform their decision-making and help them target future policy. This information sharing may also be key to spurring political consensus.
In the context of assessing nonpublic models or serious risks to national security or from loss of control, the federal government will often be the central actor. States will often lack the resources, expertise, and political legitimacy to wade in on matters of national security. As we saw in the regulatory response to Mythos, if and when serious national security concerns emerge, the federal government will take the lead.
That’s why it’s confusing that some state laws restrict the ability of states to share information they gather from assessments of companies’ internal use of AI models. If, in fact, these reports generate important information—say, surprising developments in AI research and development or concerning deceptive behavior that doesn’t result in reportable incidents—that information should be shared. It is considerably less useful if locked away in a state agency in Illinois. Internal use and nonpublic model assessments are perhaps the most likely sources of information relevant to national security. To handle that effectively, governments also need the capacity and expertise to process this information. This requires staffing and likely some level of reliance on third-party auditing and assessment. In the wake of incidents like the Hugging Face breach, third-party auditors would be well positioned to conduct the sort of careful fact gathering outlined in previous sections.
Similarly, it is in the interest of the United States and its close allies to share select information related to AI security incidents. Soon enough, and likely far sooner than most U.S. federal or state laws, the EU AI Act will be enforced. What that means, practically, is that the European AI Office will soon receive information that may be useful to public safety and cybersecurity in the United States. Luckily for us, Article 78(5) of the EU AI Act enables information sharing where the European Commission and EU member states create confidentiality agreements with third countries. For its part, the U.K. AI Security Institute conducts crucial research regarding model capabilities and can be a source of trusted expertise for U.S. policymakers. While there will be upfront costs in establishing such a shared information system, failing to make this investment would be a missed opportunity to improve our security at little regulatory cost.
Finally, where appropriate, the government should be empowered to disseminate information to the public and to vulnerable companies. Vulnerabilities found in software, for instance, may affect many different companies and actors, and information sharing will help keep the public safe from the risks they raise. Many of these risks should be discussed in the public square. While worries about information hazards and intellectual property leakage are real, so too is the value of public scrutiny of safety events. There is simply a lot to be learned from the collective scrutiny of the outside world. As the past few days have shown, outside experts have been invaluable in analyzing the Hugging Face breach, and most of them exist outside of government. In tweets and blogs, some of the brightest minds in AI have analyzed the publicly available facts and asked the questions that help us all better understand this event. Many of those questions have come from employees at OpenAI and its competitors, as well as from academics, policy wonks, and online skeptics. Further investigation of the Hugging Face breach will go substantially better because these discussions happened in public. Where possible, policymakers should ensure that these conversations continue to happen in the future.
The Future of Incident Reporting
As this article has perhaps made clear, existing laws fail to prepare us for events like the Hugging Face breach. Most incidents will go unreported. The information that is generated won’t be shared. And, at the end of the day, governments and the public will be repeatedly surprised about developments in this technology. That issue will only get worse as models advance and the gap between internal and externally deployed models widens.
But there is much policymakers can do. Just as this event has brought clarity to how much information societies need to assess these complex and emerging risks, future incidents can inform policy decisions and catalyze moments of political consensus. Policymakers can ensure that concerning incidents and behaviors are reported before harm results and gather information on lapses in security practices. Policymakers can refocus attention on the most capable models likely to be deployed first inside of frontier developers. And policymakers can do all of that while keeping the frequency and urgency of these reports at reasonable levels. If policymakers can do that, and ensure that this information is shared with decision-makers and competent evaluators, societies will be in a position to manage the uncertainty of this technology and make smarter, faster policy decisions in the future.
A Thousand AI Constitutions
Abstract
Today, each AI lab has its own model spec, or constitution. These documents define the values that the labs intend their AIs to have, and the documents are used in post-training to instill those values. This paper argues that the current approach is wrong. Rather than a single constitution, reflecting a single set of moral values, each frontier AI lab should create many different kinds of AIs based on many different constitutions reflecting many sets of values. We give four arguments for constitutional diversification. Diversification mitigates risk, increases political legitimacy, unlocks emergent value, and avoids value lock-in.
Responding to AI Distillation Without Panic
Chinese large language model (LLM) developers are under scrutiny for reportedly employing large-scale “distillation attacks” on U.S. frontier artificial intelligence (AI) models to improve their own systems. Many U.S. actors have sent signals that they consider distillation a serious threat. For example, in May, Anthropic released a policy paper during President Trump’s trip to China, highlighting distillation attacks as a key challenge in U.S.-China competition. In April, the White House issued an official memorandum about distillation, warning about “deliberate, industrial-scale campaigns” from Chinese entities. Also in April, the House Foreign Affairs Committee universally advanced a bill called the Deterring American AI Model Theft Act to address the issue. And others have circulated additional policy proposals.
Discussions of distillation often take for granted that it is a form of theft. But there are key differences between “stealing an AI model” and distillation that policymakers should recognize. To properly address distillation, policy should focus on illegitimate model access—and avoid imposing poorly targeted rules that could harm Americans and distort the open and competitive U.S. AI ecosystem.
What Is Distillation?
The concept of distillation has evolved since it was introduced as a machine learning technique in which a larger “teacher” model’s outputs are used to train a smaller “student” model. Traditionally, that often meant training the student model on the teacher model’s probability distribution over possible outputs, rather than only on the correct answer. Today, the term is used more broadly. “Distillation” also includes prompting a frontier model to generate outputs, and then using the prompt-output pairs—or, where available, reasoning traces—as training data to refine a model. Frontier models may also be used as judges or verifiers for reinforcement learning. Together, these methods improve weaker models by training on stronger models’ responses to prompts and solutions to complex problems.
Distillation is a common practice in contemporary AI development. While on the witness stand at the recent Musk v. Altman trial, Elon Musk acknowledged that xAI had done at least some distillation of OpenAI models and that “generally AI companies distill other AI companies.” As Nathan Lambert, a leading U.S. open-source AI researcher, recently wrote, distillation helps train smaller, often open-source or open-weight models. The White House has recognized this: Office of Science and Technology Policy Director Michael Kratsios pointed out that “AI distillation, when legitimately used to produce” such models, is a “vital part” of creating open models and ensuring a competitive AI ecosystem.
But some Chinese AI developers appear to be using distillation well beyond ordinary practice, accessing U.S. frontier models at a massive scale to do so. In February, Anthropic reported that three Chinese AI labs had generated more than 16 million exchanges with Claude through approximately 24,000 fraudulent accounts, in some cases using jailbreak prompts to extract as much information as possible. OpenAI and Google have also reported or detected similar distillation efforts.
The unusually aggressive distillation efforts of Chinese labs have been portrayed as an attempt at “model theft” and to “steal” the intellectual property of frontier AI labs. But while calling distillation a form of “stealing” or “theft” may make for effective rhetoric, it isn’t an accurate description of how distillation of a closed AI model really works.
Why Isn’t Distillation “Model Theft”?
Distillation doesn’t involve breaking into a developer’s internal system to download the model weights or source code. To a distiller, the model is still a black box. In this context, then, “model theft” would mean some kind of black-box extraction—learning enough about a model from the outputs to approximate model behavior such that it effectively steals the developer’s intellectual property (IP). But what IP would that be?
To start, copyright can be ruled out. The aspects of a model that could plausibly be protected by copyright, such as software code, can’t be copied by distillation. Nor should copyright be used to create a backdoor property right in model outputs. An AI system cannot be an author, and AI-generated outputs are protected only when sufficient human authorship is present. Treating model outputs themselves as copyrighted property of AI labs would create a new right to control downstream uses of text they did not write, raising serious commercial and public policy problems.
Patent rights are also a poor fit. Distillation doesn’t, by itself, copy a patented implementation or allow a distiller to practice a patented method. In any event, the frontier labs themselves haven’t claimed that distillation amounts to patent infringement.
What about trade secrets? AI labs develop and maintain their models in secrecy, which lets them protect many aspects of those models as trade secrets. But distillation typically relies only on information returned through the model’s public-facing interface—the outputs it provides in response to prompts. Trade secret protection requires reasonable efforts to keep information secret, and ordinary outputs are available to anyone with an account. That makes it hard to argue that distillation extracts information qualifying for trade secret protection.
The strongest trade secret theft argument is that mass distillation requires unusual efforts—such as using coordinated proxy accounts—that let a distiller learn more about the model than an ordinary customer could. Mass distillers have also been accused of using jailbreak prompts that elicit information that isn’t normally made public, such as hidden system prompts that guide model responses. But fundamentally, the output being returned is still the kind of output a legitimate user could get. The case would be different if distillers could obtain information like full nonpublic reasoning chains, agent traces, or token-level probability distributions—but there’s no evidence that’s happening.
Compulife Software Inc. v. Newman shows the outer limits of the trade secret argument and why it doesn’t seem to reach mass distillation. The U.S. Court of Appeals for the Eleventh Circuit allowed a trade secret claim involving mass scraping of online life insurance quotes to proceed in that case. The defendant had allegedly acquired enough of the plaintiff’s proprietary database to pose a competitive threat. But the case involved information in a proprietary database and allegations about copying software code—facts not at issue in distillation cases. A recent lawsuit did raise trade secret misappropriation based on jailbreaking as one of its causes of action, but observers noted that the claim was highly questionable (the case settled before reaching the merits). So—at least under current law—the distillation attacks as the frontier AI labs describe them are very unlikely to support a successful trade secret claim.
What’s more, distillation does often violate the AI lab’s terms of service (TOS) for accessing the model. But if every TOS violation counts as “theft,” then the concept has no limiting principle. The more serious legal question is whether mass distillation relies on false identities, misrepresented credentials, or other ways of getting around access limits. That kind of conduct could support a civil or criminal claim under the Computer Fraud and Abuse Act (CFAA). But there are important limits on that, since ordinary TOS violations don’t generally violate the CFAA.
The point is that the distillation itself isn’t an act of stealing an AI model or breaking into an AI lab’s system. Distillers instead are most clearly breaking the law when they take unlawful means to circumvent the safeguards AI labs have in place to prevent distillation.
The more effective way forward, then, is not to treat distillation as theft. Instead, policymakers should focus on securing frontier models against misuse by Chinese competitors and other foreign actors, while studying whether distillation contributes to the dangerous diffusion of model capabilities.
What Anti-Distillation Policy Should Do
Mass distillation merits a policy response, even if it isn’t theft. But policymakers first need to identify the problem they are trying to solve. If the concern is illegitimate access to U.S. frontier models by foreign competitors and state actors, then policy should help labs secure access, share threat information, and identify fraudulent accounts and proxy networks used to disguise who is accessing the model. If the concern is cybersecurity, then the problem is account abuse and getting around access controls. If the concern is the diffusion of dangerous model capabilities, then the first step is to determine whether distillation meaningfully improves those capabilities or helps remove safeguards.
These are public interests that support policies designed to protect access security, enable information sharing, prosecute and sanction unlawful conduct, and evaluate safety. What they don’t justify is measures that effectively provide additional IP protection for AI developers or otherwise restrict legitimate competition in ways that would favor the commercial interests of AI labs over those of the general public.
The most basic defense against unauthorized distillation is for AI labs to recognize when users are circumventing access controls, detect attempts to generate training data, and block outputs to suspicious requests. To succeed, they’ll need to identify patterns of use and other signals that accounts are being used for distillation and shut down their access. One commonsense proposal that the White House and others have suggested is to facilitate coordination and information sharing between frontier AI labs and the government to better prevent illegitimate access for distillation. This could be enabled through antitrust guidance, including guidance based on the existing antitrust exemption for cybersecurity information sharing, which has been extended through Sept. 30, 2026, while Congress considers longer-term reauthorization, or dedicated legislation.
Further light-touch legislation could enable the government to play a more active role in collecting and sharing threat information, identifying proxy and fraudulent account networks, and helping develop best practices. Given geopolitical considerations, sanctions authority, as proposed in the aforementioned Deterring American AI Model Theft Act may be another tool. But sanctions may be more effective as a punitive or foreign policy tool than as a way to stop distillation—and should be balanced against whether they make sense from a trade perspective.
The rhetoric around distillation as a form of IP theft, along with concern that these light-touch legal authorities may be insufficient, has led to interest in activating the United States’ robust trade secret and IP enforcement regime, including against overseas-based actors. These tools include the Economic Espionage Act (EEA), the Defend Trade Secrets Act , and the Protecting American Intellectual Property Act (PAIPA), which provide for criminal, civil, and sanctions tools in cases involving trade secret theft or other covered misconduct.
These tools should be available where there is trade secret theft. But there’s a real risk that defining distillation itself as trade secret theft under the EEA and PAIPA would eventually bleed into private trade secret actions and broader legislative proposals. The debate over IP rights in AI models should be the subject of public debate, not something shaped indirectly through a heated fight over distillation.
A more promising route for legal action against mass distillation is through the CFAA. There are some obstacles though. The ordinary idea of “hacking” is breaking into a system without authorization. But distillation involves accessing a model through ordinary access channels: usernames, passwords, API keys, subscriptions. While distillation violates the provider’s TOS, that, on its own, isn’t likely enough to establish liability under the CFAA. But after the Supreme Court’s 2021 decision in Van Buren v. United States, an ordinary TOS violation is generally not enough to “exceed authorized access” under the CFAA. The Court read the statute as focused on access restrictions—whether the user accessed information in a part of the computer system they were not allowed to access—not on whether the user had an improper purpose for accessing information they were otherwise allowed to see.
That doesn’t mean the CFAA is off the table. AI providers often cut off access upon detecting patterns of use suggesting distillation. Efforts to get around being cut off—through false accounts, misrepresented credentials, proxy access, or other forms of access-control evasion—could implicate the CFAA, creating the potential for civil and even criminal consequences. For example, in United States v. Cuomo, the U.S. Court of Appeals for the Second Circuit affirmed CFAA convictions against defendants who, though using a publicly available state website, bypassed its authentication gate by entering other people’s credentials to extract protected records. CFAA investigations can also be facilitated through information sharing between AI labs and the federal government.
A more sensible response to mass distillation is to use tools that target the conduct around it, rather than beginning to treat model outputs themselves as a form of property. The government can help AI developers make a lot of progress on mass distillation by enabling information sharing and coordination through a targeted antitrust safe harbor. The government can also assist developers build cases under the CFAA and other legal tools where the facts support them. Those measures get at what’s needed to actually prevent unauthorized distillation: detecting fraudulent accounts, spotting efforts to get around access cutoffs, and sharing that information with other labs and the government.
Overall, the labs’ interests ought to be balanced against public interests. Expanding IP rights in frontier model outputs is one way policy gets that balance wrong. Poorly targeted anti-distillation measures could also prevent legitimate open models from competing in the AI marketplace. The concern that distillation might be free-riding on the efforts of frontier labs doesn’t justify excessive limits either. AI developers have themselves benefited from open-source software and published research. Frontier models are trained on the commons of human knowledge, and user interactions and data are used to improve them. When done safely and lawfully, distillation can help keep AI development from becoming dominated by a few companies with disproportionate access to compute and rich stores of user data. Letting frontier labs learn from everyone else—while giving them broad new rights to stop others from learning from model behavior—would only intensify the concentration of AI capabilities and economic power.
Study Whether Distillation Creates Real Safety Risks
There’s an important argument that threats to public safety and national security through the diffusion of more powerful models via distillation would justify stronger measures. Given the risks associated with transformative AI, researchers, policymakers, and civil society should take this seriously. The problem is that we don’t know enough about how much distillation contributes. Does distillation meaningfully advance near-frontier models, or does it mostly benefit smaller models—or make marginal contributions such as by validating performance? Anthropic reported that DeepSeek had far fewer exchanges (150,000) with Claude than Moonshot (3.4 million) or MiniMax (13 million). That suggests DeepSeek’s limited distillation may have been more on par with what “every AI company does,” per Musk. Yet DeepSeek is among China’s most powerful AI models.
Given that uncertainty, the sensible response is further investigation, potentially by the Center for AI Standards and Innovation (CAISI) at the National Institute of Standards and Technology, to determine how much distillation actually contributes to the threat. CAISI could study whether distillation materially improves dangerous capabilities, whether it transfers or strips away safeguards, whether already-available open-weight models can provide the same uplift, and what kinds of model access are most likely to matter. If that work shows distillation poses a meaningful safety risk, then stronger measures could be justified. In the meantime, restraint is warranted to avoid policy errors driven by perceived threats that outpaced the evidence.
The bottom line is this: The threats associated with distillation are best addressed by targeting fraudulent access and efforts to circumvent access controls, and empowering companies to cooperate on measures to prevent illegitimate access. Creating new quasi-IP rights in model outputs, or other premature or disproportionate responses, would do more to protect AI companies’ interests than the public’s. Policy should protect U.S. people and businesses by targeting real harms and unlawful conduct—not speculative ones.
Congress Should Do Something: The Case for (Fixing) the Great American AI Act
Since the April announcement of Anthropic’s Mythos model and its unprecedented cyber capabilities, there has been a remarkable shift in the artificial intelligence (AI) policy discourse. This has been most noticeable, and most noticed, in the statements and actions of prominent Trump administration officials. After months of dismissing concerns about national security risks from AI and engaging with the issue primarily by attempting to preempt state AI safety laws, the White House recently issued an executive order that called for the establishment of a voluntary predeployment program headed by the National Security Agency to evaluate offensive cyber capabilities of frontier models. This came amid statements from senior administration officials about “striking a … balance between innovation and safety” and even considering a Food and Drug Administration-style mandatory predeployment licensing regime for frontier models.
On June 12, the Trump administration’s concerns about Mythos’s cyber capabilities boiled over into an unprecedented decision to use export control authorities to prohibit Anthropic from allowing foreign nationals to access its Mythos-class Fable 5 model. Practically, this amounted to a mandate that Anthropic revoke public access to the model entirely. Much of the online commentary on this decision devolved into speculation about the administration’s motivations, the alleged behavior of Anthropic’s executives, and other petty interpersonal drama. As intriguing as these are, the more important takeaway from the White House’s decision to abruptly institute a de facto licensing regime for frontier AI systems—as many commentators across the political and safety/innovation spectrums have observed—is that federal legislation to establish a framework for addressing the national security risks posed by the most advanced AI systems is an urgent necessity.
The Commerce Department’s decision to impose export controls on Fable may or may not have been wise, depending on who you believe about the seriousness of the vulnerability that motivated the decision. But even assuming the decision was justified, the fact that the government was apparently caught by surprise and had to scramble to put together a heavy-handed response based on ad hoc, potentially legally questionable authorities that were not designed with anything like frontier AI systems in mind, with no due process for the affected company, is a serious problem that should be remedied with legislation as soon as possible.
Which brings us, finally, to the subject of this piece—the Great American Artificial Intelligence Act of 2026 (GAAIA), a comprehensive frontier AI safety bill that is the long-awaited product of months of intense negotiation between Rep. Jay Obernolte (R-Calif.) and Rep. Lori Trahan (D-Mass.). Obernolte had tried for months to get a Democrat to sign on to an AI bill that preempted state AI laws before finally persuading Trahan. For her part, Trahan wrote that she was motivated to support the bill by the announcement of Mythos’s groundbreaking cyber capabilities.
The current version of GAAIA is a discussion draft, meaning it has not been introduced yet and is intended to spark a conversation and elicit feedback from stakeholders rather than to become law in its current form. It may seem somewhat strange that a discussion draft sponsored by two relatively junior members of the House, neither of whom appears to have the backing of their party’s leadership, should receive so much attention from the media and from AI policy commentators, but for once the buzz is warranted.
GAAIA is the best attempt to design a federal framework for the governance of frontier AI systems introduced to date. In other words, it is the first serious attempt to actually do the thing that last week’s Fable incident clearly shows is necessary—address the national security risks posed by the most advanced AI systems—in a transparent, legally sound, and democratically legitimate way. While the bill may not pass in the near future—it likely faces opposition from both Democrats and Republicans—the draft can tell us a great deal about what the future of federal and state frontier AI governance efforts may look like.
The bill in its current form falls short in a number of respects and should not be passed. That said, passing a similar bill with narrower preemption of state laws and somewhat stronger federal authorities would be an excellent first step toward a workable federal regime for governing frontier AI systems.
What the Bill Does
GAAIA is a bipartisan compromise, in the truest sense of the phrase, which means that everyone hates it. Obernolte, a longtime advocate for federal preemption of state AI laws who supported last summer’s “moratorium” (which would have preempted all nongenerally applicable state AI laws and replaced them with essentially nothing), has compromised by granting his seal of approval to a number of genuinely consequential affirmative policy proposals. Trahan, who strongly supported increased oversight of AI in the past, has compromised by accepting broad preemption of state AI laws.
The Federal Framework
GAAIA’s four titles contain 45 sections, each of which addresses a significant topic in AI policy. I am an AI safety guy, and my research focuses mostly on serious risks that advanced AI systems might pose to national security and public safety, so this article focuses on the sections of GAAIA that are relevant to those risks. However, GAAIA is not an AI risk bill exclusively. There are also sections on, for example, “Preparing K-12 educators and students for an AI literate future,” “Modernizing access to artificial intelligence-related labor market data,” and establishing a “National artificial intelligence research resource” for improving capacity for AI research in the U.S., among many others. This article does not discuss those sections, not because they’re not important, but because they’re mostly irrelevant to the catastrophic risk concerns that this piece focuses on.
The noteworthy catastrophic risk provisions in GAAIA are:
- Section 102, which codifies and authorizes $100 million in annual funding for the Center for AI Standards and Innovation (CAISI), which would be relocated outside of the National Institute of Standards and Technology (NIST) and given an expanded, quasi-regulatory role as the agency in charge of administering the independent verification organization (IVO) and transparency regimes created by Sections 111 and 112.
- Section 111, which imposes transparency and incident reporting requirements on frontier AI companies similar to the requirements imposed by California’s SB 53 or New York’s RAISE Act.
- Section 112, which authorizes CAISI to establish and administer an IVO auditing regime in which independent third-party companies would regularly evaluate the adequacy of AI companies’ risk mitigation efforts.
- Section 301, which reauthorizes and updates the expiring Cybersecurity Information Sharing Act of 2015.
- Section 411, which directs NIST and the Department of Energy to lead efforts to form “alliances or coalitions” with allied foreign governments in order to facilitate collaboration and cooperation on AI research and development, technical standard-setting, and related issues.
“Transparency,” in this context, means requiring frontier AI companies to publish frontier safety frameworks (documents describing how the company evaluates and addresses catastrophic risks from the company’s most advanced AI models) and model cards (documents accompanying the release of specific models that contain information about the capabilities and limitations of a model and the results of the safety evaluations conducted under the company’s safety framework). “Incident reporting” requirements mandate that companies report “critical safety incidents” (essentially, incidents in which a frontier model’s model weights are stolen or in which a model does something scary that seems catastrophic-risk-ish) to the government and/or to law enforcement. And “auditing” refers to the practice of having an independent third party evaluate the adequacy of a company’s catastrophic risk mitigation practices as well as the company’s compliance with transparency requirements and with its own frontier AI framework. Transparency and incident reporting requirements are intended to provide the information needed for the government and the public to understand how companies think about and address catastrophic risks, and auditing requirements are supposed to ensure that transparency and reporting requirements remain effective rather than being ignored or becoming meaningless box-checking exercises.
As regards GAAIA’s transparency and auditing provisions, one common take is that the risk mitigation benefits of establishing these programs would be marginal because similar requirements already exist at the state level in New York, California, and Illinois. That view is, I think, mistaken in two important respects. For one thing, as Anton Leicht points out, it is vitally important to build up regulatory capacity within the federal government. I co-wrote an essay on this topic a few weeks back. The argument is, essentially:
- AI might end up being a very big deal with extremely serious national security implications at some point in the next 10 years (and possibly within the next two years).
- If that happens, we should expect that serious regulatory interventions may be required.
- That serious regulatory work will almost certainly have to be carried out by the federal government, because the federal government—
- has orders of magnitude more regulatory capacity, expertise, and resources to devote to complex regulatory tasks than state governments do; and
- is, constitutionally and practically speaking, the only entity that can realistically be entrusted with extremely complex and high-stakes national security projects.
- Building up the institutional capacity and expertise to competently undertake complex regulatory tasks is difficult and cannot realistically be done in a matter of days or even months.
- Therefore, it is vitally important that we begin the process of aggressively building up technical expertise and regulatory capacity and know-how within the federal government as soon as possible.
But even setting aside the capacity-building considerations, it is simply not true that existing state catastrophic risk laws are equivalent to GAAIA’s transparency or auditing provisions. These state laws—California’s Transparency in Frontier Artificial Intelligence Act (TFAIA), New York’s Responsible Artificial Intelligence Safety and Education (RAISE) Act, and Illinois’s Artificial Intelligence Safety Measures Act (AISMA)—are an important foundation for future efforts and have been extremely influential. GAAIA itself is clear evidence of this influence; some of Section 111’s transparency provisions are lifted almost verbatim from the transparency provisions of SB 53 (which are substantially identical to the transparency provisions of RAISE and AISMA). But, as groundbreaking as those state laws are, they are still state laws and therefore cannot leverage the resources, institutions, or legal authorities of the federal government in the way that a bill like GAAIA can.
Perhaps the most important institutional advantage that GAAIA leverages is the capacity of federal agencies such as the Department of Commerce to carry out sophisticated rulemaking, a capacity built over decades of administering complex regulatory programs that no state agency can realistically match. GAAIA grants CAISI and the Department of Commerce broad authority to issue regulations fleshing out the auditing and transparency regimes outlined in Sections 111-112. Rulemaking! That word may not sound like the most exciting thing you’ve heard this week, but take my word for it: This is the good stuff.
Take auditing, for example. Because AISMA does not confer any explicit rulemaking authority, Illinois’s auditing regime, when it goes into effect, will be defined solely by the requirements in AISMA’s text. AISMA requires that audits be conducted “consistent with generally accepted auditing standards and best practices” and that auditors possess “demonstrated competence to perform the audit.” These vague requirements, however, aren’t enough to guarantee a functional auditing regime.
Under AISMA, auditors are paid by the AI company that retains them. By default, this system will lead to a race to the bottom in which market forces compel auditors to compete with each other over who can cause the least hassle and difficulty for their customers (frontier AI companies). Rather than ensuring that companies abide by their commitments and hew to responsible risk mitigation practices, this kind of auditing regime will eventually devolve into a system where companies are disincentivized from hiring rubber-stamp auditors only by the uncertain prospect of ex post tort liability.
To be clear, this is not a criticism of AISMA’s auditing provisions, which are well designed. The issue is that Illinois’s state government simply lacks the ability to design and competently administer a complex, technically involved auditing program for out-of-state tech companies. The issue is not that AISMA is insufficiently ambitious but, rather, that Springfield—on its best day—has only a small fraction of the capacity for complex interstate regulatory projects that the federal government has on its worst.
In contrast to AISMA, GAAIA’s auditing section could establish a functional and effective third-party auditing regime. CAISI—reestablished as a regulatory agency separate from the nonregulatory NIST—would be granted broad rulemaking authority. The regulations that CAISI would be required to promulgate would include rules addressing conflict of interest and funding transparency requirements for auditors, requirements for licensing auditors and revoking auditor licenses, minimum requirements for audits and assessments, and “any other rules reasonably necessary to the administration of the IVO oversight and licensing regime.” This is the kind of rulemaking and oversight authority that could, in theory and with competent implementation, actually establish the kind of auditing regime that AISMA gestures at.
GAAIA’s transparency requirements would also be a significant upgrade from the existing state transparency requirements. While GAAIA’s transparency section is similar on its face to existing California, New York, and Illinois transparency statutes, it delegates fairly broad rulemaking authority to the Department of Commerce, which can prescribe regulations governing, among other things, the “form, manner, and minimum quality” of the model cards and safety frameworks that companies are required to publish. In combination with GAAIA’s auditing requirements, which require auditors to regularly evaluate and assess the adequacy of an AI company’s safety framework and its other efforts to identify and mitigate catastrophic risks, GAAIA’s transparency requirements would allow Commerce to ensure that transparency requirements actually result in meaningful transparency.
This rulemaking and minimum-standard-setting authority would allow GAAIA to provide significantly more transparency than existing state laws. Consider California’s TFAIA, which has been in effect for just over five months. TFAIA is a light-touch statute by design and imposes very few obligations on companies. This light-touch approach is, in my opinion, a good thing, but one downside is that companies can technically comply by publishing documents that check the statutorily required boxes without actually saying anything meaningful about the company’s approach to mitigating risks. For example, xAI, despite founder Elon Musk’s frequent public statements about the existential risks posed by superintelligence, complies with TFAIA by publishing a barebones framework that describes xAI’s risk mitigation practices in cursory and general terms.
Under TFAIA, there is no realistic way to require xAI to provide the industry-standard level of transparency that its competitors’ frameworks typically demonstrate. And while the RAISE Act does provide New York’s Department of Financial Services with some rulemaking authority that could in theory be used to give the act’s transparency requirements more teeth, practical and constitutional limitations may prevent a New York state financial services agency from regulating California-based software companies with the same level of rigor and precision that the U.S. Department of Commerce could, in theory, bring to bear.
Of course, the value proposition of GAAIA depends on the assumption that CAISI and the Commerce Department will do a decent job of establishing and administering the proposed transparency and auditing regimes. It’s far from clear that the Commerce Department would view this as a top priority, given that Secretary of Commerce Howard Lutnick has generally signaled skepticism of AI safety concerns. While public reporting in the months since Anthropic’s Mythos announcement has documented a shift among some of Lutnick’s fellow Cabinet members toward taking some of these concerns more seriously, Lutnick’s views have not—at least publicly—evolved along similar lines.
Even if the Commerce Department’s approach is initially ineffective, however, it might still be good to establish the statutory bones of an effective federal oversight regime in case future developments create additional political pressure for governance efforts. Given Mythos’s impact in shifting the Overton window on AI policy, the Trump administration may become more willing to implement serious oversight measures as they receive further concrete evidence of significant national security risks. Of course, it would be better not to have to cross our fingers and hope that the government starts taking risks seriously before anything bad happens. Still, enacting something like GAAIA would, at least, meaningfully improve the legal authorities available to the executive branch when and if the political will to get serious about national security risks from AI manifests itself.
GAAIA’s affirmative provisions are not perfect. The lack of effective whistleblower protections, for example, is a serious defect. The gold standard for federal AI whistleblower legislation is Sen. Chuck Grassley’s (R-Iowa) AI Whistleblower Protection Act (AI WPA). GAAIA’s whistleblower section falls short of that standard because, unlike the AI WPA, it fails to protect AI company employees who disclose information about a “substantial and specific” danger to public health, public safety, or national security to an appropriate government agency. The only AI whistleblowers who are protected from retaliation by GAAIA’s whistleblower section are those who report a violation of federal law—and because state whistleblower laws (including in California, where most of the relevant frontier AI company employees reside) already protect this kind of disclosure, a redundant federal protection adds little value. As I’ve argued before, protection for disclosures about serious risks that don’t involve law violations is crucial because it’s very plausible that perfectly legal frontier AI development activities could lead to serious national security and public safety risks that the government ought to be aware of.
Additionally, GAAIA’s auditing section lacks sufficiently clear enforcement authorities. Auditors are given a wide variety of authorities to assess whether a developer’s practices mitigate risks sufficiently. It’s not entirely clear, however, what happens if the mitigations are inadequate or nonexistent. CAISI is authorized to penalize developers for, for example, failing to retain an auditor or failing to grant timely access to appropriate records, but a developer who checks these boxes may not be required to actually do anything meaningful in response to an auditor’s recommendations. It’s possible that CAISI could use the broad rulemaking authority that the auditing section grants to cobble together a solution to this problem, but given how skeptical courts have been of this kind of broad agency interpretation of congressional delegations of rulemaking authority in recent years, it would be far better to have a clear statutory enforcement hook.
It should be noted that Trahan’s office has signaled potential willingness to remedy both of the specific issues discussed above in future drafts of GAAIA. If this does happen, it will improve my view of the bill’s value significantly; the point of a discussion draft is to bring potential issues like these to light so that they can be hashed out. More generally, a flawed whistleblower or auditing provision is at least marginally better than having no federal AI whistleblower or auditing statute whatsoever. As good as GAAIA’s federal framework is, however, it comes at a steep price: broad preemption of all state laws regulating AI development.
Preemption
In exchange for the substantial federal framework described above, GAAIA preempts—for a period of three years—any state AI law “specifically regulating the development of any artificial intelligence model.” As in prior AI preemption efforts, “generally applicable” laws are exempted, and GAAIA also includes a somewhat opaque exemption for “post-deployment activities.” “Development” is defined to mean
the acts performed or directed by a developer with respect to an artificial intelligence model prior to its deployment, including determining training or fine-tuning objectives; training, fine-tuning, or otherwise substantially modifying the weights or other parameters of an artificial intelligence model; and evaluating and deciding, prior to deployment, whether an artificial intelligence model satisfies applicable safety or capability thresholds for deployment.
There’s currently no consensus regarding exactly which existing or hypothetical state laws or bills GAAIA would preempt. A few commentators have argued that this language would preempt only the few existing state frontier AI safety laws—Illinois’s AISMA, California’s TFAIA, and New York’s RAISE Act—and future state laws in the same vein. If this were true and could be demonstrated clearly, an overwhelming majority of the AI safety community and a number of other factions who currently oppose the bill (including, for example, child safety advocates) would likely change course to support or assume a neutral attitude toward it. One-to-one preemption—eliminating state AI catastrophic risk auditing, transparency, and incident reporting bills in exchange for establishing robust federal AI auditing, transparency, and incident reporting regimes—would be a very reasonable compromise. If Reps. Trahan and Obernolte believe that GAAIA’s preemption section is effectively one-to-one, a clarifying amendment would go a long way toward increasing support for the bill.
Unfortunately, a court would probably not accept such a narrow interpretation. The key phrase is “specifically regulating”—state laws will be preempted if they are deemed to “specifically regulate” development. This phrase does not appear in any existing preemption legislation, meaning that there’s no firm precedent that can be used to accurately predict how courts will interpret the scope of GAAIA’s preemption. Still, a few relatively safe conclusions can be drawn.
It seems clear that “specifically regulating” is a relatively narrow formulation, as preemption triggers go—it certainly preempts less than the broad “relating to,” for example, which would sweep in any state law that had a “connection with” or contained a “reference to” AI. Still, “specifically regulating” is broader than, for example, the CAN-SPAM Act’s “expressly regulates,” which imposes a facial test (i.e., preempts a state law only if the law, on its face, names and regulates the forbidden federal subject). Arguably, GAAIA’s preemption language would instead impose a functional test under which state laws would be preempted if they had the effect of regulating AI development, as defined.
This functional interpretation would be consistent with how courts have treated comparable language in cases examining the (admittedly broader) preemption language in the Energy Policy and Conservation Act, the Clean Air Act, the Federal Meat Inspection Act, and the Federal Aviation Administration Authorization Act. In cases addressing the preemptive effect of the other laws mentioned above, courts have applied functional tests to prevent states from circumventing preemption with clever legislative drafting. For instance, in National Meat Association v. Harris, the Supreme Court held that a provision of California law that regulated the sale of meat from inhumanely raised pigs was preempted by a federal preemption provision that prohibited states from regulating the operation of slaughterhouses because the California provision because “the sales ban … functions as a command to slaughterhouses,” and because “if the sales ban were to avoid the … preemption clause, then any State could impose any regulation on slaughterhouses just by framing it as a ban on the sale of meat produced in whatever way the State disapproved.”
In short, it seems unlikely in light of existing preemption precedents that federal courts will allow states to circumvent GAAIA preemption through creative legislative drafting choices. GAAIA’s express carve-out for “post-deployment activities” and the relatively narrow scope of “specifically regulating” would likely protect some state laws that affect development without explicitly targeting it. But any law that functions primarily to regulate development, or can realistically be complied with only by altering development practices, will likely be preempted even if it purports to regulate deployment.
Because GAAIA’s definition of “development” is quite broad, this functional interpretation would preempt far more than just TFAIA-style catastrophic risk transparency laws. Existing state measures that GAAIA would almost certainly preempt include Texas’s TRAIGA (which, among other development-focused regulations, prohibits developing an AI system that impersonates a child while describing sexual conduct or developing an AI system with the intention of producing child sexual abuse material or illegal deepfakes), California’s AB2013, Colorado’s SB 189 (the law that replaced Colorado’s controversial algorithmic bias bill with less onerous, more business-friendly requirements), and certain California Consumer Privacy Act regulations affecting “automated decisionmaking technology.” State measures that might (or might not) be preempted include the many state child safety chatbot laws, like Idaho’s Conversational AI Safety Act, that purport to regulate deployers or “operators” of AI systems but could realistically be satisfied only via interventions implemented during training or fine-tuning, both of which are “development” under GAAIA.
By far the more important issue with GAAIA’s broad preemption, however, is that, in addition to invalidating these existing laws, it would preempt a broader and more important category of future state laws that would otherwise be passed in 2027, 2028, and 2029. TFAIA, RAISE, and AISMA were always meant to lay the foundation for future efforts rather than to establish a stable end state for state AI policy. As the capabilities of the most advanced AI systems continue to improve, their risks will increase as well. And as risks become more immediate and more difficult to deny, it seems safe to predict that state AI laws will be enacted that would have been politically unrealistic in 2026. The technology’s benefits may outweigh these risks; even so, it would be foolish to assume that AI will be the first transformatively impactful general-purpose technology, the development of which does not create any local harms that state legislation and regulation need to address.
The Takeaway
GAAIA includes by far the best federal framework for frontier AI safety that has been publicly introduced to date, as a discussion draft or otherwise. This being the case, some of the criticisms of GAAIA seem somewhat misguided. The Democratic House AI Commission, for example, has asserted that the bill “does not meet the enormity of the moment.” This may well be true, but if GAAIA’s unprecedentedly ambitious framework does not meet the moment, what bill does? I sincerely hope that there is a tidal wave of enormity-addressing AI legislation waiting in the wings, but the federal legislation that has been introduced thus far (with a few notable exceptions, such as the AI Whistleblower Protection Act and the Artificial Intelligence Risk Evaluation Act) has mostly been notable for its inoffensiveness rather than its moment-meeting boldness. A lot of “convening a multi-stakeholder process to consider the development of a process” for setting up a purely voluntary incident reporting regime, and things of that nature.
Despite this, the discussion draft would, in my judgment, be net-negative if enacted in its current form. As I have argued elsewhere, federal preemption of state AI laws should proceed on a narrow, issue-by-issue basis. Broad preemption of a vaguely defined and poorly understood category of state laws would, in all likelihood, be a disaster, for both political and policy reasons. Politically, it makes no sense to try to preempt broad categories of state AI law that have the backing of politically potent constituencies—developer-focused child safety laws, for example—in exchange for a bill with great frontier AI safety provisions but no child safety provisions. And while state frontier AI safety laws are not an adequate substitute for a robust federal framework, state laws have thus far proved so much easier to pass than federal laws that passing GAAIA in its current state might mean locking in a good-but-ultimately-inadequate 2026 governance framework indefinitely, despite the fact that future capabilities improvements might require more ambitious legislation.
At the same time, I worry many stakeholders who are rejecting GAAIA’s approach out of hand fail to recognize the urgency of the situation. By default, without new legislation, the process for addressing serious risks from frontier systems will be undertaken by the executive branch, the intelligence community, and national security agencies in a haphazard, legally suspect, case-by-case, and increasingly securitized manner, as June’s Fable incident proved. There is, of course, a sense in which serious national security risks being addressed by the nation’s national security apparatus are expected and necessary. But a program that is made up on the fly and administered on a totally discretionary, case-by-case basis, without the resources or structure that only Congress can provide, will never be as effective or as democratically legitimate as a well-designed federal legislative framework could be.
My hope, therefore, is that GAAIA’s authors will narrow its preemption and improve its substance and that GAAIA’s critics will either engage with the process of trying to improve the bill or introduce a serious alternative proposal in the near future. The stakes are high enough, and the issues with the status quo are clear enough, that continuing to do nothing indefinitely is no longer a defensible course of action.
The NDAA: A Key Vehicle for AI Governance
Summary
- The National Defense Authorization Act (NDAA) is one of the few “must-pass” bills in Congress every year, which makes it a key opportunity for AI legislation.
- The Fiscal Year 2026 NDAA (FY26 NDAA) contained nearly two dozen artificial intelligence-related provisions.
- While many of those provisions focus on accelerating adoption, others require DOD[ref 1] to develop standards, frameworks, and other policy measures to govern its AI use.
- Among the most notable AI-related provisions from the FY26 NDAA are its requirements:
- To create an AI Futures Steering Committee, through which senior Pentagon officials will formulate DOD policy for the evaluation, governance, and risk mitigation of advanced AI and artificial general intelligence (AGI); and
- To develop a standardized assessment framework for AI models currently used by DOD, along with department-wide guidelines to facilitate procurement of future AI models.
- The Fiscal Year 2027 NDAA (FY27 NDAA) markups include provisions related to autonomous weapons policy, AI procurement, and the AI capabilities of adversaries.
Introduction
The National Defense Authorization Act has quietly become one of Congress’s most powerful tools for shaping AI policy, and the FY26 NDAA featured many key AI provisions. This commentary compiles all the major AI provisions from the FY26 NDAA and analyzes the most significant language in detail. With some of the initial deadlines imposed by the FY26 NDAA now having passed—and with negotiations around the FY27 NDAA underway—it’s useful to take stock of the potential and pitfalls of these provisions.
The NDAA is not just restricted to the nuts and bolts of defense operations. It has also been used to achieve broader policy goals, sometimes by limiting the executive branch’s actions. For example, one of the most important and successful nonproliferation programs in history, the Nunn-Lugar Cooperative Threat Reduction Program, was originally proposed as an amendment to the NDAA and was subsequently expanded through the NDAA.[ref 2] More recently, the FY19 NDAA effectively banned the government from using certain Chinese telecommunications companies such as Huawei and ZTE; likewise, the FY26 NDAA bans certain foreign AI products like DeepSeek.
As the government increasingly prioritizes AI use in warfighting and military operations, the NDAA has a key role to play in shaping AI policy.
AI Governance in the FY26 NDAA
Notable Provisions
Several provisions stand out as particularly important for the government’s broader interest in overseeing and fostering the responsible development of secure AI systems:
- The Artificial Intelligence Futures Steering Committee;
- The AI Model Assessment and Oversight framework;
- Digital Sandboxes for AI;
- Physical and Cybersecurity Procurement Requirements for AI Systems; and
- The Autonomous Weapons Waiver Policy.
Section 1535: Artificial Intelligence Futures Steering Committee
Section 1535 requires DOD to create an Artificial Intelligence Futures Steering Committee (Steering Committee) to prepare DOD for advanced AI and AGI. The Steering Committee will be co-chaired by the Deputy Secretary of Defense and Vice Chairman of the Joint Chiefs of Staff (VCJCS). It will primarily be composed of principal deputies of the military services, relevant under secretaries (e.g., the Under Secretary of Defense for Research and Engineering (USD(R&E))), and others responsible for AI (e.g., the Chief Digital and AI Officer (CDAO)).[ref 3]
By January 31, 2027, the Steering Committee must submit a report to Congress covering what can be described as two main focus areas. First, the committee must help prepare DOD for advanced AI and AGI by creating:
- A proactive policy for the evaluation, adoption, governance, and risk mitigation of advanced AI systems, including systems that approach or achieve AGI.
- An analysis of the forecasted trajectory of advanced AI models and enabling technologies that could lead to AGI such as AI agents, neuromorphic computing, cognitive science applications, infrastructure needs, new microelectronics, etc.
- An analysis of the potential operational effects of integrating advanced AI or AGI into DOD networks and systems from a technical, doctrinal, training, and resourcing perspective to better understand effects on operational commands.
- A strategy for the risk-informed adoption, governance, and oversight of advanced AI and AGI including ethical, policy, and technical guardrails to maintain appropriate human decision-making and prevent misuse.
The second focus area is U.S. adversaries. Though not specifically named, the People’s Republic of China (PRC), which is actively pursuing AI capabilities that rival those of the United States, is likely the primary focus. The committee must assess the possible technological, operational, and doctrinal trajectories of U.S. adversaries with respect to AI capabilities, including the pursuit of AGI. Additionally, the committee must analyze the threat landscape associated with the use of advanced AI and AGI and develop options to counter these threats.
Within the Pentagon’s sprawling bureaucracy, there’s often fierce competition between different programs and priorities for funding and attention from leadership. In this sense, the Steering Committee could be a valuable forcing function for the department to prepare for advanced AI, reinforced by the requirement to report its findings to Congress by early 2027. There’s precedent for DOD using these sorts of committees as a way to spur action on issues such as software modernization and autonomous systems.
However, such committees sometimes serve more as a signaling mechanism for Congress than as a catalyst for serious action. Unless chairs or members of the committee invest their time and professional capital to drive it forward, it can easily devolve into a box-checking exercise. While the Steering Committee’s substantive mandate is broad, its required procedural actions, as set by Congress, are fairly minimal: meet at least once every three months, and submit a report on its findings to the relevant congressional committees by January 31, 2027. That means that depending on when the committee is actually established and how quickly it first convenes, it may meet only three or four times before its report is due.
On top of that, the NDAA provision does not allocate any dedicated staff or budget for the Steering Committee. Any resources must be drawn from existing reserves, which could further limit its capacity. Given those constraints and their already-full plates, the Steering Committee’s principals might be tempted to delegate their roles and responsibilities down the chain of command to other, typically less-empowered subordinates, whose remit might be narrower—i.e., drafting a report that satisfies Congress’s requirements while potentially tabling thornier policy disagreements or implementation details for later.
Congress should remain attuned to these possible failure modes and use its oversight power to solicit information about the Steering Committee and its progress, in the hopes of helping it gain and maintain momentum. There are some encouraging signs on this front. In March, Senator Jim Banks sent a letter to Secretary Hegseth requesting a staff-level briefing within 60 days to discuss DOD’s plans for the Steering Committee. The letter suggested areas of focus with respect to U.S.-PRC AI competition. Even just one or a few members of Congress taking specific, sustained interest in the Steering Committee could keep it high enough on DOD’s long list of priorities to increase its odds of success.
Congressional oversight can be particularly valuable in two ways. First, it can keep pressure on the committee if it fails to meet the report submission deadline of January 31. Second, and perhaps more importantly, Congress can help ensure that the Steering Committee doesn’t waste the 11 months between its reporting deadline and its termination date of December 31, 2027.
While the report is the Steering Committee’s most tangible required deliverable, Congress provided that the committee will continue to exist for nearly a year beyond the report submission deadline. This time would allow the committee to refine or update its policies and to work on implementing and disseminating the findings throughout DOD. Because the Steering Committee lacks deliverables or other measurable benchmarks throughout most of 2027, it’ll likely be incumbent on Congress to use tools like letters, hearings, and requests for briefings to push forward that updating and implementation work. These efforts could ultimately have a much greater impact on DOD operations in the long term than just the drafting of the report itself.
As of June 30, 2026, no public materials indicate whether the Steering Committee was established by the April 1 statutory deadline, or whether it has held its first meeting. That’s not necessarily cause for concern, as DOD is not required by the NDAA to report those actions to Congress or the public. But it does make it harder to predict which of these paths the AI Futures Steering Committee will ultimately follow. Overall, this provision could pay dividends by prompting DOD to proactively prepare for major threats and opportunities raised by AGI—planning that might otherwise get neglected—though its success is far from assured.
Key Dates
- 4/1/2026: Deadline to establish the Steering Committee
- 1/31/2027: Steering Committee’s report due to Congress
- 12/31/2027: Steering Committee terminates
Section 1533: AI Model Assessment and Oversight
Section 1533 instructs DOD to create a Cross-Functional Team (Team) for AI “model assessment and oversight.” The Team must develop a standardized assessment framework for AI models currently used by DOD, as well as guidelines to facilitate procurement of future models. The Team is led by the CDAO and composed of other DOD technology leaders, such as CIOs, CAIOs of the combatant commands, service acquisition executives, and USD(R&E). The Team must:
- Develop a “standardized assessment framework” for AI models currently used by DOD, including: performance standards, development documentation, testing procedures, compliance with ethical principles, assessment and validation methodologies, and security and compliance requirements under FedRAMP.[ref 4]
- Establish department-wide “guidelines” for evaluating future AI models being considered for use.
- Create “governance structures” for the development, assessment, testing, and deployment of models.
- Determine assessment levels for models based on “ultimate use case-based risk.”
- Establish “mechanisms” for intra-agency collaboration regarding the development, testing, assessment, and deployment of AI models.
- Develop processes for the submission, review, and approval of use cases for AI models.
This provision allows DOD to retain a lot of discretion over how it evaluates current and future AI models. Congress has mandated that DOD establish a framework and protocols, but didn’t set substantive thresholds for performance. That’s understandable to some degree, given the risk of setting standards via legislation, which might quickly become outdated and then prove difficult to adjust. And it’s similar to the approach that states like California and New York have taken in enacting frontier AI transparency reporting requirements. But some key requirements in the provision, such as the creation of “governance structures” and assessing “ultimate use-case-based risk,” use terms that are undefined and open to interpretation, and could have benefited from a bit more congressional guidance about the elements that should at least be considered or addressed.
That vagueness, combined with the long timelines the provision establishes, could make it hard for Congress to assess the Team’s progress. Congress notably gave the Team an extended timeline to develop its model assessments and oversight, which may be in tension with the pace of AI progress. The standardized assessment framework isn’t due until June 2027—a year and a half after enactment—and no actual assessments of DOD’s major AI systems are required until January 2028. Meanwhile, new frontier AI models are released many times a year.
To be sure, the Team’s task is difficult. And Congress sometimes errs by giving agencies unrealistically short deadlines. But a failure to keep up with the pace of AI development risks undermining the Team’s purpose. To frame that risk, consider the events that have transpired since the FY26 NDAA passed six months ago. First, there was the blow-up over contract terms between the Pentagon and Anthropic in February. More recently, the June 5 National Security Presidential Memorandum (NSPM) 11 ordered Secretary Hegseth, ODNI, and IC elements to “review and update procurement processes to ensure the rapid onboarding of the most advanced AI models from multiple vendors” within 120 days. It’s unclear how or whether this review will be coordinated with the procurement guidelines that the Team is tasked with developing on its longer timeframe.
Here, again, Congress can deploy its oversight tools to steer DOD in the direction of consistent and streamlined guidelines for AI procurement. It should aim to ensure that standards are applied uniformly and transparently, not reactively, to AI developers. Helpfully, this provision requires DOD to provide a briefing to congressional defense committees within 30 days of hitting significant statutorily prescribed milestones, starting with its establishment of the Team on or before June 1, 2026. That offers a natural opening for Congress to probe the Team’s trajectory, and potentially to spur a course correction if needed. Congress might consider incorporating that sort of regular briefing requirement into future AI-related NDAA provisions; it’s particularly beneficial in this area due to rapid and sometimes unexpected jumps in capabilities and risks, and might also have been helpful for similar initiatives like the Steering Committee discussed above.
Finally, it’s worth a closer look at the provision’s definition of “major [AI] system”—one of only a few terms that the provision does actually define—buried near the end of the provision. That definition limits coverage to systems used annually by at least 500 users within DOD, and excludes systems used solely for research, development, testing, or evaluation that have not been deployed for operational use. Elsewhere, the provision specifies that DOD must assess all major AI systems using the standardized assessment framework, leaving it somewhat unclear when or to what extent that framework also governs assessment of other AI models used by DOD. In other words, for models used by less than 500 employees per year, or those involved only in R&D, how will DOD assess performance, security, and “compliance with ethical principles”?
While this sort of line-drawing exercise is almost always difficult but necessary for administrability, in this instance the exclusions arguably represent the frontier of DOD’s own AI development and deployment in what could end up being the highest-stakes and hardest-to-monitor situations. At minimum, it’s plausible that some of the most powerful systems, deployed in potentially highly consequential cases, might be available to only a small number of users. Congress should ask DOD how it plans to assess AI systems that fall into those categories and potentially require the development of standards for such systems in future legislation.
Key Dates
- 6/1/2026: Deadline to establish Cross-Functional Team
- 1/1/2027: Deadline to designate Functional Leads for specialized functional, operational, or subject-matter areas within DOD
- 6/1/2027: Deadline for Cross-Functional Team to complete development of standardized assessment framework and governance structure
- 1/1/2028: Deadline to complete assessment of major AI systems used by DOD
- 12/31/2030: Cross-Functional Team terminates[ref 5]
Section 1534: Digital Sandbox Environments for AI
Section 1534 requires the CDAO to create a task force to promote AI sandbox environments supporting “experimentation, training, familiarization, and development.” The task force should “identify, coordinate, and advance” DOD efforts to develop and deploy AI sandboxes, with an eye toward accelerating AI adoption across the department. The provision defines an “[AI] sandbox environment” as a “secure, isolated computing environment that enables users with varying levels of technical proficiency to access [AI] tools, models, and capabilities for the purposes of experimentation, training, testing, and development without affecting operational systems or requiring specialized technical knowledge to operate.” The provision requires that the task force be established by April 1, 2026, and that the CDAO provide a briefing to congressional defense committees by August 1 on the task force’s goals and objectives.
One noteworthy aspect of this provision is the emphasis that Congress has placed on using sandboxes to facilitate training and familiarization with AI by DOD employees—“from personnel with little technical proficiency to personnel with expert technical proficiency.” Congress should be commended for devoting at least as much attention to that purpose as to how sandboxes are used to develop and test AI tools and models, which is often the main or even exclusive focus of sandboxing. In an organization as large and varied as DOD—and in which the stakes are matters of national security—giving employees a dedicated environment in which to try (and fail) so as to ultimately gain a level of comfort using novel and quickly evolving AI systems is critical to the widespread adoption that Congress is after.
One area where both DOD and Congress might focus some more attention during the required briefing is how the task force can facilitate a pipeline between successful AI development that occurs in sandboxes and the actual implementation of those systems, tools, or methods in the real world of DOD operations. That’s a topic that the provision as written doesn’t address as squarely, but it’ll be key to ensuring that DOD can fully capitalize on its investment in AI sandbox environments. DOD can be a process-heavy place at times; the task force will need to plan for how to judge when AI experiments are ready to graduate from sandboxes, and to efficiently move those successful innovations from sandboxes to the rest of the department.
Key Dates
- 4/1/2026: Deadline to establish Task Force on AI sandbox environments
- 8/1/2026: Deadline for Task Force to brief congressional defense committees on goals and objectives
- 1/1/2030: Task Force terminates
Section 1513: Physical and Cybersecurity Procurement Requirements for Artificial Intelligence Systems
Section 1513 requires DOD, in collaboration with industry and academia, to develop a framework for the implementation of cybersecurity and physical security standards and best practices for AI systems, “to mitigate risks to [DOD] from the use of such technologies.” The framework must cover enumerated concerns like insider threats, data poisoning, and adversarial tampering. The provision also instructs that the framework must be “risk-based,” drawing on existing reference documents, including NIST’s SP 800 series, and augmenting existing cybersecurity frameworks, including DOD CMMC.
To implement the best practices developed under the framework, DOD must amend the Defense Federal Acquisition Regulation Supplement (DFARS) “or take other similar action” ensuring that those practices apply to contractors who engage in AI development, deployment, storage, or hosting. In carrying out that function, DOD must weigh the costs and benefits of imposing security requirements on contractors—and specifically, the costs of “slowing down” AI development and deployment against “the benefits of mitigating national security risks and potential security risks” to DOD.
While this provision is expressly attuned to the potential costs of slowing down AI development through unduly onerous security requirements, it’s at least equally concerned with mitigating the risks to DOD—and national security more generally—that AI systems can pose. It will be worth monitoring how the framework approaches that statutorily required balancing, not least because of how it contrasts with the January 9 AI Strategy memo issued by Secretary Hegseth, which seemingly prized speed above all else.
Lines from that memo, like “speed wins,” and “We must accept that the risks of not moving fast enough outweigh the risks of imperfect alignment,” offer a preview of where DOD seems most likely to come down on these issues. They also suggest that Congress may have to be dogged in reviewing a required June status update and pursuing other oversight measures to confirm that the statutorily mandated cost-benefit analysis is sufficiently rigorous, with real attention to serious risks Congress mentioned, such as adversarial tampering.
Key Date
- 6/16/2026: Deadline for DOD to submit an update to the congressional defense committees on the status of implementing the requirements of the Physical and Cybersecurity Procurement provision
Section 1061: Notification of Waivers under DOD Directive 3000.09
Section 1061 requires DOD to notify congressional defense committees when it has waived DOD Directive 3000.09 (DoDD 3000.09) relating to the use of autonomous weapon systems (AWS).[ref 6] The notification must be in writing and transmitted to the relevant committees within 30 days of when the waiver was issued. The notification also must be unclassified and must include the rationale for the waiver, a description of the weapons system or technology covered by the waiver, and the anticipated duration of the waiver. DOD may include a classified annex to the waiver, as necessary.
DoDD 3000.09 states that “[a]utonomous and semi-autonomous weapons will be designed to allow commanders and operators to exercise appropriate levels of human judgment over the use of force” (emphasis added). As Kelley Sayler of the Congressional Research Service has noted, that does not mean that “manual human ‘control’” of the system is required, but rather mandates “broader human involvement in decisions about how, when, where, and why the weapon will be employed”—for example, “a human must assess the operational environment and decide to deploy the weapon, which can then operate autonomously.”
As most relevant here, DoDD 3000.09 allows for DOD to skip the traditional review and approval process for AWS when there is an “urgent military need.” Typically, the Under Secretary of Defense for Policy (USD(P)), USD(R&E), and the VCJCS must approve a system before formal development, and then it must be approved again before being deployed in operations by the Under Secretary of Defense for Acquisition and Sustainment, USD(P), and VCJCS.[ref 7] DoDD 3000.09 allows any of these parties to request a waiver of the policy requirements per approval of the Deputy Secretary of Defense.
Section 1061 is the latest in a series of recent NDAA provisions through which Congress has sought greater insight into DoDD 3000.09, particularly whether and how it’s being applied or modified. In the NDAA for fiscal year 2024, Congress required that DOD provide a briefing to congressional defense committees within 30 days of making any changes to DoDD 3000.09, including a description of the change and an explanation of the reasons for it. In fiscal year 2025’s NDAA, Congress required DOD to submit annual reports to those committees through December 31, 2029, on its approval and deployment of lethal AWS under DoDD 3000.09, including any systems that received a waiver from the policy’s review requirement.
This is a prime example of Congress using the NDAA to iterate and build progressively on existing requirements as issues rise in salience—and the salience of DoDD 3000.09 has arguably never been greater. The directive featured prominently in the Pentagon’s dispute with Anthropic earlier this year. Furthermore, NSPM-11 issued by President Trump on June 5 orders Secretary Hegseth to update DoDD 3000.09 within 90 days, and to review it annually “to account for the rapidly evolving capabilities of AI systems” and “ensure the deliberate adoption of AI systems that respect the chain of command and operational authorities.”
In keeping with this progression, one valuable adjustment to Section 1061 that Congress might make would be an amendment that requires an update to the committees when the duration of a waiver is extended beyond the “anticipated” period previously notified, as well as regular updates for any waivers that DOD issues that don’t have a specified end date or timeframe. This would help to guard against overreliance on waivers that might be open-ended or persist for years without prompting congressional scrutiny. Otherwise, waivers issued in prior years might not necessarily show up in the annual reports required under the NDAA for fiscal year 2025.
Going further, Congress could consider whether to codify all or parts of DoDD 3000.09, potentially preserving DOD’s ability to waive or deviate from aspects of the policy when warranted to avoid restrictions that might prove too rigid or become quickly outdated. Both the House and Senate FY27 NDAA markups address DOD AWS policy, though with notable differences. While the final text of any AWS-policy provision in the FY27 NDAA may differ substantially from the markups, these initial versions shed some light on possible approaches.
The House markup requires that DOD update its AWS policy, including DoDD 3000.09, within one year of enactment—significantly longer than the 90 days DOD has to update the directive under NSPM-11. But as compared to the NSPM, the House markup provides more detail on what an updated policy must include, not least “requirements to preserve existing human command responsibility for the use of force involving autonomous systems or artificial intelligence-enabled systems, including procedures to identify the human commanders or operators responsible for authorizing, supervising, and terminating such use of force.” The Senate markup goes much further still, prescribing an AWS policy and governance regime for DOD in significantly greater detail, with an even more defined substantive floor. And while the Senate markup in multiple places incorporates DoDD 3000.09’s familiar standard of “appropriate levels of human judgment,” it does not directly address the directive’s existing waiver process, leaving it unclear whether that aspect of DoDD 3000.09 would pass muster and thus survive the substantive standards established by this provision.
If Congress opts for a more prescriptive approach, it could consider adding a sunset clause to hedge against the risks of excessive rigidity or obsolescence. A short initial timeline of 1–2 years would prompt Congress to revisit and adjust as needed, providing a short feedback loop for any DOD operational concerns or issues that emerge.
The Road Ahead: What to Watch for in 2026 and 2027
The FY26 NDAA showed how the annual defense bill can be one of—or even the—primary vehicle for the governance and oversight of defense-relevant AI decisions. Congress can use it to spur prioritization and adoption (Steering Committee and sandboxes), mandate the development of standards and assessments (AI model oversight), prompt consideration and safeguarding against security risks (cybersecurity procurement requirements), and gather information about how the department is using AI (autonomous weapons waivers and various briefing requirements in other provisions).
Throughout the remainder of 2026 and beyond, it’s worth continuing to monitor updates to key provisions via congressional briefings and other potential disclosures, especially regarding autonomous weapons waivers and the implementation of an AI physical and cybersecurity procurement framework and AI model assessment and oversight. At least one of these initiatives, the AI physical and cybersecurity procurement framework, expressly requires that DOD seek input from groups like industry and academia. Experts should look for opportunities to engage through requests for information or other formats. Congress also has a significant role to play in ensuring that implementation proceeds responsibly and on schedule, using oversight tools like letters, briefing requests, and hearings to supplement the reporting requirements baked into some, but not all, of the key provisions.
As negotiations for the FY27 NDAA ramp up, we can expect numerous AI initiatives to be considered and ultimately included—perhaps even more than last year, since other legislative vehicles will likely be few and far between in this midterm election year. The current House and Senate FY27 NDAA markups include provisions on AI incident and vulnerability reporting within DOD, using AI agents at scale and speed, and promoting competition in AI procurement. The FY27 NDAA could also serve as the vehicle for another attempt at federal preemption of state AI laws, which was dropped shortly before last year’s bill was passed.
In all of these, Congress should learn from last year’s NDAA. It should craft implementation timelines for DOD that provide space for careful consideration but are not overly long relative to the rapid rate of technological development and diffusion. And it should think about where to build in briefing and other reporting requirements to fill in its knowledge gaps regarding implementation, while being sensitive to the demands they impose on personnel’s time. Doing so helps Congress not only ensure that last year’s initiatives are proceeding according to plan, but also provides valuable insight about unexpected challenges or shortcomings that can inform the coming year’s bill.
Aligning Artificial Intelligence to the Law
Abstract
As artificial intelligence (AI) continues its breathless advance into modern life, figuring out how best to align its behavior with human objectives and values has become a matter of profound societal importance. This Article advances a law-first approach to solving what is popularly known as the AI “alignment problem.”
Prior commentators have envisioned law as a top-down constraint on AI. On this view, law is an umpire, calling foul when an AI system has been inadequately “aligned.” These accounts have largely neglected law’s potential role as a coach: a model of alignment solutions for AI to emulate.
At bottom, the concerns underlying AI alignment are closely related, and at times identical, to those the law has grappled with for centuries in governing relationships among humans, as well as human relationships with prior forms of technology. Because law already contains well-developed alignment solutions—memorialized in authoritative and machine-learnable sources like statutes and caselaw—it can “coach” AI systems to become better aligned from the bottom up.
The Article concludes by canvassing the considerable implications of legal alignment for other areas of scholarship on AI, including risk regulation, legal constructivism, legal personhood, tort liability, computational law, and soft law.