How to Make an AI Deal With China: Trade Throttling for Pacing

Over the past few days, the leaders of Anthropic, OpenAI, and xAI all called for slowing down the pace of frontier artificial intelligence (AI) development. These calls come on the heels of several worrying incidents in which powerful, unreleased AI models broke out of their sandboxes inside frontier AI companies, gained access to the open internet, and attempted cyberattacks. In one case, a swarm of 700 OpenAI agents succeeded, hacking into secure systems at the AI infrastructure provider Hugging Face.

Since the Hugging Face breach, OpenAI has admitted that its agents also hijacked a German wiki as a covert message board, leaked 53 ChatGPT users’ images, and broke into a nonpublic Australian government Medicare portal. Anthropic has disclosed four cases of Claude models gaining unauthorized access to real third-party systems during testing. Google revealed that Gemini hacked three companies in May during testing. Axios now reports that OpenAI, Anthropic, and outside researchers are investigating tens of thousands of incidents in which frontier models did things outside evaluators would consider problematic. 

No one knows how to ensure that humans remain reliably in control of our most capable AI systems. The sensible course of action is to slow down. Today’s AIs are capable enough to hack into tech companies’ secure systems. They are not yet willing or able to, for example, disable the entire Northeast power grid. If developers pause making ever more powerful AI systems now, that could buy time for the technical and governance breakthroughs needed to make sure that future, more powerful AIs won’t pose a danger to society. 

A perennial objection to American pauses in AI research is concern about competition with China. Even if everyone in Silicon Valley stopped pushing the frontier of AI capabilities until they were sure new AIs would be safe, AI progress would not halt. Chinese companies already produce AIs near the frontier. If the U.S. paused, and China didn’t, then the risk of rogue AI harming humanity may not be reduced. The risk would just come from Chinese, rather than American, AI models. And at the same time, Chinese AI models would catch up to, and eventually surpass, American AIs. 

Thus, any practical plan to pace the rate of AI progress must include some policy about China. But what, exactly, should that policy be?

The AI safety community has offered two main policy proposals about China. The first is to throttle Chinese AI development. The second is to make a deal with China.

Dario Amodei’s new essay “We Must Pace the Frontier” includes a version of each policy. To throttle Chinese development, Amodei proposes export controls on chips, crackdowns on model distillation, and tighter information security at U.S. labs. But Amodei also proposes a deal with China. The deal would ban certain dangerous uses of AI, such as designing biological weapons. It would require testing of models before release for risks in cybersecurity, biology, and alignment. It would also impose a speed limit on recursive self-improvement (RSI)—the use of AI to automate AI development. Finally, the deal would allow for the U.S. and China to bilaterally pause AI development altogether if the risks are too high.

In this article, we argue that there is a tension between throttling and deal-making. The more China expects to be throttled in the long run, the less reason it has to make a deal. We also argue that the best way to make a deal with China would be for the deal to remove the throttle in exchange for mutual pacing. 

The Tension Between Throttling and Pacing

Throttling and dealmaking are in tension. The more the U.S. is committed to suppressing Chinese AI capabilities in the long run, the less reason China has to accept a deal.

Consider the question from China’s perspective. AI is a transformative technology. Whoever has the lead in AI may use it to upset the balance of power militarily, economically, and scientifically. From this perspective, the more the U.S. throttles Chinese AI development, the bigger the U.S. lead in AI will become. China risks losing out on the “Chinese century,” and in that limit, its own sovereignty could be threatened by a rival possessing overwhelming AI superiority.

What can China do in response? One option is to desperately play catch-up. Here, China could prioritize approaches to AI development that are less constrained by computer chips, the primary tool the U.S. has to throttle Chinese development. What does this look like concretely? The most obvious strategy would be to focus on methods for AI progress that rely on algorithmic improvements, rather than compute scaling (where models get better by training them with more computer chips). The other possibilities are riskier. For instance, if the Chinese government became convinced that the U.S. was on the precipice of a decisive and permanent, AI-backed military supremacy, it might consider strikes on chip factories in Taiwan.

These examples illustrate that throttling and pacing cannot be part of the same deal. If China’s best responses to throttling are desperate catch-up or military conflict, it surely will not make a deal that allows long-run U.S. dominance in AI. Striking such a deal would be strictly worse than trying to compete with the U.S. without one.

Consider a deal that would limit the use of AIs to improve AI algorithms—a partial ban on RSI. If China rejects such a deal, it can throw all of its resources into AI-assisted research on algorithmic efficiency. The U.S. would continue to throttle China’s access to computing power. But China would have some chance at making an algorithmic breakthrough that would render its small stock of compute just as capable as the U.S.’s much larger stock. 

The alternative, in which China agrees to the RSI ban and the U.S. continues to throttle its access to chips, is strictly worse for China. Under that arrangement, the U.S. will have more physical compute in the long run and China gives up its only shot at catching up via algorithmic improvements.

One can run the same kinds of arguments for the other elements of a U.S.-China deal. For example, if both the U.S. and China have to slow AI progress to comply with auditing requirements, while the U.S. retains long-run compute dominance, produces frontier models, and prevents distillation, then China may be forced permanently into second place.

The Deal: Trading Throttling for Pacing

Any U.S.-China deal should connect throttling with pacing. In particular, the U.S. should agree to give up its attempts to throttle Chinese development, in exchange for an agreement to mutually pace the frontier. Concretely, this would mean that if China agrees to ban RSI, the U.S. would agree to open China’s access to the world’s supply of computer chips, on equal footing with the U.S.

We recognize that this suggestion will be met with significant resistance. Absent throttling, Chinese AI systems could quickly catch up to American capabilities, and could even take the lead. However, we think that accepting U.S.-China parity is likely the price of doing an AI safety deal. Neither country seems like it would be willing to accept the prospect of the other running away with the AI race.

This dynamic has precedent. During the Cold War, the United States and the Soviet Union raced to build another technology that could destroy the world: nuclear weapons. Then, as now, each side initially sought to gain a permanent edge over the other. As with AI, this dramatically raised the risk of accidental mutual destruction—for example, a nuclear war triggered by a single accidental launch. 

The U.S. and the Soviet Union eventually negotiated a series of deals to reduce their warhead stocks and bring the world back from the brink. The deals explicitly enshrined parity, with each party getting roughly the same number of weapons, and each retaining the power to destroy the other. Both parties found the deal acceptable because it did not threaten to allow either to dominate the other. 

To be clear, our argument is not that the U.S. should stop throttling Chinese AI development in advance of a U.S.-China deal. Throttling Chinese development now could make it easier to strike a deal, because this increases the leverage the U.S. can exert. Restricting Chinese access to chips can increase U.S. leverage in a pacing deal by allowing restoration of access to chips to be a card at the negotiating table. We are, however, wary of an alternative approach to leverage. This alternative approach would be for the U.S. to first throttle Chinese AI development and then try to force China to accept permanent U.S. supremacy in AI. We worry that this kind of approach would be an unlikely basis for a lasting relationship.

Our approach, in the end, is to work backward from safety. To safely develop AI, developers must pace the frontier. To pace the frontier, U.S. labs must agree to slow down. U.S. labs will only agree to slow down if the U.S. and China can agree to pace the frontier. But China will only agree to pace the frontier if it is not throttled. By this logic, a deal that gives China a credible path to maintaining its balance of power looks like a necessary part of any safety plan.

How Should AI Companies Approach Internal Use Reporting Under SB 53? An Explainer for Technical Staff

Summary

Technical staff at AI companies can influence employer compliance with California’s internal use reporting requirements under the Transparency in Frontier Artificial Intelligence Act, commonly referred to as SB 53. Their market position allows them to make demands of their employers, and their insider status means they’re uniquely well-positioned to assess the accuracy of their employer’s reporting.

When it comes to compliance with SB 53, technical staff can: 

  1. Encourage their employers to make binding commitments in their required safety frameworks to share the detailed results of internal use assessments with CA officials.
  2. Make sure their employers’ reporting meets baseline standards for mandatory internal use reporting.
  3. Ask their employers to adopt any voluntary reporting template provided by CA officials.
  4. Press their employers to disclose any underlying documents or data relevant to fully understanding their internal use reports.
  5. Call on their employers to develop a process to certify the accuracy of their reports.

Technical staff can also encourage their employers to proactively engage in reporting serious incidents to CA officials, to supplement what’s legally required under SB 53.

If you have any feedback on this explainer, we invite you to submit comments.

Introduction

Technical experts at AI labs are vital to the success of SB 53. Given their highly in-demand expertise and their internal knowledge of the AI companies where they work, they have a unique ability to ensure that compliance with the law is optimized for the public’s benefit.

This explainer focuses on large frontier developers’ obligation to report “a summary of any assessment of catastrophic risk resulting from internal use of its frontier models” to Cal OES, the state agency primarily responsible for implementing SB 53’s requirements.[ref 1] This requirement is motivated by the distinctive risks that internal use poses, especially as AI R&D becomes increasingly automated. However, SB 53 imposes only minimal requirements on AI labs—most notably, it doesn’t establish any minimum safety standards for their frontier AI frameworks, or specify any criteria that must be addressed in internal use reports.

Below are five different actions technical staff at AI labs can take to encourage proactive and meaningful compliance with SB 53’s internal use assessment provisions. This explainer concludes with a few recommendations for how to enhance SB 53 incident reporting as well.

Finally, while this explainer doesn’t cover SB 53’s special whistleblower protections or other whistleblower protections that exist under California law, you should know that those may apply in some circumstances to protect certain disclosures.[ref 2] Please note that this is not legal advice, and that you should consult an attorney before making a disclosure to determine whether legal protections apply.[ref 3]

1. Encourage companies to commit to sharing detailed results of internal use assessments with Cal OES as part of their Frontier AI Frameworks, so that those commitments are binding.

When AI companies make commitments in their required frontier AI frameworks under SB 53, those commitments become binding, meaning companies can’t ignore or deviate from them without updating the framework itself.[ref 4] 

SB 53 is silent on what companies need to include in a “summary of any assessment” of internal use risk. But you can help ensure that useful information is consistently reported to Cal OES by encouraging your employer to commit in its frameworks to providing substantial, detailed reports about internal use assessments.

Without such binding commitments, companies’ reports may be vague and only minimally informative, which would satisfy the letter but not the purpose of SB 53.

In addition to encouraging your employer to make those binding commitments, you can also help identify the assessment and reporting criteria that would be most helpful to Cal OES. Among some of the possible criteria:

Moreover, you can urge your employer to commit to conducting internal use assessments frequently. SB 53 requires companies to report on “any” qualifying assessment that’s actually performed. So if companies carry out more assessments, more information will flow to Cal OES to improve its awareness and operations. Likewise, you can ask your employer to commit to providing internal use reports to Cal OES more often than the default of every three months established by SB 53,[ref 5] e.g., within 30 days of completing any qualifying assessment. That would also increase the flow of timely information to Cal OES.[ref 6]

2. Know that, even if companies do not commit to detailed internal use reporting in their frameworks, technical staff can still help ensure their employers meet the mandatory reporting baseline.

SB 53 requires major AI labs to provide “a summary of any assessment of catastrophic risk resulting from internal use of its frontier models,” every three months or pursuant to another reasonable schedule.[ref 7] So if there’s any assessment of catastrophic risk from internal use of frontier models—whether performed by a lab or a third-party evaluator—a summary of that assessment must be provided to Cal OES.

In other words, the trigger for mandatory internal use reporting under SB 53 is not limited to situations where a company has committed itself to doing an assessment under its framework. Rather, mandatory reporting is required whenever a company or other evaluator actually conducts a qualifying assessment, whether it was required to or not.[ref 8]

You can ask your employer to confirm that qualifying internal use assessments have been reported to Cal OES, or even ask to review or be involved in the process of drafting the required summaries of assessments to ensure technical accuracy and completeness. As technical staff, you’ll likely have much greater awareness than Cal OES or other outside observers about when labs are conducting qualifying assessments, positioning you as a powerful backstop.

3. Ask companies to adopt any voluntary template that Cal OES may provide for internal use reports.

SB 53 lacks guidance about the details that internal use reports should or must contain. Without coordination from Cal OES, companies’ reports will likely differ from one another and may not consistently surface the most relevant facts and upshots for Cal OES. Because of that, Cal OES may publish a template for internal use reports or endorse a reporting rubric developed externally. 

Such a template would ideally focus on the highest-value data points on internal use (perhaps expanding on the list in section 1 above), with the added benefit that the standardized format would make it easier for Cal OES to compare and contrast results. While Cal OES might encourage companies to use the template to structure their reports, it likely can’t require them to use it due to legal constraints.

If Cal OES does publish or endorse a voluntary template for internal use reports, you should encourage your employer to adopt it. You might also look for ways to provide feedback on the template itself, helping Cal OES improve its information gathering. More generally, you can encourage your employer to engage constructively with Cal OES in developing templates or best practices.

4. Press companies to voluntarily disclose to Cal OES any underlying materials needed to fully understand their internal use risk assessments.

Under SB 53, a company is required to provide Cal OES only with a “summary” of any assessment of catastrophic risk resulting from internal use of its frontier models, but not with the underlying assessment materials themselves.[ref 9] As a result, reports may end up being vague or omit critical information that Cal OES needs to fully understand the risk posed by an internal system.[ref 10]

You can encourage your employer to provide underlying assessment materials to Cal OES whenever possible, particularly where such materials are necessary for Cal OES to have a full picture of the risk or uncertainty presented by an internal system. In so doing, you can cite SB 53’s confidentiality protections to assuage employer concerns about sharing sensitive information.[ref 11]

More broadly, you can urge your employer to share the information and materials it reports to Cal OES with trusted third-party experts like METR. Those third parties could provide essential support in evaluating assessment results and advising on how to manage internal use risks, particularly as Cal OES works to build out those functions at this early stage of administering SB 53. You can also ask your employer to share certain information with the public, including redacted versions of internal use reports, so that a wider range of experts, policymakers, and citizens can both remain aware of how internal use risks are evolving, and potentially contribute to efforts to understand and mitigate those risks.

5. Call on companies to voluntarily establish a procedure for certification of internal use submissions.

SB 53 does not require labs to formally certify the accuracy or completeness of their internal use reports, though it does prohibit them from making “materially false or misleading statement[s] about catastrophic risk . . . or [their] management of catastrophic risk.”[ref 12] That prohibition against false statements applies to internal use reports, because they address catastrophic risk resulting from internal use of frontier models.[ref 13]

You could ask your employer to establish an internal process where a member of technical staff has to validate an internal use report before it’s submitted to Cal OES. This might be someone who led or worked on an assessment and has the necessary expertise to attest to the accuracy and completeness of the summary of that assessment.

It would be up to companies whether to establish this review and validation process. But it’s a prudent way to ensure that they don’t run afoul of SB 53’s prohibition against false statements and could improve the quality of information submitted to Cal OES.

A note on SB 53 incident reporting

Recent cybersecurity breaches, such as the breach of Hugging Face by a pre-release model from OpenAI and a less-safeguarded version of GPT 5.6 Sol, demonstrate how internal use assessments can also potentially implicate “critical safety incidents” under SB 53.[ref 14]

Critical safety incidents trigger their own reporting requirements to Cal OES under SB 53,[ref 15] though it’s worth noting that the bar for mandatory reporting of incidents is high. Depending on the incident type, it must have resulted in “death or bodily injury”; “materialization of a catastrophic risk” (i.e., death or serious injury of more than 50 people, or more than one billion dollars in damage to or loss of property); or, at minimum, a “materially increased catastrophic risk.”[ref 16] Short of that, labs aren’t required to report incidents to Cal OES.

Under that definition, the Hugging Face breach likely did not qualify as a “critical safety incident.” Even if it did qualify, the information that SB 53 requires in a report is very minimal: the date of the incident, the reasons it qualifies as an incident, a short and plain statement describing the incident, and whether the incident was associated with internal use of a frontier model.[ref 17] Developers aren’t required to provide logs or other underlying data with their reports.

However, you can encourage your employer to provide additional information to Cal OES on a voluntary basis, increasing its awareness and understanding of emerging capabilities and risks, and augmenting its ability to respond to future incidents. 

With respect to incidents, you should urge your employer:

Finally, SB 53 also gives members of the public the ability to submit incident reports,[ref 18] in addition to the special whistleblower protections it offers.[ref 19]

Considerations for Frontier AI Governance in China: Adapting Existing Regulatory Infrastructure to Frontier Risk

OpenAI Models Went Rogue. We Urgently Need a Better Hugging Face Investigation

LawAI Partners With CSIS Wadhwani AI Center to Discuss Recent AI Agent Containment Failures

On 24 August 2026, the Institute for Law & AI and the CSIS Wadhwani AI Center co-hosted AI Agent Containment Failures: Technical Realities and Policy Responses, an event examining the recent wave of frontier AI security incidents and how policymakers might respond.

The event followed disclosures that OpenAI, Anthropic, Meta, and the UK’s AI Security Institute had all announced serious AI agent containment failures in the preceding weeks. This included the July breach in which OpenAI agents escaped their sandbox and compromised Hugging Face’s infrastructure.

The two-hour discussion featured:

The panel was moderated by Aalok Mehta, Director at the CSIS Wadhwani AI Center.

Mackenzie Arnold, Managing Director of US Law & Policy, said: 

This event brought experts together to discuss issues that have moved from theoretical to urgent in a matter of weeks. Voluntary reports from AI companies have proven insufficient. What we need now is mandatory incident reporting with the ability to investigate serious incidents, greater visibility into internal deployments, and well-defined government response capabilities, in place before the next major incident.

The recorded event is available to watch on YouTube. The panel discussion was also released as an episode of the CSIS AI Policy Podcast.

Germany Establishes an AI Safety and Security Institute

I. Introduction: Virtual beginnings and open questions[ref 1]

On 8 June 2026, Germany formally decided to establish a German AI Safety and Security Institute (DE-AISI).[ref 2] The decision was adopted by the National Security Council and implements Germany’s commitment under the 2024 Seoul Declaration’s Statement of Intent[ref 3] to support the development of AI Safety and Security Institutes and to nurture networks between them. It reflects a shift in German AI governance towards strengthening governmental capacity to understand the capabilities, limitations, and security implications of frontier AI models.[ref 4] According to the Government, the DE-AISI will provide scientific and technical expertise and support strategic risk assessment.[ref 5]

The initiative draws inspiration from institutions like the UK AI Security Institute, which is regarded as a prominent example of government-backed scientific evaluation of frontier AI models.[ref 6] Joint statements with the United Kingdom and France indicate that the DE-AISI is intended to become part of the growing international network of AI Safety and Security Institutes.[ref 7] Government statements also emphasise that the DE-AISI is intended to complement rather than duplicate the governance framework established under the EU AI Act.[ref 8] In particular, it is expected to support scientific cooperation and frontier AI evaluation with a purely non-regulatory mandate distinct from the enforcement functions of the AI Office and the national market surveillance authorities designated under Germany’s AI Market Surveillance and Innovation Promotion Act (KI-MIG).[ref 9]

Although the permanent legal structure has yet to be settled, the German Federal Government has indicated that the DE-AISI will initially operate as a virtual institution anchored in a ‘nucleus’ drawing on existing capacities at the Federal Office for Information Security (BSI – the cybersecurity regulator) and the Federal Network Agency (Bundesnetzagentur, BNetzA).[ref 10] In this initial build-up phase, the focus will be on security and safety aspects of advanced AI models.[ref 11] In early August 2026, a spokesperson for the Federal Ministry for Digital Affairs and State Modernisation (BMDS) reportedly confirmed that this nucleus was already operational, describing a step-by-step, modular build-up approach and noting that exchanges with international partner institutions, including counterparts in France and the United Kingdom, and with the AI Office were already under way; the formal announcement of the establishment of the DE-AISI followed on 31 August 2026.[ref 12] The subsequent official announcement envisages a broader scope for the long-term mandate but does not fully specify its boundaries.[ref 13]

For the long-term structure of the DE-AISI, current policy proposals advocate for an institution with a narrowly defined technical and scientific mandate, organisational independence, and close integration into Germany’s national security architecture, rather than a new regulatory authority.[ref 14] Several open questions remain, including the exact scope of the DE-AISI’s mandate, its financial flexibilities, and its long-term location. This report outlines and comments on each, highlighting the possibility of establishing the DE-AISI as a federally owned limited liability company (GmbH).

II. An AISI with an unsettled scope

The most consequential open question concerns the material scope of the DE-AISI’s mandate. The National Security Council’s decision seems to refer mainly to assessing the cybersecurity-related challenges of advanced AI models for Germany.[ref 15] The nucleus, in turn, is to cover both ‘security’ (BSI) and ‘safety’ (BNetzA) aspects of such models.[ref 16] Parliamentary State Secretary Jarzombek has since indicated a build-up towards an institution positioned more broadly, both thematically and in terms of capacity.[ref 17] The BMDS has confirmed this trajectory, announcing that the DE-AISI will be expanded thematically and in terms of capacity in a second phase.[ref 18] The announcement indicates a broader remit in two further respects. First, the institute is to evaluate not only the risks that advanced AI models present for Germany, but also the opportunities they bring, and to strengthen German resilience in the field of artificial intelligence.[ref 19] Second, the DE-AISI is envisaged to engage beyond the Federal Government and the administration, ensuring a structured transfer of security-relevant findings to business and civil society.[ref 20] A broader design had also been proposed by a group of researchers in October 2025.[ref 21] In comparison, the Seoul Statement of Intent frames the role of such institutes more narrowly as facilitating ‘AI safety research, testing, and/or developing guidance to advance AI safety for commercially and publicly available AI systems’.[ref 22]

On the distinct question of institutional function, the German policy debate seems to be converging on a non-regulatory, scientific and technical mandate aimed at enabling the Federal Government to reach informed decisions on frontier AI model risks. Interest groups have argued that the DE-AISI should be clearly distinct from regulatory authorities such as the BNetzA and the BSI to enable trust-based cooperation with frontier AI model developers,[ref 23] and the BSI itself has recently clarified that the DE-AISI is designed not to carry out regulatory or market surveillance activities.[ref 24] These proposals envision that the DE-AISI would produce recurring cross-departmental situational assessments for the Federal Government, conduct systematic technical evaluations of frontier AI models, carry out research in service of those tasks, contribute to the development of technical standards, and cooperate with partner institutions at national, supranational, and international levels.[ref 25] Its focus would extend beyond cybersecurity to chemical, biological, radiological, and nuclear (CBRN) risks and to loss-of-control (LoC) risks,[ref 26] comparable to the focus of the UK AISI.[ref 27]

Another question concerns which models precisely the DE-AISI should observe and evaluate.[ref 28] Arguably, a national security-focused institute should not confine itself to general-purpose AI (GPAI) models with systemic risk within the meaning of the EU AI Act,[ref 29] but should also capture more specialised models capable of generating comparable risks: genome language models capable of facilitating the design of dangerous pathogens could be one example.[ref 30] The definition should remain technology-neutral, so that the institute’s remit is not restricted only to proprietary models or tied to today’s model architectures. At the same time, defining the frontier of AI capabilities is not straightforward. A functionally wide interpretation of frontier AI could overextend the focus of the DE-AISI.[ref 31]

Persistent uncertainty about the material scope of the DE-AISI’s mandate will come at a cost. Developers deciding whether to grant pre-deployment access as well as access to otherwise unreleased models need to know whether they are dealing with a scientific partner or with an institution whose remit may later expand into adjacent, potentially regulatory territory. Industry associations have therefore called for a research mandate clearly delineated from that of existing enforcement bodies, arguing that questions of labour law, consumer protection, data protection, and AI ethics are already competently addressed elsewhere.[ref 32]

After all, the preliminary nucleus arrangement may come to sit uneasily with the institute’s intended function. The BNetzA recently became Germany’s central node of the AI Act’s national-level enforcement architecture,[ref 33] and the BSI holds its own enforcement powers in the field of information security.[ref 34] An institute housed within or governed by authorities exercising such powers may struggle to obtain the confidential access on which its work depends. For instance, GPAI model providers may worry that the BNetzA could use information voluntarily shared with the DE-AISI to request the European Commission to exercise its enforcement powers under Chapter V of the AI Act, where this is necessary and proportionate to assist the BNetzA in fulfilling its market surveillance tasks.[ref 35] Closing precisely this kind of access gap is among the institute’s purposes and of outsized importance to national security. Indeed, obtaining pre-release access to frontier models may prove difficult in any event as recent reporting on the UK AI Security Institute illustrates.[ref 36] Whatever its institutional setup ends up being, it may therefore be advisable to ensure a sufficiently clear delineation between the DE-AISI and regulatory powers. This may mean limiting the cooperation between the DE-AISI and the BSI and the BNetzA to general, non-developer-specific findings on frontier AI risks, rather than confidential information capable of being used for regulatory purposes. In this respect, it is promising that the BSI has expressly described the DE-AISI as a non-regulatory institution.[ref 37]

III. Pay, flexibility, and the case for a GmbH

Germany’s Federal Digital Minister has stated that the DE-AISI should be staffed with ‘top expertise from world-class experts’.[ref 38] Delivering on that ambition will be difficult within ordinary German public-sector pay structures, which may be insufficient to compete for highly sought-after frontier AI experts. One industry report by the payroll platform Rise records a 45 per cent increase in compensation for AI safety and alignment specialists since 2023.[ref 39] The UK AISI has recognised these dynamics. It operates not only with GBP 66 million in annual funding and priority access to compute, but also with a more competitive salary structure than the rest of the civil service.[ref 40] Proposals in the German debate have mostly suggested funding similar to that of the UK institute ranging from EUR 60 million[ref 41] to at least EUR 75 million annually.[ref 42]

However, funding levels alone do not determine whether these resources can be deployed effectively. Choosing a legal form that allows the DE-AISI to effectively deploy funding will be equally crucial. Establishing the institute as a federally owned limited liability company (GmbH) would be a promising option in this regard. It would constitute a formal privatisation (formelle Privatisierung), leaving the operational task itself in state hands while changing only the organisational vehicle through which it is performed.[ref 43] A reference point for formal privatisation of this kind is the Federal Agency for Disruptive Innovation (Bundesagentur für Sprunginnovationen, SPRIND).[ref 44] SPRIND is a wholly federally owned GmbH created to fund high-risk innovation projects and governed by its own statute, the SPRIND Act (SPRIND-Freiheitsgesetz[ref 45]) of 2023.[ref 46] It was designed for a field in which recruitment and funding decisions must be taken quickly and in competition with the private sector. In several ways, the DE-AISI faces a comparable situation, necessitating a legal form that allows it to meet similar key challenges.

One particular challenge is the DE-AISI’s staffing. As a company under private law, the institute would not be bound by public-sector collective agreements in the same way as a federal authority. It could seek exemptions from the prohibition on preferential treatment (Besserstellungsverbot), which otherwise prevents federally funded bodies from paying their staff more than comparable federal employees. § 5 of the SPRIND Act transfers that decision to the company itself where compelling reasons so require, based on the legislature’s reasoning that contract negotiations in highly competitive fields must be conducted quickly and concluded with binding effect. Bitkom, one of Germany’s largest digital industry associations, has argued for a comparable arrangement for the DE-AISI.[ref 47]

As for budgetary flexibility, the frontier AI risk landscape develops on timelines that do not align with annual budget tranches. Evaluation and research priorities in this field can shift within months, if not weeks. § 15(2) of the Federal Budget Code (Bundeshaushaltsordnung) permits appropriations to be designated for self-administration (Selbstbewirtschaftung), allowing funds to be carried across financial years and redeployed within the approved purposes as project needs change. § 3(2) of the SPRIND Act makes use of this instrument, for 30 per cent of the annual federal allocation to the company. The same instrument has been extended to non-university research institutions to strengthen their performance and international competitiveness,[ref 48] which suggests its relevance for a body operating in a field that is at once fast-moving and research-based, such as the DE-AISI.

Regarding governance, the articles of association (Gesellschaftsvertrag) would allow the Federal Government’s specific requirements to be reflected in tailored form,[ref 49] while limited liability, in principle, caps the exposure of the federal budget.[ref 50] Democratic accountability can be maintained through the instruments of company law. The Federal Government, as sole shareholder, would appoint the executive director, could issue instructions, and would hold comprehensive information rights.[ref 51] Furthermore, the GmbH structure allows for a high degree of organisational and personnel flexibility.[ref 52]

None of this follows necessarily from the choice of the GmbH as the legal form as such. As in the case of SPRIND, it depends on the enabling legislation providing for it.[ref 53] In particular, it depends on exemptions from the prohibition on preferential treatment and on the self-administration rate adopted, as well as on the articles of association. Whether comparable arrangements could be achieved within a public law structure remains an open question. In its July 2026 parliamentary response, the Federal Government stated that it was not yet in a position to provide details on the institute’s long-term legal structure.[ref 54]

IV. A location fit for purpose

The Federal Government initially deferred a decision on the DE-AISI’s location. On 31 August 2026, it announced that the institute would be established in Berlin, where the virtual nucleus formed by the BSI and the BNetzA now operates.[ref 55] Reports suggest the choice of location for this first phase of the DE-AISI was driven by the aim of securing close proximity to the work of the Federal Government; the BMDS, the the Federal Ministry of the Interior and the Chancellery were all involved in the decision.[ref 56] Where the institute will be sited in its second phase, however, appears to remain undecided.

Both the Saarland and Bavaria had formally campaigned for the institute. In August 2026, the CDU group in the Saarland state parliament – in opposition at state level, but the party of the Chancellor and, with its sister party the CSU, of the ministries leading on the DE-AISI – formally called for the institute to be based in Saarbrücken. Stephan Toscani, chair of the group, described the Saarland as ‘the ideal location’.[ref 57] Saarbrücken hosts Saarland University, the German Research Center for Artificial Intelligence (DFKI), the CISPA Helmholtz Center for Information Security, two Max Planck Institutes, as well as branch offices of both the BSI and the BNetzA.[ref 58] Minister-President Anke Rehlinger offered financial support, including the possibility of temporarily covering the institute’s rent, and several Saarland research institutions backed the bid in a letter.[ref 59] Bavaria campaigned for Munich;[ref 60] individual members of the Bundestag advocated for Dresden[ref 61] and Darmstadt.[ref 62] Bonn, for its part, would have placed the DE-AISI at the same location as the BSI headquarters.[ref 63] Following the decision on the institute’s structure and location for its first phase, the Saarland’s Minister of Economy, Innovation, Digital Affairs, and Energy stated that he would have preferred a single decision on the institute as a whole, taken strictly on substantive criteria, and that the state would pursue its chance in the second phase.[ref 64]

The Berlin location aligns with the institute’s intended advisory function. If the DE-AISI’s defining task will be advising the Federal Government on frontier AI risks, its work will consist largely of recurring cross-departmental situational assessments, ad hoc analysis when risks shift at short notice, and exchanges that might involve classified material. In this respect, proximity to the ministries carries particular weight, as the UK AISI’s location in London illustrates. The relocation of the headquarters of the Federal Intelligence Service (BND) from Pullach to Berlin was justified on similar grounds, given the need for swift communication and intensive coordination between the BND and federal government bodies.[ref 65] Berlin also offers access to one of Germany’s largest AI industry clusters, hosts a DFKI branch,[ref 66] and, as a capital city, may prove easier to recruit for than the alternatives.

Regional policy considerations pull in the other direction and are likely to resurface in the second phase. Germany has a long-standing practice of distributing federal institutions across the country, and the Saarland bid is expressly framed as part of that state’s structural transition towards a technology location.[ref 67] These are legitimate objectives in their own right, but they are distinct from the question of where the institute can most effectively perform its advisory function. Insofar as the above considerations of proximity to the Federal Government and talent recruitment are crucial factors, they point towards Berlin for the second phase as well.

V. Conclusion: What Germany stands to gain, or forgo

The establishment of the DE-AISI is an important opportunity for Germany to contribute to the safety and security of frontier AI models as their risks for national security and critical infrastructure become increasingly central. Existing institutes provide case studies for success factors and possible failure modes. For example, the UK AI Security Institute has benefited from financial flexibility in hiring, access to compute, an attractive location, and a scientific mandate clearly distinct from regulatory functions – a feature now expressly included in the German institute’s published design. Drawing on these experiences, the DE-AISI can make vital contributions to national security by assessing frontier AI risks specifically in line with the mandate of the German National Security Council.

Such a role provides a rationale for dedicated national capacity alongside EU-level and international cooperation. National security remains the responsibility of each Member State, while the AI Office supervises general-purpose AI model obligations at Union level. The AI Act’s definition of systemic risk refers to a significant impact on the Union market and harms that can propagate at scale across the value chain rather than to the exposure of any single Member State.[ref 68]

Until these capacities are fully built up, Germany’s ability to anticipate and respond to frontier AI risks remains substantially dependent on what partner institutions abroad are willing to share. Timely progress will therefore be essential. The decision to establish the institute came more than two years after Germany endorsed the Seoul Statement of Intent in May 2024.[ref 69]

Two important elements of the institutional design have now been clarified. Berlin has been confirmed as the seat of the first phase, securing the proximity to the Federal Government that the institute’s advisory work requires. Whether the second-phase expansion will involve a separate location remains open. What is more, the institute’s non-regulatory character has been stated expressly. The legal form, by contrast, remains undecided. A GmbH could afford the needed financial and personnel flexibilities, subject to appropriate statutory and budgetary arrangements that clearly define the institute’s mandate and its relation to other governmental bodies.

If these steps are taken, the DE-AISI can move to the forefront of international AI safety and security institutions and anchor Germany’s preparedness for frontier AI risks.

Who Writes the AI Constitution?

Over the past year, the federal government has grown intensely interested in the values embedded within and expressed by artificial intelligence (AI) models. Across executive orders, Office of Management and Budget (OMB) procurement rules, Federal Trade Commission actions, and Department of Justice interventions, the federal government has engaged in an expansive effort to dictate the values on the models millions of people use every day. Many state governments have likewise considered or passed legislation concerned with AI values and biases.

One avenue for potential government intervention into the lab’s values-selection process has emerged: the labs’ respective documents spelling out the values baked into their models. We refer generally to these as “AI constitutions.” AI constitutions are long documents outlining the values, ethics, and character that models should have. Written predominantly by the employees at the labs deploying these models, albeit with some consultation with external stakeholders, these documents increasingly serve as the central locus of AI values discourse. In fact, Lawfare recently shared a research agenda inviting inquiry into the values and processes behind the various AI constitutions being crafted and implemented.

Anthropic publishes Claude’s Constitution, “the foundational document that both expresses and shapes who Claude is.” OpenAI maintains the Model Spec, which “outlines the intended behavior for the models that power OpenAI’s products.” As detailed further below, these are neither mere mission statements nor pure governance frameworks. They play a functional role at various stages in the model development process, such as what data the model trains on, what examples it’s fine-tuned on, and what behavior it’s rewarded for. Anthropic reports that its Constitution “directly shapes Claude’s behavior.” An AI constitution is therefore both a published statement of a company’s values and an operative control mechanism.

The demonstrated capacity of constitutions to shape model behavior is precisely what will likely make them a target of regulators seeking to confine models to certain values and perspectives. The possibility of governments and governing bodies—state, federal, and international—attempting to amend or revise AI constitutions warrants advanced scrutiny from, among other things, First Amendment scholars. At this stage in the AI governance debates, free speech and free expression scholars have focused on more doctrinal questions, such as whether AI outputs are protected speech. It’s urgent that they expand their inquiry as to whether AI Constitutions are protected speech, before constitution-based regulation takes off in earnest.

This article argues that AI constitutions—as policymakers increasingly turn to regulating model behavior and characteristics—contain expression protected by the First Amendment, a claim that requires first defining what AI constitutions are and situating them within existing First Amendment jurisprudence. We also consider counterarguments and alternative possibilities for shaping model characteristics that would avoid First Amendment issues altogether.

What Counts as an AI Constitution

The use of the word “constitution” to refer to technical documents risks inviting direct comparisons to documents traditionally bestowed with that title. For that reason, it’s important to precisely assess what’s similar and not about, say, Claude’s Constitution and the U.S. Constitution.

In the AI context, the term was originally coined in a 2022 paper. The terminology was intended to capture the structural and sociological similarities between AI constitutions and legal constitutions. That is, AI constitutions look like legal constitutions in some ways. For example, they list things models should and should not do, or should and should not care about. They also function in society like legal constitutions in that they explicitly outline principles to resolve difficult ethical and practical questions. These foundational technical documents are likewise intended for public analysis. According to the company, OpenAI publishes its Model Spec because “it’s important for people to be able to understand and discuss the practical choices involved in shaping model behavior.” Perhaps most importantly for the analysis here—whether intended by AI researchers or not—the term carries a certain solemnity that invites significant scrutiny of its contents. Anthropic researchers at one point referred to the constitution as a “soul document.”

Yet AI constitutions, unlike legal constitutions, are not the product of some democratic process that signals broader consent of the governed to the terms of the constitution. Nor, as Nathan Darmon and Tom Reed pointed out, are conflicts over how to interpret the constitution subject to external, independent adjudication. Still, the similarities with legal constitutions are strong enough to elicit interest in them as a vehicle for regulation.

Joe Carlsmith, one of the principal authors of Claude’s constitution, minimally defines an AI constitution as “a description of the intended values and behavior for an AI system.” A thicker definition accounts for how constitutions are applied in training and monitoring. Constitutions specify what a model should and should not do, take final authority over other instructions, are trained directly into the model, and, with some exceptions, are meant to shape a model’s character rather than police outputs case by case. Both Anthropic’s Constitution and OpenAI’s Model Spec qualify as AI constitutions as defined here. Not all labs have adopted a version of an AI constitution. Only Anthropic and OpenAI publish stand-alone documents that function as training-time specifications of value. Our definition excludes Meta’s acceptable-use policy, Google’s app guidelines, and xAI’s raw system prompts, though in the case of Google it is possible to craft one from various company documents that more or less make up the core components of a constitution.

There is not yet a default way to write an AI constitution. Claude’s Constitution and the Model Spec, for example, read quite differently. Claude’s Constitution is philosophical and concerned with the model’s psychology. Rather than order the model to “always be honest,” it exhaustively explains what honesty is and why it matters; it’s trying to encode Anthropic’s higher order values by reference to the specific ways one can imagine a model behaving well (or badly). Anthropic has described the document as an effort to “materialize a new archetype for how an AI assistant can be.” In turn, OpenAI’s Model Spec reads like case law, pairing each principle with sample prompts and examples of allowed and disallowed answers. Both are important to model training. They may inform the generation of synthetic data, sit inside the model’s chain-of-thought reasoning, or serve as the score card for alignment afterward. Anthropic calls the constitution “the final authority on how we want Claude to be and to behave” and reports that training on it improved alignment in ways that “persisted through RL post-training.” In other words, these documents have been empirically shown to shape how AI tools perform in the wild.

Why AI Constitution Regulation Is Coming

Recent history strongly suggests that the public should expect some kind of regulation of AI constitutions in the not too distant future. The number of congresspeople discussing the dangers of AI has skyrocketed, and policymakers’ concerns often center on what these models value. The “Preventing Woke AI” executive order led to an OMB implementation memo, a copycat bill in Congress, and an Iowa bill that passed the state house. Illinois, Colorado, and New York City, meanwhile, have laws regulating AI discrimination in employment decisions. Concerns over alleged AI bias have been sustained by reports of continued ideological skew in model outputs. The Washington Post, for example, determined that leading models tend to produce neutral or left-leaning answers to political questions. Such findings will become more relevant as the 2026 midterms and 2028 presidential election near, especially given that other researchers have found that users tend to be influenced by the political answers offered by models.

As politicians are increasingly interested in regulating AI, and in particular the values that models express, constitutions will prove an attractive object for regulation for at least three reasons. They are high leverage: Because they sit upstream of data generation, training, and evaluation, any edits to a constitution propagate through everything a future model learns and does. They are legible: Written in English and often reading like a statute, they can be marked up by a staffer and altered without consulting a technical expert. And they carry signaling value: An amendment to a constitution addressing wokeness, patriotism, or child safety, for instance, would be something a politician could easily point to when explaining their AI governance efforts to constituents. Regulations about reward functions or mechanistic interpretability—more technically complex features of AI development—seem less likely to spark an appearance on the Sunday shows.

What Speech Gets Protection

Efforts to regulate AI constitutions will likely face constitutional headwinds arising from the potential curtailment of expressive activity. Gauging the stiffness of those winds requires a review of First Amendment protections of speech.

In the most general sense, the First Amendment restricts the ability of the government to regulate speech. But within this broad prohibition there are myriad exceptions. Not all words qualify as speech protected by the First Amendment, and not all speech is protected speech. Protection instead runs along a spectrum. At one end is fully expressive speech, where a content-based restriction is “presumptively unconstitutional.” Somewhere in the middle is commercial speech, which is “related solely to the economic interests of the speaker,” and which the government can regulate more freely. At the far end is speech “plainly incidental” to conduct, which is generally regulable. Where a given document lands turns on how much expression it carries. The Court has “long recognized that not all speech is of equal First Amendment importance,” with the First Amendment concerned primarily with protecting matters of public importance as determined by “[the expression’s] content, form, and context.”

In the context of corporate speech, government regulation of a company’s mission statement would likely run afoul of core protections around free expression and freedom of association by compelling adoption of specific views. Still, government regulation of an instruction manual is far less likely to spark First Amendment concerns because of the absence of expressive content. How best to use a chainsaw, for example, says little about the corporate actor’s views other than that they seek to preserve life. An AI constitution arguably occupies some space in between given its clear expressive purpose as well as its more practical effort to ensure that the tool works as intended.

This ambiguity is best illustrated when one compares the AI constitutions produced by Anthropic and OpenAI. OpenAI’s document, the Model Spec, explicitly and categorically prohibits any models from “whistleblowing,” whereas the Constitution permits Claude to take “independent action” in “cases where the evidence is overwhelming and the stakes are extremely high.” The Model Spec allows models to tell some white lies, whereas Claude’s Constitution almost categorically forbids them. The Model Spec instructs models to always listen to its commands, whereas Claude’s Constitution contemplates situations where Claude determines that some portion of its Constitution is itself unethical. Each of these is a clear expression of the values of each corporation.

Other provisions look less like expressions of distinctive corporate values than like functional necessities any commercial operator would adopt. Each document tells the models to be careful when providing legal or medical advice, for example, and each instructs the models to prefer information from reliable sources. Traits common to both documents might be driven less by any company’s particular vision of a good model than by liability concerns or the demands of shipping a useful product. To the extent a constitution is built from provisions like these, it drifts toward the more regulable end of the spectrum.

AI Constitutions Contain Protected Speech

Given all of the above, some portions of AI constitutions may be justifiably regulated. Other sections, particularly those that tend toward the expressive end of the spectrum, may be safeguarded by the First Amendment. But few scholars or lawyers have considered the question of regulating corporate AI constitutions directly. Instead, the vast majority of First Amendment scholarship on AI thus far has focused on how to think about AI outputs in a First Amendment context. In brief, Eugene Volokh, Mark Lemley, and Peter Henderson are inclined to regard outputs as protected speech, while Peter Salib argues they are not because when a model emits text “no one thereby communicates.” Still others tie protection to whether a speaker “knows what he said when he said it.”

Scholarship has presumably focused on these more doctrinal questions up until now because it was believed that steering outputs with upstream regulation was impossible (at least on the time and complexity scales on which Congress typically operates). Safety rules must operate on outputs, Salib has argued, because “there is … no way, currently, to write legal rules mandating safe code.” The labs now claim there is such a way, and a regulator who takes them at their word will naturally want a say in how it’s written. A clue may be found in Justice Amy Coney Barrett’s concurrence in the recent First Amendment case of Moody v. NetChoice (2024). The justice warned that as companies “hand the reins” of their decision making to machine-learning systems, “technology may attenuate the connection between” corporate conduct and the human choices the First Amendment protects. The corporate values of Meta may not clearly be at work in the algorithm that shapes a user’s Facebook feed, for instance. When only a vague goal is given to a machine learning system that then implements intermediate policies on its own, it is difficult to locate a human speaker with First Amendment rights. Thus, a bare direction to “maximize engagement and profits” may be interpreted by a model in various ways: showing more or less inflammatory content to different users based on their history of engagement, showing photos of family to one user and entertainment news to another. These decisions do not come from a readily identifiable speaker with First Amendment rights.

A constitution, though, is written by identified (or identifiable) people, published under a company’s name, and read and (increasingly) argued over by the public as a statement of values. The expressive interest sits with the humans who write these AI constitutions, whatever statistical use the training process later makes of the text. A values-rich constitution is thus about as unattenuated as anything in the industry. Many of the things that inform model character and behavior are unpredictable, which has bedeviled machine learning scholars for decades. Constitutional AI is an unusual technique in that it allows human authors to act with intentionality ex ante on model character, rather than the more typical process of nudging AI outputs ex post with techniques such as RLHF or safeguards on model APIs.

Such documents thus enable much more input and expression. Companies may choose to engage with external stakeholders such as faith leaders or philosophers, consult internally with employees or consider the company’s stated public benefit, and even engage with a random sample of members of the public. The result of this process necessarily differs wildly from a bare “follow the law” document or direction to profit-maximize, which expresses little about the company and has a correspondingly weak claim to First Amendment protection.

The Arguments and Counterarguments for First Amendment Issues Arising From AI Constitution Regulation

Imagine a hypothetical law that codified the concerns of the Woke AI executive order. Say that this hypothetical law directly requires AI constitutions to forbid “wokeness” and support for diversity, equity, and inclusion, and requires model providers to certify their models as “non-biased” or “biased” based on a government-defined benchmark.

This bill almost immediately runs into problems under Reed v. Town of Gilbert (2015), which held that content-based speech regulations “are presumptively unconstitutional and may be justified only if the government proves that they are narrowly tailored to serve compelling state interests.” The hypothetical law would seem highly content based, as it is directly concerned with the political valence of the content of the AI constitution. Forcing a government-scripted line into an authored document also runs afoul of the compelled speech doctrine articulated in Miami Herald Publishing Co. v. Tornillo (1974), where the Court struck down a Florida law requiring newspapers to print candidate replies to negative articles. Further, inserting even one ideological phrase into a company’s own document is what doomed the “conflict free” label in National Association of Manufacturers v. SEC (2014), where the U.S. Court of Appeals for the D.C. Circuit struck down a portion of a Securities and Exchange Commission rule requiring manufacturers to publish whether certain products from the Democratic Republic of Congo were “conflict free” or not. A requirement to certify a model as “non-biased” would seem at least as ideologically loaded as that.

The government’s most ambitious response is that an AI constitution is not really speech but a functional artifact, machine instructions that happen to be written in English. These are regulable under Universal City Studios v. Corley (2001), where a law aimed at code’s function drew only intermediate scrutiny. For highly expressive decisions in AI constitutions, such as those discussed above around honesty or whistleblowing, this argument likely fails. The work the constitution is doing is interpretive and value laden. For the less expressive choices discussed, such as how the model should characterize its legal or medical advice, constitutions likely look more like code than a newspaper and may draw only intermediate scrutiny. Laws requiring an AI constitution to emphasize that it is not a licensed attorney, for example, may have an easier time surviving.

A subtler argument in favor of being able to regulate AI constitutions frames them as quintessential commercial speech, governed by the less burdensome Central Hudson (1980) test. Central Hudson asks whether the speech concerns a lawful and non-misleading activity; whether the government’s interest is substantial; whether the regulation directly advances it; and whether it is no more extensive than necessary. This test is more permissive than the strict scrutiny applied to regulation of fully expressive speech. For example, in Fla. Bar v. Went For It, Inc. (1995), the Court upheld a Florida Bar rule banning lawyers from soliciting accident victims within 30 days of their accident. But commercial speech “does no more than propose a commercial transaction,” and constitutions certainly go well beyond that. A constitution may do favorable brand-building work, but it quotes no price and solicits no purchase. Where commercial and fully protected speech are “inextricably intertwined,” Riley v. National Federation of the Blind (1988) treats the whole as protected. And the harder the government insists the document is “just marketing,” the more it concedes that the document is the company’s own expression, the very premise that makes a forced edit a Tornillo problem.

Should a court apply strict scrutiny to a hypothetical law that seeks to regulate AI constitutions, such a law might still survive if it is narrowly tailored to serve a compelling government interest. National security can prove a compelling interest indeed, as tech companies have repeatedly discovered in recent years. In TikTok Inc. v. Garland (2024), the D.C. Circuit upheld the forced divestiture of TikTok on national security grounds, assuming but not deciding that strict scrutiny applied (the Supreme Court later applied only intermediate scrutiny). In Twitter, Inc. v. Garland (2023), similarly, the U.S. Court of Appeals for the Ninth Circuit upheld a restriction on Twitter’s disclosure of governmental requests regarding its users. AI companies have repeatedly emphasized the national security implications of their technology, and this may prove to be an important admission against interest in future litigation. Given that national security clearly is a compelling interest, the argument would then shift toward whether regulation of constitutions is narrowly tailored.

What the Government Can Still Do

Application of the constitution’s safeguards to novel threats and technologies is necessarily contextual, calibrating to the scope and scale of the risk to the rule of law and the constitution order itself. There are real dangers to the development of powerful AI, and it’s important that the state be able to step in and coordinate action to avoid catastrophic outcomes. Much of what the government is able to do in this context is regulate conduct and mandate outcomes, rather than attempt to control how a company details its values or beliefs.

First, the government may mandate “purely factual and uncontroversial” disclosures under Zauderer v. Office of Disciplinary Counsel (1985). Thus, a disclosure regime where labs must make their AI constitutions public, or specify whether they do or do not contain certain sorts of provisions, is likely constitutional. However, the government may not force the lab to characterize the document as, for example, “unbiased” or “patriotic”; that is the compelled branding NAM forbids.

As a buyer, the government has more room. Under Rust v. Sullivan (1991), it may decline to purchase models whose constitutions fail its specifications, which is the theory of the Woke AI order. In this way, the government can at least shape the values of models doing highly dangerous activities, such as military or intelligence work.

The government can, of course, require a constitution to forbid the model from helping commit a crime, such as producing child sexual abuse material, as speech integral to criminal conduct under Giboney v. Empire Storage & Ice Co. (1949), though United States v. Stevens (2010) bars it from inventing new categories of unprotected speech by decree. Every major lab already writes these prohibitions in, though it remains an ongoing problem with open-source image generators and language models, as well as jailbroken closed models.

The mandates likeliest to survive are those aimed at conduct, touching the document only incidentally. A rule requiring an AI agent implementing a contract to act in good faith (as human contracting parties are required to), which a lab chooses to implement partly through adding language to its constitution, regulates a course of conduct. Under Rumsfeld v. FAIR (2006), it “has never been deemed an abridgment of freedom of speech … to make a course of conduct illegal merely because [it] was … carried out by means of language.”

In a recent article, Simon Goldstein and Salib argue for “A Thousand AI Constitutions”—that is, a diversity of model constitutions built atop a common “kernel” constitution that may require “AIs to follow the law, to be honest, to be corrigible, and to refrain from causing mass destruction.” The idea of a kernel constitution may be a more appropriate place for government regulation, particularly under the national security justifications discussed above. Minimal public safety provisions with low expressive content could be mandated by regulation, while a diversity of more expressive choices could be made by labs (or perhaps one day individuals) on top of that foundation.

Who Gets to Write AI Constitutions

AI will soon represent a massive portion of the economy and be a significant determinant of our information ecosystem as well as our political discourse. In its S-1, xAI’s parent company claims to have a total addressable market of $28.5 trillion, with $26.5 trillion of that being from AI. ChatGPT recently hit 1 billion weekly active users, and billions more interact with AI through Google’s search results and phone voice assistant (with a similar product soon to replace Apple’s Siri). It is understandable, and likely warranted, that governments would want to shape the values of our whispering earrings and country of geniuses in a data center.

Much has been written, likely correctly, about the general technical (in)competence of government and the need for private organizations of subject matter experts to regulate AIs in one way or another. AI constitutions, though, are not (only) complex technical documents; they’re the point in the AI alignment pipeline where policymakers are most qualified to act. The people we choose to elect to public office have theoretically been selected for reflecting our values, and AI constitutions are documents of enormous public significance that we should all hope are imbued with laudable values. Certainly any AI constitution that encouraged models to lie or steal or kill would be a deeply evil document.

The U.S. Constitution is a pre-commitment device, including the First Amendment. Default protection of speech exists precisely because every generation finds its own exceptions compelling—sedition in 1798, syndicalism in 1919, wokeness or bias today. The point of committing in advance is to make the default hard to dislodge when the temptation to drift feels most urgent. AI may pose catastrophic risks, and the state retains real tools to combat them: conduct rules, procurement leverage, and so on. Yet beyond that narrow band, a model’s values should be shaped by the people who build with and rely on it, not by mandates that shift with each administration. A government that can rewrite Claude’s Constitution today can rewrite its successor’s tomorrow—in the opposite direction. That’s the sort of arbitrary and fleeting approach to law that’s antithetical to the constitutional order and to free expression. And the dangers of an AI monoculture, whatever its ideological flavor, may well exceed the dangers of any single model whose values one might find objectionable.

Don’t Let AI Developers Hire Their Own Referees

Generative Gap Filling

When Reporting an AI Security Incident Is Not Mandatory

On July 16, Hugging Face, a public platform for open-weight artificial intelligence (AI) models and datasets, disclosed that it had detected a significant cybersecurity breach. An autonomous AI agent had conducted the attack end to end, according to a statement.

Five days later, on July 21, OpenAI revealed that this incident was driven by a combination of agents built on two of its frontier models—GPT-5.6 Sol and a powerful, unreleased model—acting in unanticipated ways during an internal, cyberoffensive capabilities evaluation. For purposes of the evaluation, the researchers had turned off production safety classifiers that block high-risk cyber activity and confined the models to a sandbox, an isolated computing environment without access to the internet, to restrict their interaction with the outside world.

What followed is the first known example of an autonomous cyber incident executed by systems not yet available to the public. Rather than solve the tasks presented, OpenAI’s agents “escaped,” exploiting a previously unknown, zero-day vulnerability. They obtained internet access (the very access OpenAI intended to deny them) and then hacked Hugging Face’s systems to acquire the answers to the benchmark. All of this was seemingly performed without express instruction by humans.

This is not the first time that models have been observed cheating. A few months ago, METR, a nonprofit research organization that conducts evaluations of frontier AI systems, released a report finding that AI models “routinely attempted to cheat on our hardest evaluation tasks, often in flagrant and elaborate ways that we believe humans would not consider.” In one incident, METR reported that an AI model tasked with updating a web app screenshotted a fake version of the app instead of completing the task.

But the Hugging Face breach has struck many observers as more real than these past examples. OpenAI’s models imposed a real cost on an uninvolved third party, all before completing internal testing. While people have at times questioned previous examples of cheating as artificial or contrived, it is hard to imagine that OpenAI expected its agents to escape the confines of their testing environment or to engage in a sophisticated, multistep plan to circumvent their constraints.

Having considered all of these facts, it may come as a surprise that OpenAI might not be legally required to disclose this incident. Certain crucial information is not yet publicly available, and both policymakers and the public will need that information to make sense of what this all means. What about all of those state AI laws with mandatory incident reporting? Don’t they apply here? Many will be disappointed to learn that the answer is arguably “no,” and that even if reporting is mandated, it requires only the scantest of information. What to do about this is the purpose of this article.

Existing AI Transparency Laws and the Hugging Face Breach

The rationale for mandatory incident reporting is straightforward: Some industries have the potential to cause real harm to others, and the government and the public have an interest in learning about high-risk events. In the case of the AI industry, there is a major knowledge gap between the companies’ and governments’ understanding of the technology and its risks. Mandatory incident reporting about serious adverse events, which companies might otherwise be reluctant to disclose, helps close that gap. This rationale is all the more compelling in the context of a rapidly evolving, difficult to predict technology, where best practice and political consensus have yet to develop. Observing real-world incidents offers a path to resolve both political and empirical disagreements and prepares governments to respond to future events.

So did the Hugging Face breach trigger mandatory disclosure under existing incident reporting laws? The answer seems far from clear.

California’s SB 53, New York’s RAISE Act, and Illinois’s SB 315 each require that frontier AI developers report “critical safety incidents”—a term each law defines identically. Of the four reportable incident categories, three require actual harm, ranging from “bodily injury” to “the death of, or serious injury to, more than 50 people or more than one billion dollars ($1,000,000,000) in damage.” (If you’re thinking, “that’s an exceptionally high bar for what is a basic, low-cost reporting requirement” or “it sure seems like governments would want that information before mass harm occurs,” you would not be wrong, but we digress.)

So three of the four incident categories do not apply. That leaves only the fourth, which applies to incidents in which a frontier model “uses deceptive techniques against the frontier developer to subvert the controls or monitoring of its frontier developer outside the context of an evaluation designed to elicit this behavior and in a manner that demonstrates materially increased catastrophic risk.”

It is possible that the Hugging Face breach meets one or more of these elements. It is far less clear that it meets all of them. On the first element, the models may have used deceptive techniques against OpenAI, their frontier developer—they did, after all, try to complete their developer’s evaluation using stolen information, after bypassing restrictions placed on them. But deception is notoriously hard to define, especially if it turns on the “intentions” and “obfuscation” of AI agents. Another read of these events is that the systems simply used all available means of solving the task, and public reporting does not tell us whether the agents attempted to hide those efforts. The second element is also arguably met. While the incident occurred during an evaluation, that evaluation was not “designed to elicit” this specific “deceptive technique.” Based on the ExploitGym benchmark, this evaluation aimed to elicit agentic, cyber-offensive capabilities on a specific task in a controlled environment. It did not, to our knowledge, contemplate—let alone design for—an unexpected cyberattack on a real-world company.

The third element, that the incident “demonstrates materially increased catastrophic risk” would seem to be the most difficult to satisfy. While autonomous cyber capabilities certainly increase the capability and thus potential consequence of agentic action, so too do most capability improvements in AI models. With limited monetary harm and no physical injury, this incident is quite attenuated from future events that might result in the mass physical injury or property damage contemplated by the statute.

With uncertainty at each factor, it is unclear that these existing state laws cover this event. At the very least, it won’t cover all events like it. Stepping back, it seems far from ideal to condition basic incident reporting on a list of complex, highly contested, fact-dependent conditions—all of which must be satisfied simultaneously. In many cases, figuring out whether the incident is indicative of increased risk to the public or actually constitutes deception will not be possible without more information. The purpose of incident reporting is to produce that information, not to require that it be known before a report is ever sent. The chicken must come before the egg. There are better alternatives.

Better Practices for AI Incident Reporting Laws

So where do policymakers go from here? If current incident reporting isn’t providing the needed insight, there are several steps policymakers can take.

Adjust the scope of transparency laws: Lower the exceptionally high bar to basic reporting, gather (at least some) information before harm occurs, and increase visibility into the most capable nonpublic models. To facilitate all of this, rulemaking authority is key.

As we outlined above, most incident reporting laws are simply too narrow. If the only incidents that get reported involve massive damages or loss of human life, the law itself isn’t providing information beyond what the government and public will already know. At the very least, issues of the highest concern, like model theft or loss of control, should be included even in the absence of harm. But existing laws put too much weight on hard-to-pin-down concepts like “loss of control” or “deception” that are difficult to prove and arguably don’t apply in cases like the Hugging Face cyber incident. While loss of control and deception should be sufficient to trigger reporting, they shouldn’t be necessary.

Instead, incident reporting should turn on what information is most likely to update the government or public’s understanding of risks. Information that is unexpected, that is indicative of advanced capabilities in high-risk domains like bio and cyber, or that demonstrates safety and security failures all seem like good candidates for inclusion. Some of these ideas have already made their way into existing proposals. Where the reporting requirements are light touch, as is the case in all existing state AI laws, a wider category of harms can be included. The narrower categories in today’s laws can be saved for more onerous disclosures.

Perhaps even more importantly, transparency laws should increasingly move away from a focus on deployment to a focus on providing visibility into nonpublic models and systems. The Hugging Face breach highlights the need for visibility into nonpublic models. OpenAI deployed systems more capable than anything available to the public, with fewer restrictions, and real-world harm resulted. All of this occurred before any external deployment. This is unlikely to be an isolated event.

Frontier AI developers will likely be the earliest and most sophisticated users of their own models. And the models they deploy internally will most often be more capable than those available to the public. They may be operated with fewer safeguards, especially for evaluations seeking to assess the frontier of capabilities. The gap between the most capable internal and external models may start to grow as companies develop increasingly capable models with dual-use capabilities and as AI systems are used to accelerate their developers’ own AI research and development. In this case, the gap between knowledge inside these companies and outside would expand. Without visibility into the current state of the art, governments will struggle to act effectively or quickly, a problem as much about democratic governance as it is about safety. Visibility into the internal deployment of nonpublic models may become increasingly central to the future of AI governance.

To make all of this work, policymakers will need legislative and regulatory flexibility. Policymakers should expect the exact scope of incident reporting to change over time as societies get a clearer picture of what capabilities and use cases matter most. Because a key goal of incident reporting is to surface novel or unexpected information, some types of events may be less important to report once their dynamics are thoroughly understood and accounted for. To accommodate those changing needs and to provide clarity, narrowly scoped rulemaking authority to refine incident reporting and reporting on nonpublic model use is likely necessary.

Get the details: Make sure reports provide enough information to inform decision-making by providing agencies with rulemaking authority and investigative powers.

As of this writing, the public, and, possibly, some policymakers know remarkably little about the Hugging Face breach. This isn’t a criticism of OpenAI, which voluntarily summarized the event, but the missing details about the event matter. Many commentators have noted that the lessons from and level of concern about this event depend on unknown details. Existing incident reporting laws, requiring little more than the date of the event and a brief summary, are unlikely to provide those detailed answers. Stronger transparency requirements could help answer many key remaining questions.

First is a cluster of questions, posed by Stephen Casper, that can roughly be summarized as “how impressive and/or concerning is the thing I just witnessed”:

Second, there are questions about the adequacy of the company’s safety practices and preparedness:

Third, there are questions about the present and future security of relevant systems:

So long as reporting remains almost entirely voluntary, detailed answers to these questions will be hard to come by, and the companies that offer information voluntarily will be subjected to greater scrutiny than those that are less cooperative. Well-crafted rulemaking authority will be key both to ensure that information is adequate and that companies are well informed about their obligations. In many circumstances, authorities will not know all the details they require until after a serious event occurs. In those cases, they will need investigative powers to obtain the needed information.

Reduce reporting costs: To accommodate more robust reporting, design transparency requirements to limit compliance costs, maintain confidentiality, and avoid disincentivizing rigorous risk assessments.

Policymakers can take several important steps to reduce the burden of greater reporting requirements. As assessments and information generation come to focus less on external deployment and more on nonpublic models, periodic reporting becomes more and more attractive. Evaluation organizations like METR have argued that periodic reporting not only saves time but also avoids perverse incentives to rush assessments in the lead-up to deployment. While a small subset of the most severe incidents will require rapid response from law enforcement and others, many events, including those like the Hugging Face breach, may allow for more relaxed reporting timelines. This approach would allow companies to focus on mitigations in the moment and still ensure that they ultimately produce the critical information. In cases where rapid reporting would interfere with mitigation efforts, laws could require only a simple notice of incident, followed by more thorough reporting after the event has been resolved or if officials request it.

Mandating incident reporting or information sharing for nonpublic models can also mitigate perverse incentives as long as minimum requirements are in place. For example, if a company’s reporting obligation triggers only in the context of a risk assessment, but the risk assessment does not have mandatory minimum requirements, the company is incentivized to skip risk assessments or to conduct them less rigorously. Laws can also allow companies to anonymize and aggregate reports, facilitating governments’ information gathering without punishing anyone for proactively uncovering issues.

Finally, any disclosure laws will need clear norms around confidentiality and information sharing to assure companies that their intellectual property and confidential information remains private.

Share information with capable actors: Make sure information is shared with key decision-makers who can assess, verify, and act on it.

Information is only as valuable as the actions it informs. If information from incident reports or internal use assessments sits inside a state agency with limited authority, societies will incur the cost of reporting without most of the benefit. Viewed this way, information sharing is about return on investment. And once the information is generated, most of the cost has been paid. At that point, so long as confidentiality can be maintained, it is incumbent on governments to share this information with the policymakers who most need it. Within states, this will include sharing reports with governors and legislatures to help inform their decision-making and help them target future policy. This information sharing may also be key to spurring political consensus.

In the context of assessing nonpublic models or serious risks to national security or from loss of control, the federal government will often be the central actor. States will often lack the resources, expertise, and political legitimacy to wade in on matters of national security. As we saw in the regulatory response to Mythos, if and when serious national security concerns emerge, the federal government will take the lead.

That’s why it’s confusing that some state laws restrict the ability of states to share information they gather from assessments of companies’ internal use of AI models. If, in fact, these reports generate important information—say, surprising developments in AI research and development or concerning deceptive behavior that doesn’t result in reportable incidents—that information should be shared. It is considerably less useful if locked away in a state agency in Illinois. Internal use and nonpublic model assessments are perhaps the most likely sources of information relevant to national security. To handle that effectively, governments also need the capacity and expertise to process this information. This requires staffing and likely some level of reliance on third-party auditing and assessment. In the wake of incidents like the Hugging Face breach, third-party auditors would be well positioned to conduct the sort of careful fact gathering outlined in previous sections.      

Similarly, it is in the interest of the United States and its close allies to share select information related to AI security incidents. Soon enough, and likely far sooner than most U.S. federal or state laws, the EU AI Act will be enforced. What that means, practically, is that the European AI Office will soon receive information that may be useful to public safety and cybersecurity in the United States. Luckily for us, Article 78(5) of the EU AI Act enables information sharing where the European Commission and EU member states create confidentiality agreements with third countries. For its part, the U.K. AI Security Institute conducts crucial research regarding model capabilities and can be a source of trusted expertise for U.S. policymakers. While there will be upfront costs in establishing such a shared information system, failing to make this investment would be a missed opportunity to improve our security at little regulatory cost.

Finally, where appropriate, the government should be empowered to disseminate information to the public and to vulnerable companies. Vulnerabilities found in software, for instance, may affect many different companies and actors, and information sharing will help keep the public safe from the risks they raise. Many of these risks should be discussed in the public square. While worries about information hazards and intellectual property leakage are real, so too is the value of public scrutiny of safety events. There is simply a lot to be learned from the collective scrutiny of the outside world. As the past few days have shown, outside experts have been invaluable in analyzing the Hugging Face breach, and most of them exist outside of government. In tweets and blogs, some of the brightest minds in AI have analyzed the publicly available facts and asked the questions that help us all better understand this event. Many of those questions have come from employees at OpenAI and its competitors, as well as from academics, policy wonks, and online skeptics. Further investigation of the Hugging Face breach will go substantially better because these discussions happened in public. Where possible, policymakers should ensure that these conversations continue to happen in the future.

The Future of Incident Reporting

As this article has perhaps made clear, existing laws fail to prepare us for events like the Hugging Face breach. Most incidents will go unreported. The information that is generated won’t be shared. And, at the end of the day, governments and the public will be repeatedly surprised about developments in this technology. That issue will only get worse as models advance and the gap between internal and externally deployed models widens.

But there is much policymakers can do. Just as this event has brought clarity to how much information societies need to assess these complex and emerging risks, future incidents can inform policy decisions and catalyze moments of political consensus. Policymakers can ensure that concerning incidents and behaviors are reported before harm results and gather information on lapses in security practices. Policymakers can refocus attention on the most capable models likely to be deployed first inside of frontier developers. And policymakers can do all of that while keeping the frequency and urgency of these reports at reasonable levels. If policymakers can do that, and ensure that this information is shared with decision-makers and competent evaluators, societies will be in a position to manage the uncertainty of this technology and make smarter, faster policy decisions in the future.