Hero background

How California can meet the moment on frontier AI safety

Last updated 9 October 2026 Contact: info@secureaiproject.orgDownload PDFDownload with appendices (PDF)

The situation in frontier AI is serious and urgent. Without the intent of their developers, frontier AI systems have already gone rogue and hacked or attempted to hack third parties, including government websites in the United States and Australia. Even after OpenAI learned about and remediated the Hugging Face incident, OpenAI agents escaped containment and accessed the internet again. Google, Meta, and Anthropic also reported loss of control incidents this summer.

We frequently talk to rank and file researchers at the frontier AI companies. These are the people in the best position to understand the risks, and many of them are among the most concerned. They feel they don’t have the time or scientific capability they need to keep the systems they are creating safe, which is why they’ve called for efforts like pacing. And while a number of employees at frontier AI companies have resigned over safety concerns, others have chosen to stay, fearing that leaving would only make leadership free to act more recklessly.

As all this is happening, many scientists believe that AI progress could accelerate, and at the same time accelerate the magnitude of the risk. Anthropic has recently said that Claude “leads 26% of model R&D work” at Anthropic and collaborates with humans on nearly all of the rest. If AI research is fully automated, this could lead to a feedback loop, known as recursive self-improvement, that leads to advances far faster than even those we have seen in recent years. This would be a destabilizing and dangerous situation. Imagine how difficult it would be to ensure AI safety if all AI research was performed by AI and progress was happening far faster than it is today. AI developers already have trouble controlling their systems; this would make the problem much worse.

It is the duty and responsibility of the government to protect the public. Often, this involves solving collective action problems that can’t be solved by private entities alone. And while the leaders of the largest AI companies signed an accord at the White House pledging to monitor their models for dangerous capabilities and submit to independent audits, there is no enforcement mechanism, and the President himself called it only "morally binding." California has an opportunity to lead the way.

California can’t wait for more incidents to put the necessary tools in place. Many of the legislative actions we need to take are already clear, and should be codified now. If we wait to do so, it may well be too late. Those legislative actions are:

  • Scoping: SB 53’s scoping needs to be fixed. Today’s frontier models need to be uniformly covered under the law, and serious incidents must be reported.
  • Verification: We need trusted, embedded independent verification organizations (IVOs) to assess risk created by developer activities and ensure that developers are following the law. They should be selected by the state, not by developers.
  • Standards: The Governor’s Office of Emergency Services (Cal OES), California’s frontier AI incident reporting authority, should have the authority to create minimum safety and security standards that all AI companies must follow.
  • Orders: No set of standards is going to be complete as long as frontier AI development continues at an accelerated pace. In cases where standards fall short or danger is imminent, Cal OES needs to be able to issue orders to plug the holes before the public is harmed and while they fix safety problems, including by activating a company-designed shutdown procedure or “kill switch.”
  • Pacing: All of the above may all fail if recursive self-improvement produces advances that are simply too fast to control. California needs to be able to prevent runaway recursive self-improvement by pacing the frontier of AI.

The need for these measures becomes even clearer when considering the status quo. It is hard to believe that events like Hugging Face shouldn’t be reported, that AI companies should evaluate their own safety or choose the people who do so, that there should be no enforceable standards, that the government should have no authority to intervene in an emergency, or that there should be no limits on uncontrolled recursive self-improvement.

If California implements these measures now, we believe it would substantially reduce risks from catastrophic harm that may occur in the coming years. Otherwise, we could be left unprepared in the event of a disaster. Imagine, for instance, if in a year a company produced a model capable of recursive self-improvement and lost control of it in a Hugging Face like incident. Imagine it attacked critical infrastructure and copied itself to computers outside its developer’s control. California would be left scrambling to prevent even more severe follow-up disasters without clearly defined responsibilities and authority within the government, without trained staff and expertise, and without the ability to create regulations to address the problem. Fixing this would require significant time to pass new laws, hire staff, and promulgate required regulations — perhaps years. Incidents can unfold in hours. California needs to act now to implement policies that will be adequate to address these risks and keep the public safe from foreseeable harm.

SB 53 covers models trained with more than 10^26 operations. But the public has no evidence that today’s frontier models actually meet that threshold. No policy regime will work if it does not even cover the models at issue. The definition should be adjusted to ensure that it covers all frontier models. Cal OES should also be able to adjust the definition through notice-and-comment rulemaking.

In addition, under SB 53 a company is only covered for most requirements if it has $500M in revenue. But revenue is not closely related to risk, especially loss of control risk – some companies with billions in funding like Safe Superintelligence Inc. don’t have any revenue and may continue to have none for a while. We recommend fixing this by moving from the $500M revenue threshold to a $1B AI R&D spending threshold, which correlates more strongly with catastrophic risks.

Finally, as far as we know, no incident has ever clearly qualified as a “critical safety incident” under SB 53. This is a big problem because it means the only way we know about these incidents is voluntary reporting from AI companies and investigations by external researchers. The government has no way to know if the spate of reported incidents at OpenAI is because OpenAI is being especially transparent, has especially capable models, or is being especially reckless, because other players might just be choosing not to report incidents. The definition must be expanded to include the kinds of loss of control incidents that have been reported recently.

Our laws do not deem self-reporting sufficient when companies report their financials, the safety or efficacy of a drug, or even whether a restaurant is sanitary. We should regulate trillion-dollar AI companies - companies that have repeatedly warned us that their work could lead to extreme disasters or even an existential catastrophe – at least as seriously, and require even stronger third-party verification.

There is growing support for the idea of requiring frontier AI developers to embed evaluators, or IVOs, in their companies to independently assess risks. METR has already been essential to the investigation of the Hugging Face incident. California should mandate this practice in law.

Outputs

IVOs should publish two kinds of reports:

  • Annually, an assessment of whether the developer has complied with the law. This deters developers from violating the law, including any minimum standards established by Cal OES. This is an “audit.”
  • Quarterly, an assessment of catastrophic risks and security risks from the developer’s models, as well as any recommended corrective actions. This ensures that somebody checks if large risks exist, even if a developer is following every rule. This is an “evaluation.”

It is very important that these reports be published (with redactions). If they are submitted privately to a government agency, then the public and independent scientists will remain in the dark about the level of risk being imposed on them, and the risk of regulatory capture by the companies will be much higher. The state would also have to rely solely on the IVO and government experts to assess the state of risks, when most expertise will be outside both.

IVOs should also be able to notify the government and the public at any time if they come to believe that the company has violated the law or that there is an imminent risk from their activities.

For more information, see our explainer on the difference between audits and evaluations.

Assignment

In our view, letting companies select their own IVOs to perform risk assessments would introduce significant conflicts of interest, lower the expected quality of evaluations, and produce perverse incentives. In the leadup to the 2008 financial crisis, investment banks choosing and paying for their own credit rating led to severe inflation of such ratings, ultimately contributing to the crisis itself. We need to avoid a similar situation in AI, especially given the judgment-driven nature of risk assessment. If companies can select evaluators based on the conclusions they are likely to reach, that creates severe conflicts of interest.

Furthermore, if a single revenue source accounted for more than 10% of an accounting firm’s revenue, its independence would likely be called into question. But there are only a few frontier AI companies; if companies are able to choose their own IVOs, revenue concentration would blow far past that limit. Some organizations have attempted to avoid this problem by organizing as nonprofits and not taking company funding, but this can only go so far. As long as a company gets to choose which evaluator it brings in, that company has considerable leverage over the evaluator.

Instead, we believe the government should assign an evaluator for each function (for example, evaluating loss of control risk or biological risk) to companies, and choose evaluators that it deems to have capacity to take on additional engagements and that are the most qualified for that function. Evaluators should be paid by companies through a process managed by the state. This would better ensure that IVOs answer to the people of the state, rather than the developers that hire them. A government accreditation system, where companies choose from a list of accredited auditors, would in theory permit a state to set a minimum bar for IVO quality. But such a minimum bar is hard to completely articulate in such a fast moving industry, so commercial incentives would point in the direction of IVOs exploiting gaps in that minimum bar and introducing the conflicts of interest described above. Government selection would ensure that the best possible IVO was selected for this important public safety responsibility, and bypass these conflict of interest concerns.

For more information, see our explainer on why IVOs should be assigned and not selected by AI companies.

Transparency and visibility into AI, as important as they are, simply are not sufficient anymore. Many frontier AI companies have lax cybersecurity, poor agent monitoring practices, and vague policies written to avoid them being locked in to taking any particular safety mitigations. The government should hold all frontier AI developers to the best existing safety and security standards at any given time.

Cal OES should be authorized to adopt regulations to create minimum standards for frontier AI frameworks, including required mitigations (see why we believe Cal OES is the right agency to do this under Other recommendations). These standards should be informed by industry and IVO input and go through the normal notice and comment rule-making process, but Cal OES should be able to make emergency regulations on a shorter time frame when needed.

For more information, see our explainer on minimum safety standards and rulemaking.

Minimum standards help, but given how rapidly AI is advancing, no minimum standard can be written now that would be sufficient to prevent risks from materializing. The government needs a way to respond if needed, and it needs to know that it will work.

There has been much talk of a “kill switch” over the years, including in the recent executive order. There are many possible implementations, some of which don’t make sense: for example a literal switch would create a massive cybersecurity vulnerability if a hacker breached it. And it’s not technically feasible for a developer to install a kill switch for an open source copy of a model. But the idea that AI companies should be able to shut down instances of their systems that they control is important and feasible — indeed, it’s exactly what we saw OpenAI do after the Hugging Face incident. AI companies should be required to have documented shutdown procedures in place which describe how they will shut down their models if the company’s management, or a government regulator, gives the order to do so. This need not be a big red button, and could be more like an evacuation plan in the event of a fire which defines who is responsible for taking the necessary steps in the event of an emergency. There can and should be an exemption ensuring these requirements aren’t read to apply to open source models. IVOs should verify that this shutdown procedure would work in a serious emergency.

But the use of a shutdown procedure should not be left up purely to developers. Cal OES should have the power to issue orders to use the shutdown procedure, require mitigations, or adjudicate disputes. We think three types of orders make sense:

  • Corrective action: Order to require a developer to implement a corrective action or cease an activity that is creating risk. This could include activating the shutdown procedure.
  • Information access: Order to require a developer to provide certain information to the IVO or to Cal OES itself.
  • Removing redactions: Order to require an IVO to remove a redaction that was made without good justification by the IVO or developer.

By default, orders should take effect after 15 days to allow time for the developer to appeal to a court if the developer feels an order is unjustified. But in cases where there is an imminent emergency, an order should be able to take effect immediately.

We believe, as do many experts in AI safety, that all of the above may not be enough. Recursive self-improvement could make advances occur so fast that they can’t be controlled.

An analogy: right now this field is moving at about 70 mph. Recursive self-improvement could mean that we start going 140 or 210 mph by next spring. We think a government agency should set some limits on how much faster this process can go, unless and until regulations are in place which would provide a high degree of assurance of safety. For example, Cal OES could define applicable benchmarks to use to measure capability progress (such as increases in the Epoch Capabilities Index per month; this is used for illustrative purposes and details would need to be worked out separately), require companies to report on their progress on these benchmarks, and require companies to slow down when capabilities are advancing too rapidly. This could be done by, for example, limiting the computational power that AI companies are permitted to spend on running AI systems executing AI R&D tasks.

For more information, see our explainer on how California can pace the frontier.

Company risk reports: Companies should be required to produce their own quarterly reports assessing catastrophic risks, security risks, and automation of AI R&D, in addition to what the IVOs produce. We can’t have all the onus be on the IVOs, and it’s important that companies are assessing these risks themselves as well.

Preventing AI aiding in coercively taking control of government: In addition to risks from chemical, biological, radiological, and nuclear weapons, cyber, autonomous crime, and loss of control, we recommend including the risk of a model being used by a malicious actor to take control of a government or a government agency. This would respond to the growing body of research that if AI capabilities expand as quickly as many think they could, it could be used to subvert democracy or even stage a coup.

Disclosing secret loyalties: Widespread use of extremely powerful AI systems would concentrate power in a very small number of hands. For example, a company could train a model to suggest policy ideas that systematically favor the company, or an employee at the company could use powerful models to help commit crimes. If AI companies train models to provide benefits to the company or the company’s insiders that fall outside of normal corporate practices, they should be required to disclose this to the public.

Preventing security risks: We recommend requiring developers to take steps to prevent model weight access/exfiltration/etc. or theft of model training secrets (such as the amateur hack of OpenAI’s internal code), even if it doesn’t lead to an articulable catastrophic risk. Unless AI developers have effective security, US leadership in AI may be undermined if and when a foreign adversary steals the model weights and algorithmic secrets. Better security practices would simultaneously enhance both AI safety and the US’s lead in AI.

Accounting for the automation of AI R&D: SB 53 lacks any explicit requirement to assess risks resulting from the automation of AI R&D, which is critical for detecting and understanding recursive self-improvement. Every major AI company already includes something about this in their frontier AI frameworks. We recommend incorporating such a requirement and covered by minimum standards and IVO verification.

Why Cal OES? We recommend that Cal OES be given most responsibilities under the law. Cal OES’s culture, which is shaped by responding to disasters like wildfires and cyberattacks, is conducive to the scale of the risks involved in frontier AI. Cal OES already has the most important non-enforcement authorities under SB 53 and has been doing well, including by hiring leading experts in frontier AI safety. A new AI division, similar to the existing cybersecurity integration center, should be created within Cal OES, and the division should be staffed up as soon as possible by providing Cal OES with more funding in the 2027 budget. We think setting up a new agency would likely take too long given the timeframes in which risks could materialize.

Explainers on specific policies