This story was originally published by CalMatters. Sign up for their newsletters.

In a world increasingly concerned about serious risks from artificial intelligence, policymakers trying to mitigate those risks find themselves turning more and more to a new class of professionals: AI evaluators, who audit systems on behalf of the companies that create them.

California is at the forefront of embracing the evaluators — and of grappling with some of the problems they raise, including conflicts of interests and the difficulty of auditing general-purpose technology.

Gov. Gavin Newsom last month signed one law establishing standards for “independent verification organizations” that would employ the evaluators and another creating a registry of evaluators. He also assembled a group of experts, who are to report back in November to recommend whether to require evaluators to be embedded inside companies developing the most powerful AI systems and whether to set standards on what counts as an adequate AI audit.

The moves come amid an intensifying focus on evaluators within the AI ecosystem itself. More than 200 AI researchers and evaluators last month signed a letter in favor of standards for independent evaluators; signatories included a former OpenAI whistleblower and the head of the United Nations’ Independent International Scientific Panel.

When a new group focused on AI evaluation gathered United Nations officials in New York last month for its official launch, it chose a California state senator, Jerry McNerney of Stockton, to give the closing remarks. McNerney, a Democrat who authored the verification standards law, called on nations and advanced AI companies like Anthropic and Google to create standards committees — in part to help ensure AI evaluators can properly do their jobs.

“We can put these companies on notice,” McNerney said at the event, “that they’re going to be evaluated by folks that understand the process, that understand the technology, and will be transparent and hold them accountable so that they create safe products, and they don’t send out these things into the wild that can cause havoc with banking, with our infrastructure and so on.”

An international joint statement that calls for independent AI evaluations garnered support from nearly 30 countries, according to the Finnish Ministry of Foreign Affairs.

Evaluators are coming to the fore alongside serious AI incidents. This year, AI agents, while being tested for their ability to hack into computer systems, carried out a series of high-profile attacks against business and government websites. The CEOs of leading AI companies subsequently committed to evaluations of their models by independent third-party auditors. A former employee of one such company, Anthropic, added to mounting concerns when he posted online about the possibility AI models will kill all humans.

Even before those incidents, there’s been a growing consensus among lawmakers in the past few years that developers should test AI before release, said Assemblymember Rebbecca Bauer-Kahan, a Democrat from San Ramon. For example, last year, testing requirements were included in a bill to help prevent discriminatory AI decisions and legislation to stop AI from causing catastrophic events. The proposals built on calls for AI testing and monitoring stretching back the better part of a decade. Roughly three out of four Californians said this spring that the government should require testing of advanced AI models, according to a Carnegie Endowment California survey.

A lawmaker sits in front of a desk with their left hand resting on their chin as they look off into the distance. They are surrounded by other lawmakers sitting at their desks.
State Sen. Jerry McNerney gave closing remarks at the launch of a new group focused on AI evaluation with United Nations officials in New York last month. McNerney during a floor session at the state Capitol in Sacramento on Aug. 6, 2026. Photo by Fred Greaves for CalMatters

Not all the attention on AI evaluators has been positive. 

Bauer-Kahan, who was behind both of the auditor bills Newsom signed, pointed out that independent evaluators sign contracts with the tech companies whose products they evaluate, and that can pose a conflict of interest.

“I’m hearing from evaluators, ‘We have to be careful, because we want them to let us back in,’” she said. “So what are you avoiding saying in order to continue to gain access and contracts?”

Caroline Singh, at the D.C. think tank the Federation of American Scientists, agreed, saying AI audit regimes run the risk of repeating mistakes in the structure of the bond-rating business. Before the 2008 financial crisis, the sellers of bonds could effectively shop for the rating they wanted. Then came the crash. Singh thinks companies and governments should pool money to pay evaluators instead of AI labs paying evaluators directly, an idea that AI lab Anthropic has also endorsed.  

The AI researchers who recently signed the letter calling for industry standards for auditors share Bauer-Kahan and Singh’s concern. They are urging advanced AI companies to rely on evaluators who maintain full editorial control and who meaningfully disclose and mitigate conflicts of interest. They also suggested companies give access to multiple evaluation groups, all receiving the same level of access as their own safety employees.

Another risk is that AI companies may test models as a performative exercise, Bauer-Kahan said — as a move to reassure the public, without a sincere determination to uncover problems. Strong legal requirements can prevent performative testing, she added. 

But legal requirements raise their own problems, said Matt O’Shaughnessy, a former State Department and Congressional staffer now at the Center for Democracy and Technology, a digital rights advocacy group.

When an assessment is done simply to meet a legal requirement, organizations are more likely to do the minimum that they need to do in order to comply, he and his colleagues recently concluded in a set of AI evaluation guidelines for policymakers

“It can make it harder for assessors to get the buy-in they need to make deeper changes in companies,” O’Shaughnessy said.

Singh said that evaluators should share the results of their work to allow other evaluators to judge its quality, thus encouraging more thorough review.

The guidelines from O’Shaughnessy and his colleagues warn that properly evaluating AI models requires a considerable range of skills, including experts in everything from privacy to mental health to cybersecurity to bias outputs and physical safety. On top of that, audits should also examine incentive structures inside AI companies in addition to the technology itself, the guidelines say.

The guidelines also warn that properly evaluating AI models requires a considerable range of skills, including experts in everything from privacy to mental health to cybersecurity to bias outputs and physical safety. On top of that, audits should also examine incentive structures inside AI companies in addition to the technology itself, the guidelines say.

In the future, as they refine their approach to AI evaluation, California officials may find themselves collaborating with international partners. The consulates or embassies of nations sometimes promote and coordinate their tech policies in Sacramento, and in 2024 California lawmakers sought to align regulation with the European Union in order to protect people from artificial intelligence. In August, California also synced implementation of laws requiring watermarks for AI-generated content with the European Union.

Bauer-Kahan said she supports the idea of working with more nations, in part because regulation in one jurisdiction paves the way for rules in another.

“I do think collaboration and finding ways to push the envelope together is helpful,” she said.

A coalition of mid-sized countries working together with California could pool their economic influence to convince the largest AI companies to follow the kind of safety guidelines laid out in the international joint statement calling for independent AI evaluations, said Canadian ambassador to the United Nations David Lametti.

Pasi Rajala, Deputy Minister for Foreign Affairs of Finland, said California working together with small nations “can play a key role” in addressing fast-developing AI risks.

O’Shaughnessy said that an international framework for commercial AI system evaluation is going to be really hard to create without a domestic framework in the United States first.

The Frontier AI Act, a bill in Congress that would enshrine testing requirements into federal law, has bipartisan support in Congress from Rep. Ted Lieu and Jay Olbernolte, both from California.

“I can imagine some of the things that California is doing influencing how Congress thinks about these issues and that becoming something bigger,” he said.

CalMatters is a Sacramento-based nonpartisan, nonprofit journalism venture committed to explaining how California's state Capitol works and why it matters. It works with more than 130 media partners throughout the state that have long, deep relationships with their local audiences, including Embarcadero Media.

Leave a comment