- Home
- Research
- Independent AI evaluation is expanding. Who checks the answers in use?
Governance Commentary
Independent AI evaluation is expanding. Who checks the answers in use?
New California laws and embedded evaluators are expanding independent AI evaluation. Audits of controls still leave a question: were the answers in use checked?
Lawmence Wong
Chief Strategist, Lawnise

Two weeks ago we wrote that the frontier labs had accepted independent evaluators. Since then, independent AI evaluation has picked up laws, contracts and budgets.
On 9 September, California's Governor signed two bills on third-party AI evaluation. SB 813 sets up a framework for designating independent verification organisations, with criteria due by 1 January 2028. AB 1405 creates an AI Auditor Registry by 1 January 2029. From then, only registered auditors may offer or conduct an audit covered by the law. Both laws define that kind of audit as one that assesses the internal controls, processes or systems around an AI system or model that are needed to comply with state law.
On 18 September, the Governor issued an executive order asking the state to speed up both laws. One proposal it sends to an expert panel is to require frontier AI companies to host an independent verification organisation inside their labs.
The same day, Anthropic announced a partnership with Accenture on embedded evaluation, covering model evaluation, red-teaming, alignment assessments and safeguard testing. Both companies expect to invest at least US$1 billion in this work over five years.
That is real progress, and it reaches further than the frontier labs. California's framework covers AI systems in operation, and it names organisations that deploy or operate AI among the groups the state must consult.
Controls are not the same as answers
Covered audits under these laws assess the controls, processes and systems an organisation has around its AI. That matters. A bank with weak controls over its AI is exposed.
But a sound control environment does not, on its own, tell you that the answers your AI gave customers or staff last week were correct against your current fee schedule, procedure or product terms. That is a separate question, and it needs separate evidence. SB 813 also makes clear that no organisation is required to engage a verification organisation or undergo an audit simply to develop, deploy or operate AI in California.
So even as the evaluation industry grows, the specific question of whether answers in use are correct may go unasked unless an organisation asks it.
Access decides what can be checked
AB 1405 has a useful detail. From 2029, a registered auditor's report on a covered audit will have to describe its limitations, including any material gaps in the evidence, information, systems or access available to the auditor. An evaluation is only as good as what the evaluator could see.
Anthropic makes a similar point in its announcement. There are no standards yet for what embedded evaluators should be able to access or how they should report, and no settled way to fund independent evaluation. For now, Anthropic will fund Accenture directly, and it would prefer pooled or government funding in the long run.
For an organisation using AI, the records needed to check an answer are its own: the current rates, the updated complaint route, the product terms that changed last month. A model developer does not hold them. So checking answers in use has to start from the organisation's own facts.
Independence comes from structure
Who pays for evaluation will keep coming up. SB 813 takes a practical line. A verification organisation may accept payment from the party being assessed at reasonable market rates, but not on terms that tie payment to the results, and it must stay free of the assessed party's control in reaching its conclusions.
In our own work at Lawnise, we check AI answers an organisation submits or authorises us to sample, against the facts, procedures and public rules that apply to it, and we check what public AI says about it. We keep a record of the question, the answer, the reference used and what the review found. The organisation being assessed does not rewrite an adverse result.
What to ask now
If you are responsible for AI in a regulated organisation, more independent evaluation is coming, both of the models you use and of the controls around them. That is good news. A useful question to add is what evidence you have about the answers your AI actually gave, and whether anyone has checked them against your current records.
Lawnise. Independent evaluation for AI in use.
How to cite this
- Short form
- Lawmence Wong. (2026). Independent AI evaluation is expanding. Who checks the answers in use? Lawnise. https://www.lawnise.com/research/independent-ai-evaluation-expanding-answers-in-use
- Long form (APA)
- Lawmence Wong. (2026, September 28). Independent AI evaluation is expanding. Who checks the answers in use? Lawnise. https://www.lawnise.com/research/independent-ai-evaluation-expanding-answers-in-use
- BibTeX
@misc{lawnise2026independentaievaluationexpandinganswersinuse, author = {Lawmence Wong}, title = {Independent AI evaluation is expanding. Who checks the answers in use?}, year = {2026}, publisher = {Lawnise}, url = {https://www.lawnise.com/research/independent-ai-evaluation-expanding-answers-in-use} }
References
- [1]Office of the Governor of California, Governor Newsom signs first-in-the-nation AI safeguards (9 Sep 2026). Announces the signing of SB 813 and AB 1405 on independent AI verification organisations and an AI auditor registry. https://www.gov.ca.gov/2026/09/09/governor-newsom-signs-first-in-the-nation-ai-safeguards-to-protect-californians-calls-on-the-federal-government-to-do-its-part/(accessed 2026-09-28)
- [2]California SB 813 (Chapter 179, Statutes of 2026). Framework for designating independent verification organisations to assess AI systems and models. https://leginfo.legislature.ca.gov/faces/billTextClient.xhtml?bill_id=202520260SB813(accessed 2026-09-28)
- [3]California AB 1405 (Chapter 178, Statutes of 2026). Establishes an AI Auditor Registry and standards for registered AI auditors. https://leginfo.legislature.ca.gov/faces/billTextClient.xhtml?bill_id=202520260AB1405(accessed 2026-09-28)
- [4]Office of the Governor of California, Executive Order N-9-26 (18 Sep 2026). Directs accelerated implementation of SB 813 and AB 1405 and convenes experts on further AI safety measures. https://www.gov.ca.gov/2026/09/18/governor-newsom-issues-executive-order-to-accelerate-independent-oversight-and-advance-the-creation-of-an-ai-kill-switch/(accessed 2026-09-28)
- [5]Anthropic, Partnering with Accenture on embedded evaluation (18 Sep 2026). Announces Accenture's embedded evaluation work on Anthropic's frontier models. https://www.anthropic.com/news/accenture-embedded-evaluation(accessed 2026-09-28)