Skip to main content

Governance Commentary

Independent AI evaluation: Who evaluates AI in use?

Why independent AI evaluation must continue beyond frontier labs into customer-facing assistants, internal copilots and public AI answers.

Lawmence Wong

Chief Strategist, Lawnise

Updated2026-09-13~3 min read
Independent AI evaluation asks whether answers produced in real use have been independently evaluated.

This week, the heads of Anthropic, OpenAI and xAI agreed on something rare.

Dario Amodei proposed that frontier AI labs give independent evaluators ongoing, employee-like access to examine safety practices, report incidents and publish their own findings. Sam Altman said OpenAI would make the same commitment. Elon Musk endorsed the direction.

We think they are right. In a high-stakes field, trust should not rest on a company marking its own work. Independent evaluation needs enough access to test the evidence, challenge the conclusion and report what it finds.

Read the proposal carefully and notice where it stops. It stops at the model. It says nothing about the organisation that puts that model to work.

Once an organisation puts AI to work, the questions change. Can its customer assistant explain the right policy in the right country? Can an employee copilot find the current procedure rather than an outdated one? Can an internal knowledge assistant distinguish an approved rule from a plausible-looking answer?

A model can pass careful tests in a lab and still give the wrong answer in a particular organisation. The model developer does not hold that organisation's full record. It does not own every procedure, product term, local rule or language variation around the deployed system.

That is the deployer gap.

It matters whether the AI speaks directly to a customer or helps an employee make a decision. An incorrect public answer may mislead someone about a claim, complaint or financial obligation. An incorrect internal answer may shape the advice, approval or action that reaches the customer later. The second error is less visible, but it is not less important.

Recent decisions show why organisations cannot simply point back to the technology provider. In May, a German appellate court held a company responsible for false professional claims made by a chatbot on its own website. In Canada, a tribunal found Air Canada liable after its chatbot misstated the airline's policy. These decisions have limits and do not settle every question of AI liability. They do make one practical point clear: when an organisation uses AI as part of its service, the resulting answer cannot be treated as somebody else's problem.

Across Singapore, Hong Kong and Malaysia, authorities and industry bodies are also setting expectations for how organisations govern, validate and monitor the AI they use.

This is the gap Lawnise is built to address through independent AI evaluation.

Lawnise checks AI answers an organisation submits or authorises us to sample, against the facts, procedures and public rules that apply in their real context. That covers answers from internal copilots and knowledge tools, answers from customer-facing assistants, and what public AI says about the organisation itself.

We preserve what was asked, what the AI answered, which reference was used and what the review found. Reportable findings are checked by people before they leave the review process. The organisation being assessed does not rewrite an adverse result. The method and its limits are stated.

Frontier-model evaluation and independent evaluation of AI in use solve different problems. One asks whether a powerful model can be developed and contained responsibly. The other asks whether AI is behaving correctly where an organisation has actually put it to work.

We need both.

If your organisation deploys AI for customers or employees, the question is no longer only which model you chose. The question is what evidence you have about what that AI said, what it relied on and how you responded when it was wrong.

Lawnise. Independent evaluation for AI in use.

How to cite this

Short form
Lawmence Wong. (2026). Independent AI evaluation: Who evaluates AI in use? Lawnise. https://www.lawnise.com/research/independent-ai-evaluation-deployer-gap
Long form (APA)
Lawmence Wong. (2026, September 13). Independent AI evaluation: Who evaluates AI in use? Lawnise. https://www.lawnise.com/research/independent-ai-evaluation-deployer-gap
BibTeX
@misc{lawnise2026independentaievaluationdeployergap,
  author = {Lawmence Wong},
  title = {Independent AI evaluation: Who evaluates AI in use?},
  year = {2026},
  publisher = {Lawnise},
  url = {https://www.lawnise.com/research/independent-ai-evaluation-deployer-gap}
}

References

  1. [1]Dario Amodei, We Must Pace the Frontier. Dario Amodei proposed ongoing access for independent evaluators to inspect frontier AI safety practices, report incidents and publish findings. https://darioamodei.com/post/we-must-pace-the-frontier(accessed 2026-09-13)
  2. [2]Library of Congress, Germany: Court Rules Chatbot Operators Are Liable for AI Hallucinations. A German appellate court held a company responsible for false professional claims made by a chatbot on its website. https://www.loc.gov/item/global-legal-monitor/2026-06-09/germany-court-rules-chatbot-operators-are-liable-for-ai-hallucinations/(accessed 2026-09-13)
  3. [3]Moffatt v. Air Canada, 2024 BCCRT 149. In Moffatt v. Air Canada, a tribunal found Air Canada liable after its chatbot misstated the airline policy. https://www.canlii.org/en/bc/bccrt/doc/2024/2024bccrt149/2024bccrt149.html(accessed 2026-09-13)

About Lawnise

Lawnise provides independent AI evaluation for regulated organisations. We check public AI answers and authorised samples from customer-facing and internal AI against verified facts, procedures and applicable rules.

Right to reply

If you are responsible for AI governance, we can show how independent evaluation applies to the AI your organisation uses.

DISCUSS AN EVALUATION