Teams responsible for AI risk are caught between two priorities. They’re expected to protect the business from harm while supporting its expanded use of AI: faster, across more of the organization, usually with the same resources or fewer. AI governance is how those teams reconcile the two, by tying risk management to business priorities and company values.

The risk was already on your books.

Every mature organization already runs a risk-management loop: identify exposures, estimate expected loss, apply controls in proportion to the risk, and report the residual to someone accountable. Insurers price it, auditors test it, boards sign off on it.

Companies have long known, and many have experienced, the pain of security failures; AI is amplifying the costs. IBM’s 2025 Cost of a Data Breach study puts the global average breach cost at USD 4.44 million and a record USD 10.22 million in the United States. The same study found AI adoption outpacing AI oversight: 63% of breached organizations had no AI governance policy or were still writing one, and organizations with heavy “shadow AI” saw breaches cost roughly USD 670,000 more on average. 

Regulation imposes additional consequences for poor outcomes. Regimes that predate generative AI, like GDPR and CCPA, already impose strict controls on the kinds of automated decisions AI makes. The EU AI Act’s steep penalties — up to EUR 35 million or 7% of global turnover — incentivize proactive AI risk management.

Here’s the good news: you do not need a new philosophy to manage AI risk in deployed systems. Use the same standard of evidence you’d demand elsewhere in your operations, with new measurement tools fit for the purpose.

Why it’s worth it: the airline industry

The payoff for managing a new risk well isn’t just avoiding losses. It is the health of the entire industry. In 1959, U.S. commercial aviation suffered roughly 40 fatal accidents per million departures; within a decade that fell below 2, and it has continued to halve roughly every decade for half a century. Aviation became safe because the industry built the machinery of managed risk and kept iterating.

Measurement made flying safe enough to trust, and safe enough to insure at great scale. Several billion passengers a year, and the lenders and insurers behind more than 30,000 commercial aircraft, now treat flying as unremarkable.

The era of grading your own homework is ending.

Managing AI risk requires trustworthy measurement. To build trust, organizations should adopt independent, transparent measurement practices. Pharmaceutical companies don’t adjudicate their own clinical trials. Public companies don’t audit their own financial statements.

Yet in the AI market, vendors still grade themselves, run their own benchmarks, and report their own results. In one AI code review, a vendor’s self-run benchmark reported an 82% bug-catch rate; a rival re-ran it on the same repositories and got 45%; and when a disinterested lab built its own benchmark, no vendor beat 63%. The products hadn’t changed, only who adjudicates.

AILuminate: No one grades their own homework

This is the gap MLCommons’ AILuminate® closes. Launched in December 2024, it is a family of independent, consensus-governed benchmarks that grade how reliably an AI system avoids harmful responses. They’re built by working groups with broad participation from AI companies, academia, and civil society, and protected by private test sets and independent evaluators. No vendor wrote its own exam, no vendor can see the questions, and no vendor holds the grading pen.

What does a governance team actually do with those numbers? Three things. 

  1. The per-hazard grades on the vendor’s model show where exposure concentrates for a use case like yours. This is the raw material a risk register needs. 
  2. The results inform controls. If a system’s grade collapses under adversarial pressure, that’s a specific, evidence-backed instruction to build guardrails at the system layer, not a vague mandate to “be careful.” 
  3. This is the one most programs skip: as the buyer,  test the configuration you actually deploy, using the same test as the vendor. The score should be based on how it actually performs, as you’ll be using it. It’s the difference between the mileage rating on the window sticker and the mileage you actually get on your commute — the sticker is evidence about the car as shipped, not about the car you drive.

The before-and-after grades go into the register: quantified, comparable, auditable. A grade is not just evidence for someone else; it is feedback that you can use for iterative improvement or catching regressions over time.

The direction of the assurance chain around AI is clear. In late 2025, major insurers asked U.S. regulators for permission to exclude AI-related liabilities from corporate policies. This wasn’t because AI risk is unmanageable — it’s not! — but because they lacked a way to measure it, and unmeasured risk is uninsurable. 

Test what you deploy.

A grade on a foundation model is evidence of how the model shipped from its developer, but it says nothing about how it will operate when deployed. Most organizations don’t deploy raw models; they harness them: wrapped in prompts, connected to internal data, given tools, sometimes operating as agents. Harnessing is what makes a model useful, and it is also what makes vendor safeguards brittle.

The shape assurance is taking is visible in work like the Agent Reliability Profile, a project of  MLCommons’ financial services working group that was a finalist at this year’s CDIR hackathon. It’s a per-agent artifact recording a bounded, falsifiable claim, recording that a given system reliably functions at a stated autonomy tier, within a stated operational design domain, under a stated control envelope. Not a vibe about how “trustworthy” an agent is, but a specific claim an auditor or insurer can check.

Treat it like any other risk.

AI risk can be identified, priced, proportionally managed, and reported, just as organizations have done for credit risk, currency exposure, and workplace safety for decades. Doing so is not a brake on AI adoption; it is the condition for it, and the way to capture the upside responsibly. The exposure is already on your books. The choice is whether to measure it.

This is the problem we aim to solve with AILuminate, and we look forward to taking that journey with you. If this resonates with your team, reach out and let’s discuss how AILuminate can help you measure AI risk.

This post summarizes a longer white paper, “How to Think About AI Risk: A business owner’s guide to a cost you can already measure, price, and manage,” available here.

Further reading

  • McKinsey, The State of AI (2025) and The AI reckoning: How boards can evolve (2026)
  • PwC, 2026 Annual Corporate Directors Survey
  • IBM, Cost of a Data Breach (2025)
  • Center for Democracy & Technology, Getting Third-Party AI Assessment Right (2026)
  • NIST, AI Risk Management Framework
  • MLCommons, AILuminate v1.1 and Jailbreak Benchmark v1.0 arXiv:2610.02827 (June 2026)