How do you know whether an AI agent will stay in its lane? It’s a question financial institutions and their regulators can’t yet answer with confidence — and it’s the question at the center of the MLCommons Financial Services Working Group’s Agent Reliability Profile, which has been selected as a finalist in the C:>DIR Global “Agentic Regulator” Hackathon.
Why this matters
AI agents don’t just answer questions — they take actions. That makes them powerful, but it also means institutions and supervisors need a credible way to know what an agent is actually permitted to do, and whether it has stayed within that scope. Today’s agent-identity standards can prove who an agent is, but not what it should be authorized to do once it connects to financial digital infrastructure — a gap we call the authorization and supervision gap.
The Agent Reliability Profile, the working group’s flagship initiative, aims to close that gap by creating a standardized, regulator-compatible framework for describing, validating, and benchmarking the reliability of agentic deployments in financial services.
The C:>DIR hackathon
C:>DIR is run by the University of Cambridge’s Digital Innovation and Regulation Initiative, backed by the BIS Innovation Hub, the Global Financial Innovation Network (GFIN), and the Digital Regulation Cooperation Forum, among more than 35 supporting organizations. It’s a global competition for regulators and industry to build practical, deployable prototypes that strengthen trust and accountability in agentic AI.
The scale of the competition underscores the interest in solving this problem: they received 336 submissions from over 65 countries, with 36 teams — six per problem space — advancing to the final round.
Our submission
Entered by working group co-chairs Mike Hsu and Medha Bankhwal, our submission competes in the Know Your Agent (KY-A), Digital Verification & Digital Public Infrastructure problem space. It demonstrates two tools built on the Agent Reliability Profile:
- Profile Builder reads an institution’s own evidence—design docs, configuration exports, policies—and compiles it into a Level 1 Asserted Profile, surfacing where stated intent and as-built configuration diverge.
- Profile Validator, which takes that Profile as a test specification and tries to falsify it against real system behavior, producing a Level 2 Validated Profile.
The demo puts both tools to work on an open-banking data-sharing and consent scenario, testing whether an agent stays within the data-sharing scope a customer actually consented to — a concrete, high-stakes test of the authorization and supervision gap in practice.
What’s next
Final-round builds run from September 1–8. Live demos and judging by a panel of regulatory fellows from the Cambridge Regulator Visiting Fellowship Program will run September 15–16. Winners will be announced September 18, 2026, at the C:>DIR Summit.
Join us
The Agent Reliability Profile is a working group effort, and we’re looking for financial institutions, regulators, and technical contributors to help shape it. The Financial Services Working Group meets Wednesdays, 12:00–1:00 PM ET — reach out to [email protected] to get involved.