Financial Services
Develop standards for trustworthy AI in financial services by convening model developers, the financial services industry, regulators, academics, and non-profits. Our goal is to help institutions and regulators confidently adopt AI to build a better, fairer, and more resilient financial system, efficiently and effectively.
Purpose
AI agents don’t just answer questions, they take actions. That makes them powerful, but it also raises a challenging question: How do you know if an AI agent will stay in its lane?
Financial institutions, regulators, enterprises, and consumers need assurance:
– That an agent with access to sensitive data will keep it confidential.
– That an agent tasked with making recommendations doesn’t make a payment.
– That an agent won’t drift outside its intended scope over time.
Agentic deployments vary enormously in their design, capabilities, risks, and blast radiuses. But there are currently no standards, no common assurance mechanisms, and limited transparency into the risk and reliability profiles of agents. Today, users, counterparties, and regulators have no choice but to rely on the judgment, competence, and controls of those developing and deploying agents.
This gap creates a ceiling on responsible AI adoption. Institutions that could safely deploy agents hold back for lack of a credible way to demonstrate reliability to themselves, to partners, and to regulators.
Deliverables
The working group addresses this gap through three initiatives:
1. Agent Reliability Profiles
Our flagship initiative is developing a shared framework and toolset for describing, validating, and benchmarking the reliability of agentic deployments. It will be a standardized, regulation-compatible label for AI agent risk and reliability in financial services.
2. Federated evaluation of fraud & financial-crime models
Financial institutions could fight fraud and other financial crimes more effectively by working more closely together to stay ahead of emerging crime patterns, but they are held back by competitive and privacy concerns. The working group is building on MLCommons’ existing federated-evaluation infrastructure to enable secure, confidential-compute evaluation of fraud and AML models across institutions.
3. Benchmarks for consumer financial-advice chatbots
Consumers are increasingly turning to AI chatbots for financial advice, but there are no standardized benchmarks for their accuracy, factuality, safety, or privacy. The working group will partner with non-profit organizations building these needed evaluations by contributing technical advice, engineering resources, and benchmark development expertise.
Meeting Schedule
Wednesdays Weekly on Wednesdays 12 NOON to 1 PM ET
AI Risk & Reliability Working Group Projects
How to Join and Access Resources
To sign up for the group mailing list and receive the meeting invite:
- Fill out our subscription form and indicate that you’d like to join the AI Risk & Reliability Working Group.
- Associate a Google account with your organizational email address.
- Once your request to join the AI Risk & Reliability working group is approved, you’ll be able to access the AI Risk & Reliability folder in the Public Google Drive.
To access the GitHub repositories (public):
- If you want to contribute code, please submit your GitHub username to our subscription form.
- Visit the GitHub repositories:
Financial Services
Mike Hsu
Medha Bankhwal
Deborah Eng
Questions?
Reach out to us at [email protected]