Financial Services

Develop standards for trustworthy AI in financial services by convening model developers, the financial services industry, regulators, academics, and non-profits. Our goal is to help institutions and regulators confidently adopt AI to build a better, fairer, and more resilient financial system, efficiently and effectively. 

Purpose


AI agents don’t just answer questions, they take actions. That makes them powerful, but it also raises a challenging question: How do you know if an AI agent will stay in its lane?

Financial institutions, regulators, enterprises, and consumers need assurance: 

– That an agent with access to sensitive data will keep it confidential.
– That an agent tasked with making recommendations doesn’t make a payment.
– That an agent won’t drift outside its intended scope over time. 

Agentic deployments vary enormously in their design, capabilities, risks, and blast radiuses. But there are currently no standards, no common assurance mechanisms, and limited transparency into the risk and reliability profiles of agents. Today, users, counterparties, and regulators have no choice but to rely on the judgment, competence, and controls of those developing and deploying agents. 

This gap creates a ceiling on responsible AI adoption. Institutions that could safely deploy agents hold back for lack of a credible way to demonstrate reliability to themselves, to partners, and to regulators. 

Deliverables


The working group addresses this gap through three initiatives:

1. Agent Reliability Profiles 

Our flagship initiative is developing a shared framework and toolset for describing, validating, and benchmarking the reliability of agentic deployments. It will be a standardized, regulation-compatible label for AI agent risk and reliability in financial services. 

2. Federated evaluation of fraud & financial-crime models

Financial institutions could fight fraud and other financial crimes more effectively by working more closely together to stay ahead of emerging crime patterns, but they are held back by competitive and privacy concerns. The working group is building on MLCommons’ existing federated-evaluation infrastructure to enable secure, confidential-compute evaluation of fraud and AML models across institutions.

3. Benchmarks for consumer financial-advice chatbots

Consumers are increasingly turning to AI chatbots for financial advice, but there are no standardized benchmarks for their accuracy, factuality, safety, or privacy. The working group will partner with non-profit organizations building these needed evaluations by contributing technical advice, engineering resources, and benchmark development expertise.

Meeting Schedule

Wednesdays Weekly on Wednesdays 12 NOON to 1 PM ET

How to Join and Access Resources




Financial Services

Mike Hsu

Michael J. Hsu served as Acting Comptroller of the Currency (OCC) from May 2021 to February 2025. There he also served as a Director of the FDIC and member of the Financial Stability Oversight Council. Mike has worked at the Federal Reserve, Securities and Exchange Commission, U.S. Treasury Department, and IMF. In addition to co-chairing the ML Commons Financial Services Working Group, Mike is currently a fellow at the University of Cambridge Judge Business School and the Aspen Institute, Executive Advisor to FINOS, policy advisor to Stripe and Anthropic, board member of the Financial Health Network, venture partner at Core Innovation Capital, and informal advisor to a range of companies, startups, and central banks. He researches and writes on AI in financial services and financial regulation and supervision.

Medha Bankhwal

Medha Bankhwal is Co-founder & CEO of Mezuro. She was previously an Associate Partner at QuantumBlack, AI by McKinsey, where she co-led the firm’s AI Trust practice working on AI transformation and risk management in enterprises across sectors. Medha also was a Risk Leader for transformation of McKinsey’s internal Agentic Governance team that reports to the Chief Risk Officer. Medha holds an MBA from The Wharton School and served as Visiting Lecturer for AI and Societal Impact at the University of California, Berkeley. She has contributed to AI research at McKinsey and others such as Stanford’s 2025 AI Index Report, Google.org and MLCommons Agent Product Maturity.

Deborah Eng

Deborah Eng leads external engagement on AI and Data Industry Standards and Governance at JPMorganChase. Since joining the firm in 2017, she has also helped lead the Technology Cyber Policy & Partnerships group. She currently chairs the Global Standards Committee at the Cyber Risk Institute and previously chaired the U.S. Financial Services Sector Coordinating Council’s International Policy Committee, as well as serving as a U.S. representative to the G-7 Cyber Expert Group. Prior to JPMorganChase, she was Chief Operating Officer at The Chertoff Group and held roles at the U.S. Department of Homeland Security, the White House, and Wired Magazine. She holds a dual Master’s in Cybersecurity from New York University’s Schools of Law and Engineering and a B.A. in International Relations from the University of Pennsylvania.

Questions?

Reach out to us at [email protected]

Privacy Overview

This website uses cookies so that we can provide you with the best user experience possible. Cookie information is stored in your browser and performs functions such as recognising you when you return to our website and helping our team to understand which sections of the website you find most interesting and useful.