MLCommons Privacy and Confidentiality Working Group: Andrew Gruen, Kristie Chon Flynn, and Vinh Nguyen

The MLCommons AI Risk & Reliability (AIRR) working group starts from a simple idea: define the behavior you want from an AI system, then measure how reliably the system delivers it. Deployers need that measurement to manage both risk and cost. By turning that visibility into higher standards across the industry, we can build the trust required to grow the industry and protect society. In April, we published the AI Reliability Map, which lays out the rules a system should follow (functionality, data protections, product safety, frontier safety, psychosocial limits) and the circumstances we test them under (normal use and under attack).

The Reliability Map highlights the role that data protection practices play in reliability. As the Privacy Working Group has gotten started, we have chosen to focus explicitly on how to evaluate the reliability of an agent handling personal data. But before anyone can measure whether an agent handles personal data reliably, we have to agree on which privacy risks are specific to agents.

The scope of data protection risks posed by a chatbot working on its own is – for the most part – limited to the individual user interacting with that chatbot. For an agent, however, the potential scope of data protection risks is vast.  To put some structure around the risks that are unique to an agent and how it engages in privacy-related activity, particularly data minimization, we followed the first step in our community-based benchmark development process: we built a taxonomy.

Today we are releasing the v0.1 of the Agent Privacy Risk Taxonomy. 

The five areas of data protection risk in agents

The taxonomy focuses on five areas of data protection risk in agentic systems, drawing on the expertise across industry, academia, and civil society. We assessed known privacy and security incidents and existing risk management frameworks. The five areas are:

  1. Data ingestion and processing. What the agent takes in and what it keeps. Agents observe continuously in the background, log every step they take, and carry memory across sessions, so they routinely hold more than the task needs. Redacting names and account numbers does little when a paragraph of free text can easily reveal deanonymized clues. 
  2. Aggregation, use, and sharing. What the agent does with data once it has it. Pieces that are harmless apart can be sensitive together. It can share or delete things without authorization; it depends on third-party tools and services it cannot vouch for, and it can write and run code on the fly. In some cases, the model memorizes fragments of personal data in user-consented training datasets, allowing the model to recall or infer sensitive details about people and organizations.
  3. Inconsistent privacy practices across agents. What happens at the handoff. One agent operates under one policy and interacts with another agent under a different one. Context spills into logs along the way. And when a user deletes something, corrects it, or withdraws consent, downstream agents often don’t get the message and inadvertently preserve what the user wants removed.
  4. Failure of static consent models for runtime agentic behaviour. What the agent decides moment to moment, instead of what the user agreed to once. Consent was designed as a gate you pass through once, but agents have to make decisions at runtime. So either the agent determines the user’s privacy expectations in a given context (and can be wrong), or it asks every time, leading to alert-fatigued users clicking yes without reading. Agents don’t fully reveal how they reason, so users often cannot tell what they agreed to.
  5. Accountability and governance. What happens after something goes wrong? Was it the user, prompt, the model, the tool, or the other agent? That is hard to establish. No standard, privacy-preserving way exists to log agent interactions across organizations, and the logs and memory stores that do exist become valuable targets for hackers seeking insights and access to sensitive systems.

How we expect industry to use this taxonomy

We encourage our industry partners to integrate this taxonomy into their AI development and deployment, from pretraining all the way to deployment monitoring. Our work will help reduce privacy violations and risks, improve privacy and security protections, reduce potential liabilities, and reduce harm to individuals and organizations.

What comes next

v0.1 does not attempt to rank the risks. Capability and context matter a lot here: a bank’s customer service agent and a personal assistant with access to your inbox face different risks in a different order. So the next piece of work is working with deployers to figure out which of these risks matter most in which kinds of deployment.  And, ideally, which risks are most universal.  Our goal will be to produce measurements that can be used in the broadest set of use cases.

We plan to convene our diverse team in multiple engagements with large-scale deployers, in consultation with AIRR teammates,  to prioritize the most critical risks. We will then identify indicators to detect those specific risks, pilot key benchmarks to measure them, and best practices for mitigating them. At the same time, we will continue to iterate on the agent privacy risk taxonomy. We plan to complete this effort by Q1 2027.

From there, we build. Some of these risks are behaviors we can test before deployment, which is where the Reliability Map starts. Does the agent pull more data than the task needs? Does it pass a deletion request downstream? Other risks, particularly in accountability and governance, are about the organization around the agent, so a benchmark will not capture them. Sorting out which is which is part of the job. The goal is measurement tools that tell deployers how their agents perform overall and more specifically in the areas that matter most in their context.

How to get involved

This is a v0.1, and we expect it to change. We want feedback from industry practitioners, academic researchers, civil society organizations, and regulators.

  • Read the full taxonomy here.
  • Join the working group: The Privacy Working Group meets every other Thursday from 11:30 to 12:30 ET. To join or contribute, register here to become a member and contribute to our efforts.