Guidelines  ·  2026-10-04

MLCommons AI Risk & Reliability Working Group — new Privacy and Confidentiality workstream (risk taxonomy + agentic AI privacy benchmarks)

GuidelinesMedium impactGlobal
MLCommons' AI Risk & Reliability working group published/updated (page dated 30 September 2026) a Privacy and Confidentiality workstream whose flagship initiatives are (1) a shared Privacy and Confidentiality Risk Taxonomy for describing privacy risks in agentic AI deployments and (2) planned 2027 Privacy Benchmarks measuring how well an AI agent mitigates known privacy risks, covering data-minimisation by autonomous agents and memorisation/recall reduction in model training.
There is currently no standard framework for measuring privacy risk in frontier and agentic AI. MLCommons — the body behind industry-standard MLPerf benchmarks — building an open risk taxonomy and agentic privacy benchmarks gives governments and enterprises a vendor-neutral measurement baseline, complementing NIST/OWASP control frameworks.
Organisations deploying agentic AI should monitor the taxonomy and benchmark deliverables; security and privacy practitioners can join the working group and feed requirements on agent data-minimisation and memory/recall risk.
MLCommons — Privacy and Confidentiality Working Group
See this in the live feed Explore related AI security and governance findings — updated every morning.
Open the feed →