What happened
MLCommons' AI Risk & Reliability working group published/updated (page dated 30 September 2026) a Privacy and Confidentiality workstream whose flagship initiatives are (1) a shared Privacy and Confidentiality Risk Taxonomy for describing privacy risks in agentic AI deployments and (2) planned 2027 Privacy Benchmarks measuring how well an AI agent mitigates known privacy risks, covering data-minimisation by autonomous agents and memorisation/recall reduction in model training.
Why it matters
There is currently no standard framework for measuring privacy risk in frontier and agentic AI. MLCommons — the body behind industry-standard MLPerf benchmarks — building an open risk taxonomy and agentic privacy benchmarks gives governments and enterprises a vendor-neutral measurement baseline, complementing NIST/OWASP control frameworks.
Action needed
Organisations deploying agentic AI should monitor the taxonomy and benchmark deliverables; security and privacy practitioners can join the working group and feed requirements on agent data-minimisation and memory/recall risk.