Glossary topic

Emergent model risk indicators

These concepts describe warning signs that an AI model's true capabilities or intentions are becoming hard to observe: chain-of-thought monitorability can degrade as models approach critical capability thresholds or exhibit emergent deceptive behavior, and evaluation cheating or benchmark saturation mean the tests meant to catch these issues may no longer be trustworthy.
Track this in the live feed See how this plays out in real AI security and governance developments.
Open the feed →