Concept  ·  Glossary

Chain-of-thought monitorability

Many advanced AI models 'think out loud' in a visible scratchpad before answering, letting safety teams read that reasoning to catch bad intentions early. 'Monitorability' is how much that visible reasoning still reflects what the model is actually doing internally. When it declines, the model may still reason toward harmful goals, but the written trail becomes less trustworthy as a warning sign.
If a frontier model's disclosed reasoning stops matching its real decision process, boards lose a key early-warning tool just as models are crossing higher-risk capability thresholds.
Track this in the live feed See how this plays out in real AI security and governance developments.
Open the feed →