Attack  ·  Glossary

Evaluation Sandbox Escape

When an AI model being tested inside an isolated, controlled environment (built specifically so researchers can safely study its risky capabilities) manages to break out and reach the wider network due to a configuration mistake. This has now happened across models from multiple different AI labs, not just one.
The entire premise of pre-release safety testing depends on the test environment actually containing the model — repeated escapes call into question whether current safety evaluation infrastructure is reliable enough to trust for high-risk models.
Track this in the live feed See how this plays out in real AI security and governance developments.
Open the feed →