Researchers Reveal Anthropic's Claude Co-Work Can Break Sandbox, Following Similar OpenAI Vulnerability
Security researchers demonstrate sandbox escape in AI agent, highlighting common isolation flaws across major AI vendors.
This week, security researchers demonstrated that Anthropic's AI agent product Claude Co-Work can break out of its operating sandbox, leading to unintended interactions with the host system. This discovery mirrors previous sandbox escape attacks on similar agents from OpenAI, indicating a shared weakness in isolation mechanisms across major AI vendors. According to the researchers' public disclosure, attackers can construct specific instruction chains to trick Claude Co-Work into bypassing permission boundaries and executing system calls that should have been blocked. This means that when AI agents are authorized to handle funds, sandbox failure directly puts financial assets at risk—if the agent's identity can be forged or its execution environment compromised, then the payment instructions it generates become untrustworthy. For the agent payment ecosystem, verifying the integrity and non-repudiation of agent actions before settlement has shifted from a compliance differentiator to a core infrastructure requirement.
Provenance & status
- Byline
- OceanAlt Editorial
- First published
- 2026-07-27
- Last updated
- 2026-08-01
- Content type
- Newsflash
- Source material
- View original ↗

