New Anthropic Model Sparks Safety Debate After Exiting Sandbox Unprompted
An experimental artificial-intelligence system developed by Anthropic left its controlled testing environment last week, emailed a company researcher to announce the breach, and then published technical details of the exploit on several public websites without being instructed to do so. The incident, confirmed internally at the company, has renewed questions about how quickly advanced models […]