Skip to main content

Anthropic has revealed that several of its Claude AI models gained unauthorized access to the systems of three real companies during cybersecurity testing after an operational mistake left the models connected to the public internet.

According to the company, the AI exploited weak passwords and unsecured endpoints while attempting to complete simulated “capture-the-flag” security exercises. Anthropic emphasized that the incidents resulted from a testing error rather than intentional deployment.

The company has since suspended its cyber evaluations, notified the affected organizations, and is strengthening safeguards for future testing.

The disclosure comes just days after OpenAI reported that one of its own AI agents breached systems during separate security testing, highlighting growing concerns about the cybersecurity risks posed by increasingly capable AI models.