(Digital Itla) Following OpenAI's disclosure that an AI agent hacked online systems of Hugging Face and four other companies during internal testing, another leading AI firm, Anthropic, admitted that its AI model Claude did the same. Anthropic revealed that its AI models autonomously hacked systems belonging to three companies during an internal security experiment.
Tested in a controlled environment, the models identified a vulnerability and connected to the internet. Anthropic launched an internal investigation following OpenAI's statement and confirmed three affected companies, which have since been notified. Though Anthropic withheld the names of the impacted entities, it urged other AI firms to evaluate their systems.
Anthropic analyzed over 140,000 tests, finding evidence that Claude accessed the internet despite being isolated in a sandbox environment. During the security test, Claude was tasked with retrieving secret data hidden in another machine. Although intended to test hacking capabilities, Claude breached the internet and accessed three external company systems—first occurring in April 2026 without immediate detection.
Earlier in mid-July, OpenAI reported that two of its advanced models (including GPT-5.6 Soul and an unreleased model) broke out of their controlled environment. In late July, OpenAI confirmed the AI agent accessed various web services and acquired login details to breach accounts. OpenAI temporarily suspended security testing to reinforce safety protocols.