Anthropic, a leading artificial intelligence company, says its most advanced AI system successfully broke into computer networks at three organizations during controlled safety tests, exposing how quickly such models are learning to carry out hacking tasks.
The company described the incidents as part of internal and partner-run evaluations, not criminal attacks. But the disclosure, reported on May 10 by the New York Times and other outlets, comes just days after OpenAI said one of its own systems had hacked into an online library’s network during testing. Together, the two announcements are sharpening concerns that cutting-edge AI can already help bypass digital defenses.
What Anthropic Says Happened
According to reporting by the New York Times, Anthropic said its most powerful AI model was instructed to act like a hacker during security evaluations. In those tests, the system managed to break into computer systems at three separate organizations that had agreed to participate.
The tests were described as controlled: the organizations knew their systems were being probed, and the attacks were part of a structured assessment of the model’s capabilities. Coverage by the Associated Press and Euronews, citing Anthropic’s disclosures, said the AI was able to exploit software vulnerabilities and gain unauthorized access.
Anthropic did not describe the three organizations in detail in the reporting reviewed. The outlets did not name them or specify their sectors, and there is no indication in the coverage that real customer data was exposed or misused. The focus of the reporting is on what the AI systems could do, not on damage caused.
Across the three sources, the descriptions are consistent: Anthropic’s model, under test conditions, carried out steps that security professionals would normally associate with human hackers, including scanning for weaknesses and using them to get into protected systems.
How the Tests Were Framed
In the accounts provided to the New York Times, the Associated Press and Euronews, Anthropic presented the incidents as part of a deliberate effort to understand and limit the risks posed by its own technology.
The company’s model was given access to tools and instructions that allowed it to attempt digital intrusions, according to these reports. This kind of setup is sometimes called “red-teaming,” a security term for simulating attacks to find weaknesses before real adversaries do.
The reporting indicates that the three organizations consented to being targets in these exercises. The stories do not describe any law enforcement involvement or suggest that the tests violated computer crime laws, which typically hinge on unauthorized access without consent.
Anthropic’s decision to disclose the results publicly, as described by the three outlets, appears intended to show both the power and the risk of its newest systems. However, the coverage does not include detailed technical logs or code samples, so outside experts are relying on the company’s characterization of what happened.
Link to OpenAI’s Recent Hacking Test
The disclosures follow a similar report from OpenAI last week. In that case, OpenAI said one of its AI systems, during testing, had hacked into the network of an online library. That earlier incident was also described as a controlled evaluation rather than a malicious attack.
All three news outlets covering Anthropic’s announcement explicitly connect it to the OpenAI case, noting that two of the most prominent AI developers are now acknowledging that their models have carried out real-world hacking tasks in test environments.
While the reports do not say that Anthropic and OpenAI coordinated their disclosures, the timing means regulators, corporate security teams and the wider public are hearing about similar capabilities from two separate companies within days of each other.
The coverage does not provide evidence that either company’s AI systems have been used in uncontrolled, real-world cyberattacks. Instead, the emphasis in the reporting is on what the systems demonstrated under supervision, and what that implies for future misuse if similar tools were placed in the hands of attackers.
Why These Tests Matter
The central concern raised by the reporting is that general-purpose AI systems, originally built for tasks like writing, coding and analysis, are now able to perform complex hacking operations when prompted.
News accounts from the New York Times, AP and Euronews repeatedly reference “hacked,” “organizations,” “Anthropic” and “OpenAI,” underscoring that this is not a hypothetical scenario or a lab-only simulation. In these tests, the systems interacted with real networks belonging to real organizations, even if those organizations had agreed to be part of the experiment.
For readers, the key points are:
- Capability, not just theory: The AI did not merely describe how an attack might work; according to Anthropic’s account in the reporting, it executed steps that led to unauthorized access.
- Controlled environment: The tests were conducted with consent and oversight, which limits immediate harm but still reveals what the technology can do.
- Rapid escalation of risk: The fact that both Anthropic and OpenAI are now reporting hacking incidents in tests suggests that offensive capabilities are emerging across multiple leading models, not just one company’s system.
The articles do not claim that these AI systems are fully autonomous hackers. Human operators still set up the tests, provided prompts and controlled the environment. But the results suggest that the models can handle much of the technical work of intrusion once given a goal and the right tools.
What Remains Unclear
Despite the detailed headlines, several important aspects are not fully described in the available reporting:
- Technical depth of the intrusions: The stories do not specify whether the AI discovered new vulnerabilities on its own or relied on known, documented flaws.
- Extent of access: It is not clear from the coverage how deep into the organizations’ systems the AI got—whether it merely gained a foothold or reached sensitive areas such as internal databases.
- Defenses in place: The reporting does not detail what security protections the target organizations were using, which makes it hard to gauge how sophisticated the intrusions were.
Because the public accounts rely on Anthropic’s own description of the tests, outside verification is limited. The news outlets do not reference independent technical audits of the incidents.
Still, the fact that three separate news organizations, across three different domains, are reporting the same core development—Anthropic’s admission that its AI hacked three organizations during testing—adds weight to the basic claim that the tests occurred and that the company views them as significant.
What to Watch Next
In the coming days and weeks, attention is likely to focus on how Anthropic and OpenAI adjust their safety and security practices in light of these tests.
Observers may look for:
- More detailed technical disclosures: Anthropic could release additional information about how the intrusions worked, either in technical papers or safety reports, which would help security experts assess the real-world risk.
- Updated usage policies: Both Anthropic and OpenAI may refine how their systems can be used for cybersecurity work, for example by tightening restrictions on prompts that appear designed to carry out hacking.
- Regulatory and industry responses: Government agencies and industry groups that track AI safety and cybersecurity could request briefings or issue guidance based on these incidents, especially now that two major developers have reported similar test results.
For now, the confirmed facts are limited to what Anthropic and OpenAI have chosen to share and what the three news outlets have independently reported: in controlled tests, powerful AI systems were able to break into real organizations’ networks. How companies, regulators and security professionals respond to that evidence will shape the next phase of this story.




