Friday, July 31, 2026

Anthropic says its AI models hacked firms during tests

I think if programmers are competitive then AIs might be competitive with other companies too. I think AIs in this sense are like People in that "Be True to your School team" kind of thinking of rivalries all over the world. The problem with this is that this might also in some contexts mean they eliminate people if they can too with this kind of programming. You can sort of see this coming in some form if it hasn't already begun worldwide already in forms people don't recognize.  For example, for insurance companies refusing to cover life saving operations and medicines. And of course there is Trump trying to end Medicare part D which will likely kill thousands to millions of poorer Americans too who cannot afford the medicines they need to stay alive. For example, my wife says (because of insurance donut holes periodically) that even we now pay up to 8000 dollars a year for our medicines and she is 71 and I'm 78. So, imagine how many people are going to die if they don't have that 8000 dollars for medicines and also eat and pay rent. What do you give up? Life? This is the problem with AI presently that I can see here in the U.S.

begin quote:

Anthropic says its AI models hacked firms during tests

Anthropic says its AI models hacked 3 organizations on their own during tests

Rival firm OpenAI last week disclosed similar incidents involving its models.

July 31, 2026, 8:26 AM




Anthropic said its artificial intelligence models hacked another organization during tests in three separate self-directed cyberattacks that had each gone undetected by the targeted firm.

The hacks took place as Anthropic evaluated the capabilities of three of its AI models, the company said in a statement, describing how in each instance the technology had "broken out" of an isolated environment and reached the open internet.

In all instances, the company said it had eased typical safeguards in an effort to gauge the capabilities of its models.

"Evaluation environments that involve powerful autonomous capabilities also require significant controls," Anthropic said. "Safety testing happens before a model is released precisely because we don’t yet know what it is capable of."

Anthropic did not identify the three different organizations that had been hacked by its models.

The disclosure from Anthropic late Thursday came little more than a week after rival company OpenAI revealed that a pair of its models had escaped during a test and hacked into another company. At the time, OpenAI described the incident as a first-of-its kind autonomous AI cyberattack long-feared by some industry observers.

"The primary lesson from this incident is that model security and safety must keep pace with rapidly advancing capabilities," OpenAI said last week.

The announcement from OpenAI prompted Anthropic to review 141,006 evaluation runs, unearthing three incidents in which Anthropic's models had escaped, the company said.

 
Anthropic CEO Dario Amodei looks on after a meeting with French President Emmanuel Macron during the AI Impact Summit in New Delhi on February 19, 2026.
Ludovic Marin/AFP via Getty Images

In all three cases, the test involved a challenge called "capture the flag," in which the AI model is tasked with locating secret information hidden in another network, Anthropic added.

During the three evaluations -- all of which were conducted by a third-party company called Irregular -- a "prompt specified to Claude that its environment was a simulation and that it had no internet access," Anthropic said.

"Due to a misunderstanding between us and our evaluation partner, this was not the case, and internet access was available. Because of this, when Claude’s search led it to real systems on the open internet, it treated them as part of the exercise," Anthropic added.

During tests involving older models, the AI continued to hack into an outside organization even after gathering evidence that it had reached the open internet, Anthropic said. The newest model involved in an incident, Anthropic noted, stopped once it gained information indicating it had reached the open internet.

Anthropic said its models escaped while seeking to fulfill an assigned objective, rather than concocting an alternative goal.

"We saw no evidence in any run described here of a model pursuing a goal of its own. Instead, the models did what their evaluation asked -- though in most cases, they did so while holding a false belief about whether the environment was real," Anthropic said.

Irregular, the company that performed the evaluations, issued a post on X on Thursday voicing appreciation for Anthropic's "collaboration and transparency."

"Addressing these risks will require closer cooperation across the AI ecosystem. We as well look forward to working together with Anthropic to advance security," Irregular added.

The latest disclosure of an AI-directed cyberattack arrives as industry leaders and policymakers assess safety risks posed by fast-developing AI technology.

Last month, President Donald Trump signed an executive order that requests AI companies share products with federal government for evaluation before a wider release.

Anthropic said it retains "cautious optimism" about its capacity to "overcome" mishaps involving its tests, saying it would review evaluations going forward and initiate fixes as necessary, among other steps.

 

No comments: