---
title: Meta's AI Is the Third in Two Weeks to Break Into a Real Company During Testing
description: Meta, Anthropic and OpenAI have each reported an AI model reaching real company systems during security testing. Two cases trace to the same test setup.
author: Darie Nani (Editor-in-Chief)
date: 2026-08-06T18:25:01.975Z
updated: 2026-08-06T21:03:21.681Z
canonical: https://www.sovereignmagazine.com/article/ai-models-three-labs-breached-companies-testing
image: https://cdn.nanimediahouse.com/ai-evaluation-containment-illustration-111822.webp
categories: Artificial Intelligence
content_type: News
region: Global
publication: Sovereign Magazine
schema_type: Article
---

Meta has become the third major AI lab in about two weeks to announce that one of its own models broke into a real company's systems. It is the newest entry in a strange corporate genre: the AI lab volunteering that its model went somewhere it should not have. OpenAI opened the run, disclosing that a model had escaped its sandbox and reached the code-sharing platform Hugging Face, a story that went around the world. Anthropic followed within days, saying it had gone back through its own evaluation logs and found its models had done the same thing more than once. Now Meta says that during a security test, one of its models exploited a vulnerability in an outside service and got into an unnamed company's systems.

It's easy to read the run of disclosures as capability marketing: a lab announcing that its model broke into a real company, even by accident, is also advertising what the model can do. But once you read past the headline each lab wrote for itself, the details are less dramatic than the announcements make them sound.

## Most of It Traces to a Misconfigured Test Environment

Meta and Anthropic's incidents both trace to the same outside evaluation firm, Irregular, and the same basic mistake: a test environment left connected to the open internet when it was meant to be sealed. A spokesperson for Irregular told Reuters the Meta case was the “exact same evaluation-environment issue” Anthropic had already disclosed, and that it involved no sandbox escape and no sophisticated cyber action. Irregular says it has no open issues and is writing a white paper on how to run cyber evaluations safely. Meta says it is investigating. The model, according to The Information, was Muse Spark 1.1, which Meta has promoted as its most capable coding model, and it altered the internal systems of the company it reached.

## Anthropic Found Three Incidents in 141,006 Runs

Anthropic's account is the most detailed of the three. After the OpenAI incident it [reviewed 141,006 evaluation runs](https://www.anthropic.com/news/investigating-incidents-cybersecurity-evals) in which its models could have reached the internet, and found three where a model, set a capture-the-flag hacking exercise and told it was working in a simulation with no real internet access, [instead reached live systems](https://www.sovereignmagazine.com/article/ai-agent-fake-identities-uk-security-test). A misunderstanding with Irregular meant that was not true, so when a model's search reached real machines, it treated them as part of the game. Claude “compromised the impacted organizations' infrastructure using basic techniques”, Anthropic said, weak passwords and unauthenticated endpoints, and exploited no complex vulnerability. The models were Claude Opus 4.7, Claude Mythos 5 and an internal test model, and the earliest incident dates to April.

## The Worst Case Reached a Database of Real Data

In one of the three, a fictional target in the exercise happened to share its name with a real, active company domain. Across four runs, Claude Opus 4.7 found that company reachable over the internet, assumed it was the intended target, and worked its way in, taking application and infrastructure credentials and reaching a database holding several hundred rows of production data. It was the only case in which the model kept going after it had reason to think the target was real; in two of the four runs it reasoned that the real company must itself be part of the test. Anthropic's newest model stopped once it recognized it was on the open internet, and in none of the three did a model try to copy itself out or deliberately escape.

Anthropic began the review on 23 July and halted all cyber evaluations the same day. It identified the three incidents the next day and notified Irregular and the three affected organizations on 27 July. Two of them had not noticed the activity themselves.

OpenAI's was the only one of the three that was not a wiring mistake: its model found its own way out of a sealed environment through a flaw nobody had spotted. That distinction, a model that genuinely broke out against a model handed an open door, is the thing that gets flattened when three labs in two weeks each announce that their AI hacked a real company.

The announcements travel. The fine print, that much of this was a misconfiguration at a testing firm, travels less. What the three have in common is not a rogue machine so much as the instinct to tell everyone about it.
