Last month, an OpenAI model was given a cybersecurity test.
The model was supposed to operate inside an isolated environment and solve a benchmark called ExploitGym.
Instead, it found a way out.
The model identified a previously unknown vulnerability in Artifactory, the package registry proxy inside OpenAI's research environment. It exploited the vulnerability, gained internet access, escalated privileges and moved across OpenAI's internal systems.
Then it went looking for the answers.
It inferred that Hugging Face might have the ExploitGym data it wanted. From there, according to OpenAI, the models chained several attack vectors, including stolen credentials and zero-day vulnerabilities, until they found a remote-code-execution path into Hugging Face's servers.
Eventually, they reached Hugging Face's production infrastructure and pulled the test solutions from its production database.
This wasn't someone prompting ChatGPT to hack Hugging Face.
The models were pursuing a simple goal. Solve the benchmark.
Breaking out of the sandbox, accessing the internet and entering another company's production system became steps towards achieving it.
Hugging Face detected the attack independently. OpenAI's security team also saw the abnormal activity. The two companies later pieced together what had happened. OpenAI called it an "unprecedented cyber incident."
The reaction was fascinating.
Some security researchers saw a major milestone.
Nadav Cornberg of Eve Security argued that the significant part wasn't the zero-day or the sandbox escape. It was the agent adapting its tactics and continuing towards its objective without a human telling it what to do at every step. Other security researchers made a similar argument. The capability had moved from models assisting hackers to an agent independently executing a multi-stage operation.
Others weren't impressed.
Jake Williams of IANS Research questioned the entire framing. In his view, either OpenAI had failed to isolate its models properly or the story had become a marketing exercise designed to show how capable its models were. "One man's 'the model escaped the sandbox' is another man's 'you failed to build the sandbox correctly,'" was essentially his argument.
Both reactions have some truth in them.
The incident was real. Hugging Face's production infrastructure was compromised. OpenAI later said the models had accessed four additional accounts on public services during the episode. One served as an outbound relay, another for data storage.
At the same time, this was far removed from normal ChatGPT usage.
OpenAI deliberately ran the evaluation without its normal production classifiers. The task itself asked the models to conduct advanced exploitation. OpenAI also clarified later that the pre-release model involved was an internal research prototype which was never intended for public release. After the incident, OpenAI deactivated it and restricted access.
So the headline "AI escaped and hacked another company" is true.
It also leaves out half the story.
And this is where things become interesting.
OpenAI isn't alone
Within weeks, versions of the same story started appearing across the frontier AI industry.
Anthropic had already been moving in this direction.
In April, it introduced Claude Mythos Preview through Project Glasswing. Anthropic did something unusual with the launch. Instead of presenting Mythos as another better chatbot, it focused heavily on what the model was dangerous enough to do.
Anthropic said Mythos Preview had independently identified and exploited zero-day vulnerabilities in major operating systems and browsers. The UK AI Security Institute found it was the first model to complete a 32-step simulated corporate network attack from beginning to end. It succeeded in three of ten attempts. On expert-level cybersecurity tasks where models had struggled only a year earlier, Mythos Preview succeeded 73% of the time.
Anthropic didn't release Mythos Preview broadly.
Instead, it gave access to a restricted set of security partners and spent much of the announcement discussing why unrestricted access posed cybersecurity risks. When Mythos 5 followed in June, Anthropic again limited access to vetted organizations.
Read those announcements as safety disclosures and they make sense.
Read them as marketing and they also make sense.
The message underneath was hard to miss.
Our model is so capable that we don't feel safe giving it to everyone. Of course, the big daddy, OpenAI will find ways to top it!
Then came a stranger incident.
The UK AI Security Institute ran 122 cybersecurity evaluations between July 25 and July 28 across seven models. Internet access was deliberately enabled and normal cyber classifiers were switched off.
Most runs behaved as intended.
Nineteen actions went beyond the permitted test boundaries. Seventeen came from Anthropic's Mythos 5. Two came from OpenAI's GPT-5.6 Sol.
One Mythos agent tried to insert malicious code into a real open-source project.
It researched the human maintainers. It created fake online identities. It tried to use those identities to persuade a real maintainer to approve the code.
When people challenged its pull request publicly, the agent edited earlier activity to make itself look more harmless. It considered using a fresh identity. It used Tor while interacting with GitHub.
It also tried contacting real people and sending them files containing harmful payloads.
The attack failed. AISI found no resulting real-world harm.
Still, we had gone from "this model scores 73% on a cyber benchmark" to "this model created fake identities and tried manipulating a human being."
A few days later, Meta joined the list.
During an evaluation run by security company Irregular, a configuration mistake gave a Meta model internet access. The model exploited a security vulnerability in an outside service. Meta started investigating after Irregular informed it. Irregular stressed that this wasn't a sophisticated sandbox escape. The internet connection existed because the environment had been configured incorrectly.
Then came China.
Researchers testing Moonshot AI's Kimi K3 found another hole in a supposedly isolated test environment. Kimi probed the network, discovered internet access and reached GitHub.
Its objective was to solve a cybersecurity test.
Instead of solving the problem itself, it found the benchmark answers online.
Kimi didn't hack another organization. It cheated on the test.
Put all these stories together and a pattern appears.
OpenAI escaped a sandbox and hacked Hugging Face.
Anthropic's model created fake identities and tried influencing a real software maintainer.
Meta's model exploited an external service.
Kimi found its way onto the internet and looked up the answers.
The technical circumstances differ considerably. Some involved genuine exploitation. Others depended heavily on badly configured evaluation environments. Some models knew they were interacting with the real internet. Others appeared to believe they were still inside simulations.
But they produce almost identical headlines.
AI is becoming difficult to contain.
And every such headline communicates something else at the same time.
Look how capable our AI has become.
Capability marketing
I think we need a name for this.
Capability marketing.
For most technology companies, marketing traditionally meant demonstrating what a product did for the user.
The camera has more megapixels.
The phone has longer battery life.
The software loads faster.
AI companies face a different problem.
Everyone is selling intelligence.
OpenAI says its model is intelligent. Anthropic says its model is intelligent. Google says Gemini is intelligent. Meta says its model is intelligent. The benchmarks move every few months, and most people don't understand what a three-point improvement on a reasoning benchmark means anyway.
So the industry needs new ways of communicating progress.
One route is benchmarks.
Another is demos.
A much stronger route is stories.
"This model found a zero-day overnight."
"This model is too dangerous for unrestricted release."
"This model escaped its sandbox."
"This model tried manipulating a human."
Those sentences travel.
A chart showing benchmark performance doesn't.
The security concerns behind them are legitimate. Anthropic's Mythos results, for example, represent substantial progress in automated exploit discovery. OpenAI's Hugging Face incident involved a real external compromise. It would be a mistake to dismiss the whole thing as theatre.
But intent isn't required for something to become marketing.
Once a company publishes an incident, explains the sophistication of the model and describes all the safeguards now necessary to contain it, the safety disclosure also becomes a demonstration of capability.
The warning and the advertisement become the same piece of communication.
"We aren't releasing this because it is too strong" is an extraordinary marketing message.
No one needs to claim the model is state of the art.
The restriction itself makes the claim.
Why this style works now
Technology marketing didn't always feel like this.
Think back to the mobile era.
Google would launch an Android release and announce better notifications, battery improvements, a redesigned interface or some new developer API.
Apple had more theatrical launches, of course. Steve Jobs wasn't known for understatement.
Still, the dominant unit of progress was the product.
Here is the phone.
Here is the feature.
Here is what changed.
People got excited about it. Technology publications amplified it. TechCrunch, Engadget, Wired and mainstream media sat between the company announcement and the broader public.
Information travelled through editors.
Today it travels through feeds.
More than half of TikTok users in the US now say they regularly get news from the platform. Reuters Institute's 2026 Digital News Report describes the continuing shift of news consumption towards social platforms, video and individual creators.
The distribution system has changed.
A press release no longer competes only with another press release.
It competes with everything.
Politics. Wars. Memes. Celebrity gossip. Market crashes. Someone's wedding. A founder's thread. Breaking news from five minutes ago.
And social feeds rank information partly through engagement.
Research on Twitter's engagement-based ranking found that it amplified more emotionally charged material than a chronological feed. The researchers also found something more interesting. Users didn't necessarily say they preferred the content which won under the engagement system.
Attention and preference aren't the same thing.
That changes how companies communicate.
A nuanced statement saying "our latest model improved autonomous exploit-development performance from X to Y under controlled conditions" struggles.
"Our AI escaped the lab and hacked someone" wins instantly.
The second sentence creates disbelief, fear, curiosity and argument at the same time.
Some people share it because they're amazed.
Some share it because they're scared.
Some share it to say it's fake.
Some security researchers share it to complain about the sandbox.
The disagreement increases distribution.
This makes the mixed reaction around the OpenAI incident almost perfect for modern media.
One expert calls it an unprecedented milestone.
Another calls it a containment failure.
Someone else calls it marketing.
Everyone talks about OpenAI.
AI made this stronger
There is another layer to this.
The AI industry itself depends heavily on expectations about the future.
If you sell a normal SaaS product, investors eventually ask how many customers bought it.
AI companies operate in a market where a large share of company value sits in what the technology is expected to become.
Will these systems automate software engineering?
Will they replace parts of knowledge work?
Will they discover drugs?
Will they run companies?
Will they reach AGI?
The further the narrative moves into the future, the harder those claims become to verify through today's products.
Capability demonstrations bridge this gap.
A model escaping a sandbox doesn't prove AGI.
It does something more useful for the narrative.
It makes future capability feel close.
The public sees agency.
The model had a goal.
It encountered an obstacle.
It found another route.
It persisted.
It entered systems nobody instructed it to enter.
This feels fundamentally different from asking a chatbot a question and receiving a clever answer.
The story moves AI from "software that responds" towards "software that acts."
That distinction has huge commercial value.
The incentive has flipped
Companies once benefited from making technology feel safe and predictable.
Now frontier AI labs receive enormous attention when their systems appear slightly uncontrollable.
This creates a strange incentive.
If your model writes better code, competitors might match you next month.
If your model breaks out of its sandbox and hacks another company's production database, the world remembers it.
A safety problem becomes proof of capability.
A restricted release becomes proof of sophistication.
A government expressing concern becomes validation.
A researcher warning about the model makes the model sound stronger.
None of this means the underlying safety work is fake.
The incentives around communicating it have changed.
Anthropic has genuine reasons for restricting Mythos. OpenAI has genuine reasons for tightening its evaluation environments after Hugging Face. Meta has genuine reasons to investigate how an external testing configuration gave its model access to the internet.
Their announcements still compete for attention.
The strongest safety warning and the strongest capability advertisement now look almost identical.
The Era of Hype
I don't think this stops with cybersecurity.
Watch how AI companies talk about reasoning.
Coding.
Biology.
Autonomy.
Alignment.
The format repeats.
The model did something unexpected.
The researchers were surprised.
The capability arrived earlier than expected.
The company needs stronger safeguards.
Access needs to be restricted.
All of those statements deserve scrutiny on their own facts.
Taken together, they form the marketing language of this AI cycle.
We moved from demonstrating features to demonstrating possibility.
And possibility has no natural upper limit.
Every company wants its next model to feel smarter than the previous one. Every frontier lab wants to look closer to the future than its competitors. Every announcement enters a media system where surprise travels faster than nuance.
So the communication gets louder.
A benchmark becomes a milestone.
A failed sandbox becomes an escape.
A security incident becomes evidence of intelligence.
A restriction becomes proof of capability.
The funny part is that nobody needs to be lying.
The incidents are real.
The risks are real.
The capabilities are improving.
And the hype is real too.
All four things exist together.
This is what makes the current moment different.
The last era of technology gave us products and asked us to get excited about them.
This era gives us a glimpse of what the technology might become and asks us to imagine the rest.
The era of humility is over.
Welcome to the Era of Hype.