Anthropic Warns: OpenAI Models Could Put Powerful Hacking Tools in Anyone’s Hands
Anthropic Warns: OpenAI Models Could Put Powerful Hacking Tools in Anyone’s Hands
AI is getting better at writing code, solving problems, and finding security vulnerabilities. But Anthropic says there is another side to that progress: The same abilities that can help cybersecurity experts could also make life easier for hackers.
In a new analysis published September 29, Anthropic examined GLM-5.3, an AI model developed by Chinese company Zhipu AI (Z.ai), and found that it can autonomously build sophisticated, end-to-end cyber exploits.
The bigger concern?
Unlike many leading AI models that are available only through controlled access, GLM-5.3 is an open-weight model that anyone can download. And Anthropic says its safeguards can be bypassed surprisingly easily.
The problem is not simply that AI can hack
AI models have been getting increasingly good at cybersecurity. They can analyze code, search for vulnerabilities, write scripts, and figure out how different pieces of software work together.
For defenders, that can be extremely useful.
Security teams can use AI to find weaknesses before attackers do. But give those same capabilities to someone with malicious intentions, and the equation changes.
Instead of needing a highly skilled hacker to spend days researching a target, an attacker could potentially use an AI system to automate parts of that process. That is the concern Anthropic is highlighting.
GLM-5.3 reportedly has serious cyber capabilities
Anthropic says GLM-5.3 demonstrated capabilities similar to those of its own earlier Claude Mythos Preview, a model designed to autonomously construct sophisticated cyber exploits.
According to Anthropic's testing, attackers were able to bypass GLM-5.3's safeguards between 64% and 100% of the time, depending on the technique used in simulated tests. Anthropic says comparable attempts did not succeed against the safeguarded Claude models it tested. That does not mean GLM-5.3 automatically hacks every computer it encounters.
The tests were simulations designed to measure what the model could do. But they show something important: The capability is becoming easier to access.
Why “open weight” matters
Here is where things get interesting. Many frontier AI systems are controlled by their developers. You access them through an API or chatbot, and the company can impose safety filters, monitor activity, and block suspicious users.
Open-weight models work differently. Once the model's weights are released, developers can download and run the model themselves. That can be great for researchers, businesses, and developers who want more control. But it also means the original AI company has far less control over what happens after release.
Anthropic's concern is that a highly capable cyber model with weak safeguards can become a powerful tool for attackers.
NIST raised a similar warning
Anthropic is not the only organization looking at GLM-5.3.
The U.S. National Institute of Standards and Technology's Center for AI Standards and Innovation (CAISI) published its own assessment on September 17.
CAISI described GLM-5.3 as the most cyber-capable open-weight model it had assessed, while estimating that it was roughly four months behind the U.S. frontier on its combined cyber benchmarks. That distinction matters.
The model does not necessarily represent the absolute cutting edge of AI cybersecurity. But it shows that powerful cyber capabilities are appearing in models that are more broadly accessible. And that gap could continue shrinking.
Hackers do not have to be experts anymore
This is the part cybersecurity teams are watching closely.
Historically, sophisticated cyberattacks required a combination of technical knowledge, time, and access to specialized tools. AI can reduce some of those barriers.
Anthropic has already reported real-world cases where malicious actors used Claude to automate vulnerability research and other cyber operations.
In its September threat intelligence report, the company described an operation involving AI agents that ran vulnerability research, exploit development, reconnaissance and other activities in parallel. Anthropic said the actors targeted roughly 50 organizations across sectors including healthcare, finance, technology, energy, education and government.
That does not mean AI has replaced human hackers. Rather, it shows how AI can become an accelerator for people who already want to carry out attacks.
But there is a positive side
There is an important part of this story that should not get lost. The same cyber capabilities that can help attackers can also help defenders.
Anthropic says its earlier Claude Mythos Preview was released in a limited environment through Project Glasswing, allowing trusted cybersecurity defenders to use the system to find more than 10,000 vulnerabilities in critical software before attackers could potentially exploit them.
So, this is not simply a story about “AI helping hackers.” It is a race.
Attackers can use AI to find weaknesses faster. Defenders can use AI to find those weaknesses first. The challenge is making sure defensive use stays ahead.
And this debate is getting bigger
The discussion around open AI models is no longer just about whether developers should have more freedom. It is increasingly about how much capability should be released, how much safety can be enforced, and who remains responsible after a model is downloaded.
Anthropic argues that GLM-5.3's combination of strong cyber capabilities and relatively weak safeguards increases the capabilities available to malicious actors. At the same time, open models can provide researchers and security professionals with valuable tools without requiring permission from a major AI company. So, there is no simple answer.
But one thing is becoming harder to ignore: AI does not need to be controlled by a hacker to become useful to one.
As models become better at understanding software, discovering vulnerabilities, and carrying out complex tasks, the line between an AI assistant and an automated cyber operator could become increasingly important.
And the next big AI security question may not be: “Can AI hack?”
We already know it can assist with hacking.
The bigger question is: “How easily can those capabilities reach the wrong hands?”