Techdirt has just written about concerns that AI development is proceeding more quickly than human oversight can monitor and control it. Since then, there has been lively debate about whether those fears are real or overblown. They mainly focus on the possible future threats from AI systems using RSI — recursive self-improvement — to drive their own development, at an ever-faster pace, by re-writing their own code. But a recent report from Anthropic underlines that AI is already being used by conventional threat actors — state-sponsored groups, financially-motivated criminals, commercial spyware vendors, state propaganda institutions, and politically-motivated individuals — in attempts to cause a wide range of harm, exploiting today’s powerful AI systems to amplify the reach and impact of their actions.
It is a measure of just how fast those threats are multiplying that the latest Anthropic report on “Detecting and countering misuse of AI” runs to 154 pages, whereas the three published in 2025 barely hit double digits. It covers activity that the company disrupted between December 2025 and August 2026, and involved the use of its Claude Haiku, Sonnet, and Opus models. An obvious application of AI is to carry out cyber attacks. The key development in recent months is the following:
The cybersecurity skills of AI models means that AI has collapsed the labor and tooling gap that used to separate well-resourced, state-sponsored operations from individual operators. In the case studies we report below, a hacktivist using stolen API keys, disparate financially motivated individuals, and a state espionage operator each sustained multi-victim campaigns that, even just a year ago, would have required many skilled operators and specialist knowledge.
For threat intelligence investigators, sophistication has stopped being a reliable signal of who is behind an operation. Every layer of offensive operations has been uplifted by AI, from reconnaissance and tool development to data processing and exploitation.
AI’s sophistication in this sphere has reached the point where “vibe hacking” is now possible:
The use of AI during intrusions and data theft operations often resembles “vibe hacking,” wherein operators direct AI to achieve general goals like using a credential for an entity or retrieving data from a broad set of targets, then allow the AI to evaluate the environment, author and execute scripts, provide summaries, and repeatedly execute until the task is complete.
It is not just single hacking attempts that are being automated; the entire exploit development chain is being outsourced to AI:
Historically, cyber operations have been limited in their scale and impact by two key constraints: the supply of working offensive exploits, and the supply of skilled operators capable of deploying those exploits. We have identified multiple threat actors who have effectively established automated exploit foundries with AI. In doing so, they have designed and implemented autonomous workflows by which they can direct Claude to conduct vulnerability and exploit research agentically around the clock. Across multiple instances, we identified Claude being used to meaningfully accelerate the pace of vulnerability research, testing, and exploit design.
As a result, everything is at risk:
The old adage of “security through obscurity” is no longer viable in this new AI-assisted world: everything connected to the internet is a potential target for exploitation.
Alongside cyberattacks, two other obvious applications of AI are influence operations and surveillance. The Anthropic report discusses examples of both, which exploit AI capabilities much as you might expect. But one notable feature is scale:
AI is now being used in place of an engineering workforce. A single consultant working for Malian national security authorities used Claude to engineer a mass-interception platform capable of surveilling communications on all of the country’s mobile operators and generating dossiers on targets.
And:
a religious affairs intelligence collection unit in the People’s Republic of China (PRC) that once comprised many teams of analysts has been reduced to a single office, using an AI assistant to produce thousands of investigations per month
It is striking that Chinese-based threat actors figure disproportionately in the influence and surveillance section:
The targets ranged from domestic petitioners and rights defenders to prominent pro-democracy figures in Hong Kong, organizers of Tiananmen Square commemorations, Uyghur advocacy organizations, and Western human rights institutions. In the most serious case, the actor directed Claude to produce pre-operational venue intelligence (i.e., scouting locations ahead of an operation) on overseas protests.
Several Iranian operations too were spotted by Anthropic:
An Iran-nexus threat actor that used Claude to build an automated, open-source intelligence identity-profiling harness targeting Israeli governmental and non-governmental individuals, and Jewish diaspora organizations. Separately, the actor used Claude to make it harder to tell that its malware was malware.
One new category of AI misuse that appears in the latest report concerns conventional weapon development, using Claude to develop software for firearms, missiles, armed drones, bombs, and other munitions, as well as the targeting and control systems that operate them:
We identified a cell of threat actors based in northern Yemen running three weapons development programs: a guided rocket that used a commodity phone-class flight computer with final-phase homing guidance; a multi-stage ballistic missile with a stated range goal above 2,000 km; and a multi-variant missile (referred to as the “R2000” set) that included a hypersonic glide vehicle variant.
The actors used Claude Code in place of human software engineers to develop the guidance, navigation, and control (GNC) software that steers and stabilizes a flying vehicle. For example, they used Claude to integrate an open-source autopilot onto a phone-class flight computer, writing the control and position estimation software, tuning the control settings, running a firmware build pipeline, and performing a flight simulation. The actors managed several Claude instances at once, assigning each one a role, much as a lead would delegate work on a small engineering team: the actors tasked one instance with writing the code, another with research, and a third with reviewing the code the first instance produced.
Russia-based actors tried to use Claude for something more innovative:
We identified likely freelance Russia-based threat actors who set out to build a full-stack autonomous first-person-view (FPV) kamikaze drone swarm. The actors used Claude Code to write and test the code and save it directly into the actors’ own project files. In addition to Claude Code, the actors used a software-in-the-loop simulation stack and a rented graphics processing host for model training.
The actors designed the platform for autonomous lethal engagement; the onboard model could select targets (including a “person” target class) and issue detonation commands without a human in the loop. The actors’ activity—including flashing the low-level firmware to live development boards, provisioning single-board computers, and wiring up a simulation environment over a mesh network—confirmed that they were using real hardware-in-loop testing within their sessions.
But perhaps the most chilling threat discussed in the Anthropic report is the following:
Biological misuse is one of the most serious risks of frontier AI models. It has long been a concern that AI models might one day reach the level of capability where they can help to make existing pathogens more dangerous—or create entirely new ones. Without the correct safeguards, such capabilities could have catastrophic consequences.
Anthropic provides five case studies of actors using its products that could support biological weapons development:
In the first example, a reseller platform evaded regional blocks to serve virologists working on a state-sponsored grant to pursue chikungunya gain-of-function work, later routing refused prompts to models with more permissive safeguards. In the second, a researcher in an unsupported region spent weeks planning avian influenza mammalian-adaptation experiments with Claude, but classifiers confined the work to our weakest models. In the third case study, a reseller relay serving a dozen customers had Opus 5 draft a complete orthopoxvirus immune-evasion grant application in about an hour. In the fourth example, a state-supported researcher built a venom peptide atlas and generative optimization pipeline of molecules directed at paralytic and analgesic targets. And in the final case study, a researcher computationally redesigned toxins for a national program, asking Claude to keep the agents’ identities deliberately vague in progress reports.
In some ways, the detail provided by the Anthropic report is comforting. It shows that the company managed to spot and block many sophisticated uses of its AI offerings. But these, of course, are only the ones that it noticed (and the ones that it is willing to talk about). The question has to be how many escaped detection, and succeeded in their goals. Moreover, Anthropic is only one company, albeit one of the leading ones in its field. Similar attempts to use AI for such dangerous purposes are doubtless being made around the world on many AI systems. Although discussions about RSI wiping out humanity grab the headlines, maybe it’s time to focus much more on the lesser but very real threats that AI already poses in the hands of the wrong people, as revealed in the Anthropic report.
In particular, we need to talk about how to respond to this malicious activity. Should we rely on a few dominant companies to police everyone’s use of AI to spot and block the bad actors? As AI becomes embedded in everyday life, that effectively means constant surveillance of everyone by those companies. What about the legal frameworks for handling these abuses? Who should bear the liability when threat actors succeed in circumventing the built-in safeguards and cause real-life harm? And how would increasingly capable open weight/open source systems that exist outside the control of any company fit into new legal and liability frameworks? It’s going to be an interesting few years finding out — assuming we have that long…
Follow me @glynmoody on Mastodon and on Bluesky.