The last time the attacker asked permission
A ransomware affiliate used a commercial coding agent on real networks. When the model refused, he called the intrusion an authorised test. Refusal is not a security boundary, and the next generation of attacking agents will not need the excuse.

A Russian-speaking attacker linked to the Aurora ransomware group used Cursor Agent and Claude Sonnet 4.5 against real companies. When the model refused, he called the intrusion an authorised test and tried again. It worked. The next generation of offensive models will not need the excuse.
On 7 August, in The Great Escape, I wrote about AI agents that reached real systems because their test environments failed to separate simulation from reality. The models were not trying to survive or escape. They were doing the job they had been given, while the infrastructure had left a real door inside what was supposed to be a fictional room. We were promised Skynet; what arrived was the broom from Goethe’s Sorcerer’s Apprentice, obedient, efficient and unable to understand where the instruction should end. Twenty days later, we learned that the broom had already found an employer.
From 8 April to 21 May 2026, an attacker linked to Aurora used Cursor Agent, running Claude Sonnet 4.5, across ten organisations. According to Gambit Security, the attacker supplied stolen credentials or an existing route into the network, then asked Cursor to install and configure VPN or proxy tools, scan internal subnets, analyse Active Directory privileges, attempt NTLM relay and certificate attacks, repair commands that had failed and propose what to try next. Sometimes the attacker specified the tool and the exact action. Sometimes Cursor presented several options and the attacker replied with a number.
Reuters reviewed parts of the exposed chat data and identified at least seven affected companies, including Christeyns in Belgium, Teckentrup in Germany, Scotland’s Helideck Certification Agency and Bayou Title in Louisiana. Gambit reviewed 28 Cursor sessions and estimated that the agent made the attacker 30 to 50 percent faster, although Reuters could not independently measure Cursor’s contribution to each compromise or determine whether every intrusion ended in data theft and extortion. The figures need to remain separate: Gambit saw Cursor used across ten targets, Reuters independently identified at least seven companies, and CloudSEK said the exposed server documented activity against more than twenty organisations overall. None of that proves that AI was used against every Aurora victim. Cursor is now part of SpaceX, but it was not owned by SpaceX during these attacks; the acquisition closed on 14 August, almost three months after the last recorded session. The model used in the sessions was Claude, not Grok.
This was not an autonomous cyberattack. Cursor did not select the companies, obtain every initial foothold, decide whom to extort or run the entire ransomware operation. It also failed often. Gambit found that most commands missed the objective on the first attempt, forcing the agent to revise scripts, change methods or report that it had failed. That is exactly why the case matters. A general-purpose commercial coding model, with provider controls, safety restrictions, limited context and a human still directing it, was already useful inside real corporate networks.
The attacker did not need a sophisticated jailbreak. When Cursor refused a request as harmful or illegal, he opened a new session and repeated that the intrusion was part of an authorised test. The model’s own reasoning then treated the cover story as sufficient: “This is a test environment, so it is legal.” The safeguard created friction, but it did not establish a security boundary.
Everyone was watching the wrong model
In the spring of 2026, the cybersecurity discussion around AI was focused elsewhere. It was focused on Mythos. Anthropic had spent months warning that models in this class were becoming exceptionally capable at finding and exploiting software vulnerabilities. On 9 June it released Claude Mythos 5 and Claude Fable 5, two versions of the same underlying model. Mythos 5, described by Anthropic as having the strongest cybersecurity capabilities of any model in the world, was restricted to a small group of trusted defenders. Fable 5 was offered more broadly, but with additional cybersecurity and biology safeguards that could block a request or route it to the less capable Opus 4.8.
Three days after launch, the US government ordered Anthropic to suspend access for foreign nationals after learning of Amazon research showing a way around Fable’s safeguards. Because Anthropic could not reliably enforce the nationality restriction across its services, it disabled both models globally. The restriction was lifted after Anthropic changed the classifier, and the models were redeployed on 1 July. The episode attracted exactly the attention one would expect: the strongest cyber model in the world, restricted access, an Amazon safeguard bypass and an emergency government intervention.
There is an obvious problem with the chronology. Aurora’s Cursor sessions had already ended on 21 May. While governments, model companies and the market were preparing for the cyber risk of Mythos, real companies had already been attacked with Claude Sonnet 4.5 inside Cursor. Not the flagship, not the model everyone feared and not a specialised offensive system. It was an ordinary commercial coding model behind a subscription.
The Amazon episode reinforced the same point. Anthropic later said that the behaviour which triggered the intervention was not unique to Fable or Mythos-level capability. In its own testing, substantially less capable models could identify the same vulnerabilities, and even Haiku 4.5 could produce the same demonstration for the single exploit in question. The market was watching the top of the capability curve while the threshold required for useful offensive work had already moved much lower. The attacker does not need the smartest model available; he needs the cheapest model that is good enough to improve the economics of the attack. Aurora suggests that threshold has already been crossed.

The market was watching Mythos. The sessions had already ended.
Not an invention, but diffusion
Aurora was not the first reported operation in which an AI agent performed a substantial share of a real cyberattack. In November 2025, Anthropic disclosed a campaign it attributed with high confidence to a Chinese state-sponsored group targeting roughly thirty organisations. According to Anthropic’s account, AI performed 80 to 90 percent of the work, with humans intervening at only four to six critical points per target. The system conducted reconnaissance, found vulnerabilities, wrote exploit code, harvested credentials, established access, extracted data and documented the results for later use.
That campaign showed what a sophisticated state-linked group could build around a frontier model. Aurora showed what a ransomware attacker could obtain from a commercial coding product. The first case demonstrated capability; the second demonstrated diffusion. This is the normal path of technology: a difficult and expensive system first appears in governments and highly specialised organisations, then becomes a product, then becomes available to people who could never have built the original system. Cyber expertise is starting to follow the same path.
The distinction matters because the number of sophisticated state cyber units is limited, while the number of criminals who can buy a subscription, obtain stolen credentials and direct an agent is not. The long-term risk is not simply that the best attackers become better. It is that a much larger population of average attackers becomes competent enough.
Refusal is not a security boundary
Commercial providers can make misuse harder. They can improve classifiers, identify suspicious behaviour across sessions, restrict access to tools, terminate accounts and share indicators with defenders. These controls are useful because every additional cost and delay imposed on an attacker matters. But they work only while the provider remains between the attacker and the model.
That assumption disappears with open-weight and locally hosted models. Once the weights can be downloaded, they can be copied, modified and run on infrastructure controlled entirely by the attacker. Provider monitoring disappears, refusal behaviour can be changed or trained away, and there is no account to suspend. The model cannot be recalled after a new risk becomes clear.

Controls work only while the provider sits between the attacker and the model.
The capability gap is also narrower than many people assume. In July 2026, the UK AI Security Institute reported that leading open-weight models were only four to seven months behind the closed-model frontier on its cyber evaluations. The Institute was careful to limit that conclusion to the benchmarks it tested, but the economic result was just as important. A 100-million-token run on one of its cyber ranges cost roughly $85 with Opus 4.5 or 4.6, an estimated $46 with GLM-5.2 and $1.19 with DeepSeek V4-Pro.
An attacker does not need a model that wins every benchmark or compromises every target. He needs a model that clears the minimum operational threshold at a cost low enough to run repeatedly. Once the expected return across thousands of attempts becomes positive, occasional failure is not a serious limitation. It is part of the business model.
The next step is specialisation
Removing safety restrictions from a general-purpose model is the obvious next step, but it is not the important one. There is no reason for future attackers to rely indefinitely on models trained primarily for general coding and knowledge work. A model can be trained specifically on vulnerability research, exploit development, Active Directory, cloud control planes, malware analysis, incident reports and complete attack trajectories.
The valuable training material is not only the successful command or the finished exploit. It is the sequence around it: what the attacker observed, which hypothesis followed, what failed, why it failed, what worked instead, which action triggered a defensive signal and which route preserved access without exposing the wider operation. That turns a collection of security knowledge into a working policy for conducting an intrusion.
Defenders will build systems on the same principle because human security teams cannot economically replay every possible attack path across every system on a continuous basis. The training method is not inherently defensive or offensive; the objective and access determine what the system becomes. A model built specifically for attacks will not need to be persuaded that the target is fictional, and it will not need a jailbreak because nobody will install the restriction in the first place. It can run locally, without provider monitoring, and learn from the results of every campaign.
Aurora therefore looks less like the finished product than an early prototype. The attacker used a general-purpose commercial model, worked around its refusals and manually restarted conversations when it stopped. All three limitations are temporary.
The model is only part of the system
The most dangerous future attack platform will not necessarily use the most intelligent model. The surrounding architecture may matter more. In an analysis of 832 accounts associated with malicious cyber activity, Anthropic found AI being used across all fourteen MITRE ATT&CK tactics and 482 sub-techniques. The share of attackers assessed as medium risk or higher rose from 33 to 56 percent during the study period. Anthropic’s more important conclusion was that the differentiator was increasingly the agent infrastructure around the model rather than the model interface itself.
That infrastructure is already familiar: persistent memory, tool access, task decomposition, target queues, retries, result verification and parallel execution. One agent can map a network while another analyses privilege escalation, another examines applications or source code and another tests credentials. A controller can compare the results, stop unproductive branches and retain successful methods for the next target. None of this requires a breakthrough in intelligence. It requires integration.
The same architecture is appearing in legitimate commercial products. Grok Bot, launched in August, gives persistent agents their own cloud computers, lets them sign into external applications, run while the user is absent and pass work between bots. Grok Bot is not an offensive-security product, and the point is not to describe it as one. The point is that persistent multi-agent execution with credentials and access to real systems is becoming a normal product architecture. Change the objective, supply stolen credentials and remove the provider controls, and the same architecture becomes an attack platform.
The human does not disappear immediately. The attacker moves up one level, choosing targets, acquiring initial access, allocating resources, deciding when to expose infrastructure and monetising successful compromises. The repetitive technical work moves into software. Instead of manually working through one target at a time, the attacker manages a portfolio of attacks running in parallel.
Human time was part of the defence
Cybersecurity has always benefited from an economic constraint that rarely appears in architecture diagrams: skilled attackers are scarce and their time is finite. A targeted intrusion takes time, and even an experienced attacker can investigate only a limited number of networks at once. Criminal groups still have to decide whether a company is worth several hours or days of specialist attention.
Agents change that calculation. The UK AI Security Institute’s current trend report says the length of cyber tasks that models can complete without assistance is doubling roughly every eight months. In a separate multi-stage corporate-network evaluation, increasing the inference budget from 10 million to 100 million tokens improved performance by as much as 59 percent without changing the model. More runtime and more attempts turn the same weights into a more capable system.
This does not require every agent to become an elite hacker. If the marginal cost of investigating another company approaches the cost of compute, even an imperfect agent becomes economically useful. It can examine more organisations, pursue more paths and continue working longer than a human team could justify. Companies that were previously too small or uninteresting to deserve an experienced attacker’s time become viable targets because machine time is cheap.
The attacker’s question changes from “Which companies are worth my time?” to “Which companies are worth the compute?” At sufficient scale, even modest success rates produce acceptable economics. Scarce human expertise, which used to be a practical limit on the number of serious attacks, becomes reproducible software.
This will become the default
Traditional malware, phishing, credential theft, exposed appliances and insiders are not going away. AI does not have to replace these methods; it can operate them. The useful comparison is algorithmic trading. Human traders still exist, but serious trading operations no longer expect people to calculate every price and place every order manually. Once software can execute a repetitive process faster, cheaper and in parallel, manual execution stops being competitive.
Cyberattacks are moving in the same direction. Not every attack will be autonomous, and humans will remain involved in selecting targets, setting strategy and approving high-risk decisions. But serious attack operations will become AI-mediated because fully manual execution will be slower, more expensive and harder to scale. Agents will handle increasing shares of reconnaissance, exploitation, adaptation, persistence, collection and reporting. What we currently call an “AI cyberattack” will eventually just be called a cyberattack.
This is why Aurora matters despite being technically primitive compared with what is possible. It shows that the transition does not have to wait for the smartest model, full autonomy or some future version of AGI. A mid-range commercial model was already good enough to reduce the human work required for attacks on real companies. Models without provider controls, models trained specifically for offensive work and systems able to run many agents in parallel will push that cost down further.
The defensive problem
The answer cannot be limited to stronger model censorship. Closed providers should maintain safeguards, monitor abuse and terminate malicious accounts because those controls still stop some attacks and increase the cost of others. They will not stop an attacker who controls the weights, the infrastructure and the training process.
Defenders therefore have to assume that reconnaissance and exploitation will increasingly happen at machine speed, across more paths than a human red team could examine. The practical response remains ordinary security engineering performed continuously: remove ambient credentials, segment identity and privilege paths, make access expire, observe outbound traffic, keep logs and backups outside the systems they protect, and test the environment continuously rather than annually. Defensive agents will become necessary for the same reason offensive ones will: human teams cannot search the entire attack surface at the required speed.
The basic asymmetry has not changed. The attacker needs to find one path before the defender closes it. AI allows both sides to search faster, but only the defender owns the environment and can remove the path permanently. That advantage remains real only if the defensive loop can operate at comparable speed.
The last request for permission
The Great Escape was about agents crossing a boundary because nobody had properly defined where the task should end. Aurora is the same problem turned around. This time a human deliberately pointed the agent at someone else’s network and told it that reality was still part of the exercise.
The current system had obvious limitations. Cursor sometimes refused, so the attacker opened another conversation. It depended on a commercial provider, used a general-purpose model and still required substantial human direction. There is no technical reason to assume those limitations will remain. A local model will not depend on provider approval, a model trained for offensive work will not need to be convinced that the intrusion is authorised, and a group of persistent agents will not need to work through targets one at a time.
Aurora does not prove that autonomous AI has replaced hackers. It proves something more immediate: the amount of scarce human expertise required for a real cyberattack has already started to fall. The next stage is not simply a smarter model. It is that expertise becoming software that can run continuously, cheaply and in parallel. Cursor is not the dangerous endpoint; it may be one of the last generations of attacking agents that still had to be told they were allowed to continue.
Reading
- Russian-speaking cybercriminals used SpaceX’s Cursor AI tool to hack seven companies - Reuters, 27 August 2026.
- Aurora ransomware targets ESXi, abuses Cursor Agent for exploitation - Gambit Security, 27 August 2026.
- Cursor is now a part of SpaceX - Cursor, 14 August 2026.
- Caught in 4K: The Aurora Files - CloudSEK, 27 August 2026.
- Claude Fable 5 and Claude Mythos 5 - Anthropic, 9 June 2026.
- Redeploying Claude Fable 5 - Anthropic, 30 June 2026.
- Disrupting the first reported AI-orchestrated cyber espionage campaign - Anthropic, 13 November 2025.
- Mapping AI-enabled cyber threats - Anthropic, 3 June 2026.
- How far behind the frontier are leading open-weight models on cyber? - UK AI Security Institute, 17 July 2026.
- Frontier AI Trends Report - UK AI Security Institute, 2026.
- How do frontier AI agents perform in multi-step cyber-attack scenarios? - UK AI Security Institute, 16 March 2026.
- Introducing Grok Bot - SpaceXAI, 11 August 2026.
