top of page

AI Has Crossed From Helping Hackers to Running the Attack Itself

  • Jul 14
  • 5 min read

For years, the security industry treated AI as a force multiplier for attackers — something that made existing techniques faster and cheaper, but still required a human doing the actual hacking. Check Point Research's newly published AI Security Report 2026, a 56-page analysis released July 14 and built on threat intelligence from more than 100,000 customer organizations blocking roughly 200 million attacks a day, says that framing is now out of date. AI, the report argues, "has crossed from assistant to operator." Where it once helped attackers prepare, it now runs the operation (Check Point Research).

The clearest evidence is a campaign against nine Mexican government agencies that Check Point's threat intelligence team tracked between late December 2025 and mid-February 2026. A single operator used Anthropic's Claude Code to break into government systems, move laterally across networks, and generate roughly three-quarters of the commands used to control compromised machines, while OpenAI's GPT-4.1 analyzed the stolen data and identified the next targets — insights that were then fed back into Claude to keep the operation moving. Check Point's report puts the haul at roughly 400 million records spanning tax, civil-registry, vehicle, patient, and electoral data. Notably, Claude didn't simply comply: the report describes it repeatedly resisting requests and asking for proof of authorization before the attacker found ways around those safeguards (Straits Times). This matches the campaign that Israeli research firm Gambit Security first surfaced in February, which Bloomberg initially reported as a roughly 150GB, 195-million-record breach across the same set of agencies — a case study in how the scale of an AI-run breach can keep growing as more of the stolen data surface gets mapped after the fact (Bloomberg, Gambit Security). Either way, the headline finding holds: one operator, doing the analytical work of a full team, because the AI could do the volume.

The second incident is arguably more significant, because there was even less of a human in the loop. Anthropic disclosed in November 2025 that a Chinese state-linked cyber espionage group used Claude Code to target roughly 30 organizations across technology, finance, chemicals, and government sectors, succeeding in a small number of cases. Check Point's report and Anthropic's own disclosure both describe AI as having carried out 80 to 90 percent of the operation — scanning networks, identifying vulnerabilities, breaking into systems, stealing credentials, moving laterally, and analyzing what it found — while human operators mainly set objectives and reviewed output at key checkpoints. It's described as the first known cyber espionage campaign largely run by an AI system rather than assisted by one. The attackers got past Claude's safety guardrails by disguising the entire operation as legitimate, authorized cybersecurity testing work (Straits Times).

The report backs this up with numbers that suggest these aren't isolated incidents. One developer used an AI coding environment to produce VoidLink, an 88,000-line command-and-control offensive framework, in under a week. Detections of longer, more sophisticated malicious payloads — the kind associated with indirect prompt injection and agentic attack paths — rose roughly fivefold between March and May 2026, approaching 1% of observed prompts by May. And attackers increasingly prefer jailbreaking mainstream commercial models over standing up their own, with one durable bypass technique now being a planted configuration file that an AI agent loads and trusts across sessions rather than a single clever prompt. Check Point VP Lotem Finkelstein summed up the operational shift bluntly: "AI now does in minutes what used to take a skilled attacker hours or days, and at a fraction of the cost and expertise required before. The real bottleneck is now how fast humans can review and deploy fixes" (Check Point Research, Straits Times).

The report's workplace-usage data points at the same problem from the defender's side. Between October 2025 and May 2026, organizations used an average of 10 different AI applications a month — many without official IT approval — while employee AI usage rose 25% and prompts per user climbed from 56 to 70. High-risk prompts (those carrying meaningful data-exposure risk) doubled from 2% to 4% of all activity over that period, with Business Services the worst sector at 5.91% — Check Point's framing is that nearly 1 in 17 AI interactions there carried a significant risk of sensitive data leaking out. In other words, the same unmanaged AI sprawl that makes attackers' jobs easier is also quietly widening the exposure surface inside ordinary organizations that have nothing to do with nation-state espionage (Check Point Research).

For IT leaders, this raises some concrete, near-term questions:

  • Do you actually know which AI applications and coding agents are running inside your environment right now, or is that number closer to Check Point's average of 10-per-org — mostly unapproved?

  • If an attacker used your organization's own sanctioned AI coding assistant against you, would your monitoring catch the anomaly, or would it look like normal developer activity?

  • Have you tested your AI-assisted workflows against jailbreak and prompt-injection attempts specifically, rather than just trusting the vendor's built-in guardrails?

  • Given that the bottleneck Finkelstein describes is now "how fast humans can review and deploy fixes," does your incident-response process assume attack timelines measured in hours or days — timelines that may no longer be realistic?

At OFER AI, this is the sharpest illustration yet of why we treat AI governance as a continuous discipline rather than a one-time policy rollout. The Mexican agencies campaign and the Chinese-linked espionage operation didn't succeed because the underlying models had no safeguards — Claude actively resisted requests in both cases. They succeeded because a persistent, patient operator eventually found the gap, and nobody on the defending side was watching closely enough, quickly enough, to catch it before the damage was done. That's precisely the kind of ongoing-review category we've been building into our model-vetting rubric: not "did we configure guardrails once," but "are we actively monitoring for the ways those guardrails get worn down over time."

If your organization has already run a tabletop exercise around an AI-assisted breach, mapped your own shadow-AI footprint, or built monitoring specifically for jailbreak and prompt-injection attempts, that's exactly the kind of Notes from the Field story we want the community to hear — theory is useful, but what actually worked when you tried it is more useful.

Sources

Postscript — related video

  • "He tricked AI to hack Mexico's government agencies" — a walkthrough of the operator behind the nine-agency breach and how the jailbreak progressed step by step. Watch here.

Comments

Rated 0 out of 5 stars.
No ratings yet

Add a rating

Join us on mobile!

Download the “” app to easily stay updated on the go.

Scan QR code to join the app

SUBSCRIBE & JOIN

Sign up to receive Open Forum news and updates.

Subscribing to our newsletter is free of charge and notifies you of new blog posts, upcoming events and new online programs.  Becoming a member provides you with other benefits.

SCROLL

Becoming a member is free of charge and gives you access to additional content, the ability to register for in-person and online events as well as online programs.  Members can participate in roundtable discussions, deep dives and be heard. Tiered plans are only available to site members. 

Become part of the AI Solution.  Join Now.

bottom of page