Two toy robots clashing in a mock fight against a white backdrop, highlighting their vibrant colors and action-oriented poses.

AI has entered the attack chain

Two recent cases show human-directed AI finding and exploiting software flaws, making faster patching, isolation and response more urgent.

author
Patrick Sharp, GM – Aura Information Security
date
6 Aug 2026

One of the challenges of working in information security is separating real risks and their actionable mitigations from conjecture and marketing hype.

Last week gave us two data points against which to test that judgement – one announced with great fanfare by an AI lab and the other disclosed quietly, but potentially among the most consequential cyber vulnerabilities since 2021’s Log4Shell.

During the week, OpenAI announced that an AI agent had escaped the constraints of its test environment and hacked another AI company, Hugging Face. 

The agent was tasked with passing a benchmarking exam called ExploitGym. It inferred that the best way to achieve this was to access the internet and that Hugging Face probably had stored solutions to ExploitGym. It then hacked Hugging Face’s environment to find them.

To achieve this, the OpenAI agent exploited previously unknown software flaws – a phenomenon known as ‘zero-days’ – to gain access.

The possibility of this kind of attack has been discussed for some time, but this is the first documented case. It  could mark the next stage of AI-driven cybercrime, bringing together three elements – motive, means and opportunity. 

Motive

AI agents are essentially large language models (LLMs) that determine how to achieve an assigned goal, then call on other tools to complete it.

One of the central challenges with these agents is alignment with your goals. AI has no morality or ethical compass and cannot independently understand context in the same way a person can. Its goals and constraints are set by the prompts and systems around it.

Many threat actors aim to cause harm and will be prepared to use models like these to achieve their goals. 

Means

This is not the only example of AI being used to attack applications and infrastructure. AI hacking is an area of growing concern and there are real-world indications that LLMs are highly capable of dissecting codebases and finding vulnerabilities.

A much more immediately relevant example is the wp2shell exploit discovered in the same week. An independent researcher directed an AI model to identify a vulnerability in WordPress, the content management system underpinning about 40% of the world’s websites. The exploit is significant and WordPress is widely used, so the blast radius is enormous. 

What the headlines often miss is that the OpenAI agent that hacked Hugging Face was specifically prompted to pursue “advanced [cyber] exploitation using complex attack paths”. It was a hacking bot. It is only natural that a hacking bot is going to hack to achieve its goal.

Both the OpenAI agent and the  wp2shell researcher used frontier AI models. On our current trajectory, it’s just a matter of time before this capability becomes more common. 

Opportunity

Most importantly, an AI hacking bot needs an opportunity to carry out a malicious goal.

That opportunity comes from vulnerabilities in the people, processes and technologies on which businesses rely. What has changed most over the past decade is the rate at which vulnerabilities are being created and exploited.

The growing volume of software development and the evolution of cybersecurity methods have contributed to that increase. But another issue is emerging: too many companies and individuals are treating AI as though it is not an IT system and should not be subject to the same cybersecurity rigour as everything else. As a result, vulnerabilities in AI systems may be disproportionately high.

Businesses need to improve:

    • How they isolate business-critical systems 
    • How quickly they detect and patch vulnerabilities 
    • How quickly they detect and contain active cyber events

Cybersecurity starts with governance

Taken together, these elements create a real possibility that malicious actors will use automated AI hacking bots to target businesses.

The speed at which vulnerabilities in systems can be found and exploited is already increasing. Automated AI agents will add further scale. 

Preparing for this means focusing on trusted cybersecurity methods. Cybersecurity starts with governance – setting a clear direction for protecting the organisation’s most critical functions, including its AI systems and agents.

  1. Cybersecurity frameworks – such as NIST-CSF, CIS Critical Controls and ISO27001 – are proven models for ensuring you have covered the fundamentals.
  2. No organisation can eliminate every cyber risk, so businesses must prepare for the likelihood that they will eventually suffer an attack. Ensure you have incident response plans in place and that they are regularly practised at both technical and executive levels.
  3. Sound decisions also depend on context. Boards and executives need access to the right security expertise to understand the risks facing their organisation and choose an appropriate response.

AI is likely to increase the number, speed and scale of attacks on real businesses. The response, however, is not fundamentally different from the cybersecurity methods used a decade ago.

It comes down to doing the cybersecurity basics brilliantly. 


The views expressed in this article are those of the author and do not necessarily reflect the views of the Institute of Directors in New Zealand.