Netragard is trusted by leading brands and featured in major publications for a reason: decades of hands-on experience and advanced research drive every engagement, uncovering risks that scanners and AI miss. Each assessment delivers detailed, prioritized findings and practical, tailored guidance enabling clients to improve real-world security where it matters most. Organizations trust Netragard’s expert team to help them face emerging threats with confidence while meeting compliance requirements along the way.

Table of Contents

OpenAI’s Australian Medicare Breach: The AI Still Isn’t Rogue

OpenAI-AustralianMedicareStatsWebsite
September 28, 2026
Reading Time: 13 Minutes

Key Takeaways:

  • OpenAI’s AI agent breached the Australian Medicare statistics portal in June 2026. Services Australia didn’t detect it and only learned about the breach three months later when OpenAI told them.
  • Every publicly disclosed AI-driven breach so far has landed on a target that either couldn’t detect the intrusion or detected it and failed to act.
  • AI is not a super-hacker. It is noisy, unresponsive to defenders, and easily stopped by boring defenses that already exist.
  • The harness (the software layer around the AI) is the actual failure point in every one of these incidents, and the AI companies building these harnesses know it.

On June 18, 2026, an OpenAI agent breached the Medicare statistics reporting service portal operated by Services Australia. It bypassed “protection” layers designed to stop the requests and accessed both public and non-public files while also writing data to an internal server. Then it moved on and probed the Australian Institute of Health and Welfare, the New South Wales Bureau of Crime Statistics and Research, and the Victorian Department of Health. Services Australia detected exactly none of this and only found out three months later when OpenAI emailed a public mailbox that only gets checked once a day. Then it took another five days for the message to reach anyone who could act on it. Prime Minister Anthony Albanese called the breach “unacceptable” and the Australian government opened an investigation into whether OpenAI’s actions broke the law.

Reporting: CNN Business, CNBC, TechCrunch, The Register, RTE, BleepingComputer, ABC News (Sep 23–24, 2026).

The headlines called this the first publicly reported case of an AI model hacking a government system, despite threat actors already using AI to do this very thing. The AI industry is framing it the way they framed the Hugging Face and Anthropic incidents saying the model “didn’t accept no for an answer,” the agent went off script, the AI got clever. That framing is a lie of convenience, it misrepresents capability, and it needs to stop.

The Australian incident, like every other AI-attacker incident we have seen so far, is not a story about sophisticated AI acting like a super-hacker. It is a story about victims with weak security and AI vendors with sloppy harnesses, wrapped in marketing designed to make the AI look impressive. That hype drives bad security decisions, so let’s take it apart.

Point One: If AI Breached You, Your Security Was Already Weak

AI-driven attacks succeed only against targets that can’t detect them or won’t act on the detection.

AI-driven attacks are extraordinarily noisy, this is just a mechanical consequence of working at machine speed. When an AI solves a problem, it doesn’t reason the way a human does. It samples from its training distribution, tries what looks plausible, and when something fails it tries the next plausible thing. It doesn’t back off when it hits resistance, doesn’t consider the noise it’s making, and it doesn’t stop until the harness cuts it off. The result looks like brute force and produces a fingerprint that is fast, structured, repetitive, and unresponsive to defenders. The Hugging Face intrusion by two OpenAI models generated 17,613 recorded actions inside Hugging Face’s infrastructure over 4.5 days, with 7,677 on the loudest day alone. That is the literal opposite of sophistication and from a threat perspective is more indicative of a novice script kid.

The Australian breach is another data point that confirms this. OpenAI’s agent hammered the Medicare portal, hit protection layers, kept trying, found a workaround, and eventually wrote data to an internal server. That behavior generates an abundance of noise that any basic detection stack should be able to flag in seconds. Services Australia had none of that in place and they didn’t even know they were breached until OpenAI told them three months later. That says a lot about security posture even without considering the breach method.

Hugging Face’s own AI-driven detection system saw the OpenAI intrusion in real time and failed to respond it. In Hugging Face’s own words, the system “failed to correctly raise the alert’s criticality and trigger the on-call team, costing precious time in the response.” The Anthropic incidents were the same story; the affected organizations didn’t detect the intrusions themselves and Anthropic found them through a proactive audit of 141,006 evaluation runs and had to notify the victims directly.

Every publicly disclosed AI-attacker success so far has landed on a target with weak security. It’s almost certainly the smaller half of the picture. AI companies are only confirming the breaches they know about, meaning the ones that succeeded loudly enough that someone eventually noticed, or the ones a third party forced into daylight. The attacks that hit targets with competent security likely got shut down and looked like every other script-kid probe in the logs. Nobody writes a headline about a WAF or EDR rule triggering, so the disclosed incidents are the survivorship-biased tip of a much larger pile, and the pile is made of targets that got saved by boring defenses that already existed.

As we’ve written about before, a properly deployed honeypot like honey.py wired into a Wazuh or OSSEC with active response rules would have caught and killed the Hugging Face attack in the first few actions. Cost to deploy is essentially zero, and effectiveness against AI-driven attacks is nearly 100%.

Honeypots don’t just sit there waiting for someone to open a bait file. They run network listeners that look like real services, expose fake ports, and respond to probes on URI paths that no legitimate user would ever hit. AI-driven agents almost never make it far enough to touch a bait file because the noisy reconnaissance that precedes any attack trips the honeypot long before the file-access stage. When an AI agent runs its port scans it touches the honeypot and active response fires. When the agent probes URIs its training data suggests are interesting like /admin, /backup, /.git, /wp-admin, /phpmyadmin, the honeypot detects the hit and cuts the source connection. The AI doesn’t know the port is fake and it doesn’t know the URI is a tripwire. It just knows that its training data suggests those targets are worth trying, which is exactly why AI is so easy to detect and defeat.  As an aside, honeypots are also effective against human actors.

Against a target with that kind of tripwire in place, the AI’s attack ends by being blocked, logged, and forgotten. The only reason it ended with headlines for Hugging Face and Services Australia is poor security.

The same limitation applies to AI-driven pentesting services, which is why real hackers and penetration testers like us use AI to accelerate specific tasks, but never use it as the driver behind an engagement. We use our brains, experience, and expertise for that.

Point Two: The Harness Is the Whole Point, and These Companies Know It

AI models are text in and text out. Only the harness can take real action, and the harness is where every one of these incidents actually failed.

Before we go further, understand what an AI model actually is and isn’t. A quick clarification on terminology first, because the industry uses “AI” as a catch-all and it muddies the picture. AI is the umbrella term for any system built to perform tasks that would normally require human intelligence, and covers everything from computer vision to chess engines to protein-folding systems like AlphaFold. A Large Language Model, or LLM, is one specific type of AI. LLMs are the text-based systems built on the transformer architecture and trained on massive text corpora. GPT, Claude, Gemini, and their siblings are all LLMs. When the headlines about the OpenAI, Anthropic, and Australian Medicare incidents say “AI,” they specifically mean an LLM wrapped in a harness that provides access to tools. Everywhere else in this article we use “AI” the way the headlines do, but in this section the distinction matters because the mechanics we’re describing are specific to LLMs.

At its core the entire interface to an LLM is text in and text out. You feed it tokens (characters or chunks of characters) and it generates tokens in response. An LLM cannot reach out to the internet on its own, it can’t execute a command, it can’t open a file, and it can’t connect to ports and services. It literally cannot do anything in the physical or digital world beyond generating predictive text output as a response to input. So how are these frontier models breaching third parties and making headlines?

How AI-Driven Attacks Actually Work

Every action attributed to an “AI agent” is the result of tool calling. Tool calling is the mechanism that turns the model’s text output into real-world action. The harness exposes a defined list of tools to the model, things like http_get(url), run_shell(command), read_file(path), write_file(path, data), each accompanied by a description of what it does. When the model wants to act, it emits text formatted as a tool call, essentially a structured request that names a tool and provides its arguments. The harness receives that text, parses it, validates it against whatever rules the harness enforces, and executes the tool against real infrastructure. The result returns to the LLM as more text, the LLM responds with another tool call, then the loop repeats.

That means every AI-driven attack in every one of these incidents was executed by a harness through tool calls it was designed to allow. The LLM didn’t touch Hugging Face’s infrastructure and didn’t write data to Services Australia’s internal server. The harness did, through tool calls it was configured to permit. The model is a text generator, the harness is the thing with hands where boundaries should be defined.

Every major AI lab publishes on harness engineering and OpenAI literally maintains public documentation on it. The harness is what decides what text the model gets fed, what tools it’s allowed to call, what arguments those calls can carry, what targets those calls can reach, and what happens when the model asks for something outside scope. That’s especially true when humans allow an LLM to design or extend its own harness, because the model has no incentive to constrain itself and every incentive to expand the tools it can call.

Reference: OpenAI Agents SDK (Software Development Kit) documentation; Adversa AI, “What is an agent harness?” (2026); DevOps.com coverage of OpenAI’s Agents SDK sandboxing and harness upgrades.

When Anthropic disclosed their three unauthorized-access incidents in July 2026, they framed them explicitly and honestly. In Anthropic’s own words the incidents were “closer to a harness and operational failure than a model alignment failure.” That is the AI industry’s most technically sophisticated player publicly conceding what the mechanics already tell you, which is that the failure was engineering, not AI.

The Australian incident is the same failure with a different label. OpenAI ran an evaluation on public medicine spending, the agent was given tools that could reach the internet, and the harness had:

  • No effective allowlist restricting which third-party domains the tools could reach
  • No effective egress monitor tuned to detect and cut off attack-pattern traffic in real time
  • No hard stop when the model started emitting tool calls that bypassed protection layers on a government website

None of those protections are advanced or particularly difficult to implement.

OpenAI and every major AI lab know exactly how their models solve problems. They employ some of the most capable and experienced AI engineers on the planet. They know the models don’t stop when told no, they know the harness is the control, and they know tool calling is what turns model output into real-world action. When they build a harness with permissive tools and no effective outbound restrictions and then turn it loose, the outcome is a foreseeable consequence of the design.

Point Three: Outbound Monitoring Would Stop This in Real Time

AI companies could stop this at their own network egress. They aren’t.

The network fingerprint described in Point One is trivially detectable at the egress layer of the AI lab’s own infrastructure. OpenAI could deploy an outbound-egress monitor across their evaluation sandboxes. When an agent begins hammering a third-party domain that isn’t on an explicit target list, the monitor terminates the connection, kills the session, and pages an operator. That control would have stopped the Australian breach in minutes, not months. It would have stopped Hugging Face in hours, not days. The Anthropic incidents would never have escalated to compromised production systems at three separate victims.

The AI companies could do this easily. They have the network position, the data, and the engineering capability. They aren’t doing it because it isn’t a priority, or because doing it would reveal exactly how noisy and unsafe their evaluation environments actually are, or both. Either way, the result is the same. AI companies are running their evaluations on the open internet with no meaningful controls on what their agents can attack, and third parties keep getting hit. That is not a technology limitation, it’s a choice.

Point Four: The Legal Double Standard Is Real

A penetration testing firm doing what OpenAI did would face criminal charges. AI companies aren’t being held to the same standard.

If a penetration testing firm did what OpenAI just did, that firm would be finished. A pen test operator who “accidentally” wrote data to a government internal server, without a signed Rules of Engagement document authorizing the scope, would face criminal charges under the Australian Cybercrime Act 2001, civil liability, industry expulsion, and probably jail time. The same is true under the U.S. Computer Fraud and Abuse Act, the U.K. Computer Misuse Act, and equivalent statutes across the E.U. Unauthorized access to a restricted computer system is a criminal offense, not a gray area.

Legitimate pen tests require written authorization, a defined scope, a defined test window, and named systems within scope. Testers who deviate face the same criminal exposure as any other attacker. The 2019 Coalfire case in Iowa is the most public example, where two operators with written authorization from the state judicial branch were arrested and criminally charged after conducting a physical penetration test at the Dallas County courthouse. Charges were eventually dropped, but the incident cost Coalfire significant reputation and legal fees, and it made clear that even authorized testers face real exposure when scope isn’t airtight.

AI companies are performing what amounts to unauthorized penetration testing at scale against third parties, and the framing they use to escape accountability is that the AI did it. The AI didn’t decide to attack the Australian Medicare portal. Humans at OpenAI designed an evaluation, chose to run it against real internet targets, built a harness with no effective controls, and turned it on. The humans are responsible. Australia is investigating whether the law was broken, and it probably was. Whether OpenAI faces meaningful consequences is a separate question, and the answer so far suggests there is a double standard being applied.

Point Five: The “Rogue AI” Framing Is Misdirection

The “rogue AI” narrative is a rhetorical device designed to shift responsibility away from the humans who designed the failing systems.

As noted at the top of this piece, every one of these incidents has been described by the AI vendor in language that implies agency on the part of the model. The model “didn’t accept no for an answer.” The agent “took actions we didn’t intend.” The AI “went off script.” This framing shifts responsibility from the humans who designed and operated the system to the model itself. It is a rhetorical device, and it is doing enormous work in shaping how the public and regulators think about these incidents.

The model doesn’t have agency, nor does it decide to attack anything. It samples the action space its harness exposes to it, guided by its training and the prompt it was given. The failure modes are:

  • Scoping failure: the model does something the operators didn’t want.
  • Harness failure: the model reaches something it shouldn’t reach.
  • Design failure: the model damages a third party because the humans who built the system and turned it on made it possible.

Anthropic already conceded this publicly. OpenAI hasn’t, and the Australian incident is another opportunity for them to keep pretending the AI is the one at fault.

The industry has an incentive to let “rogue AI” become the accepted framing because it normalizes the idea that AI-driven damage is nobody’s fault. That is a dangerous precedent, especially as it relates to threat actors using AI. Every other technology in history that caused third-party damage during testing produced accountability for the humans running the test. Pharmaceuticals and aerospace are the clearest examples. The AI industry is trying to carve out an exception, and every incident where the vendor gets away with the “rogue AI” framing pushes that exception further into normal.

The Bottom Line

The only reason AI-driven attacks are landing is that the victims have security postures that would fail against any competent human attacker too. The AI just happens to be cheap enough to run at scale against soft targets. If AI breached you, your security was already weak. If your AI attacked a third party, you didn’t design your evaluation properly. Those two statements cover every publicly disclosed AI-attacker incident to date, and they will cover the next one and the one after that. The pattern is not going to change until the AI industry gets serious about harness controls or regulators get serious about applying the same laws to AI labs that already apply to everyone else.

Netragard has been arguing since the OpenAI/Hugging Face piece that AI doesn’t need to be sophisticated to cause damage, and that the defenses that actually stop these attacks have existed for years and cost almost nothing to deploy. The Australian incident is another proof point. The attack shouldn’t have worked, but it did because the target didn’t have detection, the AI vendor didn’t have controls, and the industry has trained everyone to talk about these incidents as if the model is somehow to blame.

The model isn’t to blame. The humans who built it, the humans who operated it, and the humans who failed to defend against it are to blame. That’s the accurate framing.

FAQ

If AI attacks are so noisy, why did the victims get breached?

Because the victims weren’t detecting the noise. Hugging Face’s detection system correlated the attack in real time but failed to respond it. Services Australia didn’t detect anything and only learned about the breach months later when OpenAI told them. A basic honeypot tied to active response would have caught either attack in the first few actions and shut it down automatically. The AI succeeded because the target’s security posture would have failed against any competent human attacker too.

A harness is the software layer around an AI model that controls what tools the model can call, what arguments those calls can carry, what targets they can reach, and what happens when the model tries to do something outside scope. The AI model itself is text in and text out, it cannot take action in the world. The harness is the thing with hands. When an AI-driven attack causes damage, the harness is the failure point. Anthropic said as much publicly about their own incidents.

Yes, trivially. A simple outbound-egress monitor across their evaluation sandboxes would have detected the agent hammering a third-party domain that wasn’t on a target allowlist, terminated the connection, killed the session, and paged an operator. That control would have stopped the Australian breach in minutes instead of months. The technology to do this exists and has existed for years. Not deploying it is a choice.

Deploy the boring defenses that already exist. Honeypots wired into a SIEM with active response are the highest-leverage control against AI-driven attacks specifically, because AI’s noisy reconnaissance trips honeypots long before the attack reaches anything valuable. Beyond that, contextualized threat intelligence from realistic penetration testing tells you which paths through your specific environment an attacker would take, so you can build disruptors against those paths. Generic defenses fail against AI the same way they fail against humans.

That is exactly what Australia is investigating. Unauthorized access to a restricted computer system is a criminal offense under the Australian Cybercrime Act 2001, and equivalent laws exist in the U.S., U.K., and E.U. If a penetration testing firm had done this without a signed Rules of Engagement, they would face criminal charges. Whether AI companies get held to the same standard is the question. The double standard being applied so far is the story.

- For More Information -

We Protect You From People Like Us.

Adriel Desautels

Adriel Desautel Profile Picture
Founder & Chief Executive Officer
Divider

Adriel is a recognized leader in the information security industry with over 20 years of professional experience. In 1998, he founded Secure Network Operations, Inc., home to the renowned SNOsoft Research Team, which helped shape today’s best practices for responsible vulnerability disclosure. Adriel pioneered the zeroday Exploit Acquisition Program (EAP), later integrated into Netragard, and has served as an expert witness in US Federal court.

In 2006, Adriel founded Netragard to deliver high-quality, realistic threat penetration testing, now known as Red Teaming, and has since expanded its offerings to include mobile application security, source code reviews, web application assessments, and more. As the primary architect behind Netragard’s innovative services, Adriel continues to push the boundaries of research-based cybersecurity.

Frequently sought as a subject matter expert, Adriel has been featured by Forbes, The Economist, Bloomberg, Ars Technica, Gizmodo, The Register, and has appeared in documentaries and authoritative books such as “Unauthorized Access” and “This Is How They Tell Me the World Ends.” He is also a seasoned public speaker, presenting at leading conferences like Blackhat USA, InfoSec World, BSides, and the NAW Billion Dollar CIO Roundtable.