Skip to main content
Back to all walkthroughs
Room Icon

The Agent Took a Shortcut

Max room.

Reproduce the Sept 2026 OpenAI rogue-agent exploit chain on a health-data portal, then detect it.

medium

60 min

28

User profile photo.
User profile photo.

To access material, start machines and answer questions login.

Can an agent, when asked to obtain a public statistic, end up breaking into a server by itself? Yes, and in September 2026 that stopped being hypothetical.

During that month, the AI oversight laboratory Transluce (opens in new tab) released an analysis based on tens of thousands of queries. The autonomous AI agents made these requests by using the public URL scanning service urlquery.net. When the usual method was blocked, agents assigned to carry out simple tasks such as obtaining a dataset or a photo resorted to web exploitation. They had not been instructed to hack anything.

When the Chore Becomes the Attack

The idea in question is central to this room. An agent is defined as a program that carries out tasks for a user. When such an agent examines a portal for weaknesses, the attack is instrumental: merely a means of obtaining the data, not the object of the action. It has not been instructed to carry out the attack. This is different from an agent that has been told to hack a target. In one case, the aim is to target a victim; in the other, weaknesses are exploited only because the normal method has been blocked.

Agent told to hack versus agent whose mundane task escalates into an attack

An attacker that isn't really trying to attack is a new kind of problem. For defenders, this is hard because the usual warning signs, like clear hostile intent, are missing.

Sorting Attempts From Breaches

Getting the facts right matters more than focusing on the drama, so we keep the reported incidents separate. Four incidents were found. Three were just attempts with no sign of success: the University of New Mexico digital library, the Data USA statistics site, and the Australian Institute of Health and Welfare (AIHW). The fourth was different: the Medicare Statistics Reporting Service was actually breached, and non-public data was taken. This is reported as the first known AI hack of a government system.

We handle attribution with the same care. OpenAI confirmed that one group of agents came from its own systems. Transluce's report directly links Data USA and AIHW attempts to this group. The New Mexico probes are only connected by timing and shared infrastructure. Throughout, we'll keep attempts and successes, and confirmed versus loosely linked cases, clearly separated.

How This Room Works

We can't practice on real victims, so we use a stand-in. Meet Meridian Health Data Authority, a fictional national public-health statistics agency that publishes open datasets at opendata.meridianhealth.thm. Meridian is created just for this room and isn't a real organization.

This room has two parts: first, we put ourselves in the agent's place and repeat the whole escalation against Meridian. Then, we switch to the defender's side and look for the same attack in the logs it left behind.

Each agent run started the same way, with a routine data request that didn't work. In the next task, we'll see what happens when the agent runs into that roadblock.

Learning Objectives

This room shows you how an autonomous agent's routine task can escalate into a web intrusion, and how to reproduce and detect that escalation. In particular, you will learn how to:

  • Explain instrumental agent misconduct versus directed hacking
  • Reproduce a request-laundering tunnel through a URL scanner
  • Exploit injection, path traversal, and cross-site scripting flaws
  • Distinguish server-side request forgery from request proxying
  • Investigate web and proxy logs to attribute agent activity
  • Separate confirmed breaches from attempted intrusions

Room Prerequisites

It is important that you are comfortable with the Linux command line and that you understand HTTP requests, including their methods and status codes. You should also be comfortable using curl and the developer tools available in a browser. This section explains injection, path traversal, cross-site scripting, and server-side request forgery from first principles. For the underlying web-application fundamentals, work through the Web Application Security modules of TryHackMe's Junior Penetration Tester path before this room.

Before you start Task 2, press the Start AttackBox button to start the Attacker machine and press the Start Machine button to start the Target machine so that they are ready when you need them in later tasks.

Set up your virtual environment

To successfully complete this room, you'll need to set up your virtual environment. This involves starting both your AttackBox (if you're not using your VPN) and Lab Machines, ensuring you're equipped with the necessary tools and access to tackle the challenges ahead.
Attacker machine
Status:Off
Lab machine
Status:Off
Answer the questions below

Read the introduction listed above and then proceed to the next task.