Skip to main content

Get back to routine: 30% off Premium and Certifications

05hr
:
20min
:
06sec
BLOG • 3 min read

How to Learn AI Red Teaming From Scratch

Scroll LinkedIn long enough and you'll find "AI Red Teamer" listed as a brand new headcount category, as if someone invented offensive security from nothing the day ChatGPT shipped. They didn't. Most of what gets badged AI red teaming is red teaming, full stop, recon, chaining exploits, thinking like the person trying to break something, pointed at a target that happens to take plain English as input instead of a port scan. Companies posting for it like it's an entirely new specialism is the same move as a company in 2005 hiring a "wireless penetration tester" and treating it as unrelated to penetration testing. It wasn't a new job then. It isn't now.

The 80% That Isn't New

Take EchoLeak, the vulnerability Aim Security disclosed in June 2025, a zero-click flaw in Microsoft 365 Copilot rated 9.3 out of 10 on severity. A single crafted email, already sitting in the inbox Copilot was reading as part of its normal job, got the model to leak internal documents without the victim clicking anything. Strip away the AI framing and what's underneath is a textbook chained exploit: evade the defence in place, find the bypass, chain it with a second and third bypass, exfiltrate through a channel nobody was watching. Evading Microsoft's cross-prompt injection classifier, slipping past link redaction with reference-style Markdown, abusing automatic image pre-fetching, proxying data out through a Teams preview API, that's four separate pieces of standard adversarial thinking stacked on top of each other. None of it required a machine learning degree. All of it required exactly the tradecraft red teamers already had.

That's the 80%. Recon, understanding a system well enough to find its blind spots, chaining small bypasses into one working exploit, none of that got reinvented for AI. It got redirected.

The 20% That's Actually New

Here's the part that is genuinely new, so don't mistake the argument above for "there's nothing to learn." The entry point EchoLeak actually used, indirect prompt injection through content the model processes as part of its normal workflow, is an attack class that didn't exist before language models did. Jailbreaking, getting a model to ignore its own safety training specifically, is new too. So is data poisoning inside a RAG system, feeding an attacker's document into a retrieval pipeline months before anyone notices what it's done. This is the layer where the old tradecraft doesn't transfer automatically, and pretending it does is how experienced red teamers end up embarrassingly stuck on a target that doesn't behave like a network.

Build the 80% First

If you already have red team fundamentals, recon, initial access, how a real engagement actually runs, skip ahead, you're most of the way there already. If you don't, don't let AI's novelty pull you past this step. TryHackMe's Red Team Fundamentals module is a genuinely free starting point for exactly this, no subscription required. Worth being straight about the ceiling here: the fuller, structured Red Teaming path moved to TryHackMe MAX for new subscribers as of June 2026. That's not a beginner tax bolted on for no reason, it's the same shape TryHackMe uses everywhere, breadth on the free tier and Premium, practitioner depth on MAX once you've outgrown the basics. You won't need MAX to start this guide. You'll probably want it once you're good enough to feel its absence.

Layer On the 20%

Once the tradecraft's solid, TryHackMe's AI Security path is where the actually-new layer lives, starting with how models and their training data work, because you cannot red team a system you don't understand, then moving into the prompt injection and jailbreaking rooms that cover the exact attack class EchoLeak used, then supply chain attack vectors and RAG data poisoning, the unglamorous half of the job that never trends on Twitter because a poisoned document sitting dormant for months doesn't demo as well as tricking a chatbot on camera.

Prove You Can Actually Do It

Reading about a four-step chained exploit and executing your own are different skills, and TryHackMe's AI1 certification is built to test the second one, hands-on and scenario-based rather than multiple choice, across the ground the steps above cover.

Who Actually Gets Hired for This

Not the person with the flashiest jailbreak screenshot. The job postings might read like AI red teaming is a shiny new title, but the people actually getting hired for it are the ones who already knew how to run an engagement and added the AI-specific layer on top, not the ones who learned a clever prompt and skipped everything that comes before it. If you're building this skill set properly, in that order, you're not chasing a trend. You're the person a hiring manager was already looking for before the job title existed.

authorNick O'Grady
Sep 4, 2026

Recommended

Get more insights, news, and assorted awesomeness around cyber training.

Join over 640 organisations upskilling their
workforce with TryHackMe