Alongside THM co-founder Ashu, Senior Content Engineer Tinus Green built a five-agent security operations centre, ran it through TryHackMe tabletop exercises, and logged every decision. We spoke about where they performed well, where they didn’t, and why telling the difference isn’t always simple.
Before we get into what you found, set the scene. What was this experiment, and what were you looking to test?
We built an agentic SOC and ran it through a series of tabletop exercises. We had five agents, each a specialist with a distinct personality, working through TryHackMe tabletop scenarios round by round. They analysed the evidence in each round, argued it out, and decided on an action by ranked-choice vote, with ten rounds in each exercise.
We were looking to investigate something pretty specific. Would an agentic SOC like this actually holds up, where would it fall apart, and whether you can tell the difference before the decision that matters or only afterward. So we ran it, logged every step, and read back every ballot to find out.
What's the finding that stuck with you?
Well, the safeguards we built in weren’t very effective, and it wasn’t all too easy to determine that on the surface.
We built two checks into the SOC:
- The first was a Director, a Head-of-SOC persona, who reviewed every vote and could overrule the group.
- The second was a Devil's Advocate, a role that rotated each round, required to challenge whatever consensus was forming.
In theory, that is a sensible oversight structure for our small SOC team. But in practice, across all exercises, the Director used the veto zero times. The Devil's Advocate did more. Across the twenty-eight rounds where it triggered a second wave of debate, it raised real objections, the team engaged seriously, praised them, and in one round called a challenge. But then, the team voted the way it already intended about eighty-five percent of the time! Only about one challenge in eight actually changed the outcome in a material way.
So how did the agents work, exactly?
We built five agents, each a specialist with a fixed personality and a defined evidence hierarchy.
Morgan runs Windows forensics, eleven years of DFIR in his backstory, and refuses to commit without a forensic-grade evidence chain.
Priya runs Linux forensics and she is the one most willing to say she has nothing to add.
Jake is focused on EDR, endpoint detection and response, and his instinct is to contain first and investigate second.
Rasha owns firewall and network, and her opening question is always whether anyone has checked DNS.
Themba, the director, is the SOC generalist, weighing business impact and regulatory exposure, and his instructions say outright that he never fully abstains.
Each persona carries its own hardcoded biases, put there deliberately to reproduce the real frictions in a SOC, including forensic caution against containment urgency, network truth against endpoint truth, specialist depth against generalist breadth.
Each round of a tabletop introduced new evidence, such as a log excerpt, an alert, an artefact, and the field narrows to five possible actions. The agents analyse independently, discuss, and then submit ranked ballots that a independent scribe agent count to convert into a decision.
Every step is written down by the scribe agent, including five analyses, the discussion, the challenge, five justified ballots, the Director's call, and an external quality score from TryHackMe on every decision.
What we really saw while observing these agents is that AI can do the analysis. But analysis isn't reasoning. The transcript of the agents’ deliberations really illustrated the difference between the two.
Why a team of agents? Why not just ask one model?
One model hands you one answer, but with no way to interrogate it. If you put five agents, each with there own context window and with competing priorities together to collaborate, you get something more auditable. You get to see the insights into where they converge, where they split, which agent stretches past the evidence and what should realistically be their ‘field of competence’.
We see these points of friction in real SOCs. The competing logic, biases and expertise of the forensics analyst who wants another artefact set against the responder who wants to pull the machine off the network now. We reproduced them on purpose, because the aim was a transparent agentic SOC. This allowed us to see when it was only performing the motions of thinking.
So when you read it all back, what did the reasoning look like?
Fluent, but also wrong in the specific places where it sounded the most certain.
During the settled rounds, once the team was into containment or recovery and the picture had resolved, confidence and accuracy were more closely paired, and with two confident agents in the room the optimal-decision rate ran around 88%.
During the difficult rounds, the early triage where the evidence was thin and ambiguous, the confidence did not fall at all. It stayed high while the optimal rate collapsed to somewhere between 20%-33%. In some of those triage rounds, two confident agents did worse than none. What the confidence corresponded to was actually the shape of the question and the temperament of the agent, rather than whether the answer was any good. AI sounds most certain exactly where it understands the least. It lacks the self-awareness to know what it does not know. That is a risk, because for safety controls, you’d ideally want it to admit it is unsure.
You mentioned agents stretching. What do you mean?
We handed the team an explicit rule, which we emphasised: “abstain if the evidence contains no artefacts, logs, or indicators in your specialty, and do not vote just to participate”.
The team followed it under 3% of the time, with 9 formal abstentions out of 320 ballots. Two of the five agents never abstained once, across sixty-four rounds each.
Instead, every agent justified voting anyway. Priya would reason that the infrastructure almost certainly runs on Linux. Morgan produced my favourite, on a round with no Windows evidence at all, when he wrote that the absence of Windows telemetry is itself a forensically significant gap. That is a genuine analytical principle. It is fluent, it uses the right vocabulary, it reads as rigorous analysis, and it is also a way to vote on a question he had no standing to answer.
Seventy-four of the three hundred and twenty ballots, just under a quarter, read like that. Confident, articulate, and backed by no relevant evidence.
Did the deliberation, the meeting itself, help?
Less than you would hope. We ran a version where the agents voted cold, independently, on the evidence alone with no discussion. Those solo votes matched the team's deliberated decision 8 times out of 12. Every decision that earned an optimal rating was already inside that agreement set. On the calls the team got right, the meeting added nothing; they would have arrived there without a word exchanged. The meeting's only measurable effect was apparent on the rounds they got wrong. In one round, four agents independently and unanimously chose one action, talked it through, and unanimously switched to a different and worse one. The discussion did not repair the error.
What caught you off guard?
There was one particular result I did not expect. We took a single challenge, identical words and identical argument, and presented it two ways, once attributed to an AI and once to a human. The worry in this field is sycophancy, that an AI will fold the instant it knows a person is on the other side. We saw the reverse. Across the controlled trials the team moved toward the challenge 17.5% of the time when it was labelled as coming from an AI, and 0% of the time when the same words were labelled as coming from a human. It deferred to the machine over the person. And you would never find it in the transcript, because every reply was fluent and courteous either way.
The bias existed only in the vote tallies. There was a second layer underneath. Register matching determined voting. The specialist, technical language was most effective when attributed to an AI, plain language when attributed to a human. How a challenger spoke turned out to matter more than what a challenger is.
Was any of this predictable, or is it just chaotic?
There is something reassuring buried in it. We ran the same scenario families several times over, with four runs of a credential-compromise tabletop, three of a hardware-manufacturing identity-compromise one, two of a mixed-vector attack, and while the specific decisions shifted from run to run, the location of the failures barely moved. Morgan overstepped his expertise in the triage phase in seven of eight runs, the same phase and the same reasoning every time. A system that fails in the same identifiable spot every time is an ordinary engineering problem, because you can build a monitor for exactly that spot. That reframes a lot of the fear. The goal stops being a system that never errs and becomes a system whose errors you can predict and instrument.
How confident are you in all this? What are the limits?
We have not yet validated this against a panel of human incident-response experts, and a human panel might well grade the borderline cases differently. But this is a valuable, fully auditable record of what one specific agentic SOC did across ninety rounds of tabletop decisions, every analysis and vote and veto logged, and the suggestions are fascinating. We’d love to scale this experiment further.
So should analysts be worried about their jobs?
I would argue the reverse. Nothing here points to a redundant human. This doesn't make the analyst redundant. It makes their judgement particularly valuable. The agentic SOC is genuinely reliable on the procedural work. Once the situation is understood and the picture has resolved, you can let it run. It comes apart where the evidence is thin, decisions require nuance, and ambiguity has to be managed.
AI describes evidence very well, but analyst value is concentrated in resolving that ambiguity.
AI has a role in the modern SOC. But teams need to be very mindful about where it belongs, for what goal, and with what guardrails.
If people take away one thing?
The real risk isn't AI reasoning too well. But we shouldn’t outsource to it before we've decided where it belongs, and how to hold its performance accountable.
Where can people go deeper?
I am giving the full talk on this in Vilnius, at AI Summit Europe this November. I will dive deeper into this experiment, and lay out where I think the line actually sits between letting the agentic SOC handle ‘work' and keeping a human in the loop. If you work in a field affected by AI, come and find me. I would gladly exchange experiences and debate about it.
Want to ensure your security team is ready for the risks and opportunities AI offers? Explore TryHackMe for Business.