T3MP3ST Turns a Coding Agent into a Red-Team Swarm

A TypeScript meta-harness that treats LLMs as interchangeable brains, routes them through a mission loop, and insists on evidence instead of vibes.

9 min read • View on GitHub • More from elder-plinius

A wide control room with a central brain socket feeding multiple operator stations, while a mission board tracks recon, scan, exploit, and report. Beneath the desk, a sealed evidence vault stores logs, screenshots, and paper trails, showing that the system values proof as much as autonomy.
T3MP3ST is less a chatbot and more an operating room for offensive work. The model can change, but the mission structure and evidence trail stay in place.
Key Takeaways

Why T3MP3ST Feels Like a Security Operating System

Most AI security tools start with a model and bolt features onto it. T3MP3ST starts with the workflow. It assumes the brain can change, but the mission structure, tool policy, and evidence trail should stay stable.

That is why the project calls itself a multi-agent offensive-security meta-harness. It is not trying to be a smarter scanner. It is trying to be the layer that coordinates the scanner, the exploit runner, the analyst, and the rules that keep them honest.

A multi-agent offensive-security framework, built to turn the AI coding agent you already run into a zero-day hunter.

The Swarm: Eight Operators, One Kill Chain

T3MP3ST decomposes the job into specialized operators: coordinator, recon, scanner, exploiter, infiltrator, exfiltrator, ghost, and analyst. That sounds theatrical, but the point is practical. Narrow roles reduce prompt drift and make it easier to decide which tools are allowed to run.

OperatorJobWhy it matters
CoordinatorAssigns tasks and keeps state movingPrevents the swarm from becoming a free-for-all
ReconCollects initial facts and scopeCreates the first evidence packets
ScannerChecks for likely weaknessesTurns loose observations into structured signals
ExploiterTests whether a finding is realSeparates possibility from proof
AnalystTurns receipts into a reportForces the final output to stay grounded
A close-up handoff between two operators at a narrow gate. One side deposits an evidence packet into a queue, while the other side receives it only after a lever is approved. A blocked side path shows how dangerous actions are paused until the system permits them.
The important move is not more autonomy. It is controlled handoff. Evidence flows forward, but risky actions have to clear a gate first.

Inside the Mission Loop

The architecture is tick-based. A mission starts with an objective, gets broken into tasks, and then advances through a loop where the coordinator assigns work, tools execute, and evidence gets written into a vault. That structure matters because it makes the system inspectable.

This diagram shows how T3MP3ST turns autonomy into a sequence of constrained state changes. The interesting part is the gate, which keeps the loop from becoming an unbounded agent blur.

The key abstraction is not the model. It is the state machine. That lets T3MP3ST pause, resume, branch, and demand receipts before the next step gets the green light.

Why the Receipts Rule Matters

The project is blunt about proof. It wants evidence kinds, tool output, and verified proof-of-concept execution, not vague claims that something probably worked. That sounds obvious until you compare it with the average agent demo, where the output is usually a confident paragraph and very little else.

In T3MP3ST, claims are supposed to land in the evidence vault. That gives the system a memory of what happened, but it also gives the operator a way to audit the chain of reasoning later. In security work, that is the difference between a story and a case file.

The Benchmark Layer Is Part of the Product

The repo does not treat evaluation as an afterthought. Its benchmark suite, claim verification scripts, and anti-fitting checks make performance a reproducible artifact. That matters because offensive AI is full of inflated demos and cherry-picked wins.

LayerT3MP3ST choiceWhy it stands out
EvaluationRe-derivable claims and benchmark scriptsThe numbers are meant to be checked, not admired
Training contaminationNo-fitting style integrity checksReduces the chance the tool is just memorizing the test
AutonomyMission loop with gatesPrevents the agent from pretending a success without proof
ToolingReal adapters with risk levelsAllows dangerous actions to be deliberate instead of accidental

That makes the benchmark layer part of the product story, not just the research appendix. The framework is trying to prove that autonomous offense can be measured with the same discipline as it is executed.

How T3MP3ST Differs from Other AI Security Tools

T3MP3ST is often grouped with pentest assistants, red-team frameworks, and commercial autonomous security platforms. That is understandable, but the comparison can hide the real distinction: some tools test models, some automate scans, and T3MP3ST tries to orchestrate a controlled offensive operation.

ProjectPrimary purposeBrain modelTarget typeEvidence standardBenchmark philosophyAutonomy
T3MP3STRun a coordinated red-team missionSwappable coding agentsReal targets and authorized scopesReceipts stored in an evidence vaultClaims should be reproducibleHigh, but gated
PentestGPTAssist pentesting with AI guidanceSingle assistant modelHuman-led pentest workflowsMostly conversational outputLightweight or absentLow to medium
DeepTeamRed-team AI systemsTest harness for modelsLLMs, agents, RAGEval scores and adversarial casesFramework-driven evaluationFocused on model testing
XBOWCommercial offensive platformProprietary systemBlack-box targetsPlatform-specific validationBenchmark as competitionHigh
PenligentCommercial autonomous red teamingSwarm-style agentsEnterprise targetsVendor-defined evidenceProductized eval stackHigh

The difference is not just technical. It is philosophical. T3MP3ST assumes the model is disposable, the process is sacred, and the proof has to survive contact with the operator.

The Creator and the Positioning

A WSJ-style hedcut portrait of elder-plinius, the creator of T3MP3ST. The portrait frames the project as the work of one visible maintainer, which matters because the repo's identity is closely tied to a single public persona.

The project is strongly associated with elder-plinius, whose public identity already carried weight in the jailbreak and agentic tooling world. That matters because T3MP3ST launched with a ready-made narrative: this is not a generic enterprise product, but a creator-led offensive system with a point of view.

That background also explains the reaction. Some readers saw a dangerous new automation layer. Others saw a cleaner way to structure work that was already happening in messy, manual form. Both readings are true.

What T3MP3ST Suggests About the Future of Offensive AI

The deeper implication is simple: offensive AI is moving from standalone tools to composable systems. The model becomes a replaceable component. The real product is the orchestration, the policy layer, and the proof trail.

That shift matters well beyond security. Any agentic system that touches real-world consequences will need the same ingredients T3MP3ST emphasizes: constrained roles, auditable handoffs, and claims that can be checked after the fact.