ARTEX: The Pentest Brain That Plans, Remembers, and Asks Permission

An open-source autonomous security system that turns reconnaissance, reasoning, and execution into a graph-driven workflow with human approval at the edge.

9 min read • View on GitHub • More from mhtsec

A wide editorial scene of a planning desk organized like a graph board. Three stations labeled Main Agent, Planner, and Worker are connected by thread-like lines, while a human hand holds a pen over approval checkboxes. Below them, a branching network of hosts, ports, and findings spreads across the page, separated from the approved path by a thin firewall-like barrier. The image explains that ARTEX is not a shell wrapper, but a control system for deciding what happens next.
ARTEX treats pentesting as a controlled decision loop, with planning, execution, and approval tied to a shared graph of facts.

ARTEX was originally designed for the purpose of learning and research. It aims to help enterprises and organizations conduct security risk tests within the scope of authorized assets and improve security protection capabilities.

Key Takeaways

Most pentest tools are built around commands. ARTEX is built around state. The difference matters because once a security workflow gets long, the hard problem is no longer running nmap or launching a check. It is remembering what was learned, deciding what to do next, and keeping custody over every action.

Why ARTEX Feels Like a Pentest Operating System

ARTEX is easiest to understand as a control plane for security work. The Main Agent handles steering, the Planner turns graph state into intents, and the Worker executes one task at a time. That structure makes the project feel less like a chatbot bolted onto Kali and more like an operating model for autonomous red teaming.

This loop is the heart of ARTEX. Human input shapes the Main Agent, the Planner emits intents, the Worker executes them, and the graph preserves lineage for the next cycle.

That loop is the real novelty. ARTEX is not trying to replace the operator, and it is not pretending the model can safely freewheel forever. It is organizing autonomy around custody, so every action has a place in the plan and every finding has a traceable source.

GitHub avatar for mhtsec, the project owner. The portrait is meant to be rendered as a WSJ hedcut-style editorial likeness from a verified reference photo, giving the article a human anchor without inventing a face.

The Planner, the Worker, and the Human in the Loop

The agent architecture is deliberately narrow. The Planner decides what matters next based on the graph. The Worker does the work. The human gate decides whether a tool action should proceed, be denied, or be steered midstream. That separation is what makes ARTEX feel credible instead of theatrical.

Human chat -> Main Agent -> Planner -> Intent queue -> Worker -> Tools
                                   ^                     |
                                   |                     v
                             Exploration graph <--- Facts and assets
                                   ^
                                   |
                           Compactor / digest

Human approval can intercept worker actions before execution.

This is also where the project moves past typical agent demos. A lot of autonomous systems can generate plans. Far fewer can maintain a persistent queue of intents, claim one task at a time, and let a human interrupt the execution path without collapsing the whole run.

Why the Exploration Graph Matters More Than the Shell

The graph is the memory model. ARTEX records hosts, ports, services, facts, and findings as linked nodes, then tracks lineage so the system can say where a conclusion came from. That matters because pentesting is full of partial truths, and partial truths become dangerous the moment their provenance gets lost.

DimensionLinear tool workflowARTEX graph workflow
Primary unitCommandIntent and fact
MemoryTerminal scrollbackLineage-rich graph
Decision styleOperator drivenPlanner driven
TraceabilityManual notesOwnership and parent links
Long runsContext driftsContext is compacted and preserved
Best fitSingle-task checksMulti-step autonomous exploration
A close-up editorial scene showing a dense bundle of graph nodes on the left being folded into a smaller sealed ledger on the right. A lantern marks the hot working set, while a storage cabinet holds older cold nodes in the background. The image explains how ARTEX compresses old context without losing the lineage needed for future reasoning.
ARTEX compacts old context instead of letting the prompt bloat. The graph shrinks, but the useful chain of custody stays intact.

That compaction step solves a practical problem. Security work generates too much state for raw prompts to carry forever. ARTEX trims the noise, keeps the lineage, and feeds the Planner enough structure to keep moving without losing the plot.

What the Self-Updating Lifecycle Says About the Project

ARTEX is designed to run like software that expects to survive contact with reality. The bootstrap and rollback logic are not decorative. They suggest an intent to support durable, unattended operation, which is exactly what you would want from a system that may spend hours exploring a target environment.

That production-mindedness changes how the project reads. It is not only an experiment in agent design. It is an attempt to make autonomous security work behave like infrastructure, with state transitions, recovery paths, and a stable record of what happened.

In view of the reality of tool abuse, the ARTEX project will no longer be updated and will be converted to a closed source. There will be no release of any version or maintenance support in the future.

ARTEX vs the Rest of the Landscape

ARTEX sits between classic tooling and autonomous exposure platforms. It is more structured than a shell-driven toolkit, more transparent than a black-box SaaS product, and more security-specific than a general multi-agent framework. That middle ground is the point.

ProjectPrimary purposeAutonomyState modelHuman oversightBest fit
Nmap / MetasploitManual scanning and exploitationLowSession or command basedOperator controlledTargeted checks and hands-on work
SliverCommand and control and post-exploitationMediumImplant and operator stateOperator controlledRed team operations
PenteraAutonomous exposure validationHighCommercial workflow stateGuided oversightEnterprise validation at scale
XBOWAutonomous pentesting platformHighSaaS task orchestrationPlatform governedCloud-first bug finding
MetaGPTGeneral multi-agent orchestrationVariableConversation and task stateDepends on implementationBroad software workflows
ARTEXAutonomous penetration testing systemHighGraph-backed intent lineageHuman-in-the-loopSecurity exploration with traceable custody

If classic tools are the arms, ARTEX is trying to be the brain. It does not just run checks. It decides what checks matter next, preserves the evidence chain, and pauses when the action should be reviewed by a person.

Sources and Notes