The Autonomous Pentester That's Forbidden From Pentesting

ARTEX is a multi-agent attack system whose own terms bar it from ever touching a live target — a preserved specimen of how far agentic offense engineering has gotten.
Somewhere between the insurance managers in Rolling Meadows and the textured ceiling coating that once contained asbestos, there is a third Artex: an open-source, LLM-driven autonomous penetration testing system, written in Go with a Next.js front end, whose repository currently on offer is — by its own description — a corpse.

The README states it plainly, in Chinese: this repository is the pure source backup of ARTEX’s final version. The Docker deployment source has expired; if you want it running, the author suggests, let an AI build it locally for you. That is not a release note. That is a headstone. And the attention the repo is attracting now is largely the attention a well-engineered machine attracts after it stops — the chance to open the hood without anyone telling you the engine might still be under warranty.
Two graphs, one brain
Strip away the screenshots (dashboards, asset maps, human-in-the-loop chat panes) and ARTEX’s genuinely interesting idea is architectural: it refuses to keep one graph, and instead runs two, connected by anchors.
The first is the asset graph — a global, cross-task store of what is true. Nodes are root domains, subdomains, IPs, services, apps, and endpoints, organized under company ownership. Crucially, the hierarchy and deduplication are computed by the program, not the model. Agents submit raw observations; the system decides what is a child of what and whether it has been seen before. This is a quietly important design choice, because the single most reliable way to make an LLM agent useless is to let it manage its own bookkeeping.
The second is the exploration graph — per-task, ephemeral, a record of thinking rather than being. Its nodes are goals, intents, facts, findings, and hints, chained by edges with names like spawns, derived from, yields, and proves. It answers the question traditional pentest reports cannot: which direction descended from which observations, and produced what.
The two graphs touch through anchor records that pin an intent or a finding to a specific asset. From one side, you can ask what a line of investigation actually hit. From the other, you can ask which assets were tested, by whom, with what results — which is what powers the coverage visualization, the force-directed asset map with tested nodes highlighted. Most agent frameworks conflate world-state with plan-state in one context window and hope for the best. ARTEX’s split is the kind of boring, structural decision that actually determines whether an autonomous system can be audited after the fact.
A planner with a memory problem, solved
The engine is event-driven: any change to the graphs wakes the planner, which reads the situation and dispatches zero or more intents into a frontier queue. Workers claim intents one at a time, execute with real tools, write facts, assets, and findings back to both graphs, and stop. The write-back triggers the next wake-up. The loop closes when a goal is proven.
Here is the wrinkle worth dwelling on: each planner wake-up is a fresh, stateless session. That is normally fatal for multi-step work. Real attack chains are serial — find the injection point, harvest credentials, move laterally, escalate — and dispatching all of that in parallel, from a planner with amnesia, produces chaos.
ARTEX’s answer is a shared todolist, persisted per task and carried across wake-ups. The planner records a dependency-ordered chain once, then in each subsequent round dispatches only the next step whose prerequisites have actually produced facts, marking earlier steps complete as their evidence lands. The chain survives the statelessness of the sessions that execute it. It is the same class of problem — and roughly the same class of solution — that any long-running agentic system eventually rediscovers: the conversation is disposable; the plan must not be.
Workers reading each other’s mail
The second autonomy mechanism is subtler. During deep exploration, the valuable observations — an odd error message, a hidden parameter, a suspicious response — often occur inside one worker’s execution process without ever being promoted to a formal fact. Later workers would normally repeat that labor.
ARTEX gives workers the ability to search the execution traces of other workers in the same task, excluding their own, then pull the full content of specific steps. Information flows between workers at the granularity of the process, not the conclusion, while the boundary holds: each worker still executes only the intent it claimed. It is the difference between reading a colleague’s lab notebook and being assigned their experiment.
The cage around it
All of this runs inside a fairly serious containment apparatus. Every Bash and HTTP action passes through a recording MITM proxy with CA verification, leaving a full traffic trail. Tool calls pass an approval gate; a human-in-the-loop chat lets a person steer mid-task. There is a dedicated retester agent that re-examines reported vulnerabilities and returns one of three verdicts — still reproducible, fixed, or unconfirmed — and a successful “fixed” verdict automatically updates the finding’s status. The system ships as a single Go binary with the frontend embedded, runs against PostgreSQL, and applies its schema idempotently on every start. It also integrates with ScopeSentry for asset import and supports remote MCP services.
And then there is the license section, which is where the project becomes genuinely strange.
AGPL, with an asterisk the size of the use case
ARTEX is licensed AGPL-3.0, which obliges network operators to share source. Fine, standard, unremarkable. But the author then appends a usage restriction: the tool is for reading, learning, and local isolated-environment verification only. It is forbidden to scan, probe, exploit, or attack any online system — regardless of authorization, and even if the assets are your own. Actual penetration testing, offensive-defensive exercises, and production use are all explicitly prohibited.
There is a legal tension here the README half-acknowledges: open-source licenses, including the AGPL, do not restrict fields of use. The prohibition is framed as a separate agreement between author and user, layered on top. Whether that layer would hold up anywhere is an open question; what is not in question is its effect on the project’s identity. ARTEX is an autonomous penetration testing system that its own author forbids from performing penetration testing. The demo mode, notably, only generates clearly labeled simulated records and never requests real targets — the software is consistent with its own epitaph.
One can read this two ways. Charitably, it is a researcher who built the thing to prove it could be built and wants no part of what people might do with it — a position with some precedent in security tooling, and arguably a rational one for a solo maintainer in a jurisdiction with active cybersecurity legislation (the README cites China’s Cybersecurity Law, Data Security Law, and Personal Information Protection Law). Less charitably, it is a demonstration that the hard part of autonomous offense was never the architecture — it was deciding who is allowed to flip the switch. Either way, the repo’s current status as a final-version backup makes the restriction feel less like a speed bump and more like the reason the project ended.
Where it sits in a crowded, mostly commercial, field
The timing of ARTEX’s preservation is not accidental. Autonomous penetration testing is having a commercial moment. FireCompass argues that traditional pentesting cannot keep pace with attack surfaces where, in large cloud estates, 20–30% of externally exposed assets exist for only days, and where micro-exposures — a privileged webhook token alive for three hours, a CI artifact leaking credentials for forty seconds — are invisible to scheduled assessments. Picus sells autonomous agents that plan, adapt, and chain techniques against live defenses, producing per-step verdicts on whether controls actually blocked anything. Synack’s Sara platform deploys agent swarms for discovery but insists on human validation of exploitability, positioning the human researcher as the differentiator. SecureLayer7 describes the same shift: continuous, self-adjusting testing versus point-in-time audits. A Manning early-access book on AI agents for offensive security is in progress, aimed squarely at working red teamers.
Against that backdrop, ARTEX is the odd one out: self-hosted, source-available, inspectable down to the planner’s prompt. You cannot read Picus’s orchestration logic. You can read ARTEX’s. For researchers and the merely curious, that is the entire value proposition — a full reference implementation of the planner/worker pattern, dual-graph state, cross-worker trace retrieval, and approval-gated execution, in one AGPL repository. What it is not, on the available evidence, is a validated one. The sources contain no benchmarks, no third-party evaluation, no named adopters, and no indication the system was ever run at scale against anything it was allowed to touch. The commercial vendors, whatever their marketing gloss, at least claim per-step verdicts against live environments; ARTEX’s demonstration mode, by design, never does.
What remains unclear
Several things the README leaves vague are worth flagging rather than papering over. The effectiveness of the planner’s todolist mechanism — whether chains actually complete without human rescue — is asserted architecturally but never demonstrated with results. The retest history is not yet included in report exports or task archives, and is not automatically linked to captured traffic, which the README admits. The Docker deployment path is dead, which means the path of least resistance into the project now runs through an AI-assisted local build — an ironic dependency for a tool whose whole thesis is that agents can do this kind of work. And the relationship to the referenced Cairn project, and to the author’s own norma agent SDK that supplies the agent runtime, is sketched but not documented in the provided material.
What is not unclear is the artifact’s value. ARTEX joins a small set of projects that matter less as tools than as texts — readable answers to the question of how you structure an autonomous offensive agent when the state is too big for a context window, the plan too long for one session, and the actions too dangerous to run unattended. The dual-graph split, the persistent todolist over stateless wake-ups, and process-level trace exchange between workers are transferable ideas, applicable well outside security.
It is just worth remembering, before anyone gets ideas, that the author has already had them for you. The tool forbids its own use. Read it, learn from it, build in an isolated lab if you must — and if you need a ceiling textured or a captive insurer managed, that is one of the other Artexes.
Sources
- Welcome to Artex Risk | Artex Risk
- AI Agents for Offensive Security - Mark Foudy
- Why Autonomous Penetration Testing Will Redefine Cyber ...
- Aviation
- how are you actually using AI in pentesting? (5 free copies)
- Automated Penetration Testing
- Artex
- Autonomous pentesting using artificial intelligence
- Artex Risk Solutions
- AI Pentesting Platform with Human Validation
- Autonomous Pentesting: How AI is Changing Offensive ...