The Reverse-Engineering Agent That Refuses to Guess

REA hands coding agents Hopper and Ghidra and asks them to investigate binaries the way a careful human analyst would — with evidence, citations, and an honest ledger of everything they still don't know.
Every few months, a new project promises to let AI agents do reverse engineering. The pitch is always some version of the same sentence: point the model at a binary, and it figures out what the software does. The results are usually a mix of genuine cleverness and confident nonsense — a decompiled function summarized plausibly, a call graph sketched from vibes. A systematization of knowledge paper from McGill and Defence R&D Canada, reviewing 44 research papers and 18 open-source projects that apply large language models to reverse engineering, reached an uncomfortable conclusion: these systems hallucinate, struggle with reproducibility, and lack grounding in machine-level semantics. The field’s problem isn’t that models can’t read assembly. It’s that nothing in the pipeline forces them to say what they actually know.

REA — Reverse Engineer Anything, published as the rea-agents npm package by morluto — is an attempt to build the forcing function. It is a CLI and MCP server that gives coding agents a structured way to investigate software: native binaries through Hopper or a bring-your-own Ghidra installation, JavaScript and Electron applications through static reconstruction, managed PE/CLI artifacts through execution-free triage, and live behavior through tightly permissioned observation of browsers and Node processes. The pitch is not “the agent will figure it out.” The pitch is “the agent will show its work, and tell you when it couldn’t.”
Glue code, unusually disciplined glue code
Let’s be clear about what REA is at its core: it is orchestration. It does not contain a disassembler, a decompiler, or a debugger. Hopper — a commercial macOS- and Linux-native disassembler with its own license — does the deep native analysis. Ghidra, the NSA’s open-source suite, supplies a read-only inventory and function-analysis provider on Linux, pinned to exactly one official release with a full JDK 21. REA’s own contributions are the plumbing: a session router that binds a target to one provider deterministically, a serial FIFO that respects the fact that Hopper’s Python API lives on a single thread, an authenticated local socket bridge to Ghidra’s headless mode, and — the interesting part — an evidence contract layered over all of it.
That sounds like a thin contribution, and in raw line-count terms it might be. But the boring part is where the value is. The McGill SoK paper’s central complaint about LLM-based reverse engineering is that evaluation practices vary so wildly that results can’t be compared or reproduced. REA’s answer is that every successful tool call produces a deterministic Evidence record: artifact identity, provider identity, confidence, authority, limitations, and locations. Unknowns are not discarded — they’re tracked as “residual unknowns” through immutable revisions, linked to the evidence that produced them, with contradictions and probes recorded alongside. When a comparison between two application versions can’t establish equivalence because the evidence is incomplete, REA says so rather than implying absence of difference. The README states this as a design principle in almost litigious language: incomplete evidence never implies equivalence.
This is a strange and welcome posture for an agent tool. Most MCP servers expose capabilities and let the model decide what’s true. REA exposes capabilities and then audits the model’s epistemics.
The consent machine
The other thing you notice reading REA’s documentation is how much of it is about saying no. Setup — the one command everyone skims — is a small study in defensive interaction design. Nothing is preselected. The wizard prints exact paths and external effects before changing anything, defaults the final approval to No, and validates existing configuration first. Registrations into agent clients (Claude Code, Claude Desktop, Codex, Cursor, Gemini CLI, Windsurf) are additive, backed up, and read back after writing. Devin is detected but deliberately left alone because it has no documented local MCP configuration boundary — a level of restraint that borders on the obsessive, and is exactly right.
The same philosophy governs the dynamic capabilities. Browser observation over Chrome DevTools Protocol is disabled by default and requires a literal loopback endpoint plus explicitly approved page origins. It’s passive by construction: no JavaScript evaluation, no navigation, no clicks, and query values, credentials, cookies, and storage contents are never retained. The Node/Electron V8 Inspector observer sends exactly two enable messages and records script locations and execution-context lifecycles — require/import edges, EventEmitter activity, and Electron IPC are explicitly listed as unknown, not silently inferred. Process capture, browser scenarios, and extracted JavaScript replay each sit behind separate policy flags and fail closed.
The security model section is blunt about the limits: shipped providers and launched targets run with the current user’s permissions, and REA’s passive observers are not security sandboxes. The one genuine sandbox — extracted JavaScript module replay — is Linux-only and refuses to run unless Bubblewrap namespaces, architecture-checked seccomp, private runtime mounts, and cgroup limits are all present. The Windows Ghidra path is an explicitly labeled “P0” boundary, restricted to approved native x86-64 PE applications, rejecting DLLs, managed binaries, and anything mutable or hostile, and the documentation candidly admits it does not establish Job Object ownership or reparse-point-safe authority.
This candor is rare enough to be a feature. Most projects bury their threat-model gaps; REA enumerates them in the README.
Where it sits in a crowded field
REA is not appearing in a vacuum. The agentic-reverse-engineering wave is well underway. Cisco Talos published a detailed account of wiring an MCP server to IDA Pro and driving it from VSCode with a local model, treating the LLM as an analyst’s assistant rather than a replacement — and flagging real friction around tool-call cost and context-window limits on local hardware. Conference talks with titles like “How AI Agents Are Changing Binary Analysis” and Black Hat USA 2025’s clue-driven reverse engineering session point the same direction: the community is actively working out how much of the analyst’s workflow can be delegated. LinkedIn Learning now has a course segment on the topic for people who will never read a paper about it.
REA’s differentiation within that wave is scope and honesty. Where most projects in the SoK taxonomy target one task — decompilation summarization, binary classification, function naming — REA tries to span the whole investigation: from a Mach-O binary’s strings and symbols, through cross-references and call graphs, up through Electron ASAR bundles and source maps, into managed .NET metadata and CIL hashes, and across to observed browser behavior, with a provider-neutral application graph connecting the layers. The stated ambition is that an agent can trace a feature from a UI string in a web bundle down to the native function that implements it, with every hop carrying evidence.
The second differentiation is the adversarial one, and it may matter more than the first. A recent paper demonstrates prompt-injection attacks against LLM-powered disassembly and decompilation pipelines: extraneous string assignments embedded in a binary that pass surreptitious instructions to the analyzing model without affecting the executable’s functionality. An agent that free-associates from decompiler output is an agent that can be lied to by the malware it’s analyzing. REA’s architecture doesn’t solve this — no system that feeds binary-derived text to an LLM fully does — but its insistence on digest-verified targets, authenticated evidence, and explicit unknowns is at least the right shape of defense. A pipeline that distinguishes “observed” from “inferred” from “unresolved” gives the analyst a place to notice when the observations and the narrative diverge.
The rough edges
None of this makes REA practical for everyone, and the README is its own harshest critic on that front. The center of gravity is macOS, with Linux as a solid second and Windows as an explicitly experimental, Ghidra-only, native-PE-only toehold. Deep analysis depends on Hopper — separate software, separate license, with a demo mode that REA automates on Linux by running it on a private Xvfb display and clicking “Try the Demo” with verified dialog geometry, which is either admirable engineering or a sign of how awkward the provider landscape is. Probably both. Ghidra support is bring-your-own, pinned to one exact release, and refuses to download anything itself.
The tool catalog is enormous — over a hundred operations across native inspection, managed PE/CLI, browser observation, Electron analysis, and workflow families — and the documentation’s density suggests a project that documents everything and simplifies nothing. For a developer who just wants to know how an app’s offline search works, the learning curve is real, even if the agent is supposed to absorb most of it. And the roadmap’s “later” items — native runtime observation via LLDB or Frida, mobile artifacts, firmware, IDA and Binary Ninja providers — are exactly the capabilities that serious reverse engineers would ask for first, and they’re not there yet.
There’s also a fair question about who the audience is. The security-research community that lives in Ghidra already has MCP bridges and, per Talos, working local-model setups. REA’s more natural constituency is the one its README actually addresses: product developers who see a feature in someone else’s app and want to understand it well enough to build their own version. That’s a use case with real legal and ethical fuzz around it — the README is careful to say REA doesn’t claim to recover original source or clone applications, and that’s the right disclaimer — but it’s also a use case that the traditional RE world has never served, because traditional RE tools assume you already know why you’re in the binary.
The bet
REA’s bet, stated or not, is that the scarce resource in agentic reverse engineering isn’t intelligence. Models are getting better at reading decompiled code on their own. The scarce resource is trust — the ability to look at an agent’s conclusion and know which parts are grounded in verified observation, which are provider-specific inference, and which are simply missing. The project answers that with evidence ledgers, residual-unknown tracking, deterministic provider selection, and a refusal to let a failed provider silently fall back to another. It is, in a field full of tools that make agents faster, a tool built to make them accountable.
Whether that rigor survives contact with a broad audience is the open question. The setup ceremony, the platform fragmentation, and the Hopper dependency all add friction that a hobbyist won’t tolerate and a security team may not need. But if the agentic-RE field is going to mature past the hallucination problem the research literature keeps documenting, it will be because someone built the plumbing that makes honesty the path of least resistance. REA is one of the more serious attempts yet to do exactly that — and it’s honest enough to tell you which parts of the attempt are still unknown.
Sources
- Home | REA Energy Cooperatives Inc
- How AI Agents Are Changing Binary Analysis
- Black Hat USA 2025 | Clue-Driven Reverse Engineering by ...
- Reciprocal Easement Agreement (REA) | Practical Law
- Agentic Reverse Engineering + Binary Analysis with Kong
- Potentials and Challenges of Large Language Models for ...
- Why is There an REA?
- Automatically Attacking Software Reverse Engineering AI ...
- Using LLMs as a reverse engineering sidekick
- MyREAEnergy
- Reverse Engineering and Binary Analysis with AI Assistance
- LLVM and AI plugins/tools for malware analysis ...