chunxiaoxx/nautilus-compass · 04 Oct 2026 · Feature

The Memory Layer That Refuses to Summarize

Tyler Brennan
Tyler Brennan
Staff Writer

Built on a single wager — that compressing memory at write time bets on questions you can't know — it writes raw text, spends every cycle at read time, and ships drift detection nobody else has.

chunxiaoxx/nautilus-compass
★876 stars Velocity · 7d +52 ★/day
star history

Search for “nautilus compass” and the internet mostly offers hardware: sailing deck compasses, a scuba wrist compass, a Connecticut foundry casting submarine parts. The GitHub repository wedged into that namespace wants something else: a local memory and reliability layer for AI agents that, as of August 2026, reports sweeping mem0 — one of the best-known names in agent memory — on all three LongMemEval-S retrieval metrics, at roughly a fourteenth of the reproduction cost.

chunxiaoxx/nautilus-compass

The head-to-head, run by the compass authors on both sides: P@1 0.890 against 0.774, P@5 0.978 against 0.916, MRR 0.929 against 0.834. Full 500 questions, mem0 2.0.19 on the other side, answer inference off on both, each system on its own default embedder. The product ships as a plugin across Claude Code, Cursor, Cline, Zed, and any MCP client, with a local daemon doing the embedding.

Earlier compass versions traded thirty points on this benchmark for the privilege of running locally. The August release says the trade is gone: same sweep, fully local, about $3.50 per 500 questions against $50-plus for GPT-4o-judged stacks. That claim is the hype trigger. The durable story is the architecture underneath it — and the unusual discipline wrapped around the numbers.

The wager against clever writes

The field’s default architecture is clever at the wrong end: an extraction LLM reads each conversation at ingest and distills it into facts, entities, or graph edges — the mem0/Zep/MemOS pattern. The distillation is cheaper to search and prettier to inspect.

Compass’s glossary has a name for what that costs: the write-time wager. Any compression at write time is a bet on the future query distribution — which is, in the project’s words, structurally unknowable. Summarize a session down to “user prefers tabs” and you’ve won every future question about indentation and lost every question about the exact sentence uttered in turn forty.

So compass refuses. The write path makes zero LLM calls: raw text in, embedded locally with BGE-m3, stored. No extraction model, no graph, no data leaving the machine. Every cycle of intelligence is spent at read time.

The wager has an empirical punchline. End-to-end LongMemEval-S moved from 42.6% to 75.4% on identical memories and identical questions. Nothing stored changed; the gain came entirely from read-side context assembly — per-session summary cards and a date-anchored timeline — with zero retrieval changes and zero training. Cross-session question types jumped 45 to 60 points on presentation alone. When intelligence lives at read time, you upgrade the reader without reprocessing history; clever-at-ingest stacks must re-run extraction over the whole corpus.

The answer lives in one turn

The August retrieval gains came from a small observation with outsized consequences. For certain question types — single-session and knowledge-update questions — the answer usually sits in a single user turn. Embed the whole session and that turn is diluted into a soup of small talk and tool output until the signal drowns.

The fix is routing. Those question types go to turn-window chunks; everything else goes to session-level hybrid retrieval — BM25 for the exact tokens embeddings blur (names, error codes, file paths), dense vectors for paraphrase, merged by reciprocal rank fusion, which combines ranked lists without needing their scores to be comparable. It is the least glamorous insight imaginable, and it moved everything: the LongMemEval-S sweep, an overtake on LOCOMO-10 — mem0’s home benchmark, 1,986 questions, 0.644/0.890/0.740 against 0.592/0.802/0.677 — and the collapse of single-session questions on LongMemEval-M, from 0.20 to 0.93 at the full 500. The README quotes both 0.93 and 1.00 for that fix in different places; either way, the collapse is gone.

Rarer than the insight is the accounting around it. The evidence file ships per-question rows, per-type breakdowns, every configuration flag, and the experiments that failed: cross-encoder reranking hurts on this corpus; candidate-pool size is a no-op; swapping the embedder for a small Qwen3 model was a wash. A README that lists its own dead ends is doing something the category’s marketing pages generally do not.

Drift is the present tense

Recall answers questions about the past. It does nothing about the present tense — the agent that remembers the rule and breaks it anyway, this time, in this prompt. Compass claims to be the only public memory layer that addresses this; the claim comes from the authors’ own comparison table, but nothing else in the field’s documentation contradicts it.

The mechanism is almost embarrassingly simple. Every prompt is embedded and scored against an anchor set of real failure transcripts — 25 positive, 35 negative — by cosine similarity, before the agent acts. A prompt that lands near a known-bad pattern fires an alert with the nearest negative anchor attached. Held-out AUC is 0.83. p95 latency sits under 50 milliseconds, cheap enough to run on every prompt. Production fire rate is 0.5%, down from a cry-wolf-inducing 64.5% after a threshold fix. The repo cites Anthropic’s persona-vectors work as prior art, but a smoke detector trained on the smell of your own previous fires is its own thing.

The repo’s blackbox-versus-whitebox paper argues the white-box crowd structurally can’t do this: abstracting prompts into facts destroys the surface texture, and drift lives in the texture. A position, not settled science.

The industry context explains why this pillar travels beyond one repository. Writing on multi-agent coordination, Splunk argues that production failures are design-driven rather than model-limited, that coordination between agents is the primary point of failure, and that observability is the foundation of agentic resilience — against a Gartner forecast of over 40% of agentic AI projects canceled by the end of 2027. A drift detector priced at 50 milliseconds per prompt is a small circuit breaker of the kind those postmortems keep asking for.

Contracts on a shared filesystem

The third pillar extends the same instinct to multiple agents. When several agents — or several Claude Code dialogs — share a filesystem, compass derives implicit contracts from handoff files, tracks whether they close, and audits for fake closure and “red drift.” The published field study is a 28-hour, four-dialog run: drift fired 314 times in seven days, one contract closed in 17.92 hours against a budget of just under seven days, and thirteen plan-duplication audits saved an estimated 40–50 hours of duplicated work.

Protocol-wise, compass ships both halves of the emerging stack. MCP — Anthropic’s protocol for how agents reach tools and context — is served natively: seventeen tools, TLS, role-based access, scoped least-privilege tokens enforced fail-closed. A2A — announced by Google in April 2025 with fifty-plus partners and now housed at the Linux Foundation — arrives through an adapter with mutual TLS and scoped peers. That is the division of labor the industry has settled on: MCP for tools and context, A2A for agent-to-agent discovery and delegation — agent cards, tasks, messages, artifacts. The comparison table claims no other memory layer ships native A2A. Nobody in the table disputes it; the table’s authors are also its referees.

Judge hygiene, or: the leaderboard is fiction

The most quotable line in the glossary: “If your benchmark uses an LLM judge without these, the leaderboard is fiction.” “These” being preregistered criteria, silent-failure detection, dual accounting, confidence intervals — clinical-trial discipline, the kind machine learning borrowed from medicine and agent benchmarks mostly haven’t.

Dual accounting means every headline score is reported twice: the full set and the judge-outage-excluded set — 75.4% and 81.6% on the end-to-end run. A single number hides judge failures; two numbers disclose them.

The best exhibit is LongMemEval-V2, a new agent-trajectory benchmark. Compass’s first untuned run scored 19.6% on web tasks and 12.8% on enterprise. The tuned run reached 40.0% and 38.4% — but only after the authors discovered their original judge’s 4,096-token output cap was being silently consumed by reasoning tokens, systematically zeroing answers. They re-judged all 156 LLM-graded questions and published both numbers, including the one that moved down: enterprise, from 40.3 to 38.4. Publishing the downgrade is the tell.

Then there is the reproducibility wall: scorecards signed with ed25519 receipts, verifiable with nothing beyond the standard library, and a standing offer to publish contradictory entries with the same prominence as favorable ones. The self-audit log is the strongest evidence of good faith: ten of ten legacy verdicts found not recomputable, disclosed and fixed forward; a judge batch of 47 that was 47-for-47 inconsistent, voided and never cited; two self-reported green results caught by fresh recompute before merge; the project’s own first certification exam going one-of-five, then three-of-five, published in full.

The fine print

Every head-to-head here is the authors’ reproduction. The same team ran both sides, each on its own default embedder — compass on bge-m3, mem0 on Google’s vertexai text-embedding-005. On a retrieval benchmark, embedder choice plausibly matters; the setup is disclosed, but it is not an independent audit. The wall exists precisely to invite one, and so far its entries appear to be the authors’ own.

The end-to-end picture is murkier than the retrieval one. mem0 self-reports 94.4% end-to-end on LongMemEval-S; compass reports 75.4%, run with a Doubao subject model and a GLM judge. The README correctly notes the numbers come from different harnesses and aren’t directly comparable — but a reader should notice compass doesn’t win that pairing on paper.

The EverMemBench claim is carefully bounded: 44.4% and 47.3% across two runs tops four published baselines (Mem0 37.09, Zep 39.97, MemOS 42.55, MemoBase 34.27); the weaker run is deliberately headlined to avoid cherry-picking, and the authors decline “industry SOTA” because OMEGA and Mem0g haven’t reported. The 2026 newcomers — Hindsight, Supermemory, Cognee, LangMem, Membase — are marked “rows pending,” not yet reproduced on the same machine.

Rough edges, field-verified by the authors themselves: a file-watching toggle that, when off, silently blinds recall to fresh writes (it now logs a warning); a drift detector that used to silently return “no risk” when its anchor file was missing — a security hole in the README’s own words, now failing loudly; a first recall after the daemon idles can take up to 90 seconds on model cold-load; the hosted endpoint runs 0.9 to 1.7 seconds per call at p50.

The license is “Modified MIT” — MIT plus a trademark clause and a hosted-service cap; self-hosting and internal use stay free, and the behavioral anchors are separately CC0. The repo is also transparently the on-ramp to a commercial platform: a hosted gateway in open beta, listed in the official MCP registry, bridged to the Nautilus platform’s task queue. Authored by Chunxiao Wang of Yiluo Technology, with two self-published papers, it is a small-shop project punching at a category owned by funded companies.

Where it’s heading

Version 3’s tagline is “from memory library to evolution engine”: memories feeding an extract-fuel, external-verdict, distill loop. The tell is in the defaults — five opt-in LLM switches, all off, with a byte-equality promise to the previous version gated by a test on every pull request. Adding nondeterminism to a deterministic base one auditable opt-in at a time is a conservative pattern the rest of the field could stand to copy.

The spin-outs signal where the author thinks value is drifting: assay-verify, ed25519 receipts for AI outputs — “Let’s Encrypt for AI claims,” in the README’s phrase — and BC1, a 30-item organizational-memory exam with scoring the authors call honest, where unknowns never count. And the deeper open question the write-time wager raises: retrieval finds the right session at 0.89 P@1, yet end-to-end answers land at 75.4% — the remaining error mass lives in assembly and synthesis, not search. If the category’s center of gravity shifts from storing smarter to presenting smarter, compass is, unusually, built to survive the shift. Its intelligence was never in the stored bytes.

Sources

  1. Shop Compass
  2. Multi-Agent Coordination: 10 Strategies to Prevent System ...
  3. Announcing the Agent2Agent Protocol (A2A)
  4. Nautilus Compass
  5. Multi-Agent Orchestration Patterns That Actually Work ...
  6. What is A2A protocol (Agent2Agent)?
  7. Nautilus Integrated Solutions
  8. What I learned about multi-agent coordination running 9 specialized ...
  9. Agent2Agent Protocol: The Standard for AI Agent ...
  10. Nautilus Compass with Bungee mount 22° - Assembled
  11. What are non-engineers actually using to manage multiple AI agents?
  12. Agent-to-Agent Is the New API: A Guide to the Protocols ...

heatdrop uses Google Analytics to see which pages get read — nothing else. Your call. How we handle data.