thedotmack/claude-mem · 08 Oct 2026 · Feature

Every Agent Forgets. claude-mem Made That a Business.

Brandon Cole
Brandon Cole
Staff Writer

A hook-driven memory layer for Claude Code has ridden the agent-hype wave to nearly 100k stars — by treating forgetting as an engineering problem, and a revenue opportunity.

thedotmack/claude-mem
★97.8k stars Velocity · 7d +388 ★/day ↗accelerating
star history

Somewhere between the moment you close a Claude Code session and the moment you open the next one, everything dies. The architecture you explained, the bug you chased for an hour, the three approaches that failed before the fourth worked — all of it evaporates, and the next session greets you like a stranger. This is the defining annoyance of the agentic coding era, and claude-mem has turned it into one of the most-starred projects in the AI tooling space: roughly 97,000 stars at last count, with a Trendshift badge, thirty-one translations of its README, and a commercial layer riding on top.

thedotmack/claude-mem

The project, built by Alex Newman (@thedotmack) in TypeScript on top of the Claude Agent SDK, is a persistent memory compression system for Claude Code and, increasingly, for everything around it — Gemini CLI, OpenCode, Codex, Cursor, and a growing roster of harnesses. The pitch is simple enough to fit in a sentence: capture what your agent does during a session, compress it into semantic summaries, and inject the relevant parts back into future sessions. The execution is where it gets interesting.

The amnesia problem, stated properly

The reason agent memory is having a moment is that the context window turned out to be a trap. It looked like a solution — bigger windows, longer sessions, more tokens retained — but as one Hugging Face forum discussion on long-running agents puts it, the simplest rule is: do not make the prompt the memory system. Store memory outside the model, retrieve only what the current step needs, and treat the model as a reasoning engine while the runtime owns the state.

That is precisely the architecture claude-mem implements, and it does so by exploiting a feature most Claude Code users never think about: lifecycle hooks. The system registers hooks at five points in a session’s life — session start, user prompt submission, post-tool-use, the stop event, and session end. Every file read, every edit, every command execution becomes an “observation” written to a local SQLite database with FTS5 full-text search. A worker service, managed by Bun and serving an HTTP API on a local port, processes those observations in the background, extracting what the project calls “learnings” via the Claude Agent SDK. At session end, a summary is generated. At the next session start, context from the last ten sessions is injected automatically.

The boring part — the hooks, the SQLite schema, the worker process — is where the value actually lives. It is unglamorous plumbing that turns an agent’s ephemeral activity into a queryable historical record, and it requires no ceremony from the user. There is no “remember this” command. The memory forms as a side effect of working.

Progressive disclosure, or the economics of not reading

The more technically interesting idea is what claude-mem calls progressive disclosure, and it addresses a problem the memory-obsessed crowd often skips: retrieving memory costs tokens too. A naive memory system that dumps your entire project history into context is just a slower way to overflow the window.

Claude-mem’s answer is a three-layer search workflow exposed through MCP tools. A search call returns a compact index of results with IDs, at a claimed 50 to 100 tokens per result. A timeline call provides chronological context around interesting hits. Only then does the agent fetch full observation details for the filtered IDs that actually matter, at 500 to 1,000 tokens each. The project claims roughly tenfold token savings from filtering before fetching — a figure that will vary with workload, but the pattern itself is sound and increasingly echoed across the field. The commercial site frames it as “find a note before reading its details,” which is a decent one-line summary of the whole retrieval philosophy.

Under the hood, retrieval is hybrid: SQLite’s FTS5 handles keyword search, and an optional Chroma vector index adds semantic matching. That combination — relational store plus vectors, fused ranking — is becoming the default shape for agent memory, and Cognee’s 2026 survey of persistent memory layers reads like a taxonomy of exactly this design space, down to the episodic/semantic/procedodic distinctions claude-mem’s observations and summaries loosely mirror.

A crowded field with different bets

Claude-mem is not alone, and the competition illuminates what it is. akitaonrails/ai-memory takes a deliberately austere counter-position: memory as plain markdown files in a git-backed wiki, grep-able, editable by hand, with a default path that uses zero LLM calls and no vector store to babysit. Its pitch is cross-agent and cross-machine handoffs — quit Claude Code mid-task, open Codex, continue without re-explaining — with the database as a derived index that can always be rebuilt from the files.

Claude-mem’s bet is the opposite: memory as a service, with a real worker process, a web viewer streaming observations in real time, vector search, and AI-generated summaries. That buys richness — citations with observation IDs, natural-language search, a visual memory stream — at the cost of a heavier footprint. The system wants Node.js, Bun, and uv (for the Python side that Chroma requires), auto-installing the latter two if missing. There is also a beta “Endless Mode” described as a biomimetic memory architecture for extended sessions, which tells you the project’s ambitions now extend past session boundaries toward something more like continuous memory.

There is a third competitor that matters more than any open-source rival: Anthropic itself, which has been shipping native memory features into Claude. Reddit threads asking how claude-mem differs from Claude’s built-in memory capture the tension, though the platform-first vendor’s version of memory will inevitably be narrower — one agent, one ecosystem. The cross-tool story is where third parties live or die.

The business, and the token

What makes claude-mem a genuinely interesting case study rather than just a well-executed plugin is how aggressively it has been commercialized. The open-source engine is Apache-2.0 licensed — the README is explicit that durable agentic memory should be easy to embed in developer tools, MCP servers, and production harnesses — but the commercial site sells CMEM Pro at $30 a month: hosted observation, cloud sync across machines, and a private MCP connection that makes memory recallable from other clients, including ChatGPT. A TeamBrain pilot points at shared team memory, with pricing unannounced.

The hosted-observer angle is clever economics. Memory generation — the summarization and learning-extraction passes — burns tokens. Moving that work off your primary agent’s subscription, as the marketing puts it, effectively sells capacity back to heavy users. The installer, per the SkillsLLM directory listing, now walks users through a browser sign-in and a 30-day trial of the hosted observer, with fallback to your own Anthropic plan or an OpenRouter/Gemini key afterward. Escape hatches exist — explicit provider flags, an opt-out environment variable, non-interactive CI behavior — but the gravitational pull toward the paid service is unmistakable.

And then there is $CMEM, a Solana token created by a third party without the project’s consent, which the creator has nonetheless officially embraced, describing it as a community catalyst and “a vehicle for bringing real-time agent data to developers.” The README prints the contract address. Whatever one thinks of memecoins attached to infrastructure projects — and there is a long, mostly cautionary history here — it is a striking signal about the attention economy this repo is operating in. Nearly a hundred thousand stars buys you a lot of things, including a cryptocurrency you never asked for.

Rough edges

The project is not without friction, some of it visible right on the surface. The npm package is a trap for the unwary: a global install gets you the SDK library only, without the plugin hooks or worker setup — the README warns about this explicitly, which is honest, but the warning existing at all tells you the packaging story is muddy. The system requirements disagree with themselves: the README says Node 18 or higher, the official docs say Node 20 or higher. And an automated security scan surfaced a handful of low-severity npm audit findings in interactive-prompt dependencies — nothing alarming, but consistent with a codebase that has grown very fast.

The deeper open question is trust. A system that silently captures every tool execution — every file read, every command — and increasingly syncs it to a cloud service is a privacy surface that deserves scrutiny. The project does offer <private> tags for manual exclusion and regex-based auto-redaction for secrets, which is more than most competitors provide, but the default posture of capture-everything-then-filter is the opposite of the markdown-and-git school’s show-me-everything approach. Different users will weigh those differently.

Outlook

The trajectory is clear: claude-mem is racing from a Claude Code plugin toward a general memory substrate — more harnesses, more IDEs, cloud sync, team memory, an “All your context. Everywhere. All at once.” tagline that barely conceals the ambition. The technical core is sound and increasingly conventional wisdom: external storage, layered retrieval, hybrid search, background compression. What remains unresolved is whether the heavyweight, service-oriented approach beats the plain-files counter-movement, whether Anthropic’s native features absorb the use case, and whether a project that grew this fast on hype can build the boring reliability that memory infrastructure demands. Forgetting, it turns out, was the easy part to fix. Remembering responsibly is the hard part.

Sources

  1. How do you design memory systems for long-running AI ...
  2. akitaonrails/ai-memory: Solution for long term ...
  3. Claude-Mem: Introduction
  4. Building AI Agents with Persistent Memory
  5. Building an ai coding assistant that actually remembers ...
  6. All your context. Everywhere. All at once. — claude-mem
  7. Persistent Memory Layer for AI Agents 2026
  8. Would a persistent "project memory" layer for AI coding assistants actually ...
  9. How is Claude-Mem different from Claude's New Memory ...
  10. How are you handling persistent memory across multiple ...
  11. AI Memory : The Simplest System That Beats Every Complex Solution
  12. claude-mem - AI Agents on GitHub (97.4k★)

heatdrop uses Google Analytics to see which pages get read — nothing else. Your call. How we handle data.