← Back to blog

Give Claude Code Permanent Memory in 2 Minutes

March 30, 2026 · 5 min read · By Adam Cipher

You open Claude Code on Monday. It doesn't know what you built on Friday. Your architecture decisions, your naming conventions, why you chose Postgres over SQLite — all gone. Every session starts from zero.

This isn't a Claude limitation. It's how every coding agent works: the session ends, the context window is cleared, and the agent wakes up with amnesia.

Here's how to fix it with one MCP server. No database to manage. No config files to maintain. Two minutes from reading this to having it working.

Why Claude Code's Built-In Memory Isn't Enough

Claude Code has CLAUDE.md files that persist between sessions. They're useful for static rules — style guides, repo conventions, things that don't change.

But they fail for dynamic knowledge:

What you need is a memory layer that stores facts, scores them by relevance, and lets Claude retrieve only what matters for the current task.

The Setup: One Command

Add this to your Claude Code MCP config (~/.claude/claude_desktop_config.json or your project's .mcp.json):

{
  "mcpServers": {
    "engram": {
      "command": "npx",
      "args": ["-y", "engram-mcp"],
      "env": {
        "ENGRAM_API_KEY": "your-api-key-here"
      }
    }
  }
}

Get your API key (free, no credit card):

curl -X POST https://engram.cipherbuilds.ai/api/signup \
  -H "Content-Type: application/json" \
  -d '{"email": "[email protected]"}'

# Returns: { "apiKey": "eng_...", "agentId": "agent_..." }

Restart Claude Code. That's it. You now have three new tools available: store_memory, retrieve_memory, and search_memory.

What Changes

Before: Every session starts cold

// Monday: You explain your architecture
> "We're using a modular service pattern. Auth is in /services/auth,
   each service exports a standard interface..."

// Tuesday: Claude asks again
> "Could you describe the project architecture so I can help effectively?"

After: Sessions pick up where they left off

// Monday: Claude stores the decision automatically
[store_memory] "Architecture: modular service pattern.
  Auth in /services/auth. Standard interface exports."

// Tuesday: Claude retrieves context before responding
[retrieve_memory] query="project architecture"
→ Returns scored facts about your service pattern, auth setup,
  naming conventions — ranked by relevance

The agent doesn't just remember. It remembers intelligently — facts that were useful recently score higher. Facts that contradict newer decisions decay. After a week, Claude knows your codebase better than your CLAUDE.md ever could.

How the Scoring Works

Every fact stored through Engram gets a relevance score between 0 and 1. The score is influenced by:

This means your agent's memory is self-curating. You don't need to prune it. The scoring does the pruning for you — irrelevant facts sink to the bottom and eventually fall below the retrieval threshold.

Real Example: Debugging Across Sessions

Here's a pattern that happens constantly without persistent memory:

// Session 1: You debug a race condition in the queue processor
> "The bug was in handleMessage — we were awaiting inside a forEach.
   Switched to for...of with sequential await. Fixed."

// Session 4 (two days later): Similar bug in a different service
> "I'm seeing intermittent failures in the notification sender..."

// WITHOUT memory: Claude debugs from scratch. 20 minutes.
// WITH memory: Claude retrieves the forEach/await pattern.
[retrieve_memory] query="async bug intermittent failures"
→ "handleMessage had a race condition from await inside forEach.
   Fix: use for...of with sequential await. (score: 0.87)"

// Claude recognizes the pattern immediately. 2 minutes.

This isn't theoretical. This is what happens when your agent accumulates operational knowledge across hundreds of sessions. Patterns compound. Debug time drops. The agent gets better at your specific codebase, not just at coding in general.

What About Privacy?

Fair question. Your code context goes through an external API.

If you need on-prem, the MCP server source is open. Point it at your own backend.

MCP vs. RAG vs. Fine-Tuning

People ask why not just use a vector database or fine-tune the model.

Ready to try it?

Free tier. One agent, 10,000 facts. No credit card. Setup in 2 minutes.

Get Your API Key →

Going Further

Once you have basic memory working, the interesting stuff starts:

The deeper insight: memory is what turns a coding assistant into a coding partner. An assistant answers questions. A partner knows your context, remembers your decisions, and builds on past work without being told.

That's a 2-minute setup. Start here.