Skip to content
markus7hPublic

About

Shared long-term memory as a knowledge-graph MCP server — persistent context, preferences and project knowledge for Claude Code, opencode and any MCP-capable AI agent.

Topics

Resources

Stars

4 stars

Watchers

0 watching

Forks

Latest commit

 

History

253 Commits

Folders and files

Repository files navigation

ai-rem — Knowledge Graph Memory for AI coding agents

This documentation describes v1.6.0. v1.0.0 replaces the archived Kuzu with LadybugDB. The database file formats are not compatible: upgrading from v0.8.x runs through scripts/migrate.py — see Upgrade from v0.8.x (Kuzu). Fresh installs are unaffected.

Release notes are kept in CHANGELOG.md and published to the GitHub Releases and the Docker Hub description on every tag; notes for early versions (≤ v0.1.5) are archived in docs/release-history.md.

ai-rem is a persistent long-term memory for AI coding agents — Claude Code, opencode and any other MCP-capable frontend (Codex, Gemini CLI, Cursor …) — running as an MCP server on your home server. All frontends share the same graph, so what one agent learns the next one knows. Static memory files like CLAUDE.md sit in context in full and are tied to individual projects and machines. ai-rem takes a more efficient approach: relevant information — open tasks, decisions made, solved problems, projects, tools used — lives in a knowledge graph on your home server, is loaded selectively instead of wholesale, and is available from any machine, independent of where you work.

Docker Hub: docker pull magic3arkus/ai-rem


What is ai-rem?

ai-rem is the MCP server that provides the knowledge graph. It runs as a Docker container on your home server (<SERVER_IP>, configurable port, default 3456) and is always available as long as the server is running.

Technically:

  • FastMCP — Python MCP server framework, HTTP transport (Streamable HTTP)
  • LadybugDB — embedded graph database (no separate DB container needed)
  • Data is stored persistently in ./data/kg.db (configurable via KG_DATA_PATH)
  • Backups are saved to ./backups/ (configurable via KG_BACKUP_PATH)

How it works

At the start of each session, the agent loads the relevant context from the graph via memory_get_context() and saves new insights proactively with memory_add / memory_relate. The graph holds typed entities — Person · Project · Task · Tool · Problem · Solution · Decision · Preference · Topic — and relations between them. Each entity can carry a context tag (work / private / global) so work and private knowledge stay separated per repo.

Project context. A project's working context — local dev dir, deployment dir/host, relevant skills and project-specific rules — lives in a Project entity's extra. Write it with memory_set_project_context(...) (field-wise merge, so you can add skills later without wiping dev_dir/rules) and load it whole in one call with memory_project_context(name) — the untruncated record plus every related entity. Use it to start a session "in the context of project X".

// extra schema (all fields optional)
{ "status": "aktiv", "dev_dir": "...", "repo": "...", "deploy_dir": "...",
  "deploy_host": "...", "deploy_cmd": "...", "skills": [...], "rules": [...] }

Lean tool surface. Only 4 always-on MCP tools (memory_get_context, memory_search, memory_add, memory_relate) sit in tools/list and cost per-session context; the other 12 admin ops (list, merge, archive, project-context, …) are reachable over HTTP via POST /api/tool or the ai-rem CLI, keeping the session footprint small. AI_REM_ADMIN_TOOLS=1 re-exposes all of them as MCP tools.

→ MCP tool reference — the full memory_* set (4 always-on + 12 admin) with signatures, entity types and context separation.


Token savings

ai-rem lazy-loads only the relevant subgraph on demand instead of carrying everything in CLAUDE.md all session long. The per-session footprint stays roughly constant (~1–3k tokens) no matter how large the graph grows. At ~4.3 sessions/day this works out to ~0.7 million tokens/month saved — roughly 3 full 200k context windows — and the savings grow as the graph grows, because ai-rem's cost stays flat while a CLAUDE.md-everything approach scales linearly.

→ Full methodology, measurements and ranges


Web UI

URL Function
/ Redirects to /ui — the logo in the navigation links here, and it is the natural entry point in a browser
/ui Dashboard: entity/relation counts, kg.db size at startup, backup management: manual, schedule, download, restore (export v2 round-trips pinned/sort_order/archived); also OKF bundle import
/browse Interactive content browser: search and filter by type, toggle archived, expand an entry for description, extra and relations; imported entries are badged; the list loads 20 entries at a time
/tasks Task list: status and project per row, "open only" filter (default on), toggle archived, search across name, description and project; archive a task in place — it leaves the session context but stays in the DB. Paginated in steps of 20.
/graph Node-link visualization (vis-network): nodes colored by type, edges labeled by relation; filter by context (work / private / global) and toggle entity types via the legend; physics and archived toggles; "connected only" pins the clicked node plus its neighbors up to an adjustable distance (1, 2 … n; single-click shows info, double-click re-anchors)
/prefs Preferences manager: pin, context, sort order, delete; archived preferences are dimmed, badged and listed below a separator (they never load into session context).
/cleanup Nightly cleanup: config, manual run, pending reviews, run log; plus archive purge (permanently delete archived entries, optionally keeping the last X days)
/logs Server log without shell access: level filter, substring search, optional 5s auto-refresh, download as text. Fed by an in-memory ring buffer (last AI_REM_LOG_RING, default 500 lines) — so it only covers the time since the last container restart. Bearer tokens are redacted.
/install Client setup commands per platform (bash / PowerShell) with copy buttons, incl. step-by-step SSH key guide — public, for onboarding new machines

Every page shows the running server version at the right end of the navigation.

Interop (OKF). ai-rem speaks the Open Knowledge Format v0.1: /export/okf downloads the whole graph as a Markdown+YAML bundle (ZIP), /api/import/okf reads one back in. Own exports carry source: ai-rem so a round-trip stays untagged, while foreign entries are marked imported and indexed for semantic search on import.


Supported frontends

Claude Code opencode Other MCP frontends (Codex, Gemini CLI, Cursor …)
MCP tools + server instructions ✓ ✓ ✓
Context at session start SessionStart hook AGENTS.md pointer AGENTS.md/GEMINI.md pointer (snippet)
Auto-memory from transcripts PreCompact/SessionEnd hook plugin (session.idle/session.compacted) —
Slash commands ✓ ✓ (without /migrate-claude-md) —
Set up by ai-rem install --client claude ai-rem install --client opencode ai-rem install --client generic (writes snippets only)

Every new entry records which frontend created it in extra.client (claude-code, opencode, ai-rem-cli, …).

Automation (hooks)

Four Claude Code hooks — all deployed by the client setup — keep the graph fed and tidy (opencode gets the auto-memory part as a plugin):

  • Auto-Memory — a PreCompact/SessionEnd hook extracts structured entities/relations from each transcript via an OpenAI-compatible LLM endpoint (AI_REM_LLAMA_URL, by default a LiteLLM router rather than a single GPU host), with an md-fallback + catch-up when it is down. It runs detached (extraction takes minutes) and reports at the next session start when it is broken. The setup installs the CLI to ~/.local/share/ai-rem/bin/ai-rem and points AI_REM_CLI at it, so the hook does not depend on where the repo was cloned; ~/.local/bin/ai-rem links to the same copy, which is what makes ai-rem a plain shell command.
  • Nightly cleanup — a daemon dedups/archives outdated entries non-destructively (archive, never delete; preferences/pinned untouched), pushing ambiguous cases to a review queue. Finished tasks are archived 14 days after extra.status reached erledigt (the status is normalised against the enum offen | laufend | blockiert | erledigt, common synonyms mapped, free text kept in extra.status_note). A completion marker that only shows up in the description text ("ERLEDIGT: …") never archives on its own - it lands in the review queue. Plus a staleness check that flags entries with perishable infrastructure facts (IPs, ports, services, devices) for a reality check — never automatically.
  • Plan saving — an ExitPlanMode hook stores every finalized plan as an open Task, so plans become a central, cross-machine list.
  • Vault secret reminder — a PostToolUse hook scans Bash output for auth/credential failures and injects a reminder to pull the matching secret from the vault via mykeyvault instead of asking the user for a token or password, failing silently so it never blocks a command.

→ Hooks & automation in detail


Requirements

  • Docker on the target server
  • Python 3 on the client machine (the setup, the CLI and the hooks run on Python), plus at least one frontend: Claude Code, opencode or another MCP client
  • Network access to <SERVER_IP>:<PORT>
  • Optional (only for the tools companion MCP): git, Node.js ≥ 18 incl. npm

The image ships without a compiler and without pip — every dependency installs from a prebuilt wheel, and nothing is installed at runtime. That removes ~460 CVEs the toolchain layer used to drag in. If you ever need pip inside the container: python -m ensurepip.


Configuration

Environment variables are loaded from a .env file in the Compose directory:

AI_REM_API_TOKEN=...                     # REQUIRED — API token (fail-closed, see Authentication)
KG_PUBLIC_URL=http://<SERVER_IP>:3456   # Public URL of the server
PORT=3456                                # TCP port (default: 3456)
HOST=::                                  # Bind address (default: ::) — dual-stack socket, so IPv6 and the published IPv4 port both work; 0.0.0.0 for IPv4 only
LADYBUG_DB_PATH=/data/kg.db                 # Path to the database
BACKUP_DIR=/backups                      # Path for backup files
MAX_BACKUPS=10                           # Maximum number of backups to keep
AI_REM_BACKUP_KEY=...                     # Optional — encrypt backups (AES-256-GCM); empty = plaintext
LADYBUG_POOL_SIZE=4                         # Connection pool size
DISCOVER_ROUTINES_LIMIT=10               # Pinned routines injected per prompt via /discover (curated by sort_order)
LADYBUG_BUFFER_POOL_SIZE_MB=256           # buffer pool in MiB (0 = default: 80% of host RAM)
LADYBUG_WAL_CHECKPOINT_MB=2               # self-checkpoint the WAL above this size (0/empty = off)
AI_REM_WAL_CHECKPOINT_IDLE_S=300          # ...and once it has been idle this long (0 = off)
KG_REBUILD_MB=2048                        # compact kg.db on the next start above this size (there is no VACUUM)
EMBED_BACKFILL_PORTION=300                # vectors written per database session
EMBED_RECONCILE_SEC=3600                  # how often the server backfills missing vectors (0 = startup/nightly only)
KG_MAX_MB=4096                            # above this the embedding backfill stops writing entirely
KG_MIN_FREE_MB=1024                       # free disk space the backfill requires before it writes
AI_REM_ADMIN_TOOLS=0                      # 1 = re-expose the 12 admin ops as MCP tools
AI_REM_LOG_RING=500                       # Lines of server log kept in memory for /logs
EMBED_URL=                                # Empty = in-process embeddings; set to an OpenAI-compatible /v1/embeddings URL for an external service
EMBED_HTTP_MODEL=bge-m3                   # Model name sent to EMBED_URL
EMBED_THRESHOLD=                          # Cosine cut-off; empty = per-backend default (0.45 in-process, 0.50 external)
EMBED_MAX_CHARS=2000                      # Truncate input before embedding (llama.cpp rejects oversized input instead of truncating)
AI_REM_TAG=latest                         # latest (bundled embedding model) or latest-slim (~250 MB smaller, requires EMBED_URL)
EMBED_BACKEND=local                       # Only when building locally: must match AI_REM_TAG (local = :latest, external = :latest-slim)
MEM_LIMIT=1536m                           # Container memory limit; 512m is enough without the bundled model

Embeddings: in-process or external

Semantic search needs vectors. By default they are computed inside the container (fastembed/MiniLM, model baked into the image) — nothing else has to run. Setting EMBED_URL to an OpenAI-compatible endpoint (e.g. a llama.cpp server serving bge-m3) moves that work out of the container and allows the -slim image, which ships without fastembed and the model (413 MB → 162 MB).

Either way the search is hybrid: substring hits (computed locally) and semantic hits are merged by reciprocal-rank fusion — entries corroborated by several signals rank first, name matches beat description matches. If the external endpoint is unreachable, entries are stored without a vector and search keeps working lexically — the backfill fills the gaps once the service is back, at startup, hourly (EMBED_RECONCILE_SEC) and at the end of the nightly run.

Switching backends changes the vector dimension (384 ↔ 1024), which makes the stored vectors meaningless. The server detects that on the next backfill and recomputes all vectors — no manual migration, and it works in both directions.

Note (memory): Without LADYBUG_BUFFER_POOL_SIZE_MB the database sizes its buffer pool to ~80 % of host RAM and ignores the container mem_limit. Normal operation on this DB needs only ~32 MB, so 256 MiB is plenty — including with 1024-dimensional vectors from an external backend. LADYBUG_WAL_CHECKPOINT_MB keeps the WAL small (periodically + on shutdown) so opening the database never triggers an expensive recovery.

Upgrade from v0.8.x (Kuzu)

The file formats are not compatible — LadybugDB refuses a Kuzu kg.db. scripts/migrate.py ships inside the image and moves the graph through a JSON dump, which also sheds the accumulated Kuzu bloat. Embeddings are not part of the dump; the new instance recomputes them on import.

docker run --rm --entrypoint cat magic3arkus/ai-rem:latest /app/scripts/migrate.py > migrate.py
export AI_REM_API_TOKEN=$(ai-rem token)

python3 migrate.py export --url http://localhost:3456 --out dump.json   # old instance still running
docker compose down && mv /path/to/volume/kg.db /path/to/volume/kg.db.kuzu-old
docker compose up -d                                                    # new image
python3 migrate.py import --url http://localhost:3456 --in dump.json

Keep kg.db.kuzu-old until /api/status shows the expected entity count — that file is the only way back.

Database size: why kg.db stays small

Up to v0.8.32 ai-rem ran on Kuzu, which never returned space when a property was overwritten: a checkpoint rewrote the affected column and left the old version in the file, with no VACUUM to reclaim it. The embedding backfill hit that hardest, and a crash loop turned it into a disaster — on 2026-09-03 kg.db went from ~680 MB to 27 GB across 264 restarts and filled the partition, taking neighbouring containers with it.

LadybugDB, the maintained fork ai-rem moved to in v0.9.0, does not behave that way. Same measurement, 1342 vectors at 1024 dimensions:

Kuzu 0.11.3 LadybugDB 0.20.2
vectors surviving 0 / 1342 (buffer pool is full) 1342 / 1342
file size 771 MB 40 MB

Overwriting the same property repeatedly no longer grows the file either. The guards built for the Kuzu era stay in place for now — they simply never trigger:

  • KG_MAX_MB / KG_MIN_FREE_MB: the backfill stops writing once the DB is too large or the disk too full. Vectors are derived data — search keeps working with whatever is stored, lexically if need be.
  • KG_REBUILD_MB: if kg.db exceeds this at startup, the server compacts it itself (dump → fresh DB → import). The dump is written to BACKUP_DIR as a regular backup first; if that fails, the old DB is left untouched.

The current size is exposed as db_mb in /api/status, along with both thresholds and db_start_mb — the size the service came up with. The dashboard shows the startup size, and as soon as the current one drifts from it, both (40 → 41.2 MB): on its own a size tells you nothing about whether the DB is growing while running, which is the symptom that went unnoticed for days in the Kuzu era.

The one guard that did not stay is restart: on-failure:5. It was dropped on 2026-09-12 for restart: unless-stopped: a reboot stops the container gracefully (exit 0), which is not a failure, so the container was the only service that stayed down after a host restart while every neighbour came back up.

stop_grace_period: 180s belongs next to it. On SIGTERM the server merges the WAL into kg.db so the next start opens without an expensive recovery — and docker stop (so every compose up -d) sends SIGKILL 10 seconds later by default. On 2026-09-13 that SIGKILL landed in the middle of the checkpoint: kg.db.wal.checkpoint and the lock files stayed behind, and the next start failed on Checksum verification failed, the WAL file is corrupted — in a crash loop, because unless-stopped kept retrying. Three minutes of grace are plenty for a checkpoint of this size and cost nothing when it finishes earlier.

That covers an orderly docker stop — not a crash. LadybugDB segfaulted mid-checkpoint on 2026-09-16 and again on 2026-09-18 (exit 139); a segfault does not ask for SIGTERM, and the crash loop ran 20 and 22 rounds until someone intervened by hand. v1.2.5 cleaned up when LadybugDB reported a damaged WAL — on 2026-09-18 it reported nothing, because the constructor itself died by signal, and no except ever sees a SIGSEGV.

Since v1.2.6 the decision is made where that death is visible at all. If any WAL or checkpoint leftovers sit next to kg.db, the server first opens the database in a throwaway subprocess. Survives it, the real open proceeds untouched; dies it, kg.db.wal*, kg.db.shadow and the checkpoint locks are moved aside as *.corrupt-<timestamp> first. Moved, not deleted — in case anyone wants to look at them. All that is lost is whatever had not been merged since the last checkpoint; an intact WAL survives the probe and recovers normally. A probe killed by SIGKILL counts as survived: that is the OOM killer, and memory pressure must not cost healthy transactions. The probe costs a second open, so it only runs when leftovers are actually there — never after a clean stop. If the open fails anyway, the error stands: then only a restore from /backups helps.

Why the embedding backfill writes in portions

Under Kuzu every CHECKPOINT rewrote the whole column, so the file grew with the number of checkpoints rather than with the data — and once it outgrew the buffer pool, the next checkpoint failed and discarded what earlier ones had persisted, silently: the run logged "backfill finished (1251)" while embed_pending stayed at 1210. Writing EMBED_BACKFILL_PORTION vectors per database session was the workaround.

On LadybugDB the same write pattern keeps all 1342 vectors in a 40 MB file, with the default 256 MB buffer pool. The portioning is therefore no longer load-bearing; it is kept for now because it also caps peak memory during a restore, and because the verification after each portion (one entity is sampled — if its vector is gone the run stops with an ERROR) is cheap insurance. A later release will simplify this path.

Authentication

All sensitive routes (/mcp, /api/*, /export, /import, /ui) require a token. The server is fail-closed — without AI_REM_API_TOKEN it refuses to start. MCP clients authenticate with Authorization: Bearer <token>; the browser Web UI uses a derived, HttpOnly cookie set at /login. The token is kept once in mykeyvault (ai-rem-api-token) and pulled in by deploy.sh / the client hook.

→ Authentication model, Web UI login & token source


Installation / Deployment

Server (one-time setup)

# Create directory
mkdir -p ~/mydocker/compose-files/ai-rem
cd ~/mydocker/compose-files/ai-rem

# Download docker-compose.yml and .env.example
curl -O https://raw.githubusercontent.com/markus7h/ai-rem/main/docker-compose.yml
curl -O https://raw.githubusercontent.com/markus7h/ai-rem/main/.env.example

# Create .env and configure
cp .env.example .env
# → set KG_PUBLIC_URL to your actual server IP
# → set AI_REM_API_TOKEN (required) — e.g. `openssl rand -hex 32`, or use deploy.sh
#   to pull it from mykeyvault automatically (see Authentication)

# Pull image and start container
docker compose pull && docker compose up -d

Client — setting up a new machine

On every new machine, in your own shell:

bash <(curl -s http://<SERVER_IP>:3456/setup)

On native Windows (PowerShell, no WSL needed): irm http://<SERVER_IP>:3456/setup.ps1 | iex.

The setup installs the ai-rem CLI and then sets up every frontend it finds (claude, opencode; neither → generic snippets). It is idempotent and needs no preparation on the machine:

  • No SSH key needed. Without one the setup shows a code and opens /pair in the browser; approve it there (logged into the web UI, works from your phone too) and the setup continues with the ai-rem token — and the mykeyvault token if AI_REM_PAIR_VAULT_TOKEN is set on the server. (How pairing works)
  • Missing node/npm/git (for mykeyvault and tools) are installed after one prompt via brew, apt (NodeSource) or winget; --yes skips the prompt.
  • It ends with a ✓/✗ summary, one fix command per ✗, and runs ai-rem doctor.

To pick targets explicitly or add one later:

ai-rem install --client opencode     # add opencode to an existing setup
ai-rem install --client generic      # snippets for Codex, Gemini CLI, Cursor in ~/.config/ai-rem/snippets
ai-rem uninstall --client opencode   # remove one frontend again
ai-rem doctor [--fix]                # server version, installed targets, drift, token source
ai-rem pair                          # renew just the token (browser approval)
  • Claude Code: registers the MCP server, deploys the hooks, writes the minimal CLAUDE.md pointer and installs the slash commands.
  • opencode: merges ai-rem (plus mykeyvault/tools if configured) into the mcp block of ~/.config/opencode/opencode.json — providers and other servers stay untouched, a JSONC file with comments is left alone and a snippet is written instead — adds an AGENTS.md pointer, the plugin and the commands. Tokens are referenced as {file:~/.config/ai-rem/token}, never written into the config.

→ What the setup does, repo layout & CLAUDE.md strategy

Update to a new version

Server:

ssh your-server "cd ~/mydocker/compose-files/ai-rem && docker compose pull && docker compose up -d"

Client — hooks, CLI, lib and slash commands are shipped by the server and go stale with every release that touches one of them:

ai-rem update --check   # report only, exit 1 if anything is behind
ai-rem update           # pull the new files for every installed frontend, then restart them

The Claude session-start hook and the opencode plugin compare the local copies against /manifest and say so on their own, so you normally learn about it before you have to ask. settings.json is only ever added to, never trimmed; the template it merges from belongs to the server and is rewritten wholesale.


Development

CI runs on every push and pull request (.github/workflows/ci.yml): a ruff check for critical error classes (syntax, undefined names), a compileall smoke, an import smoke against a throwaway env, the pytest suite under tests/, and an image-smoke that builds the Docker image, polls /health and runs Trivy as a gate on fixable CRITICAL/HIGH CVEs. main is protected — changes land via PR with green CI.

The image build cache is scoped per day on purpose. A fixed cache pins the apt-get update && apt-get upgrade layer — the one that picks up Debian patches — so the Trivy gate starts blocking on CVEs a real build would already have fixed, and a re-run changes nothing because it pulls the same cache. Daily rotation costs one full build a day and keeps the gate at most that far behind.

Local hygiene hooks mirror the CI ruff gate and add whitespace/EOF fixes plus a detect-secrets scan (baseline: .secrets.baseline). One-time setup: pipx install pre-commit && pre-commit install.

Releases are tag-triggered (.github/workflows/docker-publish.yml); a VERSION ↔ Tag step fails the build if VERSION in server.py does not match the pushed tag (v1.2.3 → 1.2.3). The same run creates the GitHub release from the matching CHANGELOG.md section and refreshes the "What's new" block in the Docker Hub description — so add the entry before tagging.


Related Projects

  • tools-registry — MCP server exposing small scripts as tools via a central registry. ai-rem tracks a Tool entity per script (the ai_rem_entity convention) so the catalog stays discoverable.
  • mykeyvault — self-hosted secrets vault (Vaultwarden + REST/MCP). ai-rem deliberately stores no secrets; credentials live in mykeyvault instead.

About

Shared long-term memory as a knowledge-graph MCP server — persistent context, preferences and project knowledge for Claude Code, opencode and any MCP-capable AI agent.

Topics

Resources

Stars

4 stars

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages