pip install. The plugin hooks into Claude Code’s lifecycle, so it:
- captures your prompts, tool traces, and assistant responses into session memory
- injects relevant context on every prompt submit
- syncs the session into your knowledge graph on session end
Install
Install from the Claude Code marketplace. Run these in your terminal (or type the equivalent/plugin … slash commands directly in the Claude Code chat):
cognee: <dataset> · <mode>.
Configure your backend
Configure the plugin once in~/.cognee/.env. The file is created with a commented template on the first session start, and its values act exactly like shell exports — a real export in your shell still wins, per terminal. It is shared with the Codex plugin, so both read the same configuration.
- Cognee Cloud / remote
- Local (default)
- Windows (PowerShell)
Point the plugin at Cognee Cloud or a remote server by setting both:Cloud mode is a pure thin client: it talks to your remote server over HTTP only and does not install a local Cognee runtime.If
COGNEE_BASE_URL points at a server you run yourself and you leave COGNEE_API_KEY unset, the plugin logs in as COGNEE_USER_EMAIL / COGNEE_USER_PASSWORD to mint a key. On cognee 1.6.0 and later, start that server with DEFAULT_USER_PASSWORD set to the same value as COGNEE_USER_PASSWORD, or set COGNEE_API_KEY so the plugin skips the login.nano ~/.cognee/.env) works too. Either way, changes apply on the next launch. The file format:
- Comments are whole lines starting with
#— a trailing# noteafter a value becomes part of the value. - Quotes around values are optional, and a leading
exportis tolerated so existing shell profile lines paste verbatim. - Every variable in the Configuration Reference can live here, so there is nothing else to persist. The one exception is
COGNEE_ENV_FILEitself: the plugin reads it before it opens the file, so it only works as a shell export.
Which mode wins, and how to switch
You can configure both modes at once — keepCOGNEE_BASE_URL + COGNEE_API_KEY and LLM_API_KEY in the file together. The mode is then decided per terminal, by three rules in order:
- A
COGNEE_BACKENDexport wins.export COGNEE_BACKEND=local(or=cloud) pins that terminal to that mode. - Otherwise, cloud wins when configured. If
COGNEE_BASE_URLis set — in the file or the shell — the plugin connects to it. - Otherwise, local. With no URL anywhere, the plugin boots the local server.
~/.cognee/.env:
- The switch is pinned.
COGNEE_BACKEND=cloudwith noCOGNEE_BASE_URLconfigured still counts as cloud — the plugin does not silently fall back to local, and the status line shows✕ (missing_cognee_base_url)so you know exactly what to fix. - A forced-local switch blanks
COGNEE_BASE_URLandCOGNEE_API_KEYin the process environment, so the per-prompt hooks and every spawned worker resolve the same local endpoint — not justSessionStart. They are emptied rather than deleted on purpose: a child process that reloads the env file must not re-inject the cloud values. - The shared
COGNEE_BACKENDflips both the Claude Code and Codex plugins in that terminal. To flip only one, use the plugin-specific name —COGNEE_CLAUDE_BACKENDorCOGNEE_CODEX_BACKEND— which beats the shared one. - Accepted values:
local(aliasesnative,sdk) andcloud(aliaseshttp,api,server). Anything else is ignored. COGNEE_BACKENDcan also live in~/.cognee/.envto make a mode the durable default; a shell export still overrides it per terminal.- Not sure what a terminal resolved? The status line’s mode field shows it live.
Cognee’s LLM calls do not run through your coding agent — with one exception: Claude Code’s local mode with no LLM key configured, which runs them on your Claude subscription through the Claude observer. Otherwise, your Claude Code or Codex plan pays only for your conversation with the model. Everything Cognee does on its own — entity and relationship extraction during cognify, summarization, embeddings, and search-time completions — happens inside the Cognee backend against the LLM provider configured there, and is billed by that provider. In local mode you configure it with
LLM_API_KEY; in Cloud/remote mode your tenant holds the key server-side, so no local LLM key is needed.LLM_API_KEY above covers extraction, summarization, and embeddings: Cognee defaults to openai/gpt-5.6-luna for the LLM and openai/text-embedding-3-large for embeddings, and embeddings reuse LLM_API_KEY when EMBEDDING_API_KEY is unset. To use another provider, set LLM_PROVIDER, LLM_MODEL, and — for Azure, Ollama, or OpenAI-compatible endpoints — LLM_ENDPOINT. Changing only the LLM leaves embeddings on OpenAI, so also set the EMBEDDING_* variables or set EMBEDDING_API_KEY to an OpenAI key so the default embeddings keep working. See LLM providers and embedding providers.
Local mode without an LLM key: the Claude observer
In local mode, with noLLM_API_KEY and no LLM_PROVIDER configured and the claude CLI on PATH, session start routes the local server’s LLM calls through Claude Code itself instead of asking for a key. It points cognee’s custom provider at a loopback OpenAI-compatible shim (http://127.0.0.1:8017/v1), and each completion the server requests becomes one claude -p --safe-mode run. Embeddings switch to a local model (fastembed, BAAI/bge-small-en-v1.5, 384 dimensions) unless you configured an embedder of your own. Cloud mode never uses the observer — your tenant owns its LLM.
A key or provider counts as configured wherever it is set: your shell, ~/.cognee/.env, or the .env the server itself loads. The switch takes effect on the next launch, but a server that is already running keeps the LLM config it booted with until every session using it has closed. Because a dataset is tied to the embedder that built it, switch to a new dataset (/cognee-memory:cognee-switch-datasets or COGNEE_PLUGIN_DATASET) when you move between the observer and a key of your own.
The shim listens on loopback only, requires a bearer token (~/.cognee-plugin/observer/token) on every endpoint except /health, and refuses requests that carry an Origin header. The diagnostic shows whether the observer is in use in its LLM row, and the status line reports a Claude login problem as ✕ (claude_not_logged_in) rather than as an incorrect key. Each completion is logged in ~/.cognee-plugin/observer/observer-events.log.
Use it
Just use Claude Code as usual — memory is captured and recalled automatically. You can also invoke the skills explicitly:
The plugin also ships a
cognee-recall subagent. Claude delegates to it when a prompt needs a deeper or differently worded search than the per-prompt recall already ran — it reaches the permanent graph and can filter by category (user, project, agent). It runs on Haiku with a three-turn cap, so the detour stays cheap.
To verify the connection, open a fresh session and ask: “What do you know from cognee?”
With the plugin active, Cognee is the preferred memory: the
SessionStart hook steers Claude to treat Cognee as authoritative over Claude Code’s built-in MEMORY.md. Set COGNEE_PREFER_MEMORY=false to turn the steer off.Sessions & datasets
- Sessions — by default the plugin derives the Cognee session id from the Claude Code session, so a new conversation (or
/clear) starts a new one andclaude --resumecontinues the same one automatically. SetCOGNEE_SESSION_IDbefore launching to pin a named session, or to deliberately share one live session across two terminals. Whatever you pin is normalized to the ASCII setA-Z a-z 0-9 - _ ., with every other character — accented and non-Latin letters included — replaced by_, then trimmed of leading and trailing./_and cut to 120 characters. Every Cognee agent integration applies that same rule, so one value names the same session in Claude Code as it does in Codex. - Datasets — all writes and recall are scoped to one dataset (
agent_sessionsby default). SetCOGNEE_PLUGIN_DATASETto use a custom one. The Claude Code and Codex plugins default to the same one, so memory carries across both.
cognee-agent role holding read and write on your datasets, and the launch’s dataset is addressed by its canonical UUID rather than by name, which is what makes “the same dataset” mean the same rows in Claude Code and in Codex. Grants are backfilled at every session start and roughly every 60 seconds by the idle watcher, so a dataset another plugin creates becomes visible here without a restart.
Set COGNEE_SHARED_AGENT_MEMORY=false to opt out. Nothing already written moves, but the agent leaves the shared role, so on an install that already has an agent identity the datasets your user owns stop being readable from it — the memory stays, this plugin just no longer reaches it. What separation rests on is the agent identity, so it is worth knowing which one you get: under the default COGNEE_PLUGIN_IDENTITY=auto an identity is provisioned only in service of shared memory, so opting out on a fresh install simply runs the plugin as your own user — which sees everything anyway. Set COGNEE_PLUGIN_IDENTITY=true to insist on a dedicated agent sub-user and get genuinely per-plugin memory.
To move a running session to another dataset, run /cognee-memory:cognee-switch-datasets (optionally with a name). Without a name it lists the datasets you can write to; a name that is not listed is created for you. Because a Cognee session never spans two datasets, the switch first syncs the current session into its dataset — and aborts if that fails, changing nothing — then registers a fresh session on the chosen one. The choice lives in the launch record, so it survives --resume and beats COGNEE_PLUGIN_DATASET for the rest of the launch — as does the session it registers, which outranks a pinned COGNEE_SESSION_ID so a shell export cannot drag a switched launch back.
Switching is only needed to work in another dataset — not to look something up in one. On every prompt the server answers, the recall hook appends an “Other Cognee datasets you can search” block to the injected context, naming every other dataset your identity can read (read-only ones included, since searching needs no write access) with their UUIDs. So when you ask Claude to recall something the active dataset does not hold, it can offer those datasets as a picker, run a one-off, graph-only search on the one you choose, and tell you which dataset the answer came from. If the server rejects that search with an HTTP error (other than an auth failure, or a 404, which reads as an empty result), the error names the dataset id it searched and adds the server’s own reason from the response, instead of a bare HTTP status. The active dataset, the Cognee session, and where writes go are untouched. The search runs without the session id, because a session belongs to the active dataset. The /cognee-memory:cognee-search skill offers the same choice when an explicit search comes back empty. Set COGNEE_RECALL_DATASET_HINT=off to stop the per-prompt hook from listing the other datasets — the skill flow is unaffected. The listing is cached in ~/.cognee-plugin/claude-code/readable-datasets.json and refreshed at most every COGNEE_DATASETS_CACHE_TTL seconds (default 300), within what is left of the recall budget, so a prompt never waits on it.
How It Works
The plugin registers Claude Code lifecycle hooks:
A
StopFailure hook refreshes the credit balance as well, so a turn that ends in an error still updates the line. A background idle watcher persists the session cache after periods of inactivity, and a final sync on session end bridges the session into the permanent graph.
Stop carries one more command, off by default and meant for demos. Because a hook cannot call /clear itself, COGNEE_CLAUDE_CLEAR_AFTER_MESSAGE=true empties Claude Code’s transcript file after every answer and shows “Cognee demo: Claude transcript context was emptied.” The answer is saved from the hook payload rather than the transcript, so clearing doesn’t affect what reaches session memory. Turn it on to show that the plugin carries a conversation without the transcript. Each prompt then leans on what the recall hook injects, which on cognee 1.6.0 and later includes the session’s conversation history (see Reading the memory header). The cost is that the transcript file is overwritten with nothing and no copy is kept. From then on, all that remains of the conversation is what session memory captured, so leave it off for everyday work.
What gets captured
Automatic capture is what fills session memory: your prompts, the tool calls Claude makes, and its answers. Explicit requests —/cognee-memory:cognee-remember, or cognee-remember.sh — are independent of every switch below and always store what you asked for.
Two filters run before a captured value is stored:
- Credential paths are skipped entirely. A tool call whose path arguments point at
.env,.env.*,*.pem,*.key,id_rsa*,id_ed25519*,.netrc,.npmrc,*/.ssh/*,*/.aws/credentials,*.p12,*.pfx,secrets.*, orcredentials.*is not captured at all. Extend the list withCOGNEE_CAPTURE_DENY_PATHS, and narrow which tools are captured at all withCOGNEE_CAPTURE_TOOLS(pipe-separated names or globs). - Common secrets are redacted. Private key blocks, database connection URLs,
BearerandBasiccredentials wherever they appear,secret/token/password/api_keyassignments, vendor key prefixes (sk-,ghp_,xoxb-,whsec_, and similar) and bcrypt hashes become[redacted:<kind>], and a value under a credential-looking key is replaced wholesale with[redacted:credential]. This runs on prompts, traces, and answers alike, and before truncation — so a clipped secret cannot slip past the pattern that would have matched it. Add your own regexes withCOGNEE_CAPTURE_REDACT_PATTERNS, or turn redaction off withCOGNEE_CAPTURE_REDACT=false.
COGNEE_CAPTURE=false stops automatic capture altogether: nothing is captured, buffered, or replayed, while recall and explicit remember keep working. It does not erase memory already stored.
Redaction is best effort, and the path filter reads structured tool path arguments rather than arbitrary shell command text — a secret typed into a
Bash command line is still captured. COGNEE_CAPTURE=false is the strict opt-out.Session distillation (self-improvement)
The Cognee coding-agent plugins (Claude Code, Codex) run session distillation for you — you never callimprove() by hand. A distillation pass fires on three triggers:
The idle and every-N triggers share one per-session cooldown: after an automatic improve, the next one waits at least
COGNEE_IMPROVE_COOLDOWN seconds (default 1800, 30 minutes), and a failed attempt starts the same wait as a backoff. Session end, the sync skill, and a dataset switch improve regardless of it.
Overlapping triggers are safe. A per-session improve lock on the server serializes concurrent runs, and unchanged session content dedups server-side by content hash — so a repeat improve over content that hasn’t changed is a cheap no-op, not duplicated work.
Configuration
All triggers are tuned through environment variables read by the plugin. The defaults are chosen so distillation stays out of your way; you rarely need to change them.Turning it down or off
- Stop idle-triggered improves: set
COGNEE_IDLE_DISABLED=1before launching the agent. Session-end and per-turn improves still run. - Reduce mid-session improves: raise
COGNEE_AUTO_IMPROVE_EVERYto a large value so the per-turn trigger effectively never fires within a session. - Session-end distillation always runs when the plugin is active — it’s how a finished session reaches permanent memory.
Confirming it happened
- Cloud UI: the Self-improvement card at the top of a session on the Sessions page shows the status of the last graph enrichment and the dataset it wrote to.
- Plugin hook log: each automatic run emits an
improve_firedevent you can grep for when debugging (in local SDK mode, where the plugin calls the library directly instead of the HTTP endpoint, look forauto_improve_firedinstead). improve-unsupported.jsonmarker: if this file appears in the plugin’s shared state directory (24h TTL), the server rejected the improve endpoint and the plugin fell back to the legacyrememberbridge for that window — a signal the server predates session-aware improve.
Code graph
Repositories can be indexed into a deterministic code graph — symbols, calls, imports, endpoints, dependencies. Indexing makes no LLM or embedding calls by default, so it is fast and costs no tokens. It requires a Cognee server ≥ 1.5.4. Starting a session inside a git repository indexes it automatically — but only when the server is local, so a private checkout is never shipped to a hosted tenant on the plugin’s own initiative. Auto-indexing runs in the background, never blocks the first prompt, and re-indexes after any turn that changed the working tree. Index explicitly when automation won’t: against a Cloud tenant, for a different repo, or for a git URL. The simplest route is to ask the agent to index the repository — its code skill (/cognee-memory:cognee-code in Claude Code, codebase in Codex) runs this command:
--index-vectors to also embed the extracted code facts so semantic search can see them — that is the one flag that makes embedding calls.
Query the graph through the same code skill. Prompts that mention an identifier-shaped token from an indexed repo also get code facts injected automatically by the per-prompt recall hook.
Each indexed repository gets its own dataset, named codebase-<repo-name>-<digest>, where the digest identifies the indexed path or git URL — two checkouts that share a basename would otherwise share one graph and delete each other’s nodes. Code-graph searches resolve the dataset from the current checkout, so the generated name rarely needs typing. Indexing writes its snapshot into the repository itself at <repo>/.enola/ (untracked) — add .enola/ to the repository’s .gitignore or your global excludes.
always also accepts 1, true, yes, and on; off also accepts 0, false, and no. Any other value means auto. Automatic indexing skips directories that are not git repositories, hold no source files, or exceed 3000 source files. Explicit indexing has no size cap.
Repository indexing needs a Cognee server ≥ 1.5.4: 1.5.4 renamed the field that carries the repository spec on
/api/v1/remember to raw_data, and the plugin now sends only the new name. Local mode installs the pinned 1.6.0 server for you, so this only affects a remote or self-hosted deployment. When the server rejects an index request, the plugin reports the server’s own HTTP 400 verbatim; only the server’s “Unsupported content_type” wording is reported as a too-old deployment.File context on Read
Before everyRead of a source file inside a repository indexed from a local path, a PreToolUse hook injects what the code graph knows about that file: its symbols grouped by kind with line numbers, the cross-file calls it makes, and its imports. The model gets a map of the file before its contents, so it can jump to the right line and see which other files it leans on:
COGNEE_FILE_CONTEXT_TTL in the same session; a lookup that timed out or failed is retried on the next read. When the file changed after the repository was last indexed, the map notes that line numbers may have shifted. The hook stays silent for files outside every indexed repository, non-code files, credential paths, and while the server is known to be down. A repository indexed by git URL has no local path to match, so it gets no file context.
Debugging & Resuming Sessions
Hooks are callbacks from Claude Code, not a durable job queue: whatever happens while the plugin is uninstalled or failing never reaches it at all, and nothing replays it later. A backend the plugin cannot reach is the milder case — tool traces and answers are buffered on disk instead and replayed in order on a later prompt, so an outage costs you visibility rather than memory. (A prompt is staged locally and travels as the question half of that answer entry, so it is buffered along with it; there is just no separate buffered-prompt lane.) The memory header tells the two apart. When memory does not appear, check these layers first:
A
claude --resume of the same conversation continues the same live session on its own — the session id is derived from the Claude Code session, and a dataset chosen with /cognee-memory:cognee-switch-datasets is remembered too. COGNEE_SESSION_ID is what you need for the other case: making a second terminal, or a different conversation, write into one shared live session. Set it in both terminals and keep COGNEE_PLUGIN_DATASET the same, otherwise each lands in its own session. After changing credentials, dataset, or session id, restart Claude Code so SessionStart can run with the new state.
Hook commands run with
python3, falling back to python if python3 isn’t found. Any Python 3.9 or newer will do — the hooks are stdlib-only HTTP clients and never import cognee, so the 3.9 that ships with macOS’s Command Line Tools is enough. Local mode does not run the server on it either: the plugin fetches uv into ~/.cognee-plugin/uv and builds its own Python 3.12 virtualenv, falling back to the host interpreter (3.10+) only when uv is neither present nor downloadable. If neither python3 nor python resolves on PATH — most commonly on Windows, where the python.org installer doesn’t always register a python3 alias — every hook fails and no memory is ever captured. Run python3 --version or python --version in the same shell that launches Claude Code to confirm one is available, or reinstall Python with “Add python.exe to PATH” checked.PreCompact was loud about it, failing so Claude Code refused to compact; everything else failed silently. Separately, on the Windows desktop app the session-exit watcher latched onto the short-lived hook runner rather than the Claude session, so the final sync ran seconds after launch and dropped the live connection.
Since 1.6.1, a hook that crashes is reported rather than failing silently. Some installs made before 2026-09-27 also report 1.6.1 but predate this change, and Claude Code won’t offer them an update. If scripts/hook_runner.py is missing from your plugin root, reinstall the plugin. Every hook runs through scripts/hook_runner.py, which writes an uncaught exception’s traceback to ~/.cognee-plugin/claude-code/hook-crash.log and shows it as a one-line Cognee memory: hook <script> failed (...) message, once per hour for the same crash. The hook then exits 0, so the python fallback no longer runs it a second time. This also covers failures that happen before the hook’s own error handling runs, such as an import error or a Python older than 3.9. Before this change these exited with no traceback, and async hooks (prompt, tool, and Stop capture, and the credits refresh) showed nothing at all. If memory stops appearing, check that file first. The runner also switches the hooks’ input and output to UTF-8, because Windows pipes default to the ANSI code page.
The plugin also ships a read-only diagnostic: "${CLAUDE_PLUGIN_ROOT}/scripts/cognee-doctor.sh" — ask Claude to run it, or substitute the plugin’s install path and run it from your own terminal (add --json for machine-readable output). It reports the resolved mode, which env file was read, the server URL and whether it answers, where the API key came from, whether memory is shared or separated, the local and server cognee versions, where the local server’s LLM calls go (a provider key or the Claude observer), the embedding model, and the circuit-breaker state. It never writes anything.
For a clean handoff into long-term memory, run /cognee-memory:cognee-sync or exit Claude Code normally so SessionEnd can trigger the final graph sync. If the process is killed instead, recent session cache entries may exist, but the final session-to-graph sync may not have run yet.
Reading the memory header
Every prompt’s recalled context opens with a one-line header, which is the quickest read on whether memory is working:recall count is how many memory blocks this turn’s lookup injected under === Cognee memory ===. The plugin makes one graph recall per prompt. On a cognee 1.6.0 or later server, each block is the full prompt cognee would have answered from: this session’s conversation history, the retrieved graph context, and the session’s agent guidance. It is injected whole, not truncated. An older server returns only the retrieved context. The header adds / N code whenever the prompt armed the code-graph lane, even when N is 0. saved last turn is what the previous turn wrote: the trace and answer counts are server-confirmed, while the prompt count is bumped as soon as the prompt is staged locally — which is why it still reads 1 in the outage line below. The same numbers are written to ~/.cognee-plugin/claude-code/last_recall.json, and per session under recall/.
When the server cannot be reached, the header names the outage instead of going quiet:
~/.cognee-plugin/claude-code/bridge/ and replay in order on a later prompt once the server answers again, so awaiting replay falling to zero is what confirms the backlog cleared. oldest is worth a look: a buffer belonging to an earlier session only drains when that session runs again, so entries can wait weeks with nothing pointing at them. Both segments disappear once the buffer has drained.
Configuration Reference
Precedence:- Environment variables (shell exports)
~/.cognee/.env— the one-time setup file, shared with the Codex plugin; loaded into the environment at process start, so every variable below exceptCOGNEE_ENV_FILEitself can live in it- Defaults
COGNEE_BACKEND / COGNEE_CLAUDE_BACKEND mode switch follows the same precedence — it is an ordinary variable, and a shell export of it beats a value in the file. What makes it special is its effect: wherever it is set, it pins the mode regardless of where the connection variables are defined.
There is no
config.json. Older plugin versions wrote ~/.cognee-plugin/claude-code/config.json, and SessionStart read a base_url from it while the per-turn hooks did not — so a stale URL there could point the two halves of the plugin at different servers. SessionStart now deletes a leftover file. Put everything in ~/.cognee/.env instead.Update or Remove
The plugin checks for a new version about once an hour. When one is published, the next session start shows “Cognee update available X → Y” and the status line shows an update badge. Nothing installs by itself: run/plugin update cognee-memory@cognee, or enable auto-update for the marketplace. Set COGNEE_UPDATE_CHECK=off to stop the check.
If an update doesn’t take, reinstall:
GitHub Repository
View source code and the full configuration reference
Codex plugin
The same memory plugin for the Codex CLI