Skip to main content
Add persistent memory to Codex with the Cognee memory plugin — no code and no pip install. It works in the Codex CLI and can also be activated through the Codex IDE plugin. The plugin hooks into Codex’s lifecycle, so it:
  • captures your prompts, tool traces, and assistant responses into session memory
  • injects relevant context on every prompt submit
  • syncs the session into your knowledge graph on session end
Sessions are disposable; your memory isn’t.

Install

The Cognee memory plugin depends on Codex lifecycle hooks. Enable hooks before installing it.
Enable hooks, then install from the Codex marketplace with the Codex CLI:
Make sure Cognee hooks are enabled for both the Codex CLI and the Codex IDE plugin. If Codex asks you to review hooks, open /hooks and allow or trust the Cognee hooks. Until hooks are enabled and trusted, Codex will not call the plugin on prompt submit, tool use, stop, compaction, or session end.
On startup the status line shows cognee: <dataset> · <mode> to confirm the plugin is active.

Configure your backend

Configure the plugin once in ~/.cognee/.env. The file is created with a commented template on the first session start, and its values act exactly like shell exports — a real export in your shell still wins, per terminal. It is shared with the Claude Code plugin, so both read the same configuration.
Point the plugin at Cognee Cloud or a remote server by setting both:
Cloud mode is a pure thin client: it talks to your remote server over HTTP only and does not install a local Cognee runtime.Pointing COGNEE_BASE_URL at a Cognee server you run yourself? Without COGNEE_API_KEY, the plugin logs in as COGNEE_USER_EMAIL / COGNEE_USER_PASSWORD (default default_user@example.com / default_password) to mint its key. From cognee 1.6.0, a server creates the default user only when it starts with DEFAULT_USER_PASSWORD set. So either start your server with DEFAULT_USER_PASSWORD set to the same value as COGNEE_USER_PASSWORD, or set COGNEE_API_KEY so no login is needed. Any other user must already exist on the server. If the login fails, the error message tells you which of these to fix.
Re-running any of these blocks is safe: when a key appears more than once the last value wins, so pasting again with a new value updates the setting. Editing the file directly (nano ~/.cognee/.env) works too. Either way, changes apply on the next launch. The file format:
  • Comments are whole lines starting with # — a trailing # note after a value becomes part of the value.
  • Quotes around values are optional, and a leading export is tolerated so existing shell profile lines paste verbatim.
  • Every variable in the Configuration Reference can live here, so there is nothing else to persist. The one exception is COGNEE_ENV_FILE itself: the plugin reads it before it opens the file, so it only works as a shell export.

Which mode wins, and how to switch

You can configure both modes at once — keep COGNEE_BASE_URL + COGNEE_API_KEY and LLM_API_KEY in the file together. The mode is then decided per terminal, by three rules in order:
  1. A COGNEE_BACKEND export wins. export COGNEE_BACKEND=local (or =cloud) pins that terminal to that mode.
  2. Otherwise, cloud wins when configured. If COGNEE_BASE_URL is set — in the file or the shell — the plugin connects to it.
  3. Otherwise, local. With no URL anywhere, the plugin boots the local server.
So with both modes in ~/.cognee/.env:
To go local, use the switch — not unset COGNEE_BASE_URL. Unsetting does not work: the env file re-injects the URL at the next launch. Export COGNEE_BACKEND=local instead. (Deleting the line from the file works too, but that changes the default for every terminal.)
Details worth knowing:
  • The switch is pinned. COGNEE_BACKEND=cloud with no COGNEE_BASE_URL configured still counts as cloud — the plugin does not silently fall back to local, and the status line shows ✕ (missing_cognee_base_url) so you know exactly what to fix.
  • A forced-local switch blanks COGNEE_BASE_URL and COGNEE_API_KEY in the process environment, so the per-prompt hooks and every spawned worker resolve the same local endpoint — not just SessionStart. They are emptied rather than deleted on purpose: a child process that reloads the env file must not re-inject the cloud values.
  • The shared COGNEE_BACKEND flips both the Claude Code and Codex plugins in that terminal. To flip only one, use the plugin-specific name — COGNEE_CLAUDE_BACKEND or COGNEE_CODEX_BACKEND — which beats the shared one.
  • Accepted values: local (aliases native, sdk) and cloud (aliases http, api, server). Anything else is ignored.
  • COGNEE_BACKEND can also live in ~/.cognee/.env to make a mode the durable default; a shell export still overrides it per terminal.
  • Not sure what a terminal resolved? The status line’s mode field shows it live.
Cognee’s LLM calls do not run through your coding agent — with one exception: Claude Code’s local mode with no LLM key configured, which runs them on your Claude subscription through the Claude observer. Otherwise, your Claude Code or Codex plan pays only for your conversation with the model. Everything Cognee does on its own — entity and relationship extraction during cognify, summarization, embeddings, and search-time completions — happens inside the Cognee backend against the LLM provider configured there, and is billed by that provider. In local mode you configure it with LLM_API_KEY; in Cloud/remote mode your tenant holds the key server-side, so no local LLM key is needed.
In local mode, the single LLM_API_KEY above covers extraction, summarization, and embeddings: Cognee defaults to openai/gpt-5.6-luna for the LLM and openai/text-embedding-3-large for embeddings, and embeddings reuse LLM_API_KEY when EMBEDDING_API_KEY is unset. To use another provider, set LLM_PROVIDER, LLM_MODEL, and — for Azure, Ollama, or OpenAI-compatible endpoints — LLM_ENDPOINT. Changing only the LLM leaves embeddings on OpenAI, so also set the EMBEDDING_* variables or set EMBEDDING_API_KEY to an OpenAI key so the default embeddings keep working. See LLM providers and embedding providers.

Use it

Use Codex as usual — memory is captured and recalled automatically. To verify, end a session with /exit (which syncs it into Cognee), then start a fresh session and ask: “What do you know from cognee?” Answering from a clean session proves it’s recalling from your memory. The plugin also ships skills for explicit requests. Ask for what you want and Codex picks the matching skill:

Sessions & datasets

  • Sessions — by default the plugin derives the Cognee session id from the Codex thread, so a new conversation starts a new one and codex resume continues the same one automatically. (Should a launch report no thread id at all, the plugin falls back to a fresh per-launch id.) Set COGNEE_SESSION_ID before launching to pin a named session, or to deliberately share one live session across two terminals. Whatever you pin is normalized to the ASCII set A-Z a-z 0-9 - _ ., with every other character — accented and non-Latin letters included — replaced by _; leading and trailing . and _ are then trimmed, and the result is capped at 120 characters. So proyecto-café pins proyecto-caf, two ids that first differ past character 120 land on one session, and a value with no ASCII alphanumerics at all normalizes to nothing and is ignored, leaving the thread-derived id in place. Every Cognee agent integration applies that same rule, so one value names the same session in Codex as it does in Claude Code.
  • Datasets — all writes and recall are scoped to one dataset (agent_sessions by default). Set COGNEE_PLUGIN_DATASET to use a custom one. The Codex and Claude Code plugins default to the same one, so memory carries across both.
Sharing a dataset takes more than agreeing on a name. Where the server supports it, each plugin authenticates as its own agent sub-user, and grants only flow child to parent: your user sees what an agent writes, but an agent sees nothing your user or another plugin’s agent owns. Shared agent memory — on by default — closes that gap. Every plugin agent of your user joins one cognee-agent role holding read and write on your datasets, and the launch’s dataset is addressed by its canonical UUID rather than by name, which is what makes “the same dataset” mean the same rows in Codex and in Claude Code. Grants are backfilled at every session start and roughly every 60 seconds by the idle watcher, so a dataset another plugin creates becomes visible here without a restart. Set COGNEE_SHARED_AGENT_MEMORY=false for separated, per-plugin memory: the agent leaves the shared role and starts on a private dataset, and nothing already written moves. That rests on there being an agent identity to separate, though — under the default COGNEE_PLUGIN_IDENTITY=auto one is provisioned only in service of shared memory, so opting out on a fresh install simply runs the plugin as your own user, which sees everything anyway. Pair it with COGNEE_PLUGIN_IDENTITY=true to insist on a dedicated agent sub-user. To move a running session to another dataset, ask Codex to switch datasets (the cognee-switch-datasets skill), optionally naming the dataset. Without a name it lists the datasets you can write to as a numbered list; a name that is not listed is created for you. Because a Cognee session never spans two datasets, the switch first syncs the current session into its dataset — and aborts if that fails, changing nothing — then registers a fresh session on the chosen one. The choice lives in the launch record, so it survives a resume and beats COGNEE_PLUGIN_DATASET for the rest of the launch — as does the session it registers, which outranks a pinned COGNEE_SESSION_ID so a shell export cannot drag a switched launch back. A switch is not needed just to look in another dataset. On every prompt the server answers, the recall hook also tells Codex which other datasets your identity can read — read-only ones included, since a search needs no write access. So when you ask Codex to recall something the active dataset did not have, it offers those datasets as a numbered list; pick one and it runs a one-off, graph-only search there, without the session id, which belongs to the active dataset, and says which dataset the answer came from. Nothing else moves: the active dataset, the Cognee session, and where writes go stay as they were. If the server rejects that search with an HTTP error (other than an auth failure, or a 404, which reads as an empty result), the error Codex reports says which dataset id it searched and adds the server’s own reason from the response, instead of a bare HTTP status. The listing is cached per plugin at ~/.cognee-plugin/codex/readable-datasets.json and refreshed at most every COGNEE_DATASETS_CACHE_TTL seconds, inside what is left of the recall budget, so the prompt path never waits on it. Set COGNEE_RECALL_DATASET_HINT=off to stop the per-prompt hook from naming the other datasets; asking for the search outright still works through the memory skill.

How It Works

The plugin registers Codex lifecycle hooks: A background idle watcher persists the session cache after periods of inactivity, and a final sync on session end bridges the session into the permanent graph.

What gets captured

Automatic capture is what fills session memory: your prompts, the tool calls Codex makes, and its answers. Explicit requests — the memory skill, or cognee-remember.sh — are independent of every switch below and always store what you asked for. Two filters run before a captured value is stored:
  • Credential paths are skipped entirely. A tool call whose path arguments point at .env, .env.*, *.pem, *.key, *.p12, *.pfx, id_rsa*, id_ed25519*, .netrc, .npmrc, */.ssh/*, */.aws/credentials, secrets.*, or credentials.* is not captured at all. Extend the list with COGNEE_CAPTURE_DENY_PATHS.
  • Common secrets are redacted. Private key blocks, database connection URLs, Authorization: Bearer … headers, secret / token / password / api_key assignments, vendor key prefixes (sk-, ghp_, xoxb-, whsec_, and similar) and bcrypt hashes become [redacted:<kind>]. A value under a credential-looking key — authorization, x-api-key, or any key ending in secret, token, password, or api_key — is replaced wholesale with [redacted:credential] rather than pattern-matched. This runs on prompts, traces, and answers alike, and before truncation — so a clipped secret cannot slip past the pattern that would have matched it.
COGNEE_CAPTURE=false stops automatic capture altogether: nothing is captured, buffered, or replayed, while recall and explicit remember keep working. It does not erase memory already stored.
Redaction is best effort, and the path filter reads structured tool path arguments rather than arbitrary shell command text — a secret typed into a command line is still captured. COGNEE_CAPTURE=false is the strict opt-out.

Session distillation (self-improvement)

The Cognee coding-agent plugins (Claude Code, Codex) run session distillation for you — you never call improve() by hand. A distillation pass fires on three triggers: The idle and every-N triggers share one per-session cooldown: after an automatic improve, the next one waits at least COGNEE_IMPROVE_COOLDOWN seconds (default 1800, 30 minutes), and a failed attempt starts the same wait as a backoff. Session end, the sync skill, and a dataset switch improve regardless of it.
Overlapping triggers are safe. A per-session improve lock on the server serializes concurrent runs, and unchanged session content dedups server-side by content hash — so a repeat improve over content that hasn’t changed is a cheap no-op, not duplicated work.

Configuration

All triggers are tuned through environment variables read by the plugin. The defaults are chosen so distillation stays out of your way; you rarely need to change them.
The plugin READMEs document additional advanced knobs — timing (poll deadlines, busy-retry intervals for a held session lock), session-sync retries, and the update-notification variables (COGNEE_UPDATE_CHECK, COGNEE_UPDATE_CHECK_INTERVAL). You almost never need them — reach for the table above first.

Turning it down or off

  • Stop idle-triggered improves: set COGNEE_IDLE_DISABLED=1 before launching the agent. Session-end and per-turn improves still run.
  • Reduce mid-session improves: raise COGNEE_AUTO_IMPROVE_EVERY to a large value so the per-turn trigger effectively never fires within a session.
  • Session-end distillation always runs when the plugin is active — it’s how a finished session reaches permanent memory.

Confirming it happened

  • Cloud UI: the Self-improvement card at the top of a session on the Sessions page shows the status of the last graph enrichment and the dataset it wrote to.
  • Plugin hook log: each automatic run emits an improve_fired event you can grep for when debugging (in local SDK mode, where the plugin calls the library directly instead of the HTTP endpoint, look for auto_improve_fired instead).
  • improve-unsupported.json marker: if this file appears in the plugin’s shared state directory (24h TTL), the server rejected the improve endpoint and the plugin fell back to the legacy remember bridge for that window — a signal the server predates session-aware improve.

Code graph

Repositories can be indexed into a deterministic code graph — symbols, calls, imports, endpoints, dependencies. Indexing makes no LLM or embedding calls by default, so it is fast and costs no tokens. It requires a Cognee server ≥ 1.5.4. Starting a session inside a git repository indexes it automatically — but only when the server is local, so a private checkout is never shipped to a hosted tenant on the plugin’s own initiative. Auto-indexing runs in the background, never blocks the first prompt, and re-indexes after any turn that changed the working tree. Index explicitly when automation won’t: against a Cloud tenant, for a different repo, or for a git URL. The simplest route is to ask the agent to index the repository — its code skill (/cognee-memory:cognee-code in Claude Code, codebase in Codex) runs this command:
The plugin-root variable is expanded only inside the plugin’s skills, so from your own terminal substitute the plugin’s install path. An explicit request is its own consent, so it bypasses every auto-index gate below. Add --index-vectors to also embed the extracted code facts so semantic search can see them — that is the one flag that makes embedding calls. Query the graph through the same code skill. Prompts that mention an identifier-shaped token from an indexed repo also get code facts injected automatically by the per-prompt recall hook. Each indexed repository gets its own dataset, named codebase-<repo-name>-<digest>, where the digest identifies the indexed path or git URL — two checkouts that share a basename would otherwise share one graph and delete each other’s nodes. Code-graph searches resolve the dataset from the current checkout, so the generated name rarely needs typing. Indexing writes its snapshot into the repository itself at <repo>/.enola/ (untracked) — add .enola/ to the repository’s .gitignore or your global excludes.
Freshness depends on where the server runs. A local server indexes the repository path on this machine, so the graph reflects your working tree, including uncommitted and untracked changes. A cloud or remote server clones a git URL — it cannot read your disk, so the graph reflects only the last pushed commit, and the plugin does not re-submit URL-indexed repositories after local edits. The output looks identical either way, so push before relying on code answers about work in progress, or use a local server for branches you are actively editing.
always also accepts 1, true, yes, and on; off also accepts 0, false, and no. Any other value means auto. Automatic indexing skips directories that are not git repositories, hold no source files, or exceed 3000 source files. Explicit indexing has no size cap.

Debugging & Resuming Sessions

Hooks are callbacks from Codex, not a durable job queue: whatever happens while hooks are disabled or untrusted never reaches the plugin at all, and nothing replays it later. A backend the plugin cannot reach is the milder case — tool traces and answers are buffered on disk instead and replayed on a later prompt, so an outage costs you visibility rather than memory. The memory header tells the two apart. When memory does not appear, check these layers first: A resume keeps both the session and any dataset you switched to on its own. COGNEE_SESSION_ID is what you need for the other case: making a second terminal, or a different conversation, write into one shared live session. Set it in both terminals and keep COGNEE_PLUGIN_DATASET the same, otherwise each lands in its own session. After changing hook trust, credentials, dataset, or session id, restart Codex so SessionStart can run with the new state.
Hook commands run with python3, falling back to python if python3 isn’t found. On Windows, Codex runs each hook through cmd.exe with a separate command that tries py -3 (the Python launcher) and then python, because python3 there is usually the Microsoft Store stub. Any Python 3.9 or newer will do — the hooks are stdlib-only HTTP clients and never import cognee, so the 3.9 that ships with macOS’s Command Line Tools is enough. Local mode does not run the server on it either: the plugin fetches uv into ~/.cognee-plugin/uv and builds its own Python 3.12 virtualenv, falling back to the host interpreter (3.10+) only when uv is neither present nor downloadable. If none of those interpreters resolves on PATH, every hook fails with a “hook failure” error and no session is ever created. Run python3 --version (on Windows, py -3 --version or python --version) in the same shell that launches Codex to confirm one is available, or reinstall Python with “Add python.exe to PATH” checked.
Every hook runs through the plugin’s scripts/hook_runner.py, so a hook that crashes — or runs on a Python older than 3.9 — is reported rather than failing silently. The traceback is appended to ~/.cognee-plugin/codex/hook-crash.log, Codex shows a Cognee memory: hook <script> failed (…) message that points at that file (once per hour for the same crash, so a failing PostToolUse hook doesn’t repeat it on every tool call), and the hook exits cleanly instead of being run a second time. Memory for that hook is skipped until the cause is fixed, so check hook-crash.log first when that message appears. The plugin also ships a read-only diagnostic: "${CODEX_PLUGIN_ROOT}/scripts/cognee-cli.sh" doctor — ask Codex to run it, or substitute the plugin’s install path and run it from your own terminal (add --json for machine-readable output). It reports the resolved mode and which backend switch forced it, the env file with the key names it defines and any a shell export is shadowing, the server URL and whether it answers, where the API key came from, whether memory is shared or separated, the local and server cognee versions, the embedding model, and the circuit-breaker state. It never writes anything, and it needs no Cognee checkout — unlike the rest of cognee-cli.sh, the doctor subcommand runs from wherever you happen to be. Exit Codex normally (for example with /exit) when you want SessionEnd to trigger the final graph sync. If the process is killed instead, the detached exit watcher that SessionStart left running is the fallback: it follows the Codex session process itself, and starts the sync once that process is gone.

Reading the memory header

Every prompt’s recalled context opens with a one-line header, which is the quickest read on whether memory is working:
1 memory hit is how many memory blocks this turn’s lookup found and injected. Each prompt makes one graph-scope recall request. Against cognee 1.6.0 or later, that request returns, for each dataset, one block that holds the session’s conversation history, the retrieved graph context, and the session guidance. The block is injected whole under === Cognee memory ===; there is no client-side length cap, and top_k limits the size on the server. Against an older server, the block holds only the retrieved graph context. On a prompt that arms the repository code lane — an identifier-shaped token in the prompt, and a working directory inside a repo you indexed — N code facts follows the hit count, and those facts are part of the total. 12/40 turns had hits this session is the running ratio, reading memory warming up (7 turns) until the first hit. saved last turn counts what the previous turn wrote, though not all of it is a server write: 1 prompt means the prompt was recorded locally, to be sent paired with that turn’s answer, while the trace and answer counts are server writes only. The same numbers are written to ~/.cognee-plugin/codex/last_recall.json. When the server cannot be reached, the header grows the outage instead of going quiet:
A buffered trace or answer is never counted as saved. Buffered entries sit in ~/.cognee-plugin/codex/bridge/ and replay in order on a later prompt once the server answers again, so awaiting replay falling to zero is what confirms the backlog cleared. oldest is worth a look: a buffer belonging to an earlier session only drains when that session runs again, so entries can wait weeks with nothing pointing at them. Both segments disappear once the buffer has drained.

Configuration Reference

Precedence:
  1. Environment variables (shell exports)
  2. ~/.cognee/.env — the one-time setup file, shared with the Claude Code plugin; loaded into the environment at process start, so every variable below except COGNEE_ENV_FILE itself can live in it
  3. Defaults
The COGNEE_BACKEND / COGNEE_CODEX_BACKEND mode switch follows the same precedence; its effect is described in Which mode wins.
There is no config.json. Older plugin versions wrote ~/.cognee-plugin/config.json, and SessionStart read a base_url from it while the per-turn hooks did not — so a stale URL there could point the two halves of the plugin at different servers. SessionStart now deletes a leftover file. Put everything in ~/.cognee/.env instead.

Update or Remove

The cognee marketplace tracks the repository’s main branch, so updates arrive as new commits and are not automatic. Pull the latest with:
If a stale cached copy persists, remove and re-add the plugin:

GitHub Repository

View source code and the full configuration reference

Claude Code plugin

The same memory plugin for Claude Code