pip install. It works in the Codex CLI and can also be activated through the Codex IDE plugin. The plugin hooks into Codex’s lifecycle, so it:
- captures your prompts, tool traces, and assistant responses into session memory
- injects relevant context on every prompt submit
- syncs the session into your knowledge graph on session end
Install
The Cognee memory plugin depends on Codex lifecycle hooks. Enable hooks before installing it.- CLI
- Manual
Enable hooks, then install from the Codex marketplace with the Codex CLI:
Make sure Cognee hooks are enabled for both the Codex CLI and the Codex IDE plugin. If Codex asks you to review hooks, open
/hooks and allow or trust the Cognee hooks. Until hooks are enabled and trusted, Codex will not call the plugin on prompt submit, tool use, stop, compaction, or session end.cognee: <dataset> · <mode> to confirm the plugin is active.
Configure your backend
Configure the plugin once in~/.cognee/.env. The file is created with a commented template on the first session start, and its values act exactly like shell exports — a real export in your shell still wins, per terminal. It is shared with the Claude Code plugin, so both read the same configuration.
- Cognee Cloud / remote
- Local (default)
- Windows (PowerShell), either mode
Point the plugin at Cognee Cloud or a remote server by setting both:Cloud mode is a pure thin client: it talks to your remote server over HTTP only and does not install a local Cognee runtime.Pointing
COGNEE_BASE_URL at a Cognee server you run yourself? Without COGNEE_API_KEY, the plugin logs in as COGNEE_USER_EMAIL / COGNEE_USER_PASSWORD (default default_user@example.com / default_password) to mint its key. From cognee 1.6.0, a server creates the default user only when it starts with DEFAULT_USER_PASSWORD set. So either start your server with DEFAULT_USER_PASSWORD set to the same value as COGNEE_USER_PASSWORD, or set COGNEE_API_KEY so no login is needed. Any other user must already exist on the server. If the login fails, the error message tells you which of these to fix.nano ~/.cognee/.env) works too. Either way, changes apply on the next launch. The file format:
- Comments are whole lines starting with
#— a trailing# noteafter a value becomes part of the value. - Quotes around values are optional, and a leading
exportis tolerated so existing shell profile lines paste verbatim. - Every variable in the Configuration Reference can live here, so there is nothing else to persist. The one exception is
COGNEE_ENV_FILEitself: the plugin reads it before it opens the file, so it only works as a shell export.
Which mode wins, and how to switch
You can configure both modes at once — keepCOGNEE_BASE_URL + COGNEE_API_KEY and LLM_API_KEY in the file together. The mode is then decided per terminal, by three rules in order:
- A
COGNEE_BACKENDexport wins.export COGNEE_BACKEND=local(or=cloud) pins that terminal to that mode. - Otherwise, cloud wins when configured. If
COGNEE_BASE_URLis set — in the file or the shell — the plugin connects to it. - Otherwise, local. With no URL anywhere, the plugin boots the local server.
~/.cognee/.env:
- The switch is pinned.
COGNEE_BACKEND=cloudwith noCOGNEE_BASE_URLconfigured still counts as cloud — the plugin does not silently fall back to local, and the status line shows✕ (missing_cognee_base_url)so you know exactly what to fix. - A forced-local switch blanks
COGNEE_BASE_URLandCOGNEE_API_KEYin the process environment, so the per-prompt hooks and every spawned worker resolve the same local endpoint — not justSessionStart. They are emptied rather than deleted on purpose: a child process that reloads the env file must not re-inject the cloud values. - The shared
COGNEE_BACKENDflips both the Claude Code and Codex plugins in that terminal. To flip only one, use the plugin-specific name —COGNEE_CLAUDE_BACKENDorCOGNEE_CODEX_BACKEND— which beats the shared one. - Accepted values:
local(aliasesnative,sdk) andcloud(aliaseshttp,api,server). Anything else is ignored. COGNEE_BACKENDcan also live in~/.cognee/.envto make a mode the durable default; a shell export still overrides it per terminal.- Not sure what a terminal resolved? The status line’s mode field shows it live.
Cognee’s LLM calls do not run through your coding agent — with one exception: Claude Code’s local mode with no LLM key configured, which runs them on your Claude subscription through the Claude observer. Otherwise, your Claude Code or Codex plan pays only for your conversation with the model. Everything Cognee does on its own — entity and relationship extraction during cognify, summarization, embeddings, and search-time completions — happens inside the Cognee backend against the LLM provider configured there, and is billed by that provider. In local mode you configure it with
LLM_API_KEY; in Cloud/remote mode your tenant holds the key server-side, so no local LLM key is needed.LLM_API_KEY above covers extraction, summarization, and embeddings: Cognee defaults to openai/gpt-5.6-luna for the LLM and openai/text-embedding-3-large for embeddings, and embeddings reuse LLM_API_KEY when EMBEDDING_API_KEY is unset. To use another provider, set LLM_PROVIDER, LLM_MODEL, and — for Azure, Ollama, or OpenAI-compatible endpoints — LLM_ENDPOINT. Changing only the LLM leaves embeddings on OpenAI, so also set the EMBEDDING_* variables or set EMBEDDING_API_KEY to an OpenAI key so the default embeddings keep working. See LLM providers and embedding providers.
Use it
Use Codex as usual — memory is captured and recalled automatically. To verify, end a session with/exit (which syncs it into Cognee), then start a fresh session and ask: “What do you know from cognee?” Answering from a clean session proves it’s recalling from your memory.
The plugin also ships skills for explicit requests. Ask for what you want and Codex picks the matching skill:
Sessions & datasets
- Sessions — by default the plugin derives the Cognee session id from the Codex thread, so a new conversation starts a new one and
codex resumecontinues the same one automatically. (Should a launch report no thread id at all, the plugin falls back to a fresh per-launch id.) SetCOGNEE_SESSION_IDbefore launching to pin a named session, or to deliberately share one live session across two terminals. Whatever you pin is normalized to the ASCII setA-Z a-z 0-9 - _ ., with every other character — accented and non-Latin letters included — replaced by_; leading and trailing.and_are then trimmed, and the result is capped at 120 characters. Soproyecto-cafépinsproyecto-caf, two ids that first differ past character 120 land on one session, and a value with no ASCII alphanumerics at all normalizes to nothing and is ignored, leaving the thread-derived id in place. Every Cognee agent integration applies that same rule, so one value names the same session in Codex as it does in Claude Code. - Datasets — all writes and recall are scoped to one dataset (
agent_sessionsby default). SetCOGNEE_PLUGIN_DATASETto use a custom one. The Codex and Claude Code plugins default to the same one, so memory carries across both.
cognee-agent role holding read and write on your datasets, and the launch’s dataset is addressed by its canonical UUID rather than by name, which is what makes “the same dataset” mean the same rows in Codex and in Claude Code. Grants are backfilled at every session start and roughly every 60 seconds by the idle watcher, so a dataset another plugin creates becomes visible here without a restart. Set COGNEE_SHARED_AGENT_MEMORY=false for separated, per-plugin memory: the agent leaves the shared role and starts on a private dataset, and nothing already written moves. That rests on there being an agent identity to separate, though — under the default COGNEE_PLUGIN_IDENTITY=auto one is provisioned only in service of shared memory, so opting out on a fresh install simply runs the plugin as your own user, which sees everything anyway. Pair it with COGNEE_PLUGIN_IDENTITY=true to insist on a dedicated agent sub-user.
To move a running session to another dataset, ask Codex to switch datasets (the cognee-switch-datasets skill), optionally naming the dataset. Without a name it lists the datasets you can write to as a numbered list; a name that is not listed is created for you. Because a Cognee session never spans two datasets, the switch first syncs the current session into its dataset — and aborts if that fails, changing nothing — then registers a fresh session on the chosen one. The choice lives in the launch record, so it survives a resume and beats COGNEE_PLUGIN_DATASET for the rest of the launch — as does the session it registers, which outranks a pinned COGNEE_SESSION_ID so a shell export cannot drag a switched launch back.
A switch is not needed just to look in another dataset. On every prompt the server answers, the recall hook also tells Codex which other datasets your identity can read — read-only ones included, since a search needs no write access. So when you ask Codex to recall something the active dataset did not have, it offers those datasets as a numbered list; pick one and it runs a one-off, graph-only search there, without the session id, which belongs to the active dataset, and says which dataset the answer came from. Nothing else moves: the active dataset, the Cognee session, and where writes go stay as they were. If the server rejects that search with an HTTP error (other than an auth failure, or a 404, which reads as an empty result), the error Codex reports says which dataset id it searched and adds the server’s own reason from the response, instead of a bare HTTP status.
The listing is cached per plugin at ~/.cognee-plugin/codex/readable-datasets.json and refreshed at most every COGNEE_DATASETS_CACHE_TTL seconds, inside what is left of the recall budget, so the prompt path never waits on it. Set COGNEE_RECALL_DATASET_HINT=off to stop the per-prompt hook from naming the other datasets; asking for the search outright still works through the memory skill.
How It Works
The plugin registers Codex lifecycle hooks:
A background idle watcher persists the session cache after periods of inactivity, and a final sync on session end bridges the session into the permanent graph.
What gets captured
Automatic capture is what fills session memory: your prompts, the tool calls Codex makes, and its answers. Explicit requests — thememory skill, or cognee-remember.sh — are independent of every switch below and always store what you asked for.
Two filters run before a captured value is stored:
- Credential paths are skipped entirely. A tool call whose path arguments point at
.env,.env.*,*.pem,*.key,*.p12,*.pfx,id_rsa*,id_ed25519*,.netrc,.npmrc,*/.ssh/*,*/.aws/credentials,secrets.*, orcredentials.*is not captured at all. Extend the list withCOGNEE_CAPTURE_DENY_PATHS. - Common secrets are redacted. Private key blocks, database connection URLs,
Authorization: Bearer …headers,secret/token/password/api_keyassignments, vendor key prefixes (sk-,ghp_,xoxb-,whsec_, and similar) and bcrypt hashes become[redacted:<kind>]. A value under a credential-looking key —authorization,x-api-key, or any key ending insecret,token,password, orapi_key— is replaced wholesale with[redacted:credential]rather than pattern-matched. This runs on prompts, traces, and answers alike, and before truncation — so a clipped secret cannot slip past the pattern that would have matched it.
COGNEE_CAPTURE=false stops automatic capture altogether: nothing is captured, buffered, or replayed, while recall and explicit remember keep working. It does not erase memory already stored.
Redaction is best effort, and the path filter reads structured tool path arguments rather than arbitrary shell command text — a secret typed into a command line is still captured.
COGNEE_CAPTURE=false is the strict opt-out.Session distillation (self-improvement)
The Cognee coding-agent plugins (Claude Code, Codex) run session distillation for you — you never callimprove() by hand. A distillation pass fires on three triggers:
The idle and every-N triggers share one per-session cooldown: after an automatic improve, the next one waits at least
COGNEE_IMPROVE_COOLDOWN seconds (default 1800, 30 minutes), and a failed attempt starts the same wait as a backoff. Session end, the sync skill, and a dataset switch improve regardless of it.
Overlapping triggers are safe. A per-session improve lock on the server serializes concurrent runs, and unchanged session content dedups server-side by content hash — so a repeat improve over content that hasn’t changed is a cheap no-op, not duplicated work.
Configuration
All triggers are tuned through environment variables read by the plugin. The defaults are chosen so distillation stays out of your way; you rarely need to change them.Turning it down or off
- Stop idle-triggered improves: set
COGNEE_IDLE_DISABLED=1before launching the agent. Session-end and per-turn improves still run. - Reduce mid-session improves: raise
COGNEE_AUTO_IMPROVE_EVERYto a large value so the per-turn trigger effectively never fires within a session. - Session-end distillation always runs when the plugin is active — it’s how a finished session reaches permanent memory.
Confirming it happened
- Cloud UI: the Self-improvement card at the top of a session on the Sessions page shows the status of the last graph enrichment and the dataset it wrote to.
- Plugin hook log: each automatic run emits an
improve_firedevent you can grep for when debugging (in local SDK mode, where the plugin calls the library directly instead of the HTTP endpoint, look forauto_improve_firedinstead). improve-unsupported.jsonmarker: if this file appears in the plugin’s shared state directory (24h TTL), the server rejected the improve endpoint and the plugin fell back to the legacyrememberbridge for that window — a signal the server predates session-aware improve.
Code graph
Repositories can be indexed into a deterministic code graph — symbols, calls, imports, endpoints, dependencies. Indexing makes no LLM or embedding calls by default, so it is fast and costs no tokens. It requires a Cognee server ≥ 1.5.4. Starting a session inside a git repository indexes it automatically — but only when the server is local, so a private checkout is never shipped to a hosted tenant on the plugin’s own initiative. Auto-indexing runs in the background, never blocks the first prompt, and re-indexes after any turn that changed the working tree. Index explicitly when automation won’t: against a Cloud tenant, for a different repo, or for a git URL. The simplest route is to ask the agent to index the repository — its code skill (/cognee-memory:cognee-code in Claude Code, codebase in Codex) runs this command:
--index-vectors to also embed the extracted code facts so semantic search can see them — that is the one flag that makes embedding calls.
Query the graph through the same code skill. Prompts that mention an identifier-shaped token from an indexed repo also get code facts injected automatically by the per-prompt recall hook.
Each indexed repository gets its own dataset, named codebase-<repo-name>-<digest>, where the digest identifies the indexed path or git URL — two checkouts that share a basename would otherwise share one graph and delete each other’s nodes. Code-graph searches resolve the dataset from the current checkout, so the generated name rarely needs typing. Indexing writes its snapshot into the repository itself at <repo>/.enola/ (untracked) — add .enola/ to the repository’s .gitignore or your global excludes.
always also accepts 1, true, yes, and on; off also accepts 0, false, and no. Any other value means auto. Automatic indexing skips directories that are not git repositories, hold no source files, or exceed 3000 source files. Explicit indexing has no size cap.
Debugging & Resuming Sessions
Hooks are callbacks from Codex, not a durable job queue: whatever happens while hooks are disabled or untrusted never reaches the plugin at all, and nothing replays it later. A backend the plugin cannot reach is the milder case — tool traces and answers are buffered on disk instead and replayed on a later prompt, so an outage costs you visibility rather than memory. The memory header tells the two apart. When memory does not appear, check these layers first:
A resume keeps both the session and any dataset you switched to on its own.
COGNEE_SESSION_ID is what you need for the other case: making a second terminal, or a different conversation, write into one shared live session. Set it in both terminals and keep COGNEE_PLUGIN_DATASET the same, otherwise each lands in its own session. After changing hook trust, credentials, dataset, or session id, restart Codex so SessionStart can run with the new state.
Hook commands run with
python3, falling back to python if python3 isn’t found. On Windows, Codex runs each hook through cmd.exe with a separate command that tries py -3 (the Python launcher) and then python, because python3 there is usually the Microsoft Store stub. Any Python 3.9 or newer will do — the hooks are stdlib-only HTTP clients and never import cognee, so the 3.9 that ships with macOS’s Command Line Tools is enough. Local mode does not run the server on it either: the plugin fetches uv into ~/.cognee-plugin/uv and builds its own Python 3.12 virtualenv, falling back to the host interpreter (3.10+) only when uv is neither present nor downloadable. If none of those interpreters resolves on PATH, every hook fails with a “hook failure” error and no session is ever created. Run python3 --version (on Windows, py -3 --version or python --version) in the same shell that launches Codex to confirm one is available, or reinstall Python with “Add python.exe to PATH” checked.scripts/hook_runner.py, so a hook that crashes — or runs on a Python older than 3.9 — is reported rather than failing silently. The traceback is appended to ~/.cognee-plugin/codex/hook-crash.log, Codex shows a Cognee memory: hook <script> failed (…) message that points at that file (once per hour for the same crash, so a failing PostToolUse hook doesn’t repeat it on every tool call), and the hook exits cleanly instead of being run a second time. Memory for that hook is skipped until the cause is fixed, so check hook-crash.log first when that message appears.
The plugin also ships a read-only diagnostic: "${CODEX_PLUGIN_ROOT}/scripts/cognee-cli.sh" doctor — ask Codex to run it, or substitute the plugin’s install path and run it from your own terminal (add --json for machine-readable output). It reports the resolved mode and which backend switch forced it, the env file with the key names it defines and any a shell export is shadowing, the server URL and whether it answers, where the API key came from, whether memory is shared or separated, the local and server cognee versions, the embedding model, and the circuit-breaker state. It never writes anything, and it needs no Cognee checkout — unlike the rest of cognee-cli.sh, the doctor subcommand runs from wherever you happen to be.
Exit Codex normally (for example with /exit) when you want SessionEnd to trigger the final graph sync. If the process is killed instead, the detached exit watcher that SessionStart left running is the fallback: it follows the Codex session process itself, and starts the sync once that process is gone.
Reading the memory header
Every prompt’s recalled context opens with a one-line header, which is the quickest read on whether memory is working:1 memory hit is how many memory blocks this turn’s lookup found and injected. Each prompt makes one graph-scope recall request. Against cognee 1.6.0 or later, that request returns, for each dataset, one block that holds the session’s conversation history, the retrieved graph context, and the session guidance. The block is injected whole under === Cognee memory ===; there is no client-side length cap, and top_k limits the size on the server. Against an older server, the block holds only the retrieved graph context. On a prompt that arms the repository code lane — an identifier-shaped token in the prompt, and a working directory inside a repo you indexed — N code facts follows the hit count, and those facts are part of the total. 12/40 turns had hits this session is the running ratio, reading memory warming up (7 turns) until the first hit. saved last turn counts what the previous turn wrote, though not all of it is a server write: 1 prompt means the prompt was recorded locally, to be sent paired with that turn’s answer, while the trace and answer counts are server writes only. The same numbers are written to ~/.cognee-plugin/codex/last_recall.json.
When the server cannot be reached, the header grows the outage instead of going quiet:
~/.cognee-plugin/codex/bridge/ and replay in order on a later prompt once the server answers again, so awaiting replay falling to zero is what confirms the backlog cleared. oldest is worth a look: a buffer belonging to an earlier session only drains when that session runs again, so entries can wait weeks with nothing pointing at them. Both segments disappear once the buffer has drained.
Configuration Reference
Precedence:- Environment variables (shell exports)
~/.cognee/.env— the one-time setup file, shared with the Claude Code plugin; loaded into the environment at process start, so every variable below exceptCOGNEE_ENV_FILEitself can live in it- Defaults
COGNEE_BACKEND / COGNEE_CODEX_BACKEND mode switch follows the same precedence; its effect is described in Which mode wins.
There is no
config.json. Older plugin versions wrote ~/.cognee-plugin/config.json, and SessionStart read a base_url from it while the per-turn hooks did not — so a stale URL there could point the two halves of the plugin at different servers. SessionStart now deletes a leftover file. Put everything in ~/.cognee/.env instead.Update or Remove
Thecognee marketplace tracks the repository’s main branch, so updates arrive as new commits and are not automatic. Pull the latest with:
GitHub Repository
View source code and the full configuration reference
Claude Code plugin
The same memory plugin for Claude Code