List datasets
GET /api/v1/datasets/ — List all datasets accessible to the authenticated user.
Create a dataset
POST /api/v1/datasets/ — Create a new dataset or return the existing one if the name already exists.
Datasets are also created implicitly when you call
add or remember with a dataset_name that does not yet exist.Dataset status
GET /api/v1/datasets/status — Get the processing status of all datasets.
Returns the pipeline state for each dataset: whether cognify is pending, running, or completed.
In-flight progress
GET /api/v1/datasets/status/progress — Get the same statuses, plus how far each running pipeline has got.
Takes the same selection parameters as /status: repeat dataset for each dataset UUID (omit it to cover every dataset you can read) and repeat pipeline to pick pipelines (omit it to default to cognify_pipeline).
{status, progress} rather than a bare status. The flat-versus-nested rule matches /status — flat for zero or one pipeline, nested per dataset and pipeline for more than one:
- completed_items / total_items: files finished out of files in the run. An item that errored still counts as finished, so
completed_itemsalways reachestotal_items. - current_stage: the task most recently entered in an item’s task chain. Results stream through the whole chain before they surface, so in practice this names the chain’s final task (
add_data_pointson the default cognify pipeline) once the first result lands, and isnullbefore that. Treat it as “the run is moving”, not as a stage-by-stage position —completed_items/total_itemsare the fields to drive a progress bar from.
progress is null when there is nothing in flight to report — before the run’s first progress tick, and once the run reaches a terminal state, since the completed or errored record carries no progress snapshot. Treat null as “no progress information”, not as an error.
Progress ticks are throttled to roughly 20 database writes per run (the first and last file always persist), so on a large batch the numbers advance in steps rather than one file at a time.
/status is unchanged — it still returns bare status values in the same shape. Use /status/progress when you want the N-of-M numbers, and the cognify WebSocket when you want live in-flight updates while a run is executing; polling this endpoint is how a client recovers granular progress after a page refresh or a dropped subscription.409 is returned if the progress cannot be retrieved — including when you ask for a dataset you do not have read permission on.
Graph summary
GET /api/v1/datasets/graph-summary — Get node and edge counts for each dataset.
Counts are computed once per dataset’s latest cognify run and cached, so this is much cheaper than the full graph endpoint when polling dataset sizes repeatedly.
dataset_ids query parameters to summarize specific datasets. Omit it to summarize every dataset you have read access to.
- datasetId: The dataset’s UUID
- pipelineRunId: The dataset’s latest cognify run, or
nullif it has never been cognified - numNodes / numEdges: Graph size for that run
- computedAt: When the counts were cached, or
nullif they were not cached on this call. Anullhere has two causes, and they mean opposite things: either the last count attempt degraded (for example, the graph store was unavailable), in which case the counts are0placeholders and the next call retries; or a concurrent caller cached the same run first, in which case the counts you were served are exact. Read the counts themselves rather than treating everynullas a failure.
409 is returned when the summary could not be built. This means the relational read failed — a single unreadable graph store does not cause it, since that dataset simply comes back with zero counts.
Live dataset updates
WS /api/v1/visualize/subscribe/{dataset_id} — Follow one dataset’s activity over a WebSocket instead of polling.
This is the push replacement for repeatedly calling GET /api/v1/visualize/live-events and refetching the graph payload just to notice that a cognify run finished. Authentication accepts the same credentials as every other endpoint — API key header, bearer header, or auth cookie, sent with the handshake — plus one WebSocket-only fallback: a browser cannot set headers when opening a WebSocket, so the same API key or bearer token can be passed as a ?token= query parameter instead. Prefer the header or cookie where you can send one — a handshake is an HTTP request, and its full URL (query string included) lands in the default access logs of common reverse proxies such as nginx or an AWS ALB, so deployments terminating WebSocket traffic behind one should redact the token parameter there. Cognee redacts it from uvicorn’s own logs. Nothing is ever read from client frames.
Every frame is a JSON object discriminated by kind:
ready—{"kind": "ready", "dataset_id": str, "cursor": str | null}, sent once the connection is authorized, echoing the cursor it starts fromlive_events—{"kind": "live_events", "events": [...], "cursor": str}, roughly every 2s and only when the delta is non-empty. The events are exactly whatGET /api/v1/visualize/live-eventsreturnsgraph_grew—{"kind": "graph_grew", "pipeline_run_id": str}, within about 5s of a cognify run for this dataset completing. A run already complete when you connected is the baseline and is not announcedheartbeat—{"kind": "heartbeat", "time": str}, every 15s or so, so a quiet stream is distinguishable from a dead one
cursor from the last live_events frame back as the since query parameter when reconnecting; omit it to start from every available event.
Close codes:
- 1008: not authenticated, or no read permission on this dataset. A retry replays the same rejection, so clients should stop. Permission is re-checked on every poll, so access revoked mid-stream closes the connection this way too
- 1011: the stream failed server-side; reconnecting is reasonable
since, or a dataset_id that is not a UUID, fails request validation before the connection is accepted. Depending on the ASGI server the client sees either a 1008 close or a plain HTTP rejection of the handshake, so treat a handshake that never opened as a client-side bug rather than a stream error.
Dataset data
GET /api/v1/datasets/{dataset_id}/data — List all data items in a dataset.
Each item also carries the label and metadata attached at upload time:
- label: The label given to this file via the
labelsfield on add or remember, ornullif none was set - externalMetadata: The stored metadata object — the
external_metadataentry you sent merged over loader-derived keys, plus anode_setkey when one was passed at ingest. A file uploaded without metadata stores an empty object rather thannull.
GET /api/v1/datasets/{dataset_id}/data/{data_id}/raw — Download the original file for a specific data item. Requires read permission on the containing dataset; ownership of the data item itself is not required, so a dataset shared with you is fully downloadable.
Delete
DELETE /api/v1/datasets/{dataset_id} — Delete a dataset and all its contents.
DELETE /api/v1/datasets/{dataset_id}/data/{data_id} — Delete a specific data item from a dataset.