Skip to main content
Datasets are the organizational unit for all data in Cognee Cloud. Each dataset maintains its own knowledge graph and vector store. See Datasets for the underlying concept.

List datasets

GET /api/v1/datasets/ — List all datasets accessible to the authenticated user.

Create a dataset

POST /api/v1/datasets/ — Create a new dataset or return the existing one if the name already exists.
Datasets are also created implicitly when you call add or remember with a dataset_name that does not yet exist.

Dataset status

GET /api/v1/datasets/status — Get the processing status of all datasets. Returns the pipeline state for each dataset: whether cognify is pending, running, or completed.

In-flight progress

GET /api/v1/datasets/status/progress — Get the same statuses, plus how far each running pipeline has got. Takes the same selection parameters as /status: repeat dataset for each dataset UUID (omit it to cover every dataset you can read) and repeat pipeline to pick pipelines (omit it to default to cognify_pipeline).
Every value is an object {status, progress} rather than a bare status. The flat-versus-nested rule matches /status — flat for zero or one pipeline, nested per dataset and pipeline for more than one:
  • completed_items / total_items: files finished out of files in the run. An item that errored still counts as finished, so completed_items always reaches total_items.
  • current_stage: the task most recently entered in an item’s task chain. Results stream through the whole chain before they surface, so in practice this names the chain’s final task (add_data_points on the default cognify pipeline) once the first result lands, and is null before that. Treat it as “the run is moving”, not as a stage-by-stage position — completed_items/total_items are the fields to drive a progress bar from.
progress is null when there is nothing in flight to report — before the run’s first progress tick, and once the run reaches a terminal state, since the completed or errored record carries no progress snapshot. Treat null as “no progress information”, not as an error. Progress ticks are throttled to roughly 20 database writes per run (the first and last file always persist), so on a large batch the numbers advance in steps rather than one file at a time.
/status is unchanged — it still returns bare status values in the same shape. Use /status/progress when you want the N-of-M numbers, and the cognify WebSocket when you want live in-flight updates while a run is executing; polling this endpoint is how a client recovers granular progress after a page refresh or a dropped subscription.
A 409 is returned if the progress cannot be retrieved — including when you ask for a dataset you do not have read permission on.

Graph summary

GET /api/v1/datasets/graph-summary — Get node and edge counts for each dataset. Counts are computed once per dataset’s latest cognify run and cached, so this is much cheaper than the full graph endpoint when polling dataset sizes repeatedly.
Pass one or more dataset_ids query parameters to summarize specific datasets. Omit it to summarize every dataset you have read access to.
Returns a list of summaries, one per dataset, each containing:
  • datasetId: The dataset’s UUID
  • pipelineRunId: The dataset’s latest cognify run, or null if it has never been cognified
  • numNodes / numEdges: Graph size for that run
  • computedAt: When the counts were cached, or null if they were not cached on this call. A null here has two causes, and they mean opposite things: either the last count attempt degraded (for example, the graph store was unavailable), in which case the counts are 0 placeholders and the next call retries; or a concurrent caller cached the same run first, in which case the counts you were served are exact. Read the counts themselves rather than treating every null as a failure.
A 409 is returned when the summary could not be built. This means the relational read failed — a single unreadable graph store does not cause it, since that dataset simply comes back with zero counts.

Live dataset updates

WS /api/v1/visualize/subscribe/{dataset_id} — Follow one dataset’s activity over a WebSocket instead of polling. This is the push replacement for repeatedly calling GET /api/v1/visualize/live-events and refetching the graph payload just to notice that a cognify run finished. Authentication accepts the same credentials as every other endpoint — API key header, bearer header, or auth cookie, sent with the handshake — plus one WebSocket-only fallback: a browser cannot set headers when opening a WebSocket, so the same API key or bearer token can be passed as a ?token= query parameter instead. Prefer the header or cookie where you can send one — a handshake is an HTTP request, and its full URL (query string included) lands in the default access logs of common reverse proxies such as nginx or an AWS ALB, so deployments terminating WebSocket traffic behind one should redact the token parameter there. Cognee redacts it from uvicorn’s own logs. Nothing is ever read from client frames. Every frame is a JSON object discriminated by kind:
  • ready{"kind": "ready", "dataset_id": str, "cursor": str | null}, sent once the connection is authorized, echoing the cursor it starts from
  • live_events{"kind": "live_events", "events": [...], "cursor": str}, roughly every 2s and only when the delta is non-empty. The events are exactly what GET /api/v1/visualize/live-events returns
  • graph_grew{"kind": "graph_grew", "pipeline_run_id": str}, within about 5s of a cognify run for this dataset completing. A run already complete when you connected is the baseline and is not announced
  • heartbeat{"kind": "heartbeat", "time": str}, every 15s or so, so a quiet stream is distinguishable from a dead one
Pass the cursor from the last live_events frame back as the since query parameter when reconnecting; omit it to start from every available event. Close codes:
  • 1008: not authenticated, or no read permission on this dataset. A retry replays the same rejection, so clients should stop. Permission is re-checked on every poll, so access revoked mid-stream closes the connection this way too
  • 1011: the stream failed server-side; reconnecting is reasonable
A malformed since, or a dataset_id that is not a UUID, fails request validation before the connection is accepted. Depending on the ASGI server the client sees either a 1008 close or a plain HTTP rejection of the handshake, so treat a handshake that never opened as a client-side bug rather than a stream error.

Dataset data

GET /api/v1/datasets/{dataset_id}/data — List all data items in a dataset. Each item also carries the label and metadata attached at upload time:
  • label: The label given to this file via the labels field on add or remember, or null if none was set
  • externalMetadata: The stored metadata object — the external_metadata entry you sent merged over loader-derived keys, plus a node_set key when one was passed at ingest. A file uploaded without metadata stores an empty object rather than null.
GET /api/v1/datasets/{dataset_id}/data/{data_id}/raw — Download the original file for a specific data item. Requires read permission on the containing dataset; ownership of the data item itself is not required, so a dataset shared with you is fully downloadable.

Delete

DELETE /api/v1/datasets/{dataset_id} — Delete a dataset and all its contents. DELETE /api/v1/datasets/{dataset_id}/data/{data_id} — Delete a specific data item from a dataset.
Deleting a dataset removes all associated documents, knowledge graph data, and embeddings. This cannot be undone.