> ## Documentation Index
> Fetch the complete documentation index at: https://docs.cognee.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Get Dataset Progress

> Get the processing status of datasets, together with in-flight progress.

Same dataset/pipeline selection as **GET /v1/datasets/status**, but each
status value is always an object {status, progress} instead of a bare
status — a dedicated endpoint rather than a flag on /status, so neither
endpoint's response shape ever depends on how it was called.

## Query Parameters
- **dataset** (List[UUID]): Dataset UUIDs to check (from GET /api/v1/datasets). Omit to get
  status for all datasets you can read.
- **pipeline** (List[str]): Pipeline names to check: 'add_pipeline', 'cognify_pipeline', or
  'code_graph_pipeline' (code ingestion via remember content_type='code'). Omit to default
  to cognify_pipeline.

## Response
- Single pipeline (default): {dataset_id: {status, progress}}
- Multiple pipelines: {dataset_id: {pipeline_name: {status, progress}}}

**progress** is `null` until the first in-flight progress tick, then an
object with `completed_items`, `total_items`, and `current_stage` —
present only while the pipeline is running; terminal runs (completed/
errored) do not carry a progress snapshot.

## Error Codes
- **409 Conflict**: Error retrieving status (e.g. requesting a dataset you don't have
  read permission for)



## OpenAPI

````yaml /cognee_openapi_spec.json get /api/v1/datasets/status/progress
openapi: 3.1.0
info:
  title: Cognee API
  description: Cognee API with Bearer token and Cookie auth
  version: 1.0.0
servers:
  - url: https://api.cognee.ai
    description: Production server (full functionality)
  - url: http://localhost:8000
    description: Local development server (requires local setup)
security:
  - BearerAuth: []
  - ApiKeyAuth: []
tags:
  - name: activity
    description: >-
      Activity endpoints for inspecting pipeline runs, traced spans, tenant
      users, agents, and dataset exports.
  - name: add
    description: Data ingestion endpoints for adding text, files, and structured data.
  - name: agent connections
    description: >-
      Endpoints for registering, unregistering, and inspecting agent connections
      to the instance.
  - name: agent management
    description: Endpoints for creating, listing, retrieving, and deleting agents.
  - name: auth
    description: >-
      Authentication endpoints for user registration, login, and token
      management.
  - name: checks
    description: >-
      Diagnostic endpoint for validating a Cognee Cloud API key supplied in the
      X-Api-Key header.
  - name: cognify
    description: >-
      Knowledge processing endpoints to transform raw data into knowledge
      graphs.
  - name: configuration
    description: >-
      Endpoints for storing, retrieving, and listing a user's saved
      configurations.
  - name: datasets
    description: Dataset management endpoints for listing, creating, and deleting datasets.
  - name: delete
    description: Data deletion endpoints (deprecated — use datasets endpoints instead).
  - name: forget
    description: Endpoint for removing data from the knowledge graph.
  - name: health
    description: Liveness, readiness, and component health checks.
  - name: improve
    description: Endpoint for enriching and improving an existing knowledge graph.
  - name: integrations
    description: >-
      Endpoints for connecting, provisioning, and disconnecting OAuth providers
      and plugins.
  - name: llm
    description: >-
      LLM-backed endpoints for inferring graph schemas and generating custom
      extraction prompts.
  - name: memify
    description: >-
      Endpoint for running enrichment pipelines over existing graphs or supplied
      data.
  - name: ontologies
    description: >-
      Endpoints for uploading, listing, and deleting ontology files used during
      cognify.
  - name: permissions
    description: Permission management for multi-user access control.
  - name: recall
    description: >-
      Endpoints for querying the knowledge graph and reviewing past recall
      history.
  - name: remember
    description: >-
      Endpoints for ingesting data into the knowledge graph and storing session
      memory entries.
  - name: responses
    description: Response generation endpoints using the knowledge graph.
  - name: schema
    description: >-
      Schema inspection endpoints for a dataset's derived schema inventory and
      the caller-wide memory provenance graph.
  - name: search
    description: Search endpoints for querying the knowledge graph.
  - name: sessions
    description: >-
      Endpoints for listing sessions and reporting usage, cost, and token
      statistics.
  - name: settings
    description: Configuration endpoints for managing Cognee settings.
  - name: skills
    description: >-
      Skill management endpoints for ingesting, listing, retrieving, and
      deleting dataset skills, plus read-only retrieval of improvement
      proposals.
  - name: slack
    description: >-
      Endpoints for listing workspace channels, setting channel allowlists, and
      linking Slack accounts.
  - name: sync
    description: Endpoints for syncing local data to Cognee Cloud and checking sync status.
  - name: update
    description: Endpoint for updating existing data in a dataset.
  - name: users
    description: User management endpoints.
  - name: validate
    description: >-
      Diagnostic endpoint for checking consistency between a dataset's graph and
      vector stores.
  - name: visualize
    description: Graph visualization endpoints.
paths:
  /api/v1/datasets/status/progress:
    get:
      tags:
        - datasets
      summary: Get Dataset Progress
      description: >-
        Get the processing status of datasets, together with in-flight progress.


        Same dataset/pipeline selection as **GET /v1/datasets/status**, but each

        status value is always an object {status, progress} instead of a bare

        status — a dedicated endpoint rather than a flag on /status, so neither

        endpoint's response shape ever depends on how it was called.


        ## Query Parameters

        - **dataset** (List[UUID]): Dataset UUIDs to check (from GET
        /api/v1/datasets). Omit to get
          status for all datasets you can read.
        - **pipeline** (List[str]): Pipeline names to check: 'add_pipeline',
        'cognify_pipeline', or
          'code_graph_pipeline' (code ingestion via remember content_type='code'). Omit to default
          to cognify_pipeline.

        ## Response

        - Single pipeline (default): {dataset_id: {status, progress}}

        - Multiple pipelines: {dataset_id: {pipeline_name: {status, progress}}}


        **progress** is `null` until the first in-flight progress tick, then an

        object with `completed_items`, `total_items`, and `current_stage` —

        present only while the pipeline is running; terminal runs (completed/

        errored) do not carry a progress snapshot.


        ## Error Codes

        - **409 Conflict**: Error retrieving status (e.g. requesting a dataset
        you don't have
          read permission for)
      operationId: get_dataset_progress_api_v1_datasets_status_progress_get
      parameters:
        - name: dataset
          in: query
          required: false
          schema:
            type: array
            items:
              type: string
              format: uuid
            description: >-
              Dataset UUIDs to check (from GET /api/v1/datasets). Omit to get
              status for all datasets you can read.
            examples:
              - - b8a7c3de-4f5a-4b6c-8d9e-0f1a2b3c4d5e
            default: []
            title: Dataset
          description: >-
            Dataset UUIDs to check (from GET /api/v1/datasets). Omit to get
            status for all datasets you can read.
        - name: pipeline
          in: query
          required: false
          schema:
            type: array
            items:
              type: string
            description: >-
              Pipeline names to check: 'add_pipeline', 'cognify_pipeline', or
              'code_graph_pipeline' (code ingestion via remember
              content_type='code'). Omit to default to cognify_pipeline.
            examples:
              - - cognify_pipeline
            default: []
            title: Pipeline
          description: >-
            Pipeline names to check: 'add_pipeline', 'cognify_pipeline', or
            'code_graph_pipeline' (code ingestion via remember
            content_type='code'). Omit to default to cognify_pipeline.
      responses:
        '200':
          description: Successful Response
          content:
            application/json:
              schema:
                anyOf:
                  - type: object
                    additionalProperties:
                      $ref: '#/components/schemas/PipelineRunStatusWithProgress'
                  - type: object
                    additionalProperties:
                      type: object
                      additionalProperties:
                        $ref: '#/components/schemas/PipelineRunStatusWithProgress'
                title: >-
                  Response Get Dataset Progress Api V1 Datasets Status Progress
                  Get
        '422':
          description: Validation Error
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/HTTPValidationError'
      security:
        - BearerAuth: []
        - ApiKeyAuth: []
components:
  schemas:
    PipelineRunStatusWithProgress:
      properties:
        status:
          $ref: '#/components/schemas/PipelineRunStatus'
        progress:
          anyOf:
            - additionalProperties: true
              type: object
            - type: 'null'
          title: Progress
          examples:
            - completed_items: 3
              current_stage: extract_graph
              total_items: 10
      type: object
      required:
        - status
      title: PipelineRunStatusWithProgress
    HTTPValidationError:
      properties:
        detail:
          items:
            $ref: '#/components/schemas/ValidationError'
          type: array
          title: Detail
      type: object
      title: HTTPValidationError
    PipelineRunStatus:
      type: string
      enum:
        - DATASET_PROCESSING_INITIATED
        - DATASET_PROCESSING_STARTED
        - DATASET_PROCESSING_COMPLETED
        - DATASET_PROCESSING_ERRORED
      title: PipelineRunStatus
    ValidationError:
      properties:
        loc:
          items:
            anyOf:
              - type: string
              - type: integer
          type: array
          title: Location
        msg:
          type: string
          title: Message
        type:
          type: string
          title: Error Type
        input:
          title: Input
        ctx:
          type: object
          title: Context
      type: object
      required:
        - loc
        - msg
        - type
      title: ValidationError
  securitySchemes:
    BearerAuth:
      type: http
      scheme: bearer
      bearerFormat: JWT
    ApiKeyAuth:
      type: apiKey
      in: header
      name: X-Api-Key

````