> ## Documentation Index
> Fetch the complete documentation index at: https://docs.cognee.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Get Dataset Status

> Get the processing status of datasets.

This endpoint retrieves the current processing status of one or more datasets,
indicating whether they are being processed, have completed processing, or
encountered errors during pipeline execution.

## Query Parameters
- **dataset** (List[UUID]): List of dataset UUIDs to check status for.
  If omitted, returns status for all datasets the user has read permission on
- **pipeline** (List[str], optional): One or more pipeline names to check.
  - If omitted, defaults to **cognify_pipeline** (backward-compatible behavior)
  - If one pipeline is provided, response is a flat map
  - If multiple pipelines are provided, response is nested per dataset and pipeline
  - **Available options: add_pipeline, cognify_pipeline, code_graph_pipeline**
  - Note: a background code ingest creates its pipeline run only once the
    repository is cloned — a dataset missing from the response means the run
    has not started yet, not that it failed

## Response
Returns status information in one of two shapes:
- Single pipeline (default): \{dataset_id: status\}
- Multiple pipelines: \{dataset_id: \{pipeline_name: status\}\}

Status values:
- **pending**: Dataset is queued for processing
- **running**: Dataset is currently being processed
- **completed**: Dataset processing completed successfully
- **failed**: Dataset processing encountered an error

For in-flight progress (files completed / total, current stage), see
**GET /v1/datasets/status/progress** — a separate endpoint with its own
fixed response shape, rather than a flag here that would change what
this endpoint returns depending on how it's called.

## Error Codes
- **403 Forbidden**: The request owner cannot read every requested dataset
- **409 Conflict**: An unexpected error occurred while retrieving status



## OpenAPI

````yaml /cognee_openapi_spec.json get /api/v1/datasets/status
openapi: 3.1.0
info:
  title: Cognee API
  description: Cognee API with Bearer token and Cookie auth
  version: 1.0.0
servers:
  - url: https://{tenant}.aws.cognee.ai
    description: 'Cognee Cloud: your tenant pod, named in the platform.cognee.ai dashboard'
    variables:
      tenant:
        default: your-tenant
        description: Your tenant name, shown in the Cognee Cloud dashboard
  - url: http://localhost:8000
    description: 'Self-hosted: a locally running cognee server'
security:
  - BearerAuth: []
  - ApiKeyAuth: []
tags:
  - name: activity
    description: >-
      Activity endpoints for inspecting pipeline runs, traced spans, tenant
      users, agents, and dataset exports.
  - name: add
    description: Data ingestion endpoints for adding text, files, and structured data.
  - name: agent connections
    description: >-
      Endpoints for registering, unregistering, and inspecting agent connections
      to the instance.
  - name: agent management
    description: Endpoints for creating, listing, retrieving, and deleting agents.
  - name: auth
    description: >-
      Authentication endpoints for user registration, login, and token
      management.
  - name: checks
    description: >-
      Diagnostic endpoint for validating a Cognee Cloud API key supplied in the
      X-Api-Key header.
  - name: cognify
    description: >-
      Knowledge processing endpoints to transform raw data into knowledge
      graphs.
  - name: configuration
    description: >-
      Endpoints for storing, retrieving, and listing a user's saved
      configurations.
  - name: datasets
    description: Dataset management endpoints for listing, creating, and deleting datasets.
  - name: delete
    description: Data deletion endpoints (deprecated — use datasets endpoints instead).
  - name: forget
    description: Endpoint for removing data from the knowledge graph.
  - name: health
    description: Liveness, readiness, and component health checks.
  - name: improve
    description: Endpoint for enriching and improving an existing knowledge graph.
  - name: integrations
    description: >-
      Endpoints for connecting, provisioning, and disconnecting OAuth providers
      and plugins.
  - name: llm
    description: >-
      LLM-backed endpoints for inferring graph schemas and generating custom
      extraction prompts.
  - name: memify
    description: >-
      Endpoint for running enrichment pipelines over existing graphs or supplied
      data.
  - name: ontologies
    description: >-
      Endpoints for uploading, listing, and deleting ontology files used during
      cognify.
  - name: permissions
    description: Permission management for multi-user access control.
  - name: recall
    description: >-
      Endpoints for querying the knowledge graph and reviewing past recall
      history.
  - name: remember
    description: >-
      Endpoints for ingesting data into the knowledge graph and storing session
      memory entries.
  - name: responses
    description: Response generation endpoints using the knowledge graph.
  - name: schema
    description: >-
      Schema inspection endpoints for a dataset's derived schema inventory and
      the caller-wide memory provenance graph.
  - name: search
    description: Search endpoints for querying the knowledge graph.
  - name: sessions
    description: >-
      Endpoints for listing sessions and reporting usage, cost, and token
      statistics.
  - name: settings
    description: Configuration endpoints for managing Cognee settings.
  - name: skills
    description: >-
      Skill management endpoints for ingesting, listing, retrieving, and
      deleting dataset skills, plus read-only retrieval of improvement
      proposals.
  - name: slack
    description: >-
      Endpoints for listing workspace channels, setting channel allowlists, and
      linking Slack accounts.
  - name: sync
    description: Endpoints for syncing local data to Cognee Cloud and checking sync status.
  - name: update
    description: Endpoint for updating existing data in a dataset.
  - name: users
    description: User management endpoints.
  - name: validate
    description: >-
      Diagnostic endpoint for checking consistency between a dataset's graph and
      vector stores.
  - name: visualize
    description: Graph visualization endpoints.
paths:
  /api/v1/datasets/status:
    get:
      tags:
        - datasets
      summary: Get Dataset Status
      description: >-
        Get the processing status of datasets.


        This endpoint retrieves the current processing status of one or more
        datasets,

        indicating whether they are being processed, have completed processing,
        or

        encountered errors during pipeline execution.


        ## Query Parameters

        - **dataset** (List[UUID]): List of dataset UUIDs to check status for.
          If omitted, returns status for all datasets the user has read permission on
        - **pipeline** (List[str], optional): One or more pipeline names to
        check.
          - If omitted, defaults to **cognify_pipeline** (backward-compatible behavior)
          - If one pipeline is provided, response is a flat map
          - If multiple pipelines are provided, response is nested per dataset and pipeline
          - **Available options: add_pipeline, cognify_pipeline, code_graph_pipeline**
          - Note: a background code ingest creates its pipeline run only once the
            repository is cloned — a dataset missing from the response means the run
            has not started yet, not that it failed

        ## Response

        Returns status information in one of two shapes:

        - Single pipeline (default): \{dataset_id: status\}

        - Multiple pipelines: \{dataset_id: \{pipeline_name: status\}\}


        Status values:

        - **pending**: Dataset is queued for processing

        - **running**: Dataset is currently being processed

        - **completed**: Dataset processing completed successfully

        - **failed**: Dataset processing encountered an error


        For in-flight progress (files completed / total, current stage), see

        **GET /v1/datasets/status/progress** — a separate endpoint with its own

        fixed response shape, rather than a flag here that would change what

        this endpoint returns depending on how it's called.


        ## Error Codes

        - **403 Forbidden**: The request owner cannot read every requested
        dataset

        - **409 Conflict**: An unexpected error occurred while retrieving status
      operationId: get_dataset_status_api_v1_datasets_status_get
      parameters:
        - name: dataset
          in: query
          required: false
          schema:
            type: array
            items:
              type: string
              format: uuid
            description: >-
              Dataset UUIDs to check (from GET /api/v1/datasets). Omit to get
              status for all datasets you can read.
            examples:
              - - b8a7c3de-4f5a-4b6c-8d9e-0f1a2b3c4d5e
            default: []
            title: Dataset
          description: >-
            Dataset UUIDs to check (from GET /api/v1/datasets). Omit to get
            status for all datasets you can read.
        - name: pipeline
          in: query
          required: false
          schema:
            type: array
            items:
              type: string
            description: >-
              Pipeline names to check: 'add_pipeline', 'cognify_pipeline', or
              'code_graph_pipeline' (code ingestion via remember
              content_type='code'). Omit to default to cognify_pipeline.
            examples:
              - - cognify_pipeline
            default: []
            title: Pipeline
          description: >-
            Pipeline names to check: 'add_pipeline', 'cognify_pipeline', or
            'code_graph_pipeline' (code ingestion via remember
            content_type='code'). Omit to default to cognify_pipeline.
      responses:
        '200':
          description: Successful Response
          content:
            application/json:
              schema:
                anyOf:
                  - type: object
                    additionalProperties:
                      $ref: '#/components/schemas/PipelineRunStatus'
                  - type: object
                    additionalProperties:
                      type: object
                      additionalProperties:
                        $ref: '#/components/schemas/PipelineRunStatus'
                title: Response Get Dataset Status Api V1 Datasets Status Get
        '422':
          description: Validation Error
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/HTTPValidationError'
      security:
        - BearerAuth: []
        - ApiKeyAuth: []
components:
  schemas:
    PipelineRunStatus:
      type: string
      enum:
        - DATASET_PROCESSING_INITIATED
        - DATASET_PROCESSING_STARTED
        - DATASET_PROCESSING_COMPLETED
        - DATASET_PROCESSING_ERRORED
      title: PipelineRunStatus
    HTTPValidationError:
      properties:
        detail:
          items:
            $ref: '#/components/schemas/ValidationError'
          type: array
          title: Detail
      type: object
      title: HTTPValidationError
    ValidationError:
      properties:
        loc:
          items:
            anyOf:
              - type: string
              - type: integer
          type: array
          title: Location
        msg:
          type: string
          title: Message
        type:
          type: string
          title: Error Type
        input:
          title: Input
        ctx:
          type: object
          title: Context
      type: object
      required:
        - loc
        - msg
        - type
      title: ValidationError
  securitySchemes:
    BearerAuth:
      type: http
      scheme: bearer
      bearerFormat: JWT
    ApiKeyAuth:
      type: apiKey
      in: header
      name: X-Api-Key

````