> ## Documentation Index
> Fetch the complete documentation index at: https://docs.cognee.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Get Dataset Processing Status

> Get item-level processing status for a dataset.

`GET /status` reports whether a pipeline *run* is in progress or done for a
dataset. This endpoint answers the finer question operators need when
triaging incremental loads: which of the dataset's data items carry the
per-item completion stamp for a pipeline, and which are still pending.

## Path Parameters
- **dataset_id** (UUID): The unique identifier of the dataset

## Query Parameters
- **pipeline** (str, optional): Pipeline name to inspect. Defaults to
  `cognify_pipeline`.

## Response
- **total**: Number of data items in the dataset
- **completed**: Items whose per-item status for the pipeline is completed
  (both the legacy string and the dict status representation are recognised)
- **pending**: `total - completed`
- **items**: `[{id, name, completed}]`, one entry per data item, in the same
  order as `GET /datasets/{id}/data`. `id` is the data_id accepted by
  `DELETE /datasets/{id}/data/{data_id}` and `forget(data_id=...)`

Per-item errored state is not persisted, so it is not reported: a pending
item may be untouched, in progress, or failed.

## Error Codes
- **404 Not Found**: Dataset doesn't exist or user doesn't have access
- **409 Conflict**: Error computing the status



## OpenAPI

````yaml /cognee_openapi_spec.json get /api/v1/datasets/{dataset_id}/processing-status
openapi: 3.1.0
info:
  title: Cognee API
  description: Cognee API with Bearer token and Cookie auth
  version: 1.0.0
servers:
  - url: https://{tenant}.aws.cognee.ai
    description: 'Cognee Cloud: your tenant pod, named in the platform.cognee.ai dashboard'
    variables:
      tenant:
        default: your-tenant
        description: Your tenant name, shown in the Cognee Cloud dashboard
  - url: http://localhost:8000
    description: 'Self-hosted: a locally running cognee server'
security:
  - BearerAuth: []
  - ApiKeyAuth: []
tags:
  - name: activity
    description: >-
      Activity endpoints for inspecting pipeline runs, traced spans, tenant
      users, agents, and dataset exports.
  - name: add
    description: Data ingestion endpoints for adding text, files, and structured data.
  - name: agent connections
    description: >-
      Endpoints for registering, unregistering, and inspecting agent connections
      to the instance.
  - name: agent management
    description: Endpoints for creating, listing, retrieving, and deleting agents.
  - name: auth
    description: >-
      Authentication endpoints for user registration, login, and token
      management.
  - name: checks
    description: >-
      Diagnostic endpoint for validating a Cognee Cloud API key supplied in the
      X-Api-Key header.
  - name: cognify
    description: >-
      Knowledge processing endpoints to transform raw data into knowledge
      graphs.
  - name: configuration
    description: >-
      Endpoints for storing, retrieving, and listing a user's saved
      configurations.
  - name: datasets
    description: Dataset management endpoints for listing, creating, and deleting datasets.
  - name: delete
    description: Data deletion endpoints (deprecated — use datasets endpoints instead).
  - name: forget
    description: Endpoint for removing data from the knowledge graph.
  - name: health
    description: Liveness, readiness, and component health checks.
  - name: improve
    description: Endpoint for enriching and improving an existing knowledge graph.
  - name: integrations
    description: >-
      Endpoints for connecting, provisioning, and disconnecting OAuth providers
      and plugins.
  - name: llm
    description: >-
      LLM-backed endpoints for inferring graph schemas and generating custom
      extraction prompts.
  - name: memify
    description: >-
      Endpoint for running enrichment pipelines over existing graphs or supplied
      data.
  - name: ontologies
    description: >-
      Endpoints for uploading, listing, and deleting ontology files used during
      cognify.
  - name: permissions
    description: Permission management for multi-user access control.
  - name: recall
    description: >-
      Endpoints for querying the knowledge graph and reviewing past recall
      history.
  - name: remember
    description: >-
      Endpoints for ingesting data into the knowledge graph and storing session
      memory entries.
  - name: responses
    description: Response generation endpoints using the knowledge graph.
  - name: schema
    description: >-
      Schema inspection endpoints for a dataset's derived schema inventory and
      the caller-wide memory provenance graph.
  - name: search
    description: Search endpoints for querying the knowledge graph.
  - name: sessions
    description: >-
      Endpoints for listing sessions and reporting usage, cost, and token
      statistics.
  - name: settings
    description: Configuration endpoints for managing Cognee settings.
  - name: skills
    description: >-
      Skill management endpoints for ingesting, listing, retrieving, and
      deleting dataset skills, plus read-only retrieval of improvement
      proposals.
  - name: slack
    description: >-
      Endpoints for listing workspace channels, setting channel allowlists, and
      linking Slack accounts.
  - name: sync
    description: Endpoints for syncing local data to Cognee Cloud and checking sync status.
  - name: update
    description: Endpoint for updating existing data in a dataset.
  - name: users
    description: User management endpoints.
  - name: validate
    description: >-
      Diagnostic endpoint for checking consistency between a dataset's graph and
      vector stores.
  - name: visualize
    description: Graph visualization endpoints.
paths:
  /api/v1/datasets/{dataset_id}/processing-status:
    get:
      tags:
        - datasets
      summary: Get Dataset Processing Status
      description: >-
        Get item-level processing status for a dataset.


        `GET /status` reports whether a pipeline *run* is in progress or done
        for a

        dataset. This endpoint answers the finer question operators need when

        triaging incremental loads: which of the dataset's data items carry the

        per-item completion stamp for a pipeline, and which are still pending.


        ## Path Parameters

        - **dataset_id** (UUID): The unique identifier of the dataset


        ## Query Parameters

        - **pipeline** (str, optional): Pipeline name to inspect. Defaults to
          `cognify_pipeline`.

        ## Response

        - **total**: Number of data items in the dataset

        - **completed**: Items whose per-item status for the pipeline is
        completed
          (both the legacy string and the dict status representation are recognised)
        - **pending**: `total - completed`

        - **items**: `[{id, name, completed}]`, one entry per data item, in the
        same
          order as `GET /datasets/{id}/data`. `id` is the data_id accepted by
          `DELETE /datasets/{id}/data/{data_id}` and `forget(data_id=...)`

        Per-item errored state is not persisted, so it is not reported: a
        pending

        item may be untouched, in progress, or failed.


        ## Error Codes

        - **404 Not Found**: Dataset doesn't exist or user doesn't have access

        - **409 Conflict**: Error computing the status
      operationId: >-
        get_dataset_processing_status_api_v1_datasets__dataset_id__processing_status_get
      parameters:
        - name: dataset_id
          in: path
          required: true
          schema:
            type: string
            format: uuid
            description: >-
              Dataset UUID, the id field from GET /api/v1/datasets (not the
              name)
            examples:
              - b8a7c3de-4f5a-4b6c-8d9e-0f1a2b3c4d5e
            title: Dataset Id
          description: Dataset UUID, the id field from GET /api/v1/datasets (not the name)
        - name: pipeline
          in: query
          required: false
          schema:
            type: string
            description: >-
              Pipeline whose per-item completion to count: 'cognify_pipeline'
              (default), 'add_pipeline', or 'code_graph_pipeline'.
            examples:
              - cognify_pipeline
            default: cognify_pipeline
            title: Pipeline
          description: >-
            Pipeline whose per-item completion to count: 'cognify_pipeline'
            (default), 'add_pipeline', or 'code_graph_pipeline'.
      responses:
        '200':
          description: Successful Response
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/DatasetProcessingStatusDTO'
        '404':
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/ErrorResponseDTO'
          description: Not Found
        '422':
          description: Validation Error
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/HTTPValidationError'
      security:
        - BearerAuth: []
        - ApiKeyAuth: []
components:
  schemas:
    DatasetProcessingStatusDTO:
      properties:
        total:
          type: integer
          title: Total
          description: Number of data items in the dataset
        completed:
          type: integer
          title: Completed
          description: Items carrying the per-item completion stamp
        pending:
          type: integer
          title: Pending
          description: Items without the stamp (total - completed)
        items:
          items:
            $ref: '#/components/schemas/DataItemProcessingStatusDTO'
          type: array
          title: Items
          description: >-
            One entry per data item, in the same order as GET
            /datasets/\{id\}/data
      type: object
      required:
        - total
        - completed
        - pending
        - items
      title: DatasetProcessingStatusDTO
      description: Item-level completion counts for one dataset and one pipeline.
    ErrorResponseDTO:
      properties:
        message:
          type: string
          title: Message
      type: object
      required:
        - message
      title: ErrorResponseDTO
    HTTPValidationError:
      properties:
        detail:
          items:
            $ref: '#/components/schemas/ValidationError'
          type: array
          title: Detail
      type: object
      title: HTTPValidationError
    DataItemProcessingStatusDTO:
      properties:
        id:
          type: string
          format: uuid
          title: Id
        name:
          type: string
          title: Name
        completed:
          type: boolean
          title: Completed
      type: object
      required:
        - id
        - name
        - completed
      title: DataItemProcessingStatusDTO
      description: One data item's completion state for the requested pipeline.
    ValidationError:
      properties:
        loc:
          items:
            anyOf:
              - type: string
              - type: integer
          type: array
          title: Location
        msg:
          type: string
          title: Message
        type:
          type: string
          title: Error Type
        input:
          title: Input
        ctx:
          type: object
          title: Context
      type: object
      required:
        - loc
        - msg
        - type
      title: ValidationError
  securitySchemes:
    BearerAuth:
      type: http
      scheme: bearer
      bearerFormat: JWT
    ApiKeyAuth:
      type: apiKey
      in: header
      name: X-Api-Key

````