Skip to main content

cognee.add()

Description

Add data to Cognee for knowledge graph processing. This is the first step in the Cognee workflow - it ingests raw data and prepares it for processing. The function accepts various data formats including text, files, urls and binary streams, then stores them in a specified dataset for further processing. Prerequisites:
  • LLM_API_KEY: Must be set in environment variables for content processing
  • Database Setup: Relational and vector databases must be configured
  • User Authentication: Uses default user if none provided (created automatically)
Supported Input Types:
  • Text strings: Direct text content (str) - any string that is not a file://, s3://, http:// or https:// URL and does not point to an existing local file. An absolute-looking string that is not an existing file (for example "/remember to call the dentist") is ingested as text content on every platform, including Windows.
  • File paths: Local file paths as strings in these formats:
    • Absolute paths: “/path/to/document.pdf” (treated as a file reference only when the file exists)
    • File URLs: “file:///path/to/document.pdf” or “file://relative/path.txt”
    • S3 paths: “s3://bucket-name/path/to/file.pdf”
  • Binary file objects: File handles/streams (BinaryIO)
  • Lists: Multiple files or text strings in a single call
Supported File Formats:
  • Text files (.txt, .md, .csv)
  • PDFs (.pdf)
  • Images (.png, .jpg, .jpeg) - transcribed via vision models, with an optional local OCR pass
  • Audio files (.mp3, .wav) - transcribed to text
  • Code files (.py, .js, .ts, etc.) - parsed for structure and content
  • Office documents (.docx, .pptx)
See the Supported File Formats table below for the full list grouped by loader, including which formats require optional extras. Workflow:
  1. Data Resolution: Resolves file paths and validates accessibility
  2. Content Extraction: Extracts text content from various file formats
  3. Dataset Storage: Stores processed content in the specified dataset
  4. Metadata Tracking: Records file metadata, timestamps, and user permissions
  5. Permission Assignment: Grants user read/write/delete/share permissions on dataset
Args: data: The data to ingest. Can be:
  • Single text string: “Your text content here”
  • Absolute file path: “/path/to/document.pdf” (must exist; otherwise the string is ingested as text)
  • File URL: “file:///absolute/path/to/document.pdf” or “file://relative/path.txt”
  • S3 path: “s3://my-bucket/documents/file.pdf”
  • List of mixed types: [“text content”, “/path/file.pdf”, “file://doc.txt”, file_handle]
  • Binary file object: open(“file.txt”, “rb”)
  • url: A web link url (https or http) dataset_name: Name of the dataset to store data in. Defaults to “main_dataset”. Create separate datasets to organize different knowledge domains. user: User object for authentication and permissions. Uses default user if None. Default user: “default_user@example.com” (created automatically on first use). Users can only access datasets they have permissions for. node_set: Optional list of node identifiers for graph organization and access control. Used for grouping related data points in the knowledge graph. vector_db_config: Optional configuration for vector database (for custom setups). graph_db_config: Optional configuration for graph database (for custom setups). dataset_id: Optional specific dataset UUID to use instead of dataset_name. extraction_rules: Optional dictionary of rules (e.g., CSS selectors, XPath) for extracting specific content from web pages using BeautifulSoup tavily_config: Optional configuration for Tavily API, including API key and extraction settings soup_crawler_config: Optional configuration for BeautifulSoup crawler, specifying concurrency, crawl delay, and extraction rules.
Returns: PipelineRunInfo: Information about the ingestion pipeline execution including:
  • Pipeline run ID for tracking
  • Dataset ID where data was stored
  • Processing status and any errors
  • Execution timestamps and metadata
Next Steps: After successfully adding data, call cognify() to process the ingested content:
Example Usage:
Environment Variables: Required:
  • LLM_API_KEY: API key for your LLM provider (OpenAI, Anthropic, etc.)
Optional:
  • LLM_PROVIDER: “openai” (default), “anthropic”, “gemini”, “ollama”, “mistral”, “bedrock”
  • LLM_MODEL: Model name (default: “gpt-5-mini”)
  • DEFAULT_USER_EMAIL: Custom default user email
  • DEFAULT_USER_PASSWORD: Custom default user password
  • VECTOR_DB_PROVIDER: “lancedb” (default), “chromadb”, “pgvector”
  • GRAPH_DATABASE_PROVIDER: “kuzu” (default), “neo4j”
  • TAVILY_API_KEY: YOUR_TAVILY_API_KEY
  • KEENABLE_API_KEY: YOUR_KEENABLE_API_KEY

Parameters

Union[BinaryIO, list[BinaryIO], str, list[str], DataItem, list[DataItem]]
required
Data to ingest. Accepts text strings, file paths (local, S3, or URLs), binary file objects, DataItem objects, or lists of any of these.DataItem is a lightweight wrapper that lets you attach per-item metadata, human-readable label, and an optional stable data_id. Import it from cognee.tasks.ingestion.data_item:
label and external_metadata are stored on the relational Data record. They are not propagated into the knowledge graph automatically and are not searchable via cognee.search(). Use node_set when you need tags that flow into the graph and can be used for scoped queries.
str
default:"'main_dataset'"
Name of the dataset to add data to.
User
default:"None"
User performing the operation. Uses default user if not provided.
Optional[List[str]]
default:"None"
List of node set names to associate with the data.
dict
default:"None"
Override vector database configuration for this operation.
dict
default:"None"
Override graph database configuration for this operation.
Optional[UUID]
default:"None"
UUID of an existing dataset to add data to. Alternative to dataset_name.
Optional[List[Union[str, dict[str, dict[str, Any]]]]]
default:"None"
Custom loader configuration for specific file types.
bool
default:"True"
If true, skip data that has already been ingested. The skip runs whenever this or data_cache is true, so re-ingesting an already-stored item requires both to be False. See Incremental loading and deduplication.
bool
default:"True"
Companion flag to incremental_loading — either one being true enables the already-processed skip for a data item.
Optional[int]
default:"20"
Number of data items to process per batch.
Optional[LLMConfig]
default:"None"
LLM settings to install into the current async context for this ingestion operation. When omitted, Cognee uses the active context config or global LLM config. Import LLMConfig from cognee.infrastructure.llm.config.
Optional[EmbeddingConfig]
default:"None"
Embedding settings to install into the current async context for this ingestion operation. When omitted, Cognee uses the active context config or global embedding config. Import EmbeddingConfig from cognee.infrastructure.databases.vector.embeddings.config.
Optional[float]
default:"0.5"
Floating-point score stored on the Data record for retrieval ranking. Applied uniformly to all items in the batch. Use a higher value to make items more likely to surface in ranked results.
str
default:"None"
Column name for primary key when ingesting dlt resources or database connection strings. Auto-detected if not specified. For CSV files, pass it through preferred_loaders=[{"dlt_csv_loader": {"primary_key": "..."}}] instead — this kwarg does not reach the CSV loader.
str
default:"'replace'"
How to handle existing data for dlt sources: “replace” (drop and recreate), “merge” (upsert on primary_key), or “append” (always insert). Any other value raises InvalidDLTArgumentError. This controls dlt’s staging snapshot only — a growing source also needs the ingestion skip turned off, see Re-ingesting a source that keeps growing.
str
default:"None"
SQL query that filters what is ingested when using a database connection string as input. Accepts SELECT ... FROM <table> [WHERE ...], where the FROM target may be schema-qualified and may carry a table alias. A WHERE clause that references the alias, and a FROM clause containing JOIN, raise a ValueError — see Database Connection String for the accepted shapes, the rewrites, and the two detection caveats.

Supported Input Types

Supported File Formats

See Loaders for how to override the default loader selection or register custom loaders.

Examples

For a complete guide on structured data ingestion with dlt, see the dlt integration page.