cognee.add()
Description
Add data to Cognee for knowledge graph processing. This is the first step in the Cognee workflow - it ingests raw data and prepares it for processing. The function accepts various data formats including text, files, urls and binary streams, then stores them in a specified dataset for further processing. Prerequisites:- LLM_API_KEY: Must be set in environment variables for content processing
- Database Setup: Relational and vector databases must be configured
- User Authentication: Uses default user if none provided (created automatically)
- Text strings: Direct text content (str) - any string not starting with ”/” or “file://”
- File paths: Local file paths as strings in these formats:
- Absolute paths: “/path/to/document.pdf”
- File URLs: “file:///path/to/document.pdf” or “file://relative/path.txt”
- S3 paths: “s3://bucket-name/path/to/file.pdf”
- Binary file objects: File handles/streams (BinaryIO)
- Lists: Multiple files or text strings in a single call
- Text files (.txt, .md, .csv)
- PDFs (.pdf)
- Images (.png, .jpg, .jpeg) - extracted via OCR/vision models
- Audio files (.mp3, .wav) - transcribed to text
- Code files (.py, .js, .ts, etc.) - parsed for structure and content
- Office documents (.docx, .pptx)
- Data Resolution: Resolves file paths and validates accessibility
- Content Extraction: Extracts text content from various file formats
- Dataset Storage: Stores processed content in the specified dataset
- Metadata Tracking: Records file metadata, timestamps, and user permissions
- Permission Assignment: Grants user read/write/delete/share permissions on dataset
- Single text string: “Your text content here”
- Absolute file path: “/path/to/document.pdf”
- File URL: “file:///absolute/path/to/document.pdf” or “file://relative/path.txt”
- S3 path: “s3://my-bucket/documents/file.pdf”
- List of mixed types: [“text content”, “/path/file.pdf”, “file://doc.txt”, file_handle]
- Binary file object: open(“file.txt”, “rb”)
- url: A web link url (https or http) dataset_name: Name of the dataset to store data in. Defaults to “main_dataset”. Create separate datasets to organize different knowledge domains. user: User object for authentication and permissions. Uses default user if None. Default user: “default_user@example.com” (created automatically on first use). Users can only access datasets they have permissions for. node_set: Optional list of node identifiers for graph organization and access control. Used for grouping related data points in the knowledge graph. vector_db_config: Optional configuration for vector database (for custom setups). graph_db_config: Optional configuration for graph database (for custom setups). dataset_id: Optional specific dataset UUID to use instead of dataset_name. extraction_rules: Optional dictionary of rules (e.g., CSS selectors, XPath) for extracting specific content from web pages using BeautifulSoup tavily_config: Optional configuration for Tavily API, including API key and extraction settings soup_crawler_config: Optional configuration for BeautifulSoup crawler, specifying concurrency, crawl delay, and extraction rules.
- Pipeline run ID for tracking
- Dataset ID where data was stored
- Processing status and any errors
- Execution timestamps and metadata
cognify() to process the ingested content:
- LLM_API_KEY: API key for your LLM provider (OpenAI, Anthropic, etc.)
- LLM_PROVIDER: “openai” (default), “anthropic”, “gemini”, “ollama”, “mistral”, “bedrock”
- LLM_MODEL: Model name (default: “gpt-5-mini”)
- DEFAULT_USER_EMAIL: Custom default user email
- DEFAULT_USER_PASSWORD: Custom default user password
- VECTOR_DB_PROVIDER: “lancedb” (default), “chromadb”, “pgvector”
- GRAPH_DATABASE_PROVIDER: “kuzu” (default), “neo4j”
- TAVILY_API_KEY: YOUR_TAVILY_API_KEY
Parameters
Union[BinaryIO, list[BinaryIO], str, list[str], DataItem, list[DataItem]]
required
Data to ingest. Accepts text strings, file paths (local, S3, or URLs), binary file objects,
DataItem objects, or lists of any of these.DataItem is a lightweight wrapper that lets you attach per-item metadata, human-readable label, and an optional stable data_id. Import it from cognee.tasks.ingestion.data_item:label and external_metadata are stored on the relational Data record. They are not propagated into the knowledge graph automatically and are not searchable via cognee.search(). Use node_set when you need tags that flow into the graph and can be used for scoped queries.str
default:"'main_dataset'"
Name of the dataset to add data to.
User
default:"None"
User performing the operation. Uses default user if not provided.
Optional[List[str]]
default:"None"
List of node set names to associate with the data.
dict
default:"None"
Override vector database configuration for this operation.
dict
default:"None"
Override graph database configuration for this operation.
Optional[UUID]
default:"None"
UUID of an existing dataset to add data to. Alternative to dataset_name.
Optional[List[Union[str, dict[str, dict[str, Any]]]]]
default:"None"
Custom loader configuration for specific file types.
bool
default:"True"
If true, skip data that has already been ingested.
Optional[int]
default:"20"
Number of data items to process per batch.
Optional[LLMConfig]
default:"None"
LLM settings to install into the current async context for this ingestion operation. When omitted, Cognee uses the active context config or global LLM config. Import
LLMConfig from cognee.infrastructure.llm.config.Optional[EmbeddingConfig]
default:"None"
Embedding settings to install into the current async context for this ingestion operation. When omitted, Cognee uses the active context config or global embedding config. Import
EmbeddingConfig from cognee.infrastructure.databases.vector.embeddings.config.Optional[float]
default:"0.5"
Floating-point score stored on the
Data record for retrieval ranking. Applied uniformly to all items in the batch. Use a higher value to make items more likely to surface in ranked results.str
default:"None"
Column name for primary key when ingesting structured data via dlt. Auto-detected if not specified.
str
default:"'merge'"
How to handle existing data for dlt sources: “merge” (upsert), “append” (always insert), or “replace” (drop and recreate).
str
default:"None"
SQL WHERE clause for filtering when using a database connection string as input.
Supported Input Types
Supported File Formats
See Loaders for how to override the default loader selection or register custom loaders.
Examples
For a complete guide on structured data ingestion with dlt, see the dlt integration page.