Skip to main content
Embedding providers convert text into vector representations that enable semantic search. These vectors capture the meaning of text, allowing Cognee to find conceptually related content even when the wording is different.
New to configuration?See the Setup Configuration Overview for the complete workflow:install extras → create .env → choose providers → handle pruning.

Supported Providers

Cognee supports multiple embedding providers:
  • OpenAI — Text embedding models via OpenAI API (default)
  • Azure OpenAI — Text embedding models via Azure OpenAI Service
  • Google Gemini — Embedding models via Google AI
  • Mistral — Embedding models via Mistral AI
  • AWS Bedrock — Embedding models via AWS Bedrock
  • Ollama — Local embedding models via Ollama
  • LM Studio — Local embedding models via LM Studio
  • Fastembed — CPU-friendly local embeddings
  • HuggingFace — Embedding models via HuggingFace Inference API or Inference Endpoints
  • vLLM — Self-hosted embedding models via vLLM
  • OpenAI-Compatible — Direct OpenAI SDK for llama.cpp, vLLM, TEI, and any /v1/embeddings server (bypasses LiteLLM)
  • Custom — OpenAI-compatible embedding endpoints routed through LiteLLM (DeepInfra, company-internal)
LLM/Embedding Configuration: If you configure only LLM or only embeddings, the other defaults to OpenAI. Ensure you have a working OpenAI API key, or configure both LLM and embeddings to avoid unexpected defaults.

Configuration

Set these environment variables in your .env file:
  • EMBEDDING_PROVIDER — The provider to use: openai, gemini, mistral, bedrock, ollama, fastembed, openai_compatible, custom. These are not the same values LLM_PROVIDER accepts — see Valid EMBEDDING_PROVIDER values and endpoint URL forms
  • EMBEDDING_MODEL — The specific embedding model to use
  • EMBEDDING_DIMENSIONS — The vector dimension size (must match your vector store)
  • EMBEDDING_API_KEY — Your API key (falls back to LLM_API_KEY if not set — with one exception when LLM_PROVIDER="custom")
  • EMBEDDING_ENDPOINT — Custom endpoint URL (for Azure, Ollama, or custom providers)
  • EMBEDDING_API_VERSION — API version (for Azure OpenAI)
  • EMBEDDING_MAX_COMPLETION_TOKENS — Maximum input tokens per embedded text; used for tokenizer-based chunk sizing (optional, default 8191). Set it to your embedding model’s real input limit — see Max Completion Tokens
  • HUGGINGFACE_TOKENIZER — HuggingFace Hub model ID that overrides the tokenizer Cognee uses for token counting when the embedding model is not itself a HuggingFace repo. Commonly used with Ollama embeddings (for example, nomic-ai/nomic-embed-text-v1.5).

Provider Setup Guides

OpenAI provides high-quality embeddings with good performance.
Use Azure OpenAI Service for embeddings with your own deployment.
If startup fails with KeyError: 'Could not automatically map text-embedding-3-large to a tokeniser.', the installed tiktoken is too old to recognise the model. Cognee strips the azure/ prefix and asks TikToken for the encoding of text-embedding-3-large, which requires tiktoken>=0.5.2 (Cognee pins >=0.8.0). Upgrade in your environment:
Use Google’s embedding models for semantic search.
Use Mistral’s embedding models for high-quality vector representations.
Installation: Install the required dependency:
Use embedding models provided by the AWS Bedrock service.
Run embedding models locally with Ollama for privacy and cost control.
HUGGINGFACE_TOKENIZER is the HuggingFace repo ID of the tokenizer used for token-length counting when sending requests to the Ollama embedding endpoint.Installation: Install Ollama from ollama.ai and pull your desired embedding model:
HUGGINGFACE_TOKENIZER is optional. It is no longer part of Cognee’s required startup validation, so setting EMBEDDING_PROVIDER, EMBEDDING_MODEL, and EMBEDDING_DIMENSIONS without it no longer raises a ValidationError on import. It is still recommended for Ollama: set it to the HuggingFace repo ID for the tokenizer that matches your embedding model so token-length counts stay accurate. See the HUGGINGFACE_TOKENIZER environment variable section below for how to find the correct value for your model.If HUGGINGFACE_TOKENIZER is unset or points at a repo that cannot be loaded, tokenizer resolution no longer raises — Cognee logs an advisory warning and falls back to the TikToken tokenizer, so ingestion continues with approximate token counts. Cognee also logs an advisory warning whenever the HUGGINGFACE_TOKENIZER value differs from the EMBEDDING_MODEL id, so a genuine mismatch is not silent. Because an Ollama tag such as nomic-embed-text:latest is never identical to its HuggingFace repo id (nomic-ai/nomic-embed-text-v1.5), you will see this advisory even for a correct setup — it is a reminder to confirm the two share a tokenizer, not an error.If a text input exceeds the model’s context window, the Ollama embedding engine automatically falls back by splitting the batch in half and retrying both halves. For a single overlong text, it splits the string into two overlapping segments and averages the resulting embeddings. Cognee no longer pre-truncates text before sending it to Ollama, so this fallback only activates when the server returns a context-length error.
Zero-API-key setup: To run fully offline with no OpenAI key, you must configure both the LLM provider and the embedding provider to use local backends. See the Local Setup guide for a complete combined .env example.
Ollama falling behind? Ollama processes requests sequentially. If it becomes unresponsive or returns errors under load, reduce EMBEDDING_BATCH_SIZE (default 36) to send fewer chunks per call — values between 1 and 10 work well for most local hardware:
Run embedding models locally with LM Studio for privacy and cost control.
Installation: Install LM Studio from lmstudio.ai and download your desired model from LM Studio’s interface. Load your model, start the LM Studio server, and Cognee will be able to connect to it.
Use Fastembed for CPU-friendly local embeddings without GPU requirements. Fastembed runs in-process via ONNX Runtime — no separate server, no API key, and no GPU needed.
Installation: Fastembed ships as an optional extra — it is not included in the base cognee install. Add it with:
This pulls in fastembed and a compatible onnxruntime build. The first call downloads and caches the model weights from Hugging Face, so the initial run needs network access and a few hundred MB of disk for the model cache.Supported models: Any model listed by fastembed’s TextEmbedding.list_supported_models(). Common choices:If EMBEDDING_DIMENSIONS is omitted, Cognee tries to auto-derive it from the fastembed model registry; set it explicitly to avoid a fallback to 3072 if lookup fails.
Context window handling: When a text input exceeds the model’s context window, Fastembed automatically splits the batch and retries. For a single overlong text, it splits the string into two overlapping segments and averages the resulting embeddings. If a single string is already too short to split further yet still exceeds the context window, Fastembed raises the terminal EmbeddingContextWindowTooSmallError instead of retrying (see Timeout and Retry Behavior). This mirrors the behavior of the OpenAI-compatible engine.
Token counting: Fastembed now counts tokens with the embedding model’s own HuggingFace tokenizer (BGE, MiniLM, and E5 are wordpiece models), resolved from the fastembed model list. Earlier releases counted every Fastembed model with the OpenAI gpt-4o BPE tokenizer, which mis-sized chunks and skewed the --dry-run token estimate. You do not need to set HUGGINGFACE_TOKENIZER for Fastembed. If a model is not in the known list, Cognee logs an advisory warning and falls back to the TikToken tokenizer.
Use embedding models from HuggingFace via the HuggingFace Inference API (serverless) or dedicated Inference Endpoints.
Installation: Install the HuggingFace extra for tokenizer support:
HUGGINGFACE_TOKENIZER with HuggingFace embeddings: When using EMBEDDING_PROVIDER="custom" with a huggingface/ model, Cognee automatically attempts to load a HuggingFace tokenizer from the model repo for token counting. If that fails, it falls back to the TikToken tokenizer. You do not need to set HUGGINGFACE_TOKENIZER manually for this provider — it is only required when using EMBEDDING_PROVIDER="ollama" (see the Ollama section above).
Use vLLM to serve local or self-hosted embedding models with an OpenAI-compatible API.Example with Qwen3-Embedding-4B on port 8001:
hosted_vllm/ prefix required: Include hosted_vllm/ at the start of the model name so LiteLLM routes requests to your vLLM server. The model name after the prefix should match the model ID returned by your vLLM server’s /v1/models endpoint.
Tokenization: Cognee automatically strips the hosted_vllm/ prefix when loading the HuggingFace tokenizer, so no separate HUGGINGFACE_TOKENIZER setting is needed as long as the model name after the prefix is a valid HuggingFace model ID.To verify the model name your vLLM server exposes, run:
See the LiteLLM vLLM documentation for more details.
Use EMBEDDING_PROVIDER="openai_compatible" for any local inference server that exposes the standard /v1/embeddings endpoint. This provider talks directly to OpenAI-compatible embedding servers via the OpenAI Python SDK, bypassing LiteLLM.Use this provider for: llama.cpp (llama-server --embedding), vLLM, Hugging Face TEI, LocalAI, Infinity, and similar servers.
Start llama.cpp with embedding support:
Endpoint normalisation: The engine automatically appends /v1 to EMBEDDING_ENDPOINT if it is missing, and strips a trailing /embeddings suffix. You can pass either http://localhost:8080 or http://localhost:8080/v1 — both work.
Tokenizer and token limits: The openai_compatible engine automatically loads a tokenizer for chunk sizing. It first tries to load a HuggingFace tokenizer matching EMBEDDING_MODEL; if that fails (for example, because the model name is a local alias not on the HuggingFace Hub), it falls back to the TikToken tokenizer. The token limit passed to the tokenizer is controlled by EMBEDDING_MAX_COMPLETION_TOKENS (default 8191). Set this to match your server’s input length limit if it differs from the default:
This provider also works with hosted third-party /v1/embeddings endpoints, which usually enforce a smaller input length than the model’s advertised context. See Max Completion Tokens for how to pick the value and how it interacts with EMBEDDING_BATCH_SIZE.
HUGGINGFACE_TOKENIZER is not needed for this provider. The engine automatically tries to load a HuggingFace tokenizer using the model name (e.g. BAAI/bge-large-en-v1.5) for token counting. If that fails — for example when using EMBEDDING_MODEL="default" with llama.cpp — Cognee logs an advisory warning and falls back to TikToken, so token counts become approximate but ingestion continues. You do not need to set HUGGINGFACE_TOKENIZER when using EMBEDDING_PROVIDER="openai_compatible".
Use OpenAI-compatible embedding endpoints from other providers such as DeepInfra, OpenRouter, or a company-internal server. These are routed through LiteLLM and require a provider prefix in the model name.Required variables: EMBEDDING_PROVIDER="custom", EMBEDDING_MODEL (with a LiteLLM provider prefix such as openrouter/, deepinfra/, or openai/), EMBEDDING_API_KEY, and EMBEDDING_DIMENSIONS (set it explicitly — custom models are not in the auto-derive registry, so it otherwise falls back to 3072 and causes a vector-store shape mismatch). EMBEDDING_ENDPOINT is optional: omit it for a named LiteLLM prefix like openrouter/ (LiteLLM supplies the base URL), and set it only when pointing at a specific api_base such as DeepInfra or a self-hosted server.
No endpoint normalisation for custom: Unlike openai_compatible, the custom provider passes EMBEDDING_ENDPOINT directly to LiteLLM as api_base with no automatic /v1 appending or /embeddings stripping — and LiteLLM appends /embeddings itself, so http://localhost:1234/v1/embeddings is requested as .../v1/embeddings/embeddings and returns 404. Set the endpoint to exactly the base URL your provider expects (e.g., https://api.deepinfra.com/v1/openai), or omit it entirely when using a named LiteLLM prefix such as openrouter/.

Additional Information

EMBEDDING_PROVIDER and LLM_PROVIDER are configured independently and do not accept the same values. LLM_PROVIDER is checked against a fixed set of providers, while EMBEDDING_PROVIDER simply selects an embedding engine:Only fastembed, ollama, and openai_compatible select a dedicated engine. Any other value — including a typo — falls through to the LiteLLM engine without an error.
  • EMBEDDING_PROVIDER="openai_compatible" has no LLM counterpart. Setting LLM_PROVIDER="openai_compatible" fails with ValueError: 'openai_compatible' is not a valid LLMProvider — use LLM_PROVIDER="custom" instead (LM Studio, vLLM).
  • anthropic, llama_cpp, and mcp-sampling are LLM-only: they provide no embeddings, so pair them with one of the values above.
Why a local EMBEDDING_MODEL can look like it is ignored: the EMBEDDING_PROVIDER string is never sent to LiteLLM — it picks the engine (and, for openrouter, enables the automatic encoding_format guard described in Custom Providers), nothing more. With EMBEDDING_PROVIDER="custom", routing is decided entirely by the EMBEDDING_MODEL prefix, so an unprefixed model id is treated as an OpenAI model and the request goes to OpenAI (or fails on the missing key) instead of your server. Either add the matching prefix, or use openai_compatible, which sends the model id to your endpoint verbatim.Endpoint URL form — the two OpenAI-compatible providers expect different things:
With LLM_PROVIDER="custom", EMBEDDING_API_KEY does not fall back to LLM_API_KEY for LiteLLM-routed embeddings — set it explicitly (use "." when your server needs no auth). The openai_compatible engine does fall back to LLM_API_KEY.
For a full local .env covering both halves, see LM Studio or the Local Setup guide.
Cognee does not enforce a fixed allow-list of embedding models. Supported models depend on the provider configured in EMBEDDING_PROVIDER; Cognee forwards the embedding request to that provider and stores the returned vectors.
  • fastembed: any model returned by TextEmbedding.list_supported_models(), such as sentence-transformers/all-MiniLM-L6-v2. See the Fastembed section for the common list.
  • ollama / LM Studio: any embedding model loaded locally, such as bge-m3:latest or all-minilm:latest.
  • openai_compatible: any model exposed by a local /v1/embeddings server, including llama.cpp, TEI, vLLM, LocalAI, and Infinity.
  • openai / gemini / mistral / bedrock / custom: any model available through the configured provider API, using the corresponding LiteLLM prefix such as openai/ or gemini/.
Common model examples and dimensions are shown below. Set EMBEDDING_DIMENSIONS to match the model output size:Dimensions must match the model’s output size. Cognee can auto-derive EMBEDDING_DIMENSIONS for models the fastembed or LiteLLM registries know; for anything else — local aliases, custom or openai_compatible endpoints — set it explicitly. See How do I determine EMBEDDING_DIMENSIONS? for how to find the value and what a mismatch does, and Important Notes.
EMBEDDING_DIMENSIONS is the length of the vector your embedding model returns. It is a property of the model, not a free choice. Determine it in this order:
  1. Read the model card or provider docs — for example text-embedding-3-large3072, nomic-embed-text-v1.5768, BAAI/bge-m31024. See the table in Which embedding models are supported? for more.
  2. Measure it — embed one string and count the floats. This works for local aliases and self-hosted models that no registry knows:
One caveat for LiteLLM-routed (custom) models: the engine sends the currently configured dimensions value as a request parameter, so an endpoint that rejects that parameter raises UnsupportedParamsError instead of returning a vector. openai_compatible and ollama never send it and always measure cleanly.If you leave it unset, Cognee derives it from the fastembed model registry (for EMBEDDING_PROVIDER="fastembed"), otherwise from LiteLLM’s model metadata (output_vector_size), trying the model id, the id without its provider prefix, and provider/model. When neither knows the model it logs a warning and falls back to 3072:
Local aliases and custom / openai_compatible models are usually unknown, so set the value explicitly for them.What a mismatch does: Cognee creates every vector collection using the configured size (Vector(N) in LanceDB, vector(N) in PGVector), so the first write of a differently sized vector fails with a shape/dimension error — the vectors are not silently truncated or padded. Because the size is baked into the collection schema, correcting the value also requires recreating the collections with await cognee.prune.prune_system() and re-ingesting. See Dimension Consistency.
OpenAI’s text-embedding-3-* models are the exception where a smaller value is legitimate: Cognee forwards EMBEDDING_DIMENSIONS as the dimensions request parameter, and OpenAI returns shortened vectors. Most other models reject that parameter — see UnsupportedParamsError.
EMBEDDING_BATCH_SIZE controls how many text chunks are grouped into a single embedding API call. Cognee splits all chunks into batches of this size and sends them concurrently to the embedding engine.Local inference (Ollama, llama.cpp, LM Studio): Local servers handle one request at a time with limited concurrency. The default 36 can overwhelm them. Reduce the batch size if you see errors or slowdowns:
Cloud providers: Larger batches reduce the number of API calls and are efficient with cloud APIs. The default 36 suits most cloud providers.Relationship to rate limiting: Each batch counts as one request toward EMBEDDING_RATE_LIMIT_REQUESTS. A single file may produce many chunks — with EMBEDDING_BATCH_SIZE=36, a document split into 360 chunks generates 10 requests.
Despite the name, embeddings have no “completion” — EMBEDDING_MAX_COMPLETION_TOKENS is the per-text input token budget Cognee’s tokenizer uses to size chunks for the embedding model. It should reflect your embedding model’s maximum input length.Observable impact:
  • Chunk size, cost and latency. Chunks are sized as min(EMBEDDING_MAX_COMPLETION_TOKENS, LLM_MAX_COMPLETION_TOKENS // 2). A lower value forces smaller chunks, so a document produces more chunks → more embedding requests (and more LLM extraction calls) → higher latency and cost. A higher value (within the model’s real limit) produces fewer, larger chunks. See Chunkers for how chunk size shapes the graph.
  • Over-length requests. Setting this above the embedding model’s real input limit can make individual texts exceed the limit. Cognee no longer fails outright — it splits the batch (or mean-pools an over-length string, see Timeout and Retry Behavior) — but that recovery adds latency, so it is not free. “Higher” is only better up to the model’s actual limit.
Tuning guidance: match this to your embedding model’s input limit. The default 8191 fits OpenAI’s text-embedding-3-* models. For models with a smaller limit, lower it so chunks fit without triggering the split-and-pool fallback; the related LLM_MAX_COMPLETION_TOKENS (LLM Providers) caps the other half of the formula. With hosted third-party /v1/embeddings endpoints, use the limit the endpoint enforces per text — often smaller than the model’s advertised context — and leave headroom: token counts for these endpoints are usually approximate, because the model id is rarely a loadable HuggingFace repo and Cognee falls back to the TikToken tokenizer (see How Cognee selects the tokenizer).

It is a chunk-sizing hint, not an enforced cap

Cognee never sends EMBEDDING_MAX_COMPLETION_TOKENS to the provider and never truncates text at request time — it only hands the value to the tokenizer that sizes chunks during cognify(). Two consequences:
  • Setting it too high silently produces over-length requests. The provider rejects them; whether Cognee recovers depends on the error message (see below).
  • Setting it above LLM_MAX_COMPLETION_TOKENS // 2 has no effect. With the default LLM_MAX_COMPLETION_TOKENS=16384, chunks are capped at 8192 regardless — EMBEDDING_MAX_COMPLETION_TOKENS=131072 and =8192 behave identically until you also raise the LLM value.

Chunk size vs. embedding request size

One embedding request carries EMBEDDING_BATCH_SIZE chunks (default 36, see Batch Size), so the payload is roughly chunk tokens × batch size. Providers that cap the whole request rather than each text will reject a request built from chunks that each fit comfortably. If you hit a length error even though your chunks are within the model’s per-text limit, lower EMBEDDING_BATCH_SIZE rather than EMBEDDING_MAX_COMPLETION_TOKENS.

Recovery only fires on recognized error messages

The split-and-pool recovery is triggered by string matching on the provider’s error, so a provider that phrases the limit differently gets no recovery — the run fails with EmbeddingException on openai_compatible, or with the provider’s original error (re-raised after the retry window) on the LiteLLM engines:For example, <400> InternalError.Algo.InvalidParameter: Range of input length should be [1, 33000] matches none of these phrases, so Cognee raises instead of splitting. Configure the limits correctly rather than relying on the fallback.
The LiteLLMEmbeddingEngine applies two layers of protection against slow or unreachable endpoints:How retries work: Failed attempts are retried with exponential back-off starting at 2 seconds, with random jitter, until the 128-second window is exhausted.Why the per-attempt timeout is larger than the retry window: The 300-second deadline is measured per attempt and starts before any network I/O, so waiting for a free connection in the local HTTP pool and event-loop scheduling delays both count against it — under high cognify concurrency a perfectly healthy request can spend most of its budget queued client-side. A shorter deadline cancelled those queued requests and turned them into retries, which added more load. Because the retry window is evaluated between attempts, the two limits interact:
  • A request that fails fast (connection refused, rate limit, 5xx) is retried repeatedly until the 128-second window is exhausted, exactly as before.
  • A request that genuinely hangs consumes the full 300 seconds on its first attempt. By the time it is cancelled the 128-second window has already elapsed, so it is not retried — the EmbeddingException is raised directly. A single hung request therefore blocks its task for up to 5 minutes.
The per-attempt deadline matches the one used by OpenAICompatibleEmbeddingEngine, so both engines now behave the same way.What is not retried:
  • 404 Not Found errors are raised immediately — they indicate a configuration problem (wrong model name or endpoint) rather than a transient failure.
  • asyncio.CancelledError is treated as terminal and re-raised without retrying, so cancelled tasks unwind promptly instead of consuming the full retry window. The same exclusion applies to the FastembedEmbeddingEngine, OllamaEmbeddingEngine, and OpenAICompatibleEmbeddingEngine retry decorators.
  • EmbeddingContextWindowTooSmallError is treated as terminal and raised immediately. It is thrown when a single embedding text still exceeds the model’s context window but can no longer be split (the string is too short to divide further), so retrying would deterministically fail again. This exclusion applies to the LiteLLMEmbeddingEngine, FastembedEmbeddingEngine, and OpenAICompatibleEmbeddingEngine retry decorators; the failure returns at once instead of consuming the full 128-second retry window. The exception subclasses EmbeddingException (default message Text is too short to split further but exceeds context window.), so code that already catches EmbeddingException continues to catch it — catch EmbeddingContextWindowTooSmallError specifically to distinguish this deterministic, non-retryable case and shorten or pre-split the offending input.
Over-length input recovery: When the provider rejects an embedding request because the input exceeds the model’s context window, LiteLLMEmbeddingEngine no longer fails the request. This covers both LiteLLM’s ContextWindowExceededError and a plain 400 BadRequestError whose message matches maximum input length (for example OpenAI’s maximum input length is 8192 tokens, which is returned as a plain 400 by the embeddings API). On either of these:
  • If the batch contains more than one text, it is split in half and each half is embedded in parallel, then the results are concatenated.
  • If a single over-length string is left, it is split into two overlapping segments (the first two-thirds and the last two-thirds of the string), each segment is embedded, and the two vectors are averaged (mean-pooled) into one.
This split-and-pool step recurses until each piece fits, so long documents that previously failed are now embedded automatically. If a single string is already too short to split further (fewer than three characters) yet still exceeds the context window, the engine raises the terminal EmbeddingContextWindowTooSmallError instead of retrying — this deterministic failure returns at once rather than consuming the full retry window (see What is not retried). Any other 400 BadRequestError (one whose message does not indicate an over-length input) is re-raised unchanged so genuinely malformed requests still fail fast. The exact error phrasings that trigger this recovery — for this engine and for openai_compatible — are listed under Max Completion Tokens.Common error messages and causes:
  • EmbeddingException: Embedding request timed out. Check EMBEDDING_ENDPOINT connectivity. — The attempt did not complete within 300 seconds. Verify that EMBEDDING_ENDPOINT is reachable from your network. Since the deadline also covers time spent queued client-side, check that your embedding concurrency is bounded before assuming the endpoint itself is at fault.
  • EmbeddingException: Cannot connect to embedding endpoint. Check EMBEDDING_ENDPOINT. — TCP connection was refused or the server closed the connection before responding. Confirm the server is running and the URL is correct.
  • EmbeddingException: Failed to index data points using model <model> — The provider returned a 404 Not Found. Common causes: wrong model name, missing hosted_vllm/ prefix for vLLM, or an unsupported model at that endpoint. Non-over-length 400 BadRequestError responses are re-raised unchanged.
  • 400 invalid_value on encoding_format from OpenRouter — Older LiteLLM releases serialize an omitted encoding_format as JSON null, which OpenRouter rejects. Cognee now forces encoding_format="float" on detected OpenRouter routes (see the Custom Providers accordion under Provider Setup Guides), so this should no longer occur; if you still see it, upgrade Cognee to a version that includes this guard.
Diagnosing slow local servers: If you see timeouts with Ollama, LM Studio, or vLLM, a large batch can still exceed the 300-second per-attempt limit — and because a stalled request now holds its task for up to 5 minutes before failing, the cost of each one is higher than a fast timeout would be. Reduce EMBEDDING_BATCH_SIZE to send fewer texts per request:
The timeout and retry values are hardcoded in LiteLLMEmbeddingEngine and cannot be changed via environment variables. To use different limits, subclass LiteLLMEmbeddingEngine and override embed_text with a custom @retry decorator.
For the LiteLLM-routed providers (openai, gemini, mistral, bedrock, and custom), Cognee sends the OpenAI-style dimensions parameter to litellm.aembedding() whenever an embedding dimension is configured. In the .env flow documented on this page, EMBEDDING_DIMENSIONS is required, so this parameter is normally present. Only some models accept dimensions (notably OpenAI’s text-embedding-3-*). Other models — and some LiteLLM proxies — reject it with:
Cognee does not expose an environment variable to suppress dimensions or to forward extra LiteLLM params. The intended fix is LiteLLM’s own drop_params switch, which makes LiteLLM silently drop any parameter the target model does not support (including dimensions) instead of raising — so you can keep routing embeddings through your LiteLLM proxy:
Set the flag before you call any Cognee operation:
drop_params silences the error but does not change the vector size your model returns. Still set EMBEDDING_DIMENSIONS to the model’s real output size so it matches your vector store — see Which embedding models are supported? and Important Notes.
If you don’t need the LiteLLM proxy, the openai_compatible provider talks to any /v1/embeddings server directly through the OpenAI SDK and never sends the dimensions parameter, so it avoids this error entirely — at the cost of bypassing LiteLLM.
Control client-side throttling for embedding calls to manage API usage and costs.
Rate limiting is disabled by default. You must explicitly set EMBEDDING_RATE_LIMIT_ENABLED="true" to activate it.
Defaults (when rate limiting is enabled):What counts as one request?One rate-limit request = one embed_text() API call = one batch of chunks (not one chunk). With the default EMBEDDING_BATCH_SIZE=36, processing 360 chunks produces 10 requests. See the Batch Size section for how to tune batch size.Sizing guidance:Set EMBEDDING_RATE_LIMIT_REQUESTS to your provider’s RPM limit and EMBEDDING_RATE_LIMIT_INTERVAL to 60. Use ~80–90% of your provider’s advertised limit to leave headroom.Example configurations for common provider tiersThese examples target embedding endpoints, such as OpenAI embedding models like text-embedding-3-large.
Always verify your exact tier limits in your provider’s dashboard — limits vary by model, tier, and region. The examples above are approximations for common tiers and may change.
The HUGGINGFACE_TOKENIZER environment variable specifies which Hugging Face tokenizer to use for counting tokens before sending text to the embedding model. It is optional — Cognee no longer requires it at startup, so omitting it does not raise a ValidationError on import — but it is recommended when using the Ollama provider for accurate token counting.Value format: The value is the Hugging Face model repository ID — the {organization}/{model-name} path that appears in the URL on huggingface.co/models. This should match the underlying model used by your Ollama embedding.For example, if the Ollama model nomic-embed-text:latest is built from nomic-ai/nomic-embed-text-v1.5 on Hugging Face, set:

Common model-to-tokenizer mappings

Finding the tokenizer for any model

  1. Look up the model on huggingface.co/models.
  2. The repository ID is the {organization}/{model-name} part of the URL (e.g., huggingface.co/BAAI/bge-m3BAAI/bge-m3).
  3. Use the repository ID that corresponds to the model your Ollama tag is built from. The Ollama model page typically links to the original Hugging Face repository.
HUGGINGFACE_TOKENIZER is only used by the Ollama embedding engine. It is not needed for OpenAI, Fastembed, openai_compatible, or other providers. For openai_compatible, the engine automatically tries the model name as a HuggingFace tokenizer ID and falls back to TikToken if unavailable — no separate HUGGINGFACE_TOKENIZER value is required.

Troubleshooting dependency errors

The HuggingFace tokenizer (HuggingFaceTokenizer, which wraps transformers.AutoTokenizer) backs token counting for the Ollama embedding engine. Cognee also attempts to use it for custom and openai_compatible chunk sizing; if it cannot load, Cognee falls back to TikToken. transformers is not part of the base cognee install — it is an optional dependency shipped by the huggingface and ollama extras (both pin transformers>=4.46.3,<5).ModuleNotFoundError: No module named 'transformers'The tokenizer was triggered (most often by EMBEDDING_PROVIDER="ollama") but transformers is missing. This typically surfaces as Connection to Embedding handler could not be established. Install the extra:
cannot import name 'is_offline_mode' from 'huggingface_hub'This is a version mismatch: an old transformers (older than 4.46) is paired with a newer huggingface_hub that no longer exports is_offline_mode. Cognee’s transformers>=4.46.3,<5 pin is compatible with current huggingface_hub (Cognee resolves huggingface_hub 0.36.x). If you previously installed transformers manually, upgrade it into Cognee’s supported range:
Installing or reinstalling the huggingface / ollama extra pulls in a compatible huggingface_hub automatically, so prefer the extra over pinning huggingface_hub by hand.

Important Notes

  • Dimension Consistency: EMBEDDING_DIMENSIONS must match your vector store collection schema
  • API Key Fallback: If EMBEDDING_API_KEY is not set, Cognee uses LLM_API_KEY (except for custom providers)
  • Tokenization: HUGGINGFACE_TOKENIZER is optional and no longer enforced by Cognee’s startup validation — but it is recommended for the Ollama provider; set it to the HuggingFace model repo ID that matches your embedding model for accurate token counting
  • Performance: Local providers (Ollama, Fastembed) are slower but offer privacy and cost benefits
Token counts drive chunk sizing and the --dry-run token estimate, so Cognee auto-selects a tokenizer that matches your embedding model:
  • openai — the model’s TikToken (BPE) encoding.
  • gemini — the default TikToken encoding (Gemini has no local tokenizer, so counts are approximate).
  • mistral — the Mistral tokenizer.
  • fastembed — the model’s own HuggingFace (wordpiece) tokenizer, resolved from the fastembed model list.
  • ollama — the tokenizer named by HUGGINGFACE_TOKENIZER.
  • openai_compatible / custom — the embedding model id used as a HuggingFace repo (e.g. BAAI/bge-large-en-v1.5); when that is not a loadable repo — such as a local alias or a LiteLLM-prefixed id — Cognee falls back to TikToken.
If a matching tokenizer cannot be loaded, Cognee logs an advisory warning and falls back to the TikToken tokenizer. This resolution never raises: a mismatched or missing tokenizer only degrades token-count accuracy (mis-sized chunks and a skewed --dry-run estimate), it does not stop ingestion.

The “could not load a matching tokenizer” warning is benign

A repeated log line such as:
means the model id is not loadable as a HuggingFace repo. That is expected for LiteLLM-prefixed ids (openai/..., lm_studio/..., openrouter/...), Ollama tags, and local aliases such as default. Embedding and ingestion continue normally — only chunk sizing and the --dry-run token estimate become approximate.To silence it:
  • ollama — set HUGGINGFACE_TOKENIZER to the HuggingFace repo your Ollama model is built from (see the mappings earlier in this section).
  • openai_compatible — set EMBEDDING_MODEL to the served model’s real HuggingFace repo id (for example Qwen/Qwen3-Embedding-4B instead of default). This provider needs no LiteLLM prefix, so the id can be the repo id directly.
  • custom with hosted_vllm/ — the prefix is stripped automatically before tokenizer resolution, so the warning does not appear as long as the model name after the prefix is a valid HuggingFace repo id (see the vLLM accordion).
  • custom with any other prefix (openai/, lm_studio/, openrouter/, …) — the prefix is required for routing and is passed to tokenizer resolution whole, so the warning cannot be avoided; treat it as informational. If your endpoint is OpenAI-compatible, switching to openai_compatible with the plain repo id removes it.
The warning text suggests setting HUGGINGFACE_TOKENIZER, but that variable is only consulted by the Ollama engine — custom, openai_compatible, fastembed, and the LiteLLM providers do not pass it to tokenizer resolution, so setting it there has no effect.

LLM Providers

Configure LLM providers for text generation

Vector Stores

Set up vector databases for embedding storage

Overview

Return to setup configuration overview