Skip to main content

cognee.run_startup_migrations

Applies all pending database schema migrations before the rest of your application starts.
It runs two steps in sequence:
  1. Relational schema β€” executes alembic upgrade head against your configured relational database (SQLite by default, or Postgres).
  2. Vector schema β€” runs the vector adapter’s run_migrations method for every database that needs it:
    • Single-user mode (ENABLE_BACKEND_ACCESS_CONTROL=False): migrates the single default vector engine.
    • Multi-user mode (ENABLE_BACKEND_ACCESS_CONTROL=True, the default): iterates over every dataset database and migrates each one individually. A failure for one dataset is logged and skipped; the remaining datasets continue to migrate.
    • If the active vector engine has no run_migrations method, Cognee logs a warning and skips that engine.
    • If the dataset_database table does not exist yet (a fresh database), the vector migration step is skipped with a warning instead of raising. This is handled on both SQLite (OperationalError, β€œno such table”) and PostgreSQL/pgvector (ProgrammingError / UndefinedTableError).

Entrypoints

All migration functions live in the cognee.run_migrations module. run_startup_migrations is also re-exported at the top level as cognee.run_startup_migrations, so it is the recommended entrypoint for most applications.

When to call it

For the default local setup (SQLite + LanceDB), Cognee handles migrations automatically when the API server starts. You only need to call run_startup_migrations() explicitly in server deployments or CI pipelines where you manage database lifecycle yourself.

Example

Kubernetes init container

Run migrations as a one-shot init container so the main pod only starts after the schema is ready:

Concurrency

Every migration flow runs under a single cross-process lock, so a host performs at most one migration of any kind at a time. If several processes start at once β€” multiple workers of the same server, parallel SDK runs, or several init containers β€” only one acquires the lock and migrates; the others block until it finishes, then re-read the stored revision and skip work that is already done. Nothing runs migrations in parallel. Because of this, startup can block (and time-to-ready can increase) while another process holds the lock and migrates. This is expected under the Kubernetes init container and multi-worker scenarios above β€” the wait is the coordination working as intended, not a hang. The lock backend depends on your relational database:
  • Postgres β€” a session-scoped advisory lock, which also serializes migrations across hosts. Use Postgres metadata when multiple hosts may start and migrate at the same time.
  • SQLite β€” an OS advisory file lock placed next to the database file. It serializes multiple processes on a single host (multi-worker servers, parallel SDK runs) but not across hosts or over NFS.

Errors

Set LOG_LEVEL=DEBUG to see the full Alembic output when diagnosing migration failures.