Skip to main content

How Cognee tracks time

Cognee builds a temporal knowledge graph by default: dates in your data become part of the graph (see How Dates Are Stored). Use this guide when you want to ask time-scoped questions — before, after, or between two points in time — and have them answered from the content dated inside that window rather than by embedding similarity alone. To track how facts change over time, see Conflicting and Outdated Memories.

Before You Start

  • Complete Quickstart to understand basic operations
  • Ensure you have LLM Providers configured
  • Read Recall for how querying memory works
  • Any built-in graph store works. The default Ladybug/Kuzu store, Neo4j, and the Postgres demo graph answer the time-window lookups with a native query; Turso, Neptune, and Neptune Analytics use a generic fallback that scans the graph and walks each candidate’s neighbourhood, which is slower on large graphs. A community adapter, such as FalkorDB or Memgraph, needs get_graph_data and get_neighborhood for that fallback, or its own get_timestamps_in_range and get_temporal_anchors
  • No data is required up front — the script ingests its own dated sample text, but it starts with cognee.forget(everything=True), which wipes all existing Cognee data; run it against a setup you can afford to reset

Code in Action

What Just Happened

Step 1: Remember Dated Data

The script starts from a clean state, then ingests TEXT into the timeline_demo dataset. No ingestion flag is needed: the default remember() pipeline extracts the dates in the text as Timestamp nodes, which SearchType.TEMPORAL reads, so there is no separate cognify() step. This example uses one string treated as a single document; multiple documents, files, or entire datasets are processed the same way.

Step 2: Ask Time-aware Questions

The loop runs the three query shapes time-aware search is built for — a before query, an after query, and one bounded by a pair of dates — using SearchType.TEMPORAL, imported from the cognee package at the top of the script. Each call is scoped with datasets=["timeline_demo"] so it only searches the timeline it just ingested; drop the argument to search every dataset you have access to. The answer for each query is results[0].text.
  • If the query names a time, the retriever pulls in content dated inside that window and moves it to the front of the results. Dates at least as precise as the question rank equally; only coarser dates rank lower
  • If no time window can be read from the query, or nothing among the results is dated inside it, it logs a warning and answers with plain hybrid search
  • Increase top_k to inspect more results; the candidate pool is four times top_k

How Dates Are Stored

The default pipeline asks the extraction LLM to model every date it finds as a Timestamp node, attached as a leaf to the entity it dates — for example marie_curie --born_at--> 1867. The chunk that mentions the date also links to the Timestamp through its contains edge, just as it does for entities.
  • A date covers everything it states and no more: "In 2001 version 1.0 shipped" becomes 2001 with precision year, covering all of 2001, so it stays distinguishable from 2001-01-01.
  • A Timestamp node’s id is derived from timestamp_str, so every mention of the same date across chunks and documents shares one node. Timestamp nodes are not embedded.
  • A date without a stated year (“on 27 April”, “later that spring”) is resolved against the last year stated earlier in the same document and handed to the extraction prompt as a hint; the chunk text itself is not changed.
  • A date the LLM leaves in a form that cannot be normalized stays an ordinary Entity node.
At query time the retriever does three things:
  1. Reads a time window from the question. An LLM call turns the question into one UTC window, start inclusive and end exclusive. A named year, month, or day covers that whole unit; “before X”, “until X”, “after X”, and “since X” leave one side open (see How boundary words are read).
  2. Gathers candidates. It runs a hybrid search for four times top_k candidates and adds the chunks attached to Timestamp nodes inside the window, which similarity search alone tends to miss.
  3. Reranks. Candidates attached to an in-window timestamp — directly, or through an entity that is — move to the front, and the rest follow in similarity order. The list is then cut to top_k.
Anchored candidates are ordered by how well their most precise matching timestamp fits the question’s window:
  • As precise as the question, or more precise → these all rank the same and keep their embedding-similarity order. For “in 2020”, a paragraph dated September 2020 is not pushed below a row dated 2020-09-14.
  • Coarser than the question → ranked lower, and the wider the timestamp the lower it goes. For “on 18 March 1965”, a chunk dated 1965-03-18 comes before one that only says 1965.
  • Open-ended window (“since 2022”, “before 1900”) → no precision preference. Every anchored candidate ranks the same and keeps its similarity order.

How boundary words are read

Whether the named point X is inside the window depends on the word in front of it: A timestamp matches when its period overlaps the window, so “What changed since 1995?” includes anything dated in 1995, and a period such as 1990/1996 as well.

Relative and vague time expressions

Extraction only creates a Timestamp for a date it can place on a calendar. A phrase such as “when I was young, I loved running” or “when I was at primary 2, I transferred school” still produces its entities and chunk links, but no Timestamp node:
  • The content remains fully retrievable through hybrid search, including the fallback inside SearchType.TEMPORAL.
  • It is invisible to the time window, so a “before 2005” query will not move it to the front.
On the query side, relative references such as “last year” or “today”, a question with no time in it, and a question naming several separate windows all yield no window. The retriever then logs a warning and returns the plain hybrid result, which is also what happens when a window resolves but nothing among the candidates is dated inside it.
To make a relative phrase queryable by time, give the calendar date in the ingested text — "In 2003, when I was in primary 2, I transferred school" produces a Timestamp for 2003, while the bare phrase does not.

Using the HTTP API

If your server is running, you can run temporal search via the API by setting search_type to "TEMPORAL":
The temporal_cognify option has been removed. remember() and cognify() still accept it for compatibility but ignore it without a warning, and the default pipeline runs. Datasets built earlier with temporal_cognify=True hold Event nodes that the current SearchType.TEMPORAL path does not read. To rebuild such a dataset, drop its graph and vectors and cognify it again:

Full Examples

Additional examples about temporal awareness are available on our GitHub.
  • An advanced script running temporal search over real documents is on our GitHub. Instead of the inlined four-sentence timeline above, it ingests two bundled biographies as separate documents, then mixes before / after / between range queries with person-centric questions that carry no dates — exercising the hybrid-search fallback described in the tip above.

Conflicting and Outdated Memories

Close superseded facts and check staleness with is_valid()

Core Concepts Overview

Understand how Cognee builds and stores knowledge graphs.

API Reference

Explore the search endpoint behind temporal queries.