Sep 21, 2026

Give Your Agents Memory That Outlives the Session With Cognee

A 31k-star open-source memory platform that turns documents, code, and conversations into a self-hosted knowledge graph — with a local mode that runs without any LLM API key.

#tutorial#ai-agents#context-engineering

Every new agent session starts from zero: the decisions, fixes, and rules you settled last week are gone. Cognee, at 31k stars, is an open-source AI memory platform that attacks exactly this — it turns text, code, and conversations into a self-hosted knowledge graph your agents can search across sessions.

Why This Skill Matters

Cognee exposes four operations that cover a memory lifecycle: remember stores content or code in permanent memory (or a session, when you pass a session ID), recall retrieves context and answers, improve enriches memory and bridges session knowledge into the graph, and forget removes an item or dataset. Text becomes entities, relationships, and searchable chunks; code becomes a graph of symbols and dependencies.

The part that changes adoption math: since v1.6.0 (September 18, 2026), you can build and search text memory with local extraction and embedding models — no OpenAI or Anthropic key required. Text ingestion, retrieval, and session storage all work without an LLM; LLM-dependent improvement stages skip automatically. Adding a hosted LLM later unlocks generated answers and richer processing.

Installation

Cognee needs Python 3.10–3.14. The README's local-first install adds the GLiNER extra:

uv pip install "cognee[gliner]"

An LLM key is optional. To add one later in your shell — the README sets the same LLM_API_KEY via os.environ or a .env file:

export LLM_API_KEY="YOUR OPENAI_API_KEY"

Real Workflow: Build Memory From Text — No LLM

Save this as quickstart.py and run it — this is the README's own local quickstart, which extracts a graph and embeds with local models:

import asyncio

import cognee


async def main():
    # Extract a knowledge graph and embed the text with local models.
    await cognee.remember(
        "Marie Curie was born in Warsaw and worked at the University of Paris.",
        dataset_name="local_quickstart",
    )

    # Retrieve the matching source text; no LLM generates an answer.
    results = await cognee.recall(
        "Where was Marie Curie born?",
        datasets=["local_quickstart"],
    )
    for result in results:
        print(result)


if __name__ == "__main__":
    asyncio.run(main())

The same two steps exist on the CLI:

cognee-cli remember "Marie Curie was born in Warsaw." -d local_quickstart
cognee-cli recall "Where was Marie Curie born?" -d local_quickstart

If you would rather explore than build, cognee-cli demo loads bundled sample data and runs keyword search — no API key, no model downloads.

Real Workflow: Hook It Into Claude Code

To give a coding agent the same memory, install the Claude Code plugin from the integrations repository:

claude plugin marketplace add topoteretes/cognee-integrations
claude plugin install cognee-memory@cognee

A Codex plugin exists too — it requires enabling hooks in ~/.codex/config.toml first. Cursor, Cline, and other MCP clients connect through the Cognee MCP server instead.

Tips

  • Run cognee-cli -ui to inspect a local installation in a UI — the launcher needs Node.js/npm, and Docker for its MCP service.
  • Migration is a first-class path: the README documents importing memory from Mem0, Letta, Zep, or Graphiti via the COGX exchange format.
  • The whole memory layer can run on a single Postgres instance — graph, vectors, and metadata together. Note the README's own caveat: that Postgres-as-graph-store mode is released as a demo feature; the production-ready version is a licensed product.

When Not to Use This

If your context only matters inside one session, a markdown file the agent reads at startup is simpler — Cognee pays off when knowledge has to persist and be retrieved across sessions, agents, or machines. And recall without a configured LLM returns matching source text, not generated answers; if you want synthesized replies from your memory, you will need that optional LLM configuration.


See the leaderboard for more skills.