Sep 21, 2026
Give Your Agents Memory That Outlives the Session With Cognee
A 31k-star open-source memory platform that turns documents, code, and conversations into a self-hosted knowledge graph — with a local mode that runs without any LLM API key.
Every new agent session starts from zero: the decisions, fixes, and rules you settled last week are gone. Cognee, at 31k stars, is an open-source AI memory platform that attacks exactly this — it turns text, code, and conversations into a self-hosted knowledge graph your agents can search across sessions.
Why This Skill Matters
Cognee exposes four operations that cover a memory lifecycle: remember stores content or code in permanent memory (or a session, when you pass a session ID), recall retrieves context and answers, improve enriches memory and bridges session knowledge into the graph, and forget removes an item or dataset. Text becomes entities, relationships, and searchable chunks; code becomes a graph of symbols and dependencies.
The part that changes adoption math: since v1.6.0 (September 18, 2026), you can build and search text memory with local extraction and embedding models — no OpenAI or Anthropic key required. Text ingestion, retrieval, and session storage all work without an LLM; LLM-dependent improvement stages skip automatically. Adding a hosted LLM later unlocks generated answers and richer processing.
Installation
Cognee needs Python 3.10–3.14. The README's local-first install adds the GLiNER extra:
uv pip install "cognee[gliner]"
An LLM key is optional. To add one later in your shell — the README sets the same LLM_API_KEY via os.environ or a .env file:
export LLM_API_KEY="YOUR OPENAI_API_KEY"
Real Workflow: Build Memory From Text — No LLM
Save this as quickstart.py and run it — this is the README's own local quickstart, which extracts a graph and embeds with local models:
import asyncio
import cognee
async def main():
# Extract a knowledge graph and embed the text with local models.
await cognee.remember(
"Marie Curie was born in Warsaw and worked at the University of Paris.",
dataset_name="local_quickstart",
)
# Retrieve the matching source text; no LLM generates an answer.
results = await cognee.recall(
"Where was Marie Curie born?",
datasets=["local_quickstart"],
)
for result in results:
print(result)
if __name__ == "__main__":
asyncio.run(main())
The same two steps exist on the CLI:
cognee-cli remember "Marie Curie was born in Warsaw." -d local_quickstart
cognee-cli recall "Where was Marie Curie born?" -d local_quickstart
If you would rather explore than build, cognee-cli demo loads bundled sample data and runs keyword search — no API key, no model downloads.
Real Workflow: Hook It Into Claude Code
To give a coding agent the same memory, install the Claude Code plugin from the integrations repository:
claude plugin marketplace add topoteretes/cognee-integrations
claude plugin install cognee-memory@cognee
A Codex plugin exists too — it requires enabling hooks in ~/.codex/config.toml first. Cursor, Cline, and other MCP clients connect through the Cognee MCP server instead.
Tips
- Run
cognee-cli -uito inspect a local installation in a UI — the launcher needs Node.js/npm, and Docker for its MCP service. - Migration is a first-class path: the README documents importing memory from Mem0, Letta, Zep, or Graphiti via the COGX exchange format.
- The whole memory layer can run on a single Postgres instance — graph, vectors, and metadata together. Note the README's own caveat: that Postgres-as-graph-store mode is released as a demo feature; the production-ready version is a licensed product.
When Not to Use This
If your context only matters inside one session, a markdown file the agent reads at startup is simpler — Cognee pays off when knowledge has to persist and be retrieved across sessions, agents, or machines. And recall without a configured LLM returns matching source text, not generated answers; if you want synthesized replies from your memory, you will need that optional LLM configuration.
See the leaderboard for more skills.