Skip to content

Comparison

hippo vs Zep: a local Zep alternative

Zep is a managed context service for agents, built on Graphiti, its open-source temporal knowledge graph. Its self-hosted Community Edition is deprecated, so Zep today means Zep Cloud, or your own cloud on the Enterprise plan. Hippo runs on your machine with no account, and wires itself into the coding agents it finds.

Both keep an old fact when a newer one replaces it. Graphiti gives facts validity windows and invalidates old ones rather than deleting them. Hippo links a superseded memory to its replacement and leaves it out of recall, and hippo recall --as-of <date> shows what was current on a given date.

Feature by feature

Every row of the README comparison table, hippo against Zep.

hippo and Zep, row by row from the README comparison table
Feature hippo Zep
Decay by default Yes No
Retrieval strengthening Yes No
Reward-proportional decay Yes No
Hybrid search (BM25 + embeddings) Yes Yes (graph + vec)
Schema acceleration / knowledge graph Yes (schema) Yes (temporal KG)
Conflict detection + resolution Yes Yes (auto-invalidate stale facts)
Multi-agent shared memory Yes Yes
Transfer scoring Yes No
Outcome tracking Yes No
Confidence tiers Yes No
Spatial organization No No
Lossless compression No No
Cross-tool import (ChatGPT/Claude/Cursor) Yes ?
Auto-hook install Yes No
MCP server Yes Yes (hosted, needs an account)
Zero runtime deps Yes No (managed service)
LongMemEval (best published) 98.0% any / 88.5% all R@5* (local MiniLM; 99.8% any-evidence with voyage-3-large; s_cleaned, per-haystack) 90.2% accuracy** (LoCoMo 94.7%)
Git-friendly Yes No
Framework agnostic Yes Yes
License MIT Proprietary cloud (Graphiti: Apache-2.0)

* Hippo's figures are on longmemeval_s_cleaned with a per-question haystack, each the best of five retrieval settings in the benchmark scripts, not hippo recall. Any-evidence R@5 counts a hit when any answer session is in the top 5, over all 500 questions: 98.0% with the free local MiniLM embedder (an optional install) and 99.8% with voyage-3-large (measured 2026-06-09, not re-run). All-evidence R@5 counts a hit only when every answer session is in the top 5, over the 470 questions that have an answer: 86.8 to 88.5% with MiniLM. gbrain first published 97.6%, an any-evidence score over all 500; its report (opens in new tab) now leads with all-evidence, 95.53% (449 of 470) with the paid Voyage rerank-2.5 reranker and 93.19% without it. On all-evidence recall gbrain is ahead. The June 2026 build scored 98.6 any-evidence; docs/evals/2026-09-23-longmemeval-reproduction.md (opens in new tab) has both runs. An older hippo number, 86.8% R@5 on longmemeval_oracle under pooled (non-per-haystack) retrieval, is not comparable to per-haystack figures.

** Different metric: these are end-to-end answer scores, not retrieval R@5. Mem0's 94.4 comes from its hosted platform, which its README says includes optimizations the open-source SDK lacks. Zep's 90.2% and 94.7% are accuracy figures from its homepage. Memoria's 88.78% and EverMind's 83% are overall accuracy with a reader LLM. Higher denominator + LLM helps. Not directly comparable to retrieval-only R@5 numbers above. The Mem0, Zep and Letta columns were last checked against each vendor's own pages on 2026-09-28.

What Zep ships today (checked 2026-09-28)

  • Zep Cloud, which in Zep's words "unifies business data, documents, and conversations into shared, governed context so agents can complete tasks correctly" (getzep.com (opens in new tab)). Enterprise customers can run it in their own cloud.
  • No self-hosted Zep: Community Edition is deprecated and no longer supported (Zep FAQ (opens in new tab)).
  • Graphiti, Apache-2.0, for building your own: it needs a graph database (Neo4j, FalkorDB or Amazon Neptune) and defaults to OpenAI for inference and embeddings (Graphiti README (opens in new tab)).
  • A hosted Memory MCP Server that needs sign-in through an identity provider and a Zep project, with MCP seats per account (docs (opens in new tab)).
  • Plans from a free tier of 10,000 credits a month to Flex at $125 and Flex Plus at $375 a month, then Enterprise (pricing (opens in new tab)).

When to pick which

Pick hippo

The memory is for coding agents on your machine. You want no account, graph database or model to run, lessons you can mark wrong, and markdown mirrors you can read and commit.

Pick Zep

You want a managed service that turns business data, documents and conversations into a knowledge graph for your agents, with someone else running it.

FAQ

hippo and Zep, asked directly

Can I still self-host Zep?

Not Zep itself. Zep's FAQ says Zep Community Edition, the version you could host locally, is deprecated and no longer supported (checked 2026-09-28); Enterprise customers can run Zep in their own cloud. Graphiti, the open-source framework under Zep, is Apache-2.0 and runs anywhere, but it needs a graph database such as Neo4j or FalkorDB and a language model key, OpenAI by default.

Is hippo a Zep alternative?

Yes, if the memory is for coding agents on your own machine. Hippo is MIT-licensed, stores memories in SQLite in your project, needs no account, graph database or model, and hippo init wires it into Claude Code, Codex, Cursor, OpenClaw, OpenCode and Pi. Zep fits when you want a managed service that builds a knowledge graph from business data, documents and conversations.

Is Zep's 90.2% on LongMemEval comparable to hippo's 98.0%?

No. Zep's homepage reports 90.2% accuracy on LongMemEval, which scores whether answers were right. Hippo's figure is retrieval: on LongMemEval-S with a per-question haystack, using the benchmark scripts and a free local embedder rather than hippo recall, an answer session was in the top five results 98.0% of the time. Different metrics, so neither number beats the other.