Comparison
hippo vs Zep: a local Zep alternative
Zep is a managed context service for agents, built on Graphiti, its open-source temporal knowledge graph. Its self-hosted Community Edition is deprecated, so Zep today means Zep Cloud, or your own cloud on the Enterprise plan. Hippo runs on your machine with no account, and wires itself into the coding agents it finds.
Both keep an old fact when a newer one replaces it. Graphiti gives facts validity windows and invalidates old ones
rather than deleting them. Hippo links a superseded memory to its replacement and leaves it out of recall, and
hippo recall --as-of <date> shows what was current on a given date.
Feature by feature
Every row of the README comparison table, hippo against Zep.
| Feature | hippo | Zep |
|---|---|---|
| Decay by default | Yes | No |
| Retrieval strengthening | Yes | No |
| Reward-proportional decay | Yes | No |
| Hybrid search (BM25 + embeddings) | Yes | Yes (graph + vec) |
| Schema acceleration / knowledge graph | Yes (schema) | Yes (temporal KG) |
| Conflict detection + resolution | Yes | Yes (auto-invalidate stale facts) |
| Multi-agent shared memory | Yes | Yes |
| Transfer scoring | Yes | No |
| Outcome tracking | Yes | No |
| Confidence tiers | Yes | No |
| Spatial organization | No | No |
| Lossless compression | No | No |
| Cross-tool import (ChatGPT/Claude/Cursor) | Yes | ? |
| Auto-hook install | Yes | No |
| MCP server | Yes | Yes (hosted, needs an account) |
| Zero runtime deps | Yes | No (managed service) |
| LongMemEval (best published) | 98.0% any / 88.5% all R@5* (local MiniLM; 99.8% any-evidence with voyage-3-large; s_cleaned, per-haystack) | 90.2% accuracy** (LoCoMo 94.7%) |
| Git-friendly | Yes | No |
| Framework agnostic | Yes | Yes |
| License | MIT | Proprietary cloud (Graphiti: Apache-2.0) |
* Hippo's figures are on longmemeval_s_cleaned with a per-question haystack, each the best of five retrieval settings in the benchmark scripts, not hippo recall. Any-evidence R@5 counts a hit when any answer session is in the top 5, over all 500 questions: 98.0% with the free local MiniLM embedder (an optional install) and 99.8% with voyage-3-large (measured 2026-06-09, not re-run). All-evidence R@5 counts a hit only when every answer session is in the top 5, over the 470 questions that have an answer: 86.8 to 88.5% with MiniLM. gbrain first published 97.6%, an any-evidence score over all 500; its report (opens in new tab) now leads with all-evidence, 95.53% (449 of 470) with the paid Voyage rerank-2.5 reranker and 93.19% without it. On all-evidence recall gbrain is ahead. The June 2026 build scored 98.6 any-evidence; docs/evals/2026-09-23-longmemeval-reproduction.md (opens in new tab) has both runs. An older hippo number, 86.8% R@5 on longmemeval_oracle under pooled (non-per-haystack) retrieval, is not comparable to per-haystack figures.
** Different metric: these are end-to-end answer scores, not retrieval R@5. Mem0's 94.4 comes from its hosted platform, which its README says includes optimizations the open-source SDK lacks. Zep's 90.2% and 94.7% are accuracy figures from its homepage. Memoria's 88.78% and EverMind's 83% are overall accuracy with a reader LLM. Higher denominator + LLM helps. Not directly comparable to retrieval-only R@5 numbers above. The Mem0, Zep and Letta columns were last checked against each vendor's own pages on 2026-09-28.
What Zep ships today (checked 2026-09-28)
- Zep Cloud, which in Zep's words "unifies business data, documents, and conversations into shared, governed context so agents can complete tasks correctly" (getzep.com (opens in new tab)). Enterprise customers can run it in their own cloud.
- No self-hosted Zep: Community Edition is deprecated and no longer supported (Zep FAQ (opens in new tab)).
- Graphiti, Apache-2.0, for building your own: it needs a graph database (Neo4j, FalkorDB or Amazon Neptune) and defaults to OpenAI for inference and embeddings (Graphiti README (opens in new tab)).
- A hosted Memory MCP Server that needs sign-in through an identity provider and a Zep project, with MCP seats per account (docs (opens in new tab)).
- Plans from a free tier of 10,000 credits a month to Flex at $125 and Flex Plus at $375 a month, then Enterprise (pricing (opens in new tab)).
When to pick which
Pick hippo
The memory is for coding agents on your machine. You want no account, graph database or model to run, lessons you can mark wrong, and markdown mirrors you can read and commit.
Pick Zep
You want a managed service that turns business data, documents and conversations into a knowledge graph for your agents, with someone else running it.
FAQ
hippo and Zep, asked directly
Can I still self-host Zep?
Not Zep itself. Zep's FAQ says Zep Community Edition, the version you could host locally, is deprecated and no longer supported (checked 2026-09-28); Enterprise customers can run Zep in their own cloud. Graphiti, the open-source framework under Zep, is Apache-2.0 and runs anywhere, but it needs a graph database such as Neo4j or FalkorDB and a language model key, OpenAI by default.
Is hippo a Zep alternative?
Yes, if the memory is for coding agents on your own machine. Hippo is MIT-licensed, stores memories in SQLite in your project, needs no account, graph database or model, and hippo init wires it into Claude Code, Codex, Cursor, OpenClaw, OpenCode and Pi. Zep fits when you want a managed service that builds a knowledge graph from business data, documents and conversations.
Is Zep's 90.2% on LongMemEval comparable to hippo's 98.0%?
No. Zep's homepage reports 90.2% accuracy on LongMemEval, which scores whether answers were right. Hippo's figure is retrieval: on LongMemEval-S with a per-question haystack, using the benchmark scripts and a free local embedder rather than hippo recall, an answer session was in the top five results 98.0% of the time. Different metrics, so neither number beats the other.