A biologically-inspired memory system for LLM agents — episodic storage, associative graphs, and human-like forgetting.
Give your agent a memory that behaves like a mind, not a filing cabinet.
Most agent "memory" today is a flat pile of text chunks in a vector database.
Every fact is equally important, equally permanent, and equally isolated. You
embed text, you top_k it back, you paste it into a prompt. It works — until the
pile gets big, contradictory, and full of things that mattered once and don't
anymore.
Human memory doesn't work like that. We forget on purpose. We reinforce what we revisit. We remember by association — one thing pulls up another. We keep future intentions separate from past facts. And our memories subtly change every time we recall them.
agent-memory-engine ports those mechanisms into a small, dependency-light Python
library you can drop into any agent — with a plain vector database still living
underneath as just one of the pluggable parts.
store.remember(summary="Adopted a puppy named Rex", keywords=["dog", "pet"])
store.remember(summary="Bought a leash and kibble", keywords=["supplies"], related_to=["dog"])
store.recall("Rex")
# → returns the puppy memory (direct match)
# AND the leash memory — which shares no words with "Rex",
# surfaced by spreading activation across the associative graph.| ⏳ Exponential decay + reinforcement | Memories fade over time (lazy decay — computed on read, no background jobs). Recalling one restores and strengthens it. Use it or lose it. |
| 🕸️ Associative graph + spreading activation | Memories are nodes joined by typed, weighted edges. A query lights up direct matches, then activation spreads to related memories that share no words with it. |
| ❤️ Emotional tagging | Memories carry valence & intensity; emotionally charged ones decay slower and are dampened during the "dream" pass (negative first — like overnight affective down-regulation). |
| 🔮 Prospective memory | Future intentions ("call the dentist when…") are stored apart from episodic facts, with triggers and expiry. |
| 🔁 Reconsolidation | Recall is a write, not just a read. Access reinforces; corrections supersede (demote, never delete) the memories they revise — keeping an auditable history. |
| 🌙 Nightly consolidation | A "sleep" pass matures old episodes into semantic facts, ages memories through warm → cold → archive, forgets stale detail, and (optionally, with an LLM) forms cross-context hypotheses. |
| 🔌 Pluggable everything | Embedder & vector store sit behind tiny interfaces. Ships zero-dependency (HashingEmbedder + InMemoryVectorStore) and production-ready (SentenceTransformerEmbedder + ChromaVectorStore). |
Retrievability decays exponentially, but the rate is modulated by how well the
memory is consolidated (storage strength S):
R_eff(t) = R_anchor * e^(-λ_eff * Δt) # Δt = hours since last access λ_eff = decay_rate * (1 - S * 0.75) # strong memories decay slower
Emotionally charged memories resist forgetting via a parallel protective effect. The two effects combine multiplicatively — both storage strength and emotional intensity independently suppress decay:
λ_eff = decay_rate * (1 - 0.75 * S) * (1 - 0.8 * emotional_intensity)
On a successful recall the memory is reinforced (the spacing effect). Storage strength is updated first, then retrievability is restored from the new strength value:
ΔS = β * (1 - R) * (1 - S) # the more forgotten it was, the more it gains S ← S + ΔS # 1. update storage strength first R ← 0.5 + 0.5 * S # 2. then restore retrievability from the NEW S
β ∈ [0.1, 0.3] # controls how fast S grows per successful recall # higher β = faster consolidation, fewer reviews needed
On a failed recall attempt, the memory is penalized. The penalty scales with how consolidated and expected the memory was — failing to recall something that should have been strong is a stronger error signal than failing on something already weak:
ΔS = -γ * R * S # the stronger/more expected the memory was, the more it hurts to fail S ← max(S + ΔS, 0) # storage strength can't go negative R ← R * 0.5 # retrievability takes an immediate hit, but isn't zeroed out
γ ∈ [0.2, 0.4] # penalty coefficient, not yet fit
⚠️ Calibration status: The modulation coefficients above (0.75for storage strength,0.8for emotional intensity — seedecay.py) are principled starting values, not yet fit against real recall data. They're inspired by dual-process forgetting models (Bjork's storage/retrieval strength) but haven't been optimized the way FSRS fits its 19+ weights against millions of real reviews. See Roadmap — empirical calibration is a planned milestone, not an afterthought.
The core engine needs nothing but the Python standard library (sqlite3).
git clone https://github.com/<your-username>/agent-memory-engine.git
cd agent-memory-engine
# Run the demo right away — no dependencies required:
python examples/demo.pyInstall as a package, with optional production extras:
pip install -e . # core only
pip install -e ".[chroma]" # + ChromaDB vector store
pip install -e ".[embeddings]"# + sentence-transformers (semantic embeddings)
pip install -e ".[all]" # everything, incl. test toolingRequirements: Python 3.10+. Optional: chromadb>=0.5, sentence-transformers>=3.0.
from src import Database, EpisodicMemoryStore, HashingEmbedder, InMemoryVectorStore
# 1. Set up storage (SQLite for records + graph, a vector store for similarity)
db = Database("memory.sqlite")
db.run_migrations()
store = EpisodicMemoryStore(
db=db,
vector_store=InMemoryVectorStore(), # or: ChromaVectorStore(path="./chroma")
embedder=HashingEmbedder(), # or: SentenceTransformerEmbedder()
)
# 2. Encode memories — 'related_to' wires them into the associative graph
store.remember(summary="Started learning classical guitar", keywords=["guitar", "hobby"],
emotional_valence=0.6, emotional_intensity=0.5)
store.remember(summary="Practices every morning for 30 min", keywords=["routine"],
related_to=["guitar", "hobby"], memory_type="procedural")
# 3. Recall — vector familiarity + spreading activation, ranked & token-budgeted
result = store.recall("morning guitar practice")
for m in result.memories:
print(f"{m.final_score:.3f} {m.summary} (cascade={m.cascade_score:.2f})")
# 4. A future intention
store.remember_intention(action="Book a guitar setup", trigger="when strings buzz")
# 5. The nightly "sleep" pass (maturation / aging / dreaming)
from src import run_consolidation
run_consolidation(db) # pass llm=... to also generate cross-context hypothesesWant semantic recall and a persistent vector DB? Swap two lines:
from src import ChromaVectorStore, SentenceTransformerEmbedder
store = EpisodicMemoryStore(
db=db,
vector_store=ChromaVectorStore(path="./chroma"),
embedder=SentenceTransformerEmbedder(), # multilingual, 768-dim
)python examples/demo.pyThe demo walks through every feature end-to-end with synthetic data — encoding and linking, recall via spreading activation, decay reversed by recall, prospective memory, corrections that supersede, and a full consolidation pass. Sample output:
2. RECALL — 'morning guitar practice routine'
[0.424] Practices guitar for 30 minutes every morning sim=0.38 cascade=0.00
[0.410] Signed up for a weekend photography course sim=0.00 cascade=0.50 ← via spreading
[0.395] Started learning to play the classical guitar sim=0.19 cascade=0.00
3. DECAY + RECONSOLIDATION
effective R after 20 idle days : 0.006 (S=0.10)
after recall #1: R=0.570 S=0.141 access_count=2
after recall #3: R=0.606 S=0.211 access_count=4
Plain RAG is stateless similarity search over an append-only store. This engine adds the parts of memory RAG leaves out:
| Plain RAG | agent-memory-engine | |
|---|---|---|
| Retrieval | vector similarity only | similarity + spreading activation over a graph |
| Time | none — a 2-year-old chunk ranks like today's | decay: retrievability fades, recall reinforces |
| Recall effect | read-only | reconsolidation: recall strengthens; corrections supersede |
| Associations | implicit in embedding space | explicit typed graph (associative / causal / supersedes / …) |
| Importance | all chunks equal | storage strength, hub score, emotional weight |
| Forgetting | never (or manual deletes) | layered aging warm → cold → archive + detail forgetting |
| Future | not represented | prospective memory for intentions |
RAG answers "which stored text is most similar to this query?" This answers "what would an agent with a memory like ours actually recall right now?"
remember(): text → embedding → episodic record → associative links
recall(): query → vector familiarity → lazy decay → spreading activation
→ composite scoring → token budget → reinforced recall
Full component map, sequence diagrams, and module reference live in
docs/architecture.md.
agent-memory-engine/
├── src/
│ ├── episodic_store.py # the facade: remember() / recall()
│ ├── associative_graph.py # links + spreading activation
│ ├── decay.py # decay + reinforcement math
│ ├── reconsolidation.py # recall reinforces; corrections supersede
│ ├── consolidation.py # the nightly "sleep" pass
│ ├── vector_store.py # VectorStore interface (Chroma / in-memory)
│ ├── embeddings.py # Embedder interface (hashing / sentence-transformers)
│ ├── llm.py # optional LLMProvider (extraction / dreaming)
│ ├── models.py db.py # dataclasses + SQLite layer
├── schema/migrations/ # SQL schema
├── examples/demo.py # runnable end-to-end demo
├── tests/ # pytest suite
└── docs/architecture.md # diagrams + deep dive
pip install pytest
pytest # 12 tests, runs offline in ~0.1sEverything runs with the built-in offline providers, so the test suite needs no model downloads and no network.
- Async API for high-throughput agents
- Additional VectorStore backends (pgvector, Qdrant, LanceDB)
- Reference LLMProvider implementations (OpenAI / Anthropic / local)
- Benchmarks vs. plain RAG on long-horizon recall
- Configurable consolidation schedules & policies
- Empirical calibration of decay/reinforcement coefficients against real recall logs
Ideas worth exploring but not yet scoped — no timeline, no committed design:
- Multi-agent memory sync — how would multiple agents share or reconcile overlapping memory graphs? CRDT-based merge is one candidate, but semantics for decay/reinforcement conflicts across replicas haven't been worked out.
This is a young project — here's where it currently falls short, honestly:
- Spreading activation cost at scale. Graph traversal has no depth/breadth cap yet. On graphs with thousands of nodes, a single
recall()could walk a large portion of the graph before converging. Untested beyond small/medium graphs. - Conflicting memories without explicit supersession. The
supersedeslink type handles clean corrections ("I moved to Berlin" replacing "I live in Lisbon"), but two valid, context-dependent memories that contradict each other (e.g., a preference that's true in one context and false in another) aren't disambiguated — both surface with no conflict resolution. - Uncalibrated decay coefficients. The modulation constants in the decay formula (storage strength, emotional intensity weighting) are principled starting points, not empirically fit — see Roadmap.
- No concurrency guarantees. Concurrent writes (
remember()calls from multiple processes) haven't been stress-tested against the SQLite + graph layer. Single-writer usage is the only validated pattern so far. - Consolidation pass scales linearly with memory count. The nightly "sleep" pass processes the full memory set each run; no incremental/batched consolidation yet, so very large stores may see growing consolidation time.
Ideas and issues welcome — see below.
Contributions are very welcome!
- Fork the repo and create a feature branch.
- Keep the core dependency-free — new backends go behind the existing
Embedder/VectorStore/LLMProviderinterfaces. - Add tests (
pytest) for anything you change. - Open a PR with a clear description.
Found a bug or have an idea? Open an issue.
Released under the MIT License — free to use, modify, and build on.