Goal
Memory slots recall the newest items, or items that share a word with the query. With this ticket a slot can recall by meaning: an EmbeddingProvider turns text into vectors, and two new memory providers (inMemoryVectorMemory() and sqliteVectorMemory(store)) rank items by cosine similarity to the query, so "what does the user like to eat" finds "Sam is vegetarian". No new dependency: embeddings come from any AI SDK embedding model through the existing ai peer, and vectors live in the agent's existing SQLite file. AUDIT-2 section 5, idea 10 "Semantic recall / observational memory" (Mastra ships it; we have "memory slots, newest-first").
Current state
MemoryProvider (src/memory/defineMemory.ts:17-23): list(scopeKey, { limit?, query? }) ("returns the newest items first"), add(scopeKey, { text, metadata? }), remove(scopeKey, id). MemoryItem { id, text, createdAt, metadata? } (lines 9-15).
defineMemory() (line 79): recall?: { onSessionStart?, maxItems? (10), query?: 'last-input' | 'none' } (line 51); with 'last-input' the last user message is passed to list() as query.
- Recall paths in
src/memory/withMemory.ts: the recall_<name> tool takes { query?, limit? } and calls slot.provider.list(key, { query, limit }), with the description "Search the ... memory, newest items first." (lines 40-47); the memory-recall preGenerate hook puts provider.list(key, { limit: maxItems, query }) into a <memory name="..."> block on the first model call.
- Built-in providers (
src/memory/providers.ts): itemsProvider(store, { maxItems = 1000 }) with select() (lines 16-21): query words of 3+ letters matched with includes, newest first. inMemoryMemory(), fileMemory({ dir }) (src/memory/fileMemory.ts), sqliteMemory(store) (src/storage/sqlite/sqliteMemory.ts, one JSON array per scope key in memory_items), exported from @lousho/build-ai-agent/sqlite (src/storage/sqlite/index.ts:7).
- SQLite: built-in
node:sqlite (src/storage/sqlite/driver.ts:33-56), migrations appended to MIGRATIONS (src/storage/sqlite/migrations.ts:11, "never edit shipped ones"; three entries today).
- Contract test helper:
describeMemoryProviderContract(name, make) (src/memory/providerContract.ts:6), whose "filters by query" case (line 17) expects substring filtering.
- No embedding concept anywhere: grep
-i embed in non-test src/ finds only unrelated names; embed / embedMany from ai are not used. ai is a required peer (^4.3.19 || ^6.0.0 || ^7.0.0) and is imported in src/providers/aiSdkProvider.ts:17.
- Docs:
docs/memory.md (## Providers line 101, ### SQLite line 113, ## Testing memory line 150, ## Limitations line 170).
Scope
In:
src/memory/embeddings.ts, exported from the root:
export interface EmbeddingProvider {
/** Stable id of the model and settings, stored with each vector (e.g. 'openai:text-embedding-3-small'). */
readonly id: string;
embed(texts: readonly string[], options?: { signal?: AbortSignal }): Promise<number[][]>;
}
/** Any AI SDK embedding model (e.g. openai.embedding('text-embedding-3-small')), through `embedMany` from 'ai'. */
export function aiSdkEmbedder(model: unknown, options?: { id?: string; maxBatch?: number /* default 96 */ }): EmbeddingProvider;
aiSdkEmbedder imports embedMany from ai lazily, checks that model looks like an AI SDK embedding model (has modelId and doEmbed), and uses `${model.provider}:${model.modelId}` as the default id. Works on ai 4, 6 and 7 (embedMany({ model, values }) returns { embeddings } on all three; verify against ai and the ai-v7 dev dependency in tests).
- A deterministic offline embedder for tests and docs:
hashEmbedder({ dimensions = 256 }?) exported from @lousho/build-ai-agent/testing (hashed bag of lower-cased word stems, L2-normalized; id 'hash:<dimensions>'). Good enough to make "vegetarian" close to "vegetarian food" in tests, not a real semantic model; document that.
- Vector providers (core in
src/memory/vectorProvider.ts, over a small VectorItemStore interface: load(scopeKey), insert(scopeKey, row), remove(scopeKey, id), trim(scopeKey, maxItems)):
export interface VectorMemoryOptions {
embedder: EmbeddingProvider;
maxItems?: number; // per scope key, default 5000; the oldest are dropped
minScore?: number; // cosine similarity floor for query results, default 0.2
}
export function inMemoryVectorMemory(options: VectorMemoryOptions): MemoryProvider & { reindex(scopeKey?: string): Promise<number> };
// @lousho/build-ai-agent/sqlite
export function sqliteVectorMemory(store: SqliteStore, options: VectorMemoryOptions): MemoryProvider & { reindex(scopeKey?: string): Promise<number> };
add: embed the text (one call), store text, metadata, createdAt, the embedder id, dimensions and the vector.
list(key, { query, limit }): without a query (or a blank one) return newest first, as every provider does; with a query embed it, score every item of that scope key whose embedder id matches by cosine similarity, drop scores under minScore, sort by score (ties: newer first), return limit. The score is not added to MemoryItem (keep the shape); put it in metadata.score of the returned copies.
- Items embedded with a different embedder id are not scored (they cannot be compared) but still appear in no-query listings;
reindex() re-embeds them in batches and returns how many it changed.
- Brute-force scoring in JavaScript (no vector extension): fine up to
maxItems per scope key; document the limit.
- SQLite storage: append a migration creating
memory_vectors (scope_key TEXT NOT NULL, id TEXT NOT NULL, text TEXT NOT NULL, metadata TEXT, created_at TEXT NOT NULL, embedder TEXT NOT NULL, dimensions INTEGER NOT NULL, embedding BLOB NOT NULL, PRIMARY KEY (scope_key, id)) and an index on (scope_key, created_at). Vectors stored as little-endian Float32Array bytes. sqliteMemory() and memory_items are unchanged.
- Tool wording: add an optional
readonly ranking?: 'newest' | 'relevance' to MemoryProvider; the vector providers set 'relevance', and withMemory.ts then describes recall_<name> as "Search the ... memory by meaning, most relevant first." (default wording unchanged for other providers). Recall at run start already passes the last user message when recall.query is 'last-input'; document that this is how a slot recalls by meaning at the start of a run.
- Contract: give
describeMemoryProviderContract a third parameter { query: 'filter' | 'rank' } (default 'filter', so existing callers are unchanged); in 'rank' mode the query case asserts the matching item comes first instead of asserting the others are absent. Run it against both new providers with hashEmbedder().
- Docs (
docs/memory.md): append a new ## Semantic recall section at the end of the page, after ## Limitations (a heading inside ## Providers would shift ## Testing memory and ## Limitations in the Arabic page, which is matched by heading position), with: the idea, aiSdkEmbedder(openai.embedding('text-embedding-3-small')), both providers, recall: { query: 'last-input' }, minScore, reindex() after changing models, hashEmbedder() for tests, the size limit and the cost note (one embedding call per add and per query). One sentence in ### SQLite pointing to it.
- CHANGELOG entry;
npm run docs:llms.
Out:
- Observational memory (summarizing conversations into memories automatically), chunking long documents, RAG over files, rerankers, a
sqlite-vec or other native vector index, hybrid keyword plus vector search, vector support in fileMemory() or the Cloudflare KV store.
Acceptance criteria
Live test
One real embedding round trip, only if OpenRouter serves embeddings: first check curl -s https://openrouter.ai/api/v1/embeddings with the key and body {"model":"openai/text-embedding-3-small","input":["test"]}. If it answers with an embedding array, add src/memory/embeddings.live.test.ts that builds aiSdkEmbedder(createOpenAI({ baseURL: 'https://openrouter.ai/api/v1', apiKey: process.env.OPENROUTER_API_KEY }).embedding('openai/text-embedding-3-small')), adds three short items to inMemoryVectorMemory, and asserts the expected item ranks first for a paraphrased query. It runs only when OPENROUTER_API_KEY is set (it.skipIf), never in CI; there is no cassette (cassettes record chat providers, not embedders), so commit the test but not any output. If OpenRouter does not serve embeddings, write "Live test: not possible, OpenRouter has no embeddings endpoint for this model" in the pull request and spend nothing. Maximum spend: 0.02 USD.
Dependencies
None. Conflicts: N9a also appends a migration to MIGRATIONS and touches SqliteStore exports; whichever merges second renumbers its new entry (never edit a shipped one). src/memory/withMemory.ts is not touched by other wave 3 tickets.
Notes for the implementer
- Normalize vectors once on write and on query, then cosine similarity is a dot product.
node:sqlite returns BLOBs as Uint8Array; build the Float32Array over a copy aligned to 4 bytes.
embed can fail (network, quota): add must not store an item without its vector; let the error reach the remember_<name> tool call as a normal tool error.
- Keep
src/memory/ free of node:* imports apart from what fileMemory.ts already uses; the vector core must stay usable from the in-memory provider in a Worker.
Round 2 ticket N15. Before starting, read the agent brief (worktree rules, verification list, live-test budget) and the plan. One ticket is one pull request; put Closes #<this issue> in it.
Goal
Memory slots recall the newest items, or items that share a word with the query. With this ticket a slot can recall by meaning: an
EmbeddingProviderturns text into vectors, and two new memory providers (inMemoryVectorMemory()andsqliteVectorMemory(store)) rank items by cosine similarity to the query, so "what does the user like to eat" finds "Sam is vegetarian". No new dependency: embeddings come from any AI SDK embedding model through the existingaipeer, and vectors live in the agent's existing SQLite file. AUDIT-2 section 5, idea 10 "Semantic recall / observational memory" (Mastra ships it; we have "memory slots, newest-first").Current state
MemoryProvider(src/memory/defineMemory.ts:17-23):list(scopeKey, { limit?, query? })("returns the newest items first"),add(scopeKey, { text, metadata? }),remove(scopeKey, id).MemoryItem { id, text, createdAt, metadata? }(lines 9-15).defineMemory()(line 79):recall?: { onSessionStart?, maxItems? (10), query?: 'last-input' | 'none' }(line 51); with'last-input'the last user message is passed tolist()asquery.src/memory/withMemory.ts: therecall_<name>tool takes{ query?, limit? }and callsslot.provider.list(key, { query, limit }), with the description "Search the ... memory, newest items first." (lines 40-47); thememory-recallpreGeneratehook putsprovider.list(key, { limit: maxItems, query })into a<memory name="...">block on the first model call.src/memory/providers.ts):itemsProvider(store, { maxItems = 1000 })withselect()(lines 16-21): query words of 3+ letters matched withincludes, newest first.inMemoryMemory(),fileMemory({ dir })(src/memory/fileMemory.ts),sqliteMemory(store)(src/storage/sqlite/sqliteMemory.ts, one JSON array per scope key inmemory_items), exported from@lousho/build-ai-agent/sqlite(src/storage/sqlite/index.ts:7).node:sqlite(src/storage/sqlite/driver.ts:33-56), migrations appended toMIGRATIONS(src/storage/sqlite/migrations.ts:11, "never edit shipped ones"; three entries today).describeMemoryProviderContract(name, make)(src/memory/providerContract.ts:6), whose "filters by query" case (line 17) expects substring filtering.-i embedin non-testsrc/finds only unrelated names;embed/embedManyfromaiare not used.aiis a required peer (^4.3.19 || ^6.0.0 || ^7.0.0) and is imported insrc/providers/aiSdkProvider.ts:17.docs/memory.md(## Providersline 101,### SQLiteline 113,## Testing memoryline 150,## Limitationsline 170).Scope
In:
src/memory/embeddings.ts, exported from the root:aiSdkEmbedderimportsembedManyfromailazily, checks thatmodellooks like an AI SDK embedding model (hasmodelIdanddoEmbed), and uses`${model.provider}:${model.modelId}`as the default id. Works onai4, 6 and 7 (embedMany({ model, values })returns{ embeddings }on all three; verify againstaiand theai-v7dev dependency in tests).hashEmbedder({ dimensions = 256 }?)exported from@lousho/build-ai-agent/testing(hashed bag of lower-cased word stems, L2-normalized; id'hash:<dimensions>'). Good enough to make "vegetarian" close to "vegetarian food" in tests, not a real semantic model; document that.src/memory/vectorProvider.ts, over a smallVectorItemStoreinterface:load(scopeKey),insert(scopeKey, row),remove(scopeKey, id),trim(scopeKey, maxItems)):add: embed the text (one call), store text, metadata,createdAt, the embedder id, dimensions and the vector.list(key, { query, limit }): without a query (or a blank one) return newest first, as every provider does; with a query embed it, score every item of that scope key whose embedder id matches by cosine similarity, drop scores underminScore, sort by score (ties: newer first), returnlimit. The score is not added toMemoryItem(keep the shape); put it inmetadata.scoreof the returned copies.reindex()re-embeds them in batches and returns how many it changed.maxItemsper scope key; document the limit.memory_vectors (scope_key TEXT NOT NULL, id TEXT NOT NULL, text TEXT NOT NULL, metadata TEXT, created_at TEXT NOT NULL, embedder TEXT NOT NULL, dimensions INTEGER NOT NULL, embedding BLOB NOT NULL, PRIMARY KEY (scope_key, id))and an index on(scope_key, created_at). Vectors stored as little-endianFloat32Arraybytes.sqliteMemory()andmemory_itemsare unchanged.readonly ranking?: 'newest' | 'relevance'toMemoryProvider; the vector providers set'relevance', andwithMemory.tsthen describesrecall_<name>as "Search the ... memory by meaning, most relevant first." (default wording unchanged for other providers). Recall at run start already passes the last user message whenrecall.queryis'last-input'; document that this is how a slot recalls by meaning at the start of a run.describeMemoryProviderContracta third parameter{ query: 'filter' | 'rank' }(default'filter', so existing callers are unchanged); in'rank'mode the query case asserts the matching item comes first instead of asserting the others are absent. Run it against both new providers withhashEmbedder().docs/memory.md): append a new## Semantic recallsection at the end of the page, after## Limitations(a heading inside## Providerswould shift## Testing memoryand## Limitationsin the Arabic page, which is matched by heading position), with: the idea,aiSdkEmbedder(openai.embedding('text-embedding-3-small')), both providers,recall: { query: 'last-input' },minScore,reindex()after changing models,hashEmbedder()for tests, the size limit and the cost note (one embedding call peraddand per query). One sentence in### SQLitepointing to it.npm run docs:llms.Out:
sqlite-vecor other native vector index, hybrid keyword plus vector search, vector support infileMemory()or the Cloudflare KV store.Acceptance criteria
src/memory/embeddings.test.ts:aiSdkEmbedderover a fake AI SDK embedding model (an object withprovider,modelId,doEmbed) batches bymaxBatch, preserves order, uses the default id; rejects a non-embedding model with a clearConfigurationError; runs against the installedaiand theai-v7dev dependency (follow howsrc/providers/aiMajorPeers.test.tsswitches majors).src/memory/vectorMemory.test.ts: withhashEmbedder(), a query ranks the related item first;minScoredrops unrelated items; no query lists newest first;maxItemstrims the oldest; items from another embedder id are skipped for ranking and fixed byreindex(); the contract passes in'rank'mode forinMemoryVectorMemory.src/storage/sqlite/sqliteVectorMemory.test.ts: the same contract on a temp-fileSqliteStore; vectors survive reopening the database; a database at the previoususer_versionmigrates;memory_itemsandsqliteMemoryare unaffected.createAgent({ memory: [slot] })run withmockModelandrecall: { query: 'last-input' }puts the most relevant item first in the<memory>block, and therecall_<name>tool description says "most relevant first".'filter').npm run docs:verify-snippets -- --skip-build; CHANGELOG;npm run docs:llms; every check in BRIEF-2.md "Verification" passes.## Semantic recallsection at the end ofmemory.Live test
One real embedding round trip, only if OpenRouter serves embeddings: first check
curl -s https://openrouter.ai/api/v1/embeddingswith the key and body{"model":"openai/text-embedding-3-small","input":["test"]}. If it answers with anembeddingarray, addsrc/memory/embeddings.live.test.tsthat buildsaiSdkEmbedder(createOpenAI({ baseURL: 'https://openrouter.ai/api/v1', apiKey: process.env.OPENROUTER_API_KEY }).embedding('openai/text-embedding-3-small')), adds three short items toinMemoryVectorMemory, and asserts the expected item ranks first for a paraphrased query. It runs only whenOPENROUTER_API_KEYis set (it.skipIf), never in CI; there is no cassette (cassettes record chat providers, not embedders), so commit the test but not any output. If OpenRouter does not serve embeddings, write "Live test: not possible, OpenRouter has no embeddings endpoint for this model" in the pull request and spend nothing. Maximum spend: 0.02 USD.Dependencies
None. Conflicts: N9a also appends a migration to
MIGRATIONSand touchesSqliteStoreexports; whichever merges second renumbers its new entry (never edit a shipped one).src/memory/withMemory.tsis not touched by other wave 3 tickets.Notes for the implementer
node:sqlitereturns BLOBs asUint8Array; build theFloat32Arrayover a copy aligned to 4 bytes.embedcan fail (network, quota):addmust not store an item without its vector; let the error reach theremember_<name>tool call as a normal tool error.src/memory/free ofnode:*imports apart from whatfileMemory.tsalready uses; the vector core must stay usable from the in-memory provider in a Worker.Round 2 ticket
N15. Before starting, read the agent brief (worktree rules, verification list, live-test budget) and the plan. One ticket is one pull request; putCloses #<this issue>in it.