Skip to content

[N15] Semantic recall: an embedding provider and a SQLite vector memory #258

Description

@LinuxDevil

Goal

Memory slots recall the newest items, or items that share a word with the query. With this ticket a slot can recall by meaning: an EmbeddingProvider turns text into vectors, and two new memory providers (inMemoryVectorMemory() and sqliteVectorMemory(store)) rank items by cosine similarity to the query, so "what does the user like to eat" finds "Sam is vegetarian". No new dependency: embeddings come from any AI SDK embedding model through the existing ai peer, and vectors live in the agent's existing SQLite file. AUDIT-2 section 5, idea 10 "Semantic recall / observational memory" (Mastra ships it; we have "memory slots, newest-first").

Current state

  • MemoryProvider (src/memory/defineMemory.ts:17-23): list(scopeKey, { limit?, query? }) ("returns the newest items first"), add(scopeKey, { text, metadata? }), remove(scopeKey, id). MemoryItem { id, text, createdAt, metadata? } (lines 9-15).
  • defineMemory() (line 79): recall?: { onSessionStart?, maxItems? (10), query?: 'last-input' | 'none' } (line 51); with 'last-input' the last user message is passed to list() as query.
  • Recall paths in src/memory/withMemory.ts: the recall_<name> tool takes { query?, limit? } and calls slot.provider.list(key, { query, limit }), with the description "Search the ... memory, newest items first." (lines 40-47); the memory-recall preGenerate hook puts provider.list(key, { limit: maxItems, query }) into a <memory name="..."> block on the first model call.
  • Built-in providers (src/memory/providers.ts): itemsProvider(store, { maxItems = 1000 }) with select() (lines 16-21): query words of 3+ letters matched with includes, newest first. inMemoryMemory(), fileMemory({ dir }) (src/memory/fileMemory.ts), sqliteMemory(store) (src/storage/sqlite/sqliteMemory.ts, one JSON array per scope key in memory_items), exported from @lousho/build-ai-agent/sqlite (src/storage/sqlite/index.ts:7).
  • SQLite: built-in node:sqlite (src/storage/sqlite/driver.ts:33-56), migrations appended to MIGRATIONS (src/storage/sqlite/migrations.ts:11, "never edit shipped ones"; three entries today).
  • Contract test helper: describeMemoryProviderContract(name, make) (src/memory/providerContract.ts:6), whose "filters by query" case (line 17) expects substring filtering.
  • No embedding concept anywhere: grep -i embed in non-test src/ finds only unrelated names; embed / embedMany from ai are not used. ai is a required peer (^4.3.19 || ^6.0.0 || ^7.0.0) and is imported in src/providers/aiSdkProvider.ts:17.
  • Docs: docs/memory.md (## Providers line 101, ### SQLite line 113, ## Testing memory line 150, ## Limitations line 170).

Scope

In:

  • src/memory/embeddings.ts, exported from the root:
    export interface EmbeddingProvider {
      /** Stable id of the model and settings, stored with each vector (e.g. 'openai:text-embedding-3-small'). */
      readonly id: string;
      embed(texts: readonly string[], options?: { signal?: AbortSignal }): Promise<number[][]>;
    }
    /** Any AI SDK embedding model (e.g. openai.embedding('text-embedding-3-small')), through `embedMany` from 'ai'. */
    export function aiSdkEmbedder(model: unknown, options?: { id?: string; maxBatch?: number /* default 96 */ }): EmbeddingProvider;
    aiSdkEmbedder imports embedMany from ai lazily, checks that model looks like an AI SDK embedding model (has modelId and doEmbed), and uses `${model.provider}:${model.modelId}` as the default id. Works on ai 4, 6 and 7 (embedMany({ model, values }) returns { embeddings } on all three; verify against ai and the ai-v7 dev dependency in tests).
  • A deterministic offline embedder for tests and docs: hashEmbedder({ dimensions = 256 }?) exported from @lousho/build-ai-agent/testing (hashed bag of lower-cased word stems, L2-normalized; id 'hash:<dimensions>'). Good enough to make "vegetarian" close to "vegetarian food" in tests, not a real semantic model; document that.
  • Vector providers (core in src/memory/vectorProvider.ts, over a small VectorItemStore interface: load(scopeKey), insert(scopeKey, row), remove(scopeKey, id), trim(scopeKey, maxItems)):
    export interface VectorMemoryOptions {
      embedder: EmbeddingProvider;
      maxItems?: number;     // per scope key, default 5000; the oldest are dropped
      minScore?: number;     // cosine similarity floor for query results, default 0.2
    }
    export function inMemoryVectorMemory(options: VectorMemoryOptions): MemoryProvider & { reindex(scopeKey?: string): Promise<number> };
    // @lousho/build-ai-agent/sqlite
    export function sqliteVectorMemory(store: SqliteStore, options: VectorMemoryOptions): MemoryProvider & { reindex(scopeKey?: string): Promise<number> };
    • add: embed the text (one call), store text, metadata, createdAt, the embedder id, dimensions and the vector.
    • list(key, { query, limit }): without a query (or a blank one) return newest first, as every provider does; with a query embed it, score every item of that scope key whose embedder id matches by cosine similarity, drop scores under minScore, sort by score (ties: newer first), return limit. The score is not added to MemoryItem (keep the shape); put it in metadata.score of the returned copies.
    • Items embedded with a different embedder id are not scored (they cannot be compared) but still appear in no-query listings; reindex() re-embeds them in batches and returns how many it changed.
    • Brute-force scoring in JavaScript (no vector extension): fine up to maxItems per scope key; document the limit.
  • SQLite storage: append a migration creating memory_vectors (scope_key TEXT NOT NULL, id TEXT NOT NULL, text TEXT NOT NULL, metadata TEXT, created_at TEXT NOT NULL, embedder TEXT NOT NULL, dimensions INTEGER NOT NULL, embedding BLOB NOT NULL, PRIMARY KEY (scope_key, id)) and an index on (scope_key, created_at). Vectors stored as little-endian Float32Array bytes. sqliteMemory() and memory_items are unchanged.
  • Tool wording: add an optional readonly ranking?: 'newest' | 'relevance' to MemoryProvider; the vector providers set 'relevance', and withMemory.ts then describes recall_<name> as "Search the ... memory by meaning, most relevant first." (default wording unchanged for other providers). Recall at run start already passes the last user message when recall.query is 'last-input'; document that this is how a slot recalls by meaning at the start of a run.
  • Contract: give describeMemoryProviderContract a third parameter { query: 'filter' | 'rank' } (default 'filter', so existing callers are unchanged); in 'rank' mode the query case asserts the matching item comes first instead of asserting the others are absent. Run it against both new providers with hashEmbedder().
  • Docs (docs/memory.md): append a new ## Semantic recall section at the end of the page, after ## Limitations (a heading inside ## Providers would shift ## Testing memory and ## Limitations in the Arabic page, which is matched by heading position), with: the idea, aiSdkEmbedder(openai.embedding('text-embedding-3-small')), both providers, recall: { query: 'last-input' }, minScore, reindex() after changing models, hashEmbedder() for tests, the size limit and the cost note (one embedding call per add and per query). One sentence in ### SQLite pointing to it.
  • CHANGELOG entry; npm run docs:llms.

Out:

  • Observational memory (summarizing conversations into memories automatically), chunking long documents, RAG over files, rerankers, a sqlite-vec or other native vector index, hybrid keyword plus vector search, vector support in fileMemory() or the Cloudflare KV store.

Acceptance criteria

  • src/memory/embeddings.test.ts: aiSdkEmbedder over a fake AI SDK embedding model (an object with provider, modelId, doEmbed) batches by maxBatch, preserves order, uses the default id; rejects a non-embedding model with a clear ConfigurationError; runs against the installed ai and the ai-v7 dev dependency (follow how src/providers/aiMajorPeers.test.ts switches majors).
  • src/memory/vectorMemory.test.ts: with hashEmbedder(), a query ranks the related item first; minScore drops unrelated items; no query lists newest first; maxItems trims the oldest; items from another embedder id are skipped for ranking and fixed by reindex(); the contract passes in 'rank' mode for inMemoryVectorMemory.
  • src/storage/sqlite/sqliteVectorMemory.test.ts: the same contract on a temp-file SqliteStore; vectors survive reopening the database; a database at the previous user_version migrates; memory_items and sqliteMemory are unaffected.
  • A createAgent({ memory: [slot] }) run with mockModel and recall: { query: 'last-input' } puts the most relevant item first in the <memory> block, and the recall_<name> tool description says "most relevant first".
  • Existing memory tests pass unchanged (default contract mode 'filter').
  • Docs as listed; snippets pass npm run docs:verify-snippets -- --skip-build; CHANGELOG; npm run docs:llms; every check in BRIEF-2.md "Verification" passes.
  • The pull request tells the docs site (G9) about the new ## Semantic recall section at the end of memory.

Live test

One real embedding round trip, only if OpenRouter serves embeddings: first check curl -s https://openrouter.ai/api/v1/embeddings with the key and body {"model":"openai/text-embedding-3-small","input":["test"]}. If it answers with an embedding array, add src/memory/embeddings.live.test.ts that builds aiSdkEmbedder(createOpenAI({ baseURL: 'https://openrouter.ai/api/v1', apiKey: process.env.OPENROUTER_API_KEY }).embedding('openai/text-embedding-3-small')), adds three short items to inMemoryVectorMemory, and asserts the expected item ranks first for a paraphrased query. It runs only when OPENROUTER_API_KEY is set (it.skipIf), never in CI; there is no cassette (cassettes record chat providers, not embedders), so commit the test but not any output. If OpenRouter does not serve embeddings, write "Live test: not possible, OpenRouter has no embeddings endpoint for this model" in the pull request and spend nothing. Maximum spend: 0.02 USD.

Dependencies

None. Conflicts: N9a also appends a migration to MIGRATIONS and touches SqliteStore exports; whichever merges second renumbers its new entry (never edit a shipped one). src/memory/withMemory.ts is not touched by other wave 3 tickets.

Notes for the implementer

  • Normalize vectors once on write and on query, then cosine similarity is a dot product.
  • node:sqlite returns BLOBs as Uint8Array; build the Float32Array over a copy aligned to 4 bytes.
  • embed can fail (network, quota): add must not store an item without its vector; let the error reach the remember_<name> tool call as a normal tool error.
  • Keep src/memory/ free of node:* imports apart from what fileMemory.ts already uses; the vector core must stay usable from the in-memory provider in a Worker.

Round 2 ticket N15. Before starting, read the agent brief (worktree rules, verification list, live-test budget) and the plan. One ticket is one pull request; put Closes #<this issue> in it.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    live-testHas a live-model test with a budgetmodel:sonnetWell specified; a Sonnet agent can take itround-2Round 2 plan ticketwave-3Round 2, wave 3

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions