Skip to content

Repository files navigation

ContextHub

An open-source, MCP-accessible context layer for AI agents and LLMs.

ContextHub gives every agent you use — Claude Code, Cursor, your own scripts — a single, self-updating memory drawn from the apps you actually live in. You decide, per-call, what each agent is allowed to see.

It is built around three ideas:

  1. One store, many agents. Connectors pull from your sources into a single local SQLite database. Every MCP client talks to the same memory.
  2. Ship lines, not documents. A line-level scorer assembles only the spans that answer the query, instead of dumping whole chunks into the prompt.
  3. Permissions are part of retrieval. Scope policies filter the candidate set before scoring, so blocked content never reaches the model.

Status: working MVP. Local-only, no auth. Ships with an interactive web console, a CLI, an MCP server, four example connectors (incl. paste-to-ingest), hybrid FTS5 retrieval, and a full unit-test suite. Use it as a starting point — not a hosted product.


Architecture

┌──────────────┐    ┌──────────────────┐    ┌──────────────┐
│  Connectors  │ →  │ Ingestion + tag- │ →  │  SQLite db   │
│ (gh, md, …)  │    │ ger + segmenter  │    │ items, lines │
└──────────────┘    └──────────────────┘    └──────┬───────┘
                                                   │
                                ┌──────────────────┴──────────────────┐
                                │     Permission policy (scopes)      │
                                └──────────────────┬──────────────────┘
                                                   │
                                ┌──────────────────┴──────────────────┐
                                │   Line-level retrieval scorer       │
                                └──────────────────┬──────────────────┘
                                                   │
                                ┌──────────────────┴──────────────────┐
                                │   MCP server (stdio / http)         │
                                │   tools: search_context, get_context│
                                │   list_sources, ingest_context, stats│
                                └──────────────────┬──────────────────┘
                                                   │
                    ┌──────────────────────────────┴───────────────────────────┐
                    │  Clients:  Web console (Next.js)  ·  CLI  ·  any MCP host  │
                    └────────────────────────────────────────────────────────────┘

Retrieval is hybrid: when SQLite is built with FTS5, an FTS5 BM25 prefix match narrows the candidate set, and the line-level scorer re-ranks within it. If FTS5 is unavailable it falls back to scanning all lines — same results, slower. Permission filtering is applied identically on both paths.

Repo layout

apps/
  web/           Next.js landing page + interactive console at /dashboard
  cli/           `contexthub` CLI (search, list, sources, stats, get, ingest)
  mcp-server/    MCP server (stdio)
packages/
  core/          Types, SQLite store (FTS5), tagger, segmenter, ingest pipeline
  retrieval/     Hybrid FTS5 + line-level scorer + prompt formatter
  permissions/   Scope policies + filter
  connectors/    Markdown vault, GitHub fixture, Slack export, generic text
fixtures/        Demo data so the project works end-to-end out of the box

Quickstart

Requires Node 20+ and pnpm.

pnpm install
pnpm build
pnpm ingest:demo   # seeds ./data/contexthub.db from ./fixtures
pnpm dev:mcp       # starts the MCP server over stdio

Open the web app locally — the landing page links straight to the console:

pnpm dev:web       # http://localhost:3000  (console at /dashboard)

The console (/dashboard) is a full UI over the same store: run search_context queries, toggle permission scopes live and watch the token savings vs. baseline RAG, browse every item by source, inspect lines, and paste-to-ingest arbitrary text that gets tagged + segmented on the fly.

Run the test suite

pnpm test          # 92 unit/integration tests (tagger, segmenter,
                   # permissions, retrieval + db) via Vitest

Use the CLI

export CONTEXTHUB_DB=$(pwd)/data/contexthub.db
pnpm cli stats
pnpm cli search "stripe webhooks failing" --scopes exclude_private --max 5
pnpm cli sources
pnpm cli list
pnpm cli get <id-prefix>
pnpm cli ingest:md ./path/to/vault

Wire it into Claude Code

claude mcp add contexthub \
  --command "node $(pwd)/apps/mcp-server/dist/index.js" \
  --env CONTEXTHUB_DB=$(pwd)/data/contexthub.db \
  --env CONTEXTHUB_SCOPES=exclude_private

Wire it into Cursor or any MCP client

{
  "mcpServers": {
    "contexthub": {
      "command": "node",
      "args": ["/abs/path/contexthub/apps/mcp-server/dist/index.js"],
      "env": {
        "CONTEXTHUB_DB": "/abs/path/contexthub/data/contexthub.db",
        "CONTEXTHUB_SCOPES": "exclude_private,exclude_confidential"
      }
    }
  }
}

MCP tools

Tool What it does
search_context Returns the top-scored lines for a query, grouped by item, after policy.
get_context Fetches a single item by id (use this after search_context).
list_sources Lists every source app currently represented in the store.
ingest_context Adds a blob of text to the store; auto-tagged + segmented, then searchable.
stats Store-wide summary: items, sources, FTS availability, sensitivity, topics.

search_context accepts a per-call scopes override.


Scope policies

Set defaults in CONTEXTHUB_SCOPES, comma-separated. Override per call via the scopes argument on search_context.

Scope Effect
no_restriction Default. Everything in the store is retrievable.
exclude_private Drops items tagged private (personal notes, etc.).
exclude_confidential Drops items tagged confidential (secrets, comp).
exclude_source:<name> Drops an entire source app (e.g. slack).
min_confidence:<0..1> Drops items whose tagger confidence is below the bar.

Filters are applied before scoring. Blocked items never enter the candidate set, so there is no partial leak.


Writing a connector

A connector is anything that produces RawContext[]. The pipeline handles tagging, segmentation, and storage:

import type { RawContext } from "@contexthub/core/types";

export async function loadMyApp(): Promise<RawContext[]> {
  return [
    {
      source: "my-app",
      externalId: "doc-42",
      title: "Q3 plan",
      body: "...",
      updatedAt: new Date().toISOString(),
    },
  ];
}

Then call ingest(db, raw) (one-shot) or ingestBatch(db, raws).


Roadmap

  • FTS5 hybrid retrieval (BM25 candidate set + line-level re-rank).
  • A small CLI to add sources, run ingestion, and inspect the store.
  • Generic text connector + paste-to-ingest in the web console.
  • ingest_context + stats MCP tools.
  • Unit + integration test suite.
  • Embeddings layer on top of the FTS5 candidate set.
  • HTTP transport for the MCP server (multi-user, with per-token scopes).
  • Real connectors for Notion, GitHub API, Linear, Gmail.

Acknowledgements

ContextHub is independently developed and not affiliated with any other context or memory product. It exists because the same problem — fragmented memory across AI tools — is worth solving in the open, with a transparent, local-first design.

License

MIT. See LICENSE.

About

Open-source, MCP-accessible context layer for AI agents and LLMs.

Resources

Stars

2 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages