Turn a corpus of markdown and structured data into a knowledge graph you can ship as a static file.
Built for the case described in Build-Time Knowledge Graphs: a few hundred to a few thousand documents, where the whole graph fits in a browser tab and a graph database would be an operational cost with no payoff. docs/DESIGN.md has the reasoning condensed for this repo — the size table, why entity resolution is the real work, and what's actually flexible here versus fixed on purpose. New here? docs/quickstart.md gets a graph out of a handful of throwaway notes in about five minutes.
Two rules keep this a module rather than a framework:
- Sources are adapters. Every adapter yields the same
Documentshape. The core never learns the name of your blog, your book, or your CMS. Ships with eleven built-in adapters: markdown notes (obsidian,logseq), rendered HTML (html_blog), Notion workspaces (notion), Confluence spaces (confluence), GitHub issues & PRs (github), tabular DataFrames / Arrow / dicts (dataframe), curated entity packs (entity_packs), Architecture Radar scorecards and entries (radar_scorecards,radar_entries), and a list of URLs (web). Seecorpusatlas/adapters/and docs/adapters.md for what each one handles. A new source is a class with onedocuments()method; you can copy one or ship yours as an external package via thecorpusatlas.adaptersentry-point group without forking this repo. - Output is files.
corpusatlaswritesgraph.jsonand exits. It owns no process, serves no requests, and has no opinion about what reads the output.
adapters → resolve → extract (deterministic, then curated) → merge → graph.json
No third-party dependencies. Python 3.11+.
pip install git+https://github.com/anothernoise/corpusatlas@v0.12.0
# Or with optional high-performance columnar extras:
pip install "corpusatlas[all]"
corpusatlas init --dir my-corpus # a starter config + entity registry
# edit my-corpus/corpusatlas.toml to point `path` at your notes, then:
corpusatlas build --config my-corpus/corpusatlas.toml --out graph.json --dry-run # counts, nothing written
corpusatlas build --config my-corpus/corpusatlas.toml --out graph.json
corpusatlas stats --graph graph.json
corpusatlas validate --graph graph.jsoninit scaffolds an obsidian source by default — point it at any directory
of markdown and it runs immediately, wikilinks or not. example.toml (below)
shows every adapter's shape but points at a real corpus's own directories, so
it isn't runnable as-is the way init's output is.
Iterating on a custom ontology or a registry doesn't need a corpus at all:
corpusatlas ontology-check --schema ontology.toml
corpusatlas ontology-check --schema ontology.toml --dot | dot -Tsvg -o ontology.svg
corpusatlas registry-check --entities entities.toml--dot renders the schema's own type/relation graph as Graphviz DOT — the
vocabulary itself, not any data built against it, useful while actually
designing one rather than only checking it's internally consistent.
On a large corpus, build --cache .corpusatlas_cache.json skips re-parsing
any html_blog/obsidian/logseq file whose content hasn't changed since
the last build at that cache path — opt-in, and never changes what gets
built, only how long it takes.
corpusatlas diff --old OLD.json --new NEW.json compares two built graphs —
nodes and edges added, removed or changed — and --fail-on-change turns
that into a gate: exit 1 if they differ at all, for a CI step asserting a
rebuild is a no-op (exactly the check this project's own release process
runs by hand before every release). stats, validate and diff all take
--json for machine-readable output instead of the default text.
Working on the module itself:
pip install -e . # installs the `corpusatlas` console script
python tests/run.py # invariant tests, stdlib onlypython3 -m corpusatlas … works too, without installing — the console script
and the module entry point are the same main().
[[sources]]
type = "html_blog" # adapter name
path = "../blog"
url_base = "/blog/"
[[sources]]
type = "radar_scorecards"
path = "../architecture-radar/scorecards.json"
[[sources]]
type = "radar_entries"
path = "../architecture-radar/radar.json"
[[sources]]
type = "entity_packs"
path = "entities"
url_base = "/knowledge-base/entities/"
[ontology]
entities = "entities.toml"
# Packs cite their sources by URL; the graph joins on document id. This says
# how to get from one to the other, and defaults to the values below minus
# site_url — a corpus published under /blog/ can leave the section out.
[packs]
site_url = "https://example.org"
blog_prefix = "/blog/"
assessment_prefix = "/architecture-radar/"
assessment_id_prefix = "assessment:"Every path resolves relative to the config file, so the config lives with the
corpus and this module stays portable. See example.toml.
Two kinds of node, and the split is the design:
- Entities are what the graph is about.
DEFAULT(this package's own, used unless a build says otherwise) has 14 types:Concept,ArchitecturePattern,Technology,Component,Language,API,Protocol,Standard,FileFormat,TableFormat,Product,CloudService,CompanyandUseCase. Ids are type-neutral (entity:<slug>), so reclassifying something never changes its id. - Context is where claims were published:
Document(an article),Assessment,RadarEntryandTopic.
Semantic relations (IMPLEMENTS, HAS_COMPONENT, READS, RUNS_ON,
MANAGED_BY, ALTERNATIVE_TO, ...) join entities. Context relations
(COVERS, HAS_RADAR_ENTRY, ASSESSES, ...) join entities to their
evidence. corpusatlas validate refuses a semantic edge that touches a
context node. The artifact carries entity_types and inverse_labels, so a
renderer never hard-codes the split.
The vocabulary itself is an Ontology value, not a fixed set of
constants. DEFAULT is a data-and-infrastructure-architecture ontology,
built for this package's own use, and every extractor, the resolver and
write_graph take an ontology= parameter and use it instead of reaching
for a module global — which is what lets a build actually swap it out. A
corpus about something else entirely — cooking, case law, whatever — gets its
own types and relations from a schema file:
# in your corpus's config
[ontology]
entities = "entities.toml" # unchanged: named instances of the types below
schema = "ontology.toml" # new: the types and relations themselvesSee docs/ontology-schema.md for the file format and a complete worked example, and docs/DESIGN.md for why the id scheme and the three extraction tiers stay fixed regardless — those are structural, not vocabulary.
The entity registry (entities.toml) is the vocabulary. It gives each
entity a name, type, aliases and links. It also lists the scorecard options,
radar entries and blog tags that name the entity, so one option can name two
entities ("Druid / Pinot"):
[[entity]]
id = "apache-druid"
name = "Apache Druid"
type = "Technology"
options = ["Apache Druid", "Druid / Pinot"]
case_sensitive = true # text matching respects caseThree tiers run in order, and merge keeps the first writer. A later tier can add links, aliases and descriptions to an entity, but never change its label or type.
Deterministic. Links, tags, scorecards and radar entries, resolved through the registry. Exact, free and instant.
Curated: entity packs. One JSON file per subject: typed entities, and
typed relationships, each with a confidence and an explanation. Packs are
drafted offline and machine-checked. Nothing in the build calls a model. A
pack is yielded as a Document, so every curated edge names its pack, and
deleting the file retracts all of it. An ill-typed relationship, or one that
contradicts the registry, fails the build.
Extracted: mentions. Word-boundary matching of every entity's name and
aliases over article text, with a minimum-occurrence bar. Companies are not
matched. It produces COVERS edges tagged extracted. No model; the output
is byte-identical on every run.
Every edge carries where it came from:
{
"src": "entity:clickhouse", "rel": "COMPARES_TO", "dst": "entity:starrocks",
"prov": { "doc": "starrocks-vs-clickhouse-vs-doris", "tier": "deterministic",
"extractor": "deterministic@2", "via": "scorecard" }
}An edge is a claim made by a document. When the document goes, the claim goes
with it — merge.py handles retraction, not just append.
viewer/index.html provides a zero-build, dependency-free interactive WebGL knowledge graph viewer designed for exploratory analysis and documentation embeds:
- Full-Text Fuzzy Search & Typeahead: Subsequence fuzzy matching with typo tolerance, prefix boosts, and word-boundary ranking across entities, descriptions, and aliases.
- 3D Force-Directed Graph Mode: Seamlessly switch between 2D canvas and 3D WebGL force-directed space for complex topology exploration.
- Vector & Subgraph Export: One-click export to scalable vector graphics (SVG), high-res PNG, Gephi GEXF 1.2 XML, GraphML XML, and filtered JSON subgraphs.
- Temporal Graph Evolution Player: Time-travel scrubber animating graph growth and entity arrivals across historical dates and snapshots.
- Topological Clusters & Centrality Insights: Community overlays (Louvain modularity) and on-demand hub rankings (Degree, PageRank, Betweenness).
- Smooth Navigation & Minimap: Camera history (back/forward), calibrated minimap with pan tracking, and directional particle flow animations.
corpusatlas build --config your-corpus.toml --out viewer/graph.json
python3 -m http.server 8000 --directory viewer
# open http://localhost:8000/(example.toml's paths point at a real corpus's directories, so it won't
build on its own inside a clone of this repo — swap in your own config, or
pass ?graph=<url> to point the viewer at any already-built graph.json,
shirokoff.ca's included.)
How far this scales. Measured against synthetic graphs (dense, uniformly random edges — worse than a real corpus's natural clustering): 1,000 entities loads in ~2s, 2,000 in ~4s, 5,000 in ~12s, 10,000 in ~29s. Search and click-to-focus stay responsive at every size tested — it's only the one-time ForceAtlas2 layout that gets slower, not runtime interaction afterward. Shirokoff.ca's own graph (214 entities) is nowhere near where this starts to matter. A real fix for a much larger graph would be precomputing layout coordinates at build time instead of laying out in the browser on every load; not done here since nothing this module's own consumer has needs it yet.
Ship any graph as a self-contained HTML custom element — no build step, no framework:
<script type="module" src="https://unpkg.com/@corpusatlas/graph@0.12.0/embed.js"></script>
<corpusatlas-graph data-src="graph.json"></corpusatlas-graph>Or install via NPM:
npm install @corpusatlas/graphSee npm/corpusatlas-graph/README.md for React, Next.js, and framework integration examples.
Embed interactive graphs directly in MkDocs documentation using fenced code blocks:
```corpusatlas
src: graph.json
height: 600
theme: dark
```The plugin auto-transforms fences into <corpusatlas-graph> web components. Add to mkdocs.yml:
plugins:
- corpusatlasExport your knowledge graph as a native Obsidian vault or canvas:
# Interactive JSON Canvas (.canvas) for Obsidian Canvas view:
corpusatlas export-obsidian --graph graph.json --format canvas --out graph.canvas
# Markdown vault with [[wikilinks]] and YAML frontmatter:
corpusatlas export-obsidian --graph graph.json --format vault --out-dir obsidian-vault/Automatically extract entities and relationships from unstructured text using any OpenAI-compatible or Ollama endpoint:
corpusatlas extract-llm --input docs/ --endpoint http://localhost:11434 --model llama3 --out extracted.tomlThe output is a validated CorpusAtlas TOML semantic pack, ready to feed into the build pipeline.
Project vector embeddings to 2D/3D coordinates using pure-Python PCA (no numpy required):
corpusatlas project-embeddings --graph graph.json --dimensions 2 --out projected.jsonThe viewer's "Semantic Space" layout renders these coordinates directly.
Enforce role-based visibility and redact sensitive nodes before publishing:
corpusatlas redact --graph graph.json --role public --policy rbac.toml --out public_graph.jsonDrops unauthorized nodes, sanitizes dangling edges, masks private attributes, and writes an audit log.
For corpora with hundreds of thousands of claims, corpusatlas build supports out-of-core deduplication and staging engines:
# In-memory deduplication (default):
corpusatlas build --config corpusatlas.toml --out graph.json --store memory
# Disk-backed SQLite staging (bounded memory footprint):
corpusatlas build --config corpusatlas.toml --out graph.json --store sqlite --db-path staging.db
# DuckDB-backed analytical staging:
corpusatlas build --config corpusatlas.toml --out graph.json --store duckdb --db-path staging.duckdbCorpusAtlas provides first-class, zero-copy interoperability with Polars, Apache Arrow, and Pandas (pip install corpusatlas[all]):
import corpusatlas as ca
# Load graph into tabular dataframes
graph_data = ca.load_graph("graph.json")
# Zero required dependencies — standard dict records
records = graph_data.to_dict_records()
# Optional high-performance columnar formats:
nodes_pl, edges_pl = graph_data.to_polars()
nodes_pa, edges_pa = graph_data.to_arrow()
nodes_df, edges_df = graph_data.to_pandas()Tabular sources can also be ingested directly via DataFrameAdapter:
from corpusatlas.adapters.dataframe import DataFrameAdapter
adapter = DataFrameAdapter({"nodes": my_nodes_df, "edges": my_edges_df})
docs = list(adapter.documents())Beyond building static graphs, CorpusAtlas provides a comprehensive zero-dependency backend toolkit:
Extract
# Markdown prompt context for LLMs (default):
corpusatlas rag --graph graph.json --query "Apache Spark" --k-hops 2 --token-budget 2000
# JSON structured context:
corpusatlas rag --graph graph.json --query "Spark" --format jsonTranslate plain-English questions into graph traversals, shortest paths, comparisons, dependencies, and hub rankings:
corpusatlas query --graph graph.json --ask "What technologies connect to Apache Spark?"
corpusatlas query --graph graph.json --ask "How is Airflow connected to Kafka?"
corpusatlas query --graph graph.json --ask "Compare PostgreSQL and Spark"Query entities, types, and relationships using standard SPARQL pattern matching directly over the graph:
corpusatlas sparql --graph graph.json --query "SELECT ?tech ?db WHERE { ?tech a :Technology . ?tech :writes_to ?db }"Compute network topology metrics, component sizes, articulation points (single points of failure / SPOF via Hopcroft-Tarjan), and an overall health score (0–100):
corpusatlas analyze --graph graph.json
corpusatlas analyze --graph graph.json --jsonEnforce architectural boundaries, forbidden cross-tier edges (e.g. Presentation
corpusatlas lint --graph graph.json --rules architecture_rules.tomlExpose your knowledge graph to AI coding assistants (Claude Desktop, Cursor, Gemini Antigravity) over standard I/O:
corpusatlas serve-mcp --graph graph.jsonAvailable tools:
search_nodes(query, type_filter): Substring and alias search.extract_context_ppr(entity, top_k, alpha): Personalized PageRank ego-network retrieval.traverse_subgraph(start_node_id, depth, allowed_rels): Bounded BFS expansion.audit_graph(): Topological health, bridge detection, and cycle metrics.get_provenance(edge_id): Source document audit trail and confidence.
Combines lexical keyword relevance (Okapi BM25) with topological graph centrality (Personalized PageRank) using Reciprocal Rank Fusion:
corpusatlas context --graph graph.json --entity "Apache Spark" --algorithm hybrid --top-k 20Clusters the knowledge graph into thematic communities by maximizing modularity (
corpusatlas cluster --graph graph.json --resolution 1.0 --out clustered_graph.jsonFilter graphs to point-in-time states using valid-time (valid_from, valid_to) intervals and transaction time:
corpusatlas as-of --graph graph.json --date "2024-01-01" --out snapshot_2024.jsonIn the interactive viewer, a timeline scrubber slider provides real-time playback of graph evolution over time.
Measure core engine throughput across Barnes-Hut layout, Aho-Corasick scanning, and Datalog fixpoint evaluation:
corpusatlas benchmark --quick
corpusatlas benchmark --jsonRun topological graph health checks to identify cycle anomalies, bridge bottlenecks, hub concentration, and contradictory relations:
corpusatlas audit --graph graph.jsonComputes:
- Degree Gini Coefficient: Measure network centralization and power-law distribution.
- Bridge Edge Detection: Single points of topological failure using Tarjan's bridge algorithm.
- Hierarchical Cycle Detection: Detects invalid circular taxonomic/hierarchical relations.
- Contradiction Detection: Flags conflicting reciprocal relationships between identical entities.
Refactor entity types and relationship names on an existing graph.json without re-running long corpus builds:
corpusatlas migrate --graph graph.json --migration migration.toml --out migrated.jsonFor massive graphs, partition 2D layout coordinates into a multi-resolution quadtree pyramid for client-side streaming:
corpusatlas tile --graph graph.json --out-dir tiles/ --max-zoom 3Emits tiles/{z}/{x}_{y}.json and manifest.json. The web viewer automatically streams visible bounding-box tiles on zoom and pan.
Export graph.json for analytical SQL engines, columnar storage, and graph databases:
# Export to DuckDB database, Apache Parquet, CSV, and SQL import script:
corpusatlas export --graph graph.json --format duckdb --out-dir duckdb_export/
# Export to Cypher statements for Neo4j / AWS Neptune:
corpusatlas export --graph graph.json --format cypher --out graph.cql
# Export to W3C RDF Turtle for SPARQL triple-stores:
corpusatlas export --graph graph.json --format turtle --out graph.ttlWhen DuckDB CLI is installed on the host system, export --format duckdb automatically invokes DuckDB to build native .duckdb and .parquet files directly without requiring any third-party Python pip dependencies!
Deduce implicit transitive, symmetric, or custom relational rules using forward-chaining fixpoint evaluation:
from corpusatlas.datalog import DatalogEngine, Rule
engine = DatalogEngine(rules=[
Rule(head_rel="PART_OF", body_rels=("PART_OF", "PART_OF"), confidence_factor=0.9)
])
inferred_edges = engine.evaluate(edges)Cluster and deduplicate near-identical entities using character
from corpusatlas.minhash import MinHashLSH
lsh = MinHashLSH(threshold=0.8)
clusters = lsh.cluster_entities(nodes)Enforce domain/range constraints, required properties, regex patterns, and cardinality bounds on knowledge graphs:
from corpusatlas.shacl import SHACLValidator
validator = SHACLValidator.from_toml("shapes.toml")
results = validator.validate(nodes, edges)Resolve local entity nodes to Wikidata QIDs, concept descriptions, and Wikipedia links with offline fallback caching:
from corpusatlas.link import link_entity_to_wikidata
link_info = link_entity_to_wikidata("Apache Spark")
# {'qid': 'Q3011409', 'label': 'Apache Spark', 'url': 'https://www.wikidata.org/wiki/Q3011409', ...}Pure Python embedding vector store with nearest neighbor search and semantic similarity edge materialization (SIMILAR_TO):
from corpusatlas.vector import VectorIndex
vindex = VectorIndex()
vindex.add("entity:spark", [0.12, 0.85, ...])
similar = vindex.nearest_neighbors("entity:spark", top_k=5)Serve your knowledge graph over standard HTTP with multi-threaded request processing, CORS preflight headers, and ETag caching (304 Not Modified):
corpusatlas serve-api --graph graph.json --port 8080Endpoints:
GET /: Health check, API version, and graph metadata.GET /stats: Breakdown of node and edge types.GET /nodes?query=...&type=...&limit=...: Filter and paginate entities.GET /nodes/<node_id>: Node details with inbound and outbound relationships.GET /edges?src=...&dst=...&rel=...: Filter relationships.GET /context?entity=...&algorithm=ppr|bfs|hybrid&format=json|markdown: Graph RAG ego-network context extraction.GET /audit: Topological health checks (Gini coefficient, bridges, cycles).
graph.json stays the artifact build writes — it's what the browser reads
with zero transformation, and every other format here is a converter over
it, not a second thing the pipeline produces. convert never re-runs
extraction, so it works on any graph.json this module ever wrote:
corpusatlas convert --graph graph.json --format graphml --out graph.graphml
corpusatlas convert --graph graph.json --format csv --out-dir csv/
corpusatlas convert --graph graph.json --format neo4j --out-dir neo4j/
corpusatlas convert --graph graph.json --format turtle --out graph.ttlGraphML opens directly in Gephi, yEd, NetworkX (nx.read_graphml) or igraph
— real graph analysis tools, not a browser view. The CSV pair
(nodes.csv/edges.csv) is for reach rather than fidelity: pandas, a
spreadsheet, anything with a CSV reader. neo4j writes the same fields under
Neo4j's own :ID/:LABEL/:START_ID/:END_ID/:TYPE header convention —
loads directly with neo4j-admin database import full, no column-mapping
step first; a node's type becomes its Neo4j label, an edge's rel becomes
its relationship type. All three keep the same lean column set — a node's
label, type, degree and url; an edge's relation, confidence, explanation and
scope — rather than trying to be the complete record; graph.json still is
that, provenance and all.
turtle is the odd one out: unlike the three above, it isn't a reformat of
the same fields — RDF has no single obvious mapping for a typed, provenanced
graph like this one, so it's a real modelling decision, not a converter.
Every node becomes a resource under --base (a placeholder by default —
pick your own for anything meant to be dereferenced); a node's ontology type
becomes its rdf:type, and an edge's relation becomes the predicate
directly (ca:IMPLEMENTS, not a generic "relatedTo"). Confidence,
explanation, scope and provenance can't ride on a plain triple, so an edge
carrying any of them is written twice — once as the bare triple a simple
SPARQL query expects, once as a standard rdf:Statement reification
carrying the metadata — chosen over RDF-star or singleton properties for
being the more broadly compatible option. See corpusatlas/rdf_export.py
for the full reasoning. Worth doing because something now needs to SPARQL
this graph; wasn't worth doing speculatively before that was true.
-
Barnes-Hut
$O(N \log N)$ Quadtree Layout: Precalculates 2D force-directed node coordinates using adaptive quadtree spatial decomposition (--layout), scaling to tens of thousands of nodes in seconds. -
Aho-Corasick Linear-Time Mention Scanner: Deterministic string automaton matching across vocabulary keywords with word-boundary checks in
$O(L + M)$ time. -
SQLite Out-of-Core Pipeline Store: Provides memory-bounded relational ingestion, deduplication, document retraction, and JSON streaming via
SQLitePipelineStore. - Merkle-DAG Incremental Cache: Content-addressable SHA-256 tree over extractor tier claims, invalidating only touched documents on rebuilds.
shirokoff.ca/knowledge-base builds its
graph with this module in CI on every push: ~290 articles, 16 scored
assessments and a set of curated entity packs become one graph.json, rendered
as an interactive canvas and as a
plain-HTML list. The site holds
the corpus, the config and the entity registry; this repo holds the engine.
It grew inside that site's repository and was split out with
git subtree split, so the history below predates this repository.
See CONTRIBUTING.md. CHANGELOG.md has the version history; docs/adapters.md lists all eleven sources side by side.
MIT.
