An Interactive Mechanistic Laboratory for Understanding Recurrent Memory and the Dragon Hatchling (BDH) Architecture
A Submission for DataForge 2026 – Pathway Track: "Explain the Frontier"
Approved Concept: In-Context Learning with Recurrent Memory
"A fixed-size recurrent state can carry task-relevant information forward without growing a token-by-token memory, but compressing information into that state creates interference and forgetting."
Every computation, interactive control, differential matrix view, and failure curve in this laboratory is implemented to allow learners and reviewers to directly observe, test, and verify this mathematical trade-off through live, deterministic execution.
Modern large language models achieve in-context learning by appending keys and values for every processed token into an expanding KV-cache. This design yields an unbounded spatial complexity of
Recurrent memory architectures replace the growing token cache with a fixed-size internal state space
Because the physical dimensionality
In standard transformer KV-caches, retrieval accuracy across long prompts is protected because tokens are stored uncompressed; however, memory footprint grows linearly without bound ($O(T)$). Recurrent memory makes the opposite trade-off: it achieves strict
The Dragon Hatchling (BDH) architecture, introduced by Pathway Research (arXiv:2509.26507), offers a brain-inspired resolution to this tension. BDH mathematically transforms transformer-style attention into recurrent synaptic fast weights (
BDH-CQ (arXiv:2608.09888) extends this architecture to recurrent multi-step reasoning. In standard transformer models, multi-step problem solving requires emitting intermediate tokens (Chain of Thought), further exacerbating the KV-cache bottleneck. BDH-CQ conducts multi-round reasoning directly within its fixed-size latent state without generating textual scratchpads, preserving bounded spatial memory while executing complex multi-hop inferences.
-
Advantages: Strict
$O(1)$ memory footprint; constant inference latency; zero VRAM cache growth; suitability for continuous streaming and embedded deployments. - Disadvantages: Irreversible lossy compression; vulnerability to distractor interference; exponential state decay under aggressive retention discounting; inability to reliably reconstruct exact verbatim transcripts when capacity limits are breached.
This laboratory demonstrates these mathematical principles via deterministic, client-side linear algebra. Learners observe exact matrix updates (
- Target Audience: Machine learning engineers, computer science students, computational neuroscientists, and AI researchers investigating post-transformer memory architectures and in-context learning.
- Prerequisites:
- Linear algebra fundamentals: vectors, inner products, outer products, matrix-vector multiplication, cosine similarity.
- Familiarity with the Transformer KV-cache and autoregressive inference.
- Key Question Addressed: Can an AI model maintain working memory without storing every historical token indefinitely?
Upon completing this interactive laboratory, learners will be able to:
- Differentiate Memory Complexities: Contrast the unbounded spatial footprint of growing token contexts ($O(T)$) with fixed-size recurrent states ($O(1)$).
-
Trace the Mechanistic Pipeline: Step through the end-to-end flow:
$\text{INPUT} \to \text{ENCODE} \to \text{WRITE} \to \text{MEMORY STATE} \to \text{QUERY} \to \text{READ} \to \text{RETRIEVE} \to \text{PREDICTION}$ . -
Analyze Matrix Updates: Inspect real cell-by-cell matrix differences (
$\Delta M = M_{t+1} - M_t$ ) and evaluate mean/max$|\Delta M|$ statistics. -
Induce and Explain Memory Failure: Systematically provoke associative interference by reducing dimensionality (
$D=4$ ) or increasing distractor updates, measuring the collapse in Top-1 Margin. - Distinguish BDH from SSMs: Articulate why BDH is not a continuous 1D SSM (such as Mamba), but a discrete graph of non-negative sparse activations with synaptic fast-weight plasticity.
- Understand Recurrent Latent Reasoning (BDH-CQ): Explain how recurrent internal relaxation cycles replace external autoregressive token generation for complex multi-hop queries.
The application is structured as a client-side, browser-native research workspace built with React 18, TypeScript, and Tailwind CSS. It requires no external backend, API keys, or remote database connections.
src/
├── components/
│ ├── MemoryWriteRead.tsx # 7-stage end-to-end mechanistic pipeline inspector
│ ├── MatrixDiff.tsx # Cell-by-cell before/after/delta matrix difference view
│ ├── MemoryOverwriteExperiment.tsx # Sequential write collision experiment (Japan->Tokyo vs Osaka)
│ ├── ContextOrderExperiment.tsx # Order-permutation ablation (Sequence A vs Sequence B)
│ ├── MemorySurgery.tsx # Controlled counterfactual causal probe (omit single write)
│ ├── RepresentationInspector.tsx # Latent representation vector inspector (min, max, L2 norm)
│ ├── StatePersistenceExperiment.tsx# Persistence vs intervening distractor decay curve
│ ├── CapacityExperiment.tsx # Dimensionality vs accuracy capacity boundary sweep
│ ├── InterferenceMap.tsx # 2D cross-talk interference heatmaps
│ ├── EvidenceLadder.tsx # 4-level formal epistemic hierarchy drawer
│ ├── Section01Hero.tsx # Break the Memory: interactive live toy preview
│ ├── Section02TransformerProblem.tsx# O(T) KV-cache explosion analysis
│ ├── Section03RecurrentMemory.tsx # Core recurrent state mechanics & memory inspector
│ ├── Section04InterferenceLab.tsx # Experimental suite: pipeline, collisions, decay, sweeps
│ ├── Section05FindTheFailure.tsx # Interactive challenge: systematically break the memory
│ ├── Section06RecoveryStrategies.tsx# Capacity scaling, orthogonalization, and sparsity
│ ├── Section07BiologicalMemory.tsx # Neuroscience parallels: synaptic plasticity & working memory
│ ├── Section08BDHArchitecture.tsx # BDH microscope, synaptic plasticity, sparsity inspector
│ ├── Section09BDHCQReasoning.tsx # BDH-CQ recurrent latent reasoning vs Chain-of-Thought
│ ├── Section10FinalChallenge.tsx # Synthesis challenge & parameter tuning sandbox
│ ├── Section11JudgeMode.tsx # Comprehensive criteria evaluation dashboard
│ └── navigation/ResearchNav.tsx # Accessible top navigation bar & index drawer
├── lib/
│ ├── associativeMemory.ts # Deterministic linear algebra engine & PRNG substrate
│ └── math.ts # Frobenius norm, cosine similarity, vector arithmetic
The educational simulation implements an associative vector-symbolic recurrent memory substrate:
-
Deterministic Representation: Each discrete symbol (country, capital) is mapped to a unit-norm vector in
$\mathbb{R}^D$ using a seeded pseudo-random basis generator (mulberry32), ensuring 100% deterministic, reproducible latent representations. -
Associative Outer-Product Write:
$$M_{t+1} = \lambda M_t + \eta , (v_t k_t^T)$$ where$M \in \mathbb{R}^{D \times D}$ ,$\lambda \in [0.1, 1.0]$ is retention, and$\eta \in [0.1, 1.0]$ is write strength. -
Linear Readout:
$$r = M_t q$$ where$q \in \mathbb{R}^D$ is the query vector and$r \in \mathbb{R}^D$ is the reconstructed value vector. -
Candidate Retrieval & Scoring:
The retrieved vector
$r$ is scored against all vocabulary candidates$c_i \in \mathbb{R}^D$ via cosine similarity:$$\text{Retrieval Score}(c_i) = \frac{r \cdot c_i}{|r|_2 |c_i|_2}$$ - Top-1 Margin: $$\text{Top-1 Margin} = \text{Score}{(1)} - \text{Score}{(2)}$$ reflecting the separation between the top predicted candidate and the nearest competitor.
To maintain scientific integrity, all statements and visualizations in this project are explicitly labeled according to the following 4-level taxonomy:
| Level | Epistemic Label | Definition & Boundary |
|---|---|---|
| Level 1 | LIVE TOY COMPUTATION |
Calculated on-device in real time by the browser's JavaScript engine using deterministic linear algebra. |
| Level 2 | PUBLISHED RESEARCH |
Primary architectural derivations, equations, and specifications from published literature (Pathway Research / arXiv:2509.26507). |
| Level 3 | PUBLISHED BENCHMARK |
Empirical results reported in published papers on standardized benchmarks. |
| Level 4 | EDUCATIONAL INTERPRETATION |
Pedagogical schematics, conceptual frameworks, and simplified models designed to develop intuitive understanding. |
- No Production Claims: The educational recurrent memory toy is an illustrative pedagogical abstraction. It is not a trained multi-billion-parameter language model and does not run production checkpoints.
- No Calibrated Probabilities: Cosine similarities are strictly labeled as Retrieval Scores, accompanied by tooltips stating that scores measure geometric representation similarity, not calibrated statistical probabilities.
-
No Semantic Coordinates: Individual matrix cells
$M[i][j]$ and latent vector dimensions are described strictly as computational coordinate axes, not as human-interpretable semantic features. - No Fabricated Benchmarks: Toy observations are never claimed to reproduce published empirical benchmark scores from Pathway Research.
A central pedagogical requirement of the Pathway Track is clarifying the distinction between the Dragon Hatchling architecture and Mamba-style State Space Models:
| Feature | State Space Models (Mamba / S4 / S6) | Dragon Hatchling (BDH) |
|---|---|---|
| Mathematical Basis | Continuous 1D linear time-invariant ODEs: |
Discrete particle interaction across a scale-free graph |
| Discretization | Zero-order hold or bilinear transform along the sequence dimension | Discrete multi-round internal synaptic relaxation |
| Activation Space | Unconstrained signed real numbers ( |
Strictly non-negative activations ( |
| Sparsity | Dense or weakly gated 1D hidden states | High structural sparsity (e.g., small fraction of active neurons per round) |
| Memory Locus | Latent 1D state vectors updated via input-dependent |
Synaptic matrix ( |
| Multi-Hop Reasoning | Requires autoregressive token generation | Recurrent internal latent passes (BDH-CQ) |
The laboratory provides full interactive control over the mathematical parameters:
-
State Dimension (
$D$ ): Toggle between$D=4, 8, 16, 32$ . Low dimensions induce immediate geometric interference; higher dimensions restore subspace separability. -
Retention Rate (
$\lambda$ ): Continuous slider from$0.10$ to$1.00$ . Controls how rapidly earlier matrix entries decay when new updates occur. -
Write Strength (
$\eta$ ): Adjusts the magnitude of incoming outer-product updates. - Interference Ratio: Scales the magnitude and frequency of intervening distractor facts.
-
Deterministic PRNG Seeds: Presets for
seed=42,seed=77, andseed=12345guarantee that reviewers on different machines reproduce identical matrix values and retrieval curves.
-
Pedagogical Scale: The browser toy operates on small dimensions (
$D \le 32$ ) and synthetic associative tuples. It is designed to illustrate mathematical principles of interference and capacity, not to perform real-world natural language translation. - Linear Hebbian Formulation: The toy's write rule uses a standard Hebbian outer-product update. While this captures the foundational principles of synaptic fast weights, the published BDH architecture employs a four-round loop involving memory reads, synaptic reweighting, non-negative neuron activations, and inhibitory/excitatory graphs.
- No Direct Benchmark Equivalence: Observations made within the browser toy reflect the properties of this specific educational substrate and should not be generalized as universal empirical assertions about all recurrent architectures.
To reproduce key findings:
-
Observe the Interference Failure:
- Navigate to Section 04: Interference Laboratory.
- Select the 1. WHEN MEMORY COLLIDES tab.
- Set Dimension to
$D=4$ and Retention to$0.90$ . - Observe that writing 8 factual associations drives the Top-1 Margin below zero, resulting in incorrect retrievals due to subspace saturation.
-
Inspect the Matrix Difference:
- Switch to the 0. WRITE / READ PIPELINE tab.
- Click STEP THROUGH WRITE to view the exact cell-by-cell matrix update (
$\Delta M = M_{t+1} - M_t$ ). - Verify that the reported mean
$|\Delta M|$ and max$|\Delta M|$ match the visual heatmaps.
-
Test Context-Order Equivalence:
- Run the Context Order Experiment to test whether permuting the sequence of identical facts alters retrieval accuracy under the current retention parameter.
-
Demonstrate Recovery via Dimension Scaling:
- Navigate to Section 06: Recovery Strategies.
- Increase Dimension from
$D=4$ to$D=32$ . Observe how expanded orthogonal capacity restores the Top-1 Margin and eliminates retrieval errors.
This project is built with standard Vite and React 18 in TypeScript.
# 1. Clone the repository
git clone <repository-url>
cd memory-in-motion
# 2. Install dependencies
npm install
# 3. Start the local development server (runs on port 3000)
npm run dev
# 4. Typecheck and lint
npm run lint
# 5. Build for production
npm run build"This project was developed with AI assistance (Google AI Studio / DeepMind Gemini). The author is entirely responsible for understanding, verifying, mathematically auditing, and defending every component, equation, and claim in this submission."
- Codebase: 100% custom-written TypeScript and React code. No proprietary or closed-source libraries.
- Icons: Lucide React icons.
- Styling: Tailwind CSS with custom high-contrast dark palette.
- Dataset: Deterministic synthetic entity-relationship associative tuples (e.g., countries, capitals, elements) generated on-device via PRNG.
- License: Apache License 2.0.
- The Dragon Hatchling (BDH) Architecture:
Pathway Research Team, arXiv:2509.26507
https://arxiv.org/abs/2509.26507 - Official BDH GitHub Repository:
Pathway Research,pathwaycom/bdh
https://github.com/pathwaycom/bdh - BDH-CQ: In-Context Recurrent Latent Reasoning:
Pathway Research Team, arXiv:2608.09888
https://arxiv.org/abs/2608.09888 - Pathway Research Portal:
https://pathway.com/research/ - The Equations of Reasoning:
Pathway Research Explainer
https://pathway.com/research/the-equations-of-reasoning - NeurIPS Educational Resource Guidelines:
https://neurips.cc/Conferences/2026/CallforEducationalResources