Author: Bhuvan Rajpoot, Independent Researcher, United Kingdom
Contact: bhuvanrajpoot.codes@gmail.com
A CPU-first research implementation of adaptive-depth inference. Small models learn to stop early on easy examples and continue computing on harder examples, while all baselines are compared with matched analytical compute and measured latency.
The supported publication pipeline contains two experiments:
- a leakage-free synthetic pointer task with hidden reasoning depth;
- a four-exit BERT-mini model fine-tuned on BoolQ.
The result is an accuracy-versus-compute improvement. It is not a claim that dynamic execution is always faster than a static model.
Hidden-depth pointer reasoning
The primary synthetic result averages independent seeds 42, 43, and 44. Each test split has 12,000 examples. Every input requests 12 recurrent steps, so true path length is unavailable to the model and router.
| Split | Policy | Accuracy (mean +/- SD) | Mean depth | Total MFLOPs | Reduction vs depth 12 |
|---|---|---|---|---|---|
| ID | Fixed depth 12 | 100.00 +/- 0.00% | 12.00 | 0.880 | -- |
| ID | Sequential halt | 100.00 +/- 0.00% | 5.50 | 0.630 | 28.38% |
| OOD | Fixed depth 12 | 100.00 +/- 0.00% | 12.00 | 0.880 | -- |
| OOD | Sequential halt | 100.00 +/- 0.00% | 7.42 | 0.706 | 19.77% |
The OOD split extends path length from 1-8 to 1-12, increases active nodes, and uses held-out distractor cycles. Selected depth correlates with hidden path length at 1.000 Spearman ID and 0.998 OOD.
Training uses four deterministic curriculum stages containing 6,000, 16,000, 32,000, and 64,000 generated examples. The shared model has 29,401 parameters and the sequential router has 201 parameters. A 28,548-parameter untied control and an answer-loss-only ablation retain the main routing effect.
prajjwal1/bert-mini is fine-tuned for four epochs with a shared classification head at encoder depths 1-4. Model training, model validation, router training, calibration, and locked testing are disjoint.
| Policy | Accuracy | Mean depth | Total MFLOPs | Batch-1 ms | Batch-32 examples/s |
|---|---|---|---|---|---|
| Majority class | 62.17% | 0.00 | 0.00 | -- | -- |
| Fixed depth 3 | 65.20% | 3.00 | 654.31 | 5.233 | 252.8 |
| Fixed depth 4 | 65.29% | 4.00 | 872.42 | 6.920 | 192.1 |
| Random, matched compute | 63.91% | 2.94 | 641.77 | -- | -- |
| Confidence, matched compute | 65.20% | 2.95 | 644.37 | 5.714 | 228.9 |
| Dynamics-only halt | 62.17% | 1.00 | 218.10 | 2.330 | 707.3 |
| Uncertainty-aware halt | 65.29% | 2.94 | 642.24 | 7.192 | 236.6 |
The uncertainty-aware router matches full-depth accuracy on all 3,270 locked test examples while reducing analytical model-plus-router compute by 26.38%. Its paired accuracy difference from depth 4 is 0.00 percentage points, 95% CI [-0.55, +0.52], McNemar p=1.0. It beats random matched routing by 1.38 points (p=0.0177) but is statistically tied with confidence routing and static depth 3.
Dynamic routing is 1.04x slower than depth 4 at batch size one and 1.23x faster in batch-32 throughput. Static depth 3 remains the strongest measured systems baseline.
On 15 July 2026, all six experiments were rerun from scratch after upgrading to the audited dependency stack. Every headline accuracy and compute value, primary route choice, and primary correctness value remained unchanged. Independent secure-stack checkpoint replays for synthetic seed 42 and BoolQ then matched 29 deterministic artifacts byte-for-byte. Exact BoolQ hash replay uses the configured four-thread runtime; externally forcing OMP_NUM_THREADS=1 or MKL_NUM_THREADS=1 preserves headline metrics and routes but changes low-order confidence values. Migrating from the historical stack introduced sub-micro floating-point differences and small changes in auxiliary policies, so byte equality is claimed only within the pinned stack. Across three timing sessions, the sequential/full batch-one ratio ranges from 0.99x to 1.37x and batch-32 throughput from 1.07x to 1.27x; static depth 3 remains faster throughout. See reproducibility_check.json.
The frozen, audited publication environment uses:
- 64-bit Windows or Linux;
- Python
3.11(3.11.6used for reported runs); - CPU execution with four PyTorch threads; no GPU is required;
- at least 8 GB RAM and approximately 3 GB free disk space;
- internet access only for the first BoolQ/model download.
Direct package versions are pinned in requirements.txt. The main versions are PyTorch 2.13.0+cpu, NumPy 2.4.6, Transformers 5.13.0, Tokenizers 0.22.2, and Datasets 5.0.0. On 15 July 2026, pip-audit 2.10.1 reported no known vulnerabilities for the resolved project environment against both PyPI and OSV.
python -m venv .venv
.\.venv\Scripts\python.exe -m pip install --upgrade pip==26.1.2
.\.venv\Scripts\python.exe -m pip install -r requirements.txt
.\.venv\Scripts\python.exe -m pip install -e . --no-deps --no-build-isolation
.\.venv\Scripts\python.exe -m pip check
.\.venv\Scripts\python.exe -m unittest discover -s tests -vpython3.11 -m venv .venv
.venv/bin/python -m pip install --upgrade pip==26.1.2
.venv/bin/python -m pip install -r requirements.txt
.venv/bin/python -m pip install -e . --no-deps --no-build-isolation
.venv/bin/python -m pip check
.venv/bin/python -m unittest discover -s tests -vCreate the environment without --system-site-packages; global packages can
otherwise introduce dependency conflicts that are unrelated to this project.
The pinned torch==2.13.0+cpu wheel targets the official PyTorch CPU index. macOS users should install the platform's CPU build of PyTorch 2.13.0 before installing the remaining requirements.
The synthetic smoke test is offline after dependency installation and completes quickly:
.\.venv\Scripts\python.exe -m compute_adaptive.latent_experiment --smokeThe BoolQ smoke test downloads the pinned data and backbone revisions on first use:
.\.venv\Scripts\python.exe -m compute_adaptive.boolq_experiment --smokeThe following command trains all three primary synthetic seeds, both synthetic ablations, and BoolQ from scratch; regenerates the aggregate; and checks deterministic CSV/JSON artifacts byte for byte against results/:
.\.venv\Scripts\python.exe -m compute_adaptive.reproduce `
--output-root reproduced_results `
--artifact-dir reproduced_artifacts `
--reference-root resultsOn the measured CPU, the complete suite takes roughly 90-120 minutes. Use --skip-boolq for the five synthetic runs only. The output directory must be empty or absent; use --resume after an interrupted run.
Accuracy, predictions, route choices, training histories, policy sweeps, paired statistics, and manifests must match exactly under the pinned environment. Latency tables, plots, wall-clock training time, resolved output paths, and system metadata are intentionally excluded from byte-level comparison because they are machine dependent.
For individual commands, seed schedules, split hashes, and statistical details, see docs/REPRODUCIBILITY.md.
The packaging command creates a minimal source archive containing only the paper source, bibliography, and required figure:
.\.venv\Scripts\python.exe -m compute_adaptive.arxiv_package --root ..venv/bin/python -m compute_adaptive.arxiv_package --root .The upload-ready archive is written to
dist/compute-adaptive-routing-arxiv.zip. Build and inspect it with arXiv's
pdflatex processor before completing the submission. The exact form metadata
are recorded in paper/ARXIV_METADATA.md.
configs/ Frozen experiment configurations
docs/ Data, literature, audit, and reproduction notes
paper/ LaTeX source and compiled paper
src/compute_adaptive/ Models, routers, training, metrics, and CLIs
tests/ Unit and live/offline equivalence tests
results/ Frozen publication tables, reports, and plots
artifacts/ Local model checkpoints; ignored by Git
data/cache/ Downloaded/tokenized data; ignored by Git
dist/ Generated wheels/arXiv archive; ignored by Git
Each frozen run includes its resolved configuration, data manifest, model/router histories, policy sweeps, per-example predictions, paired comparisons, latency observations, plots, and result report. Headline values are generated from the run directories into results/publication_summary/RESULTS.md.
- The pointer transition is deliberately task aligned; the synthetic benchmark tests routing and stopping, not general algorithm discovery.
- BoolQ is one English binary reading-comprehension task and has one full model-training seed.
- The 371-parameter natural router does not outperform confidence routing or static depth 3 significantly.
- Analytical FLOPs omit common embedding and small-head costs.
- Dynamic batching and tensor slicing can erase analytical savings on real hardware.
- The pretrained backbone inherits limitations and biases from its training data.
This is a research prototype, not a production decision system. Repository
source code is licensed under the MIT License; third-party and paper
licensing are documented in NOTICE.md. The planned arXiv paper
license is CC BY 4.0 and must be confirmed by the author in the submission
portal. After announcement, add the arXiv identifier and DataCite DOI to
CITATION.cff and this README.
Until an arXiv identifier is assigned, cite the software release as:
@software{rajpoot2026computeadaptive,
author = {Bhuvan Rajpoot},
title = {Compute-Adaptive Transformer Routing on Limited Hardware},
year = {2026},
version = {0.2.0},
url = {https://github.com/bhuvan-01/compute-adaptive-routing}
}Machine-readable citation metadata are available in CITATION.cff.

