DecPort tested whether input-specific probabilistic decision behavior can be transferred across
heterogeneous frozen language-model backbones without target labels, and whether a particular
learned frozen DecisionCore contributes reusable functionality after the backbone change.
Final scientific outcome: NO-SHIP for the original reusable learned-core hypothesis.
The preregistered nonlinear final gate did not support reusable learned
DecisionCoreportability in the tested setting. Separately, label-free input-specific decision-behavior distillation remains supported within that setting. Choice and Score carry most of the positive transfer result; Noul did not recover input-specific source behavior in either accepted stage. OOD effects are smaller and calibration is weaker.
Distinguishing Cross-Backbone Decision-Behavior Transfer from Frozen-Core Portability
Lorenzo Suffritti
Technical research article, 24 September 2026
Research conducted with the support of artificial intelligence (AI).
The article covers the full experimental progression through decision 0007. Citation metadata is
available in CITATION.cff, and the paper-specific overview is in
paper/README.md.
| Question | Result |
|---|---|
| Input-specific cross-backbone behavior transfer (RQ1) | Supported within the tested setting |
Reusable learned DecisionCore portability (RQ2) |
Not supported |
| Final preregistered ship gate | NO-SHIP |
| Audit | Passed |
| Same-environment reproduction | Bit-exact |
| Independent or cross-hardware replication | Not performed |
RQ1 compares correct-teacher transfer with a decision-type- and width-matched mismatched teacher. RQ2 asks the stronger question: whether the particular learned frozen core contributes reusable functionality beyond a matched arbitrary frozen core under the same teacher signal. Evidence for RQ1 does not imply evidence for RQ2.
The first accepted stage used the real released Open-Jev 2B source system, its scalar
Linear(2048 -> 1) core, five seeds, and three heterogeneous frozen targets. Label-free behavior
transfer was positive overall and was carried by Choice and Score. Noul failed. The matched random
core comparison also exposed the linear-core identifiability problem: the experiment provided no
evidence that the particular pretrained rank-one core mattered.
The preregistered final gate used a learned nonlinear 1,118,977-parameter source DecisionCore, a
weak rank-128 target adapter, a scale-matched random nonlinear core, mismatched-teacher and
untrained controls, and a target-specific distilled baseline. Criteria A–E tested whether the
learned core mattered, transfer was input-specific, effects survived distribution shift, external
behavior cleared the JevBench public-subset threshold, and the shared-core route had practical
utility. Criteria A, D, and E failed, producing NO-SHIP. The audit passed and a second execution
in the same hardware/software environment was bit-exact.
- Decision record
- Final report
- Accepted five-seed archive
- Mechanical verdict
- Audit
- Reproduction comparison
- Run provenance
The run and reproduction provenance both record the frozen protocol commit
d0aa96f4defc3ce202996be1f75d669b791d57fc and git_worktree_dirty:true. The accepted archive
verifies the frozen repository configuration and data hashes, but it is not described as a
clean-checkout guarantee.
RESEARCH.md maps decisions 0001–0007 and makes the progression from supervised
transfer to the final falsification gate explicit. The repository preserves three evidence levels:
benchmarks/pilots/— bounded pipeline checks, not benchmark evidence;benchmarks/diagnostics/— controlled studies used to refine the question, not accepted benchmark claims;benchmarks/accepted/— meaningful-scale multi-seed runs that passed audit and reproduction before acceptance.
Acceptance certifies the archived measurement, including its negative findings; it does not extend the claim beyond the recorded models, datasets, tasks, and training budget.
Install the development and Open-Jev dependencies and run the ordinary package checks:
uv sync --extra dev --extra openjev
uv run ruff check .
uv run pytestReal-model tests are opt-in because they download checkpoints:
DECPORT_RUN_REAL_MODELS=1 uv run pytest tests/test_real_backbones.py tests/test_real_openjev.pyThe final report and verdict can be inspected without rerunning the expensive experiment. The accepted final-gate archive README records the exact data preparation, full run, audit, comparison, JevBench public-subset, and mechanical-verdict commands. The archived evidence includes configs, data manifests, per-seed results, aggregate tables, artifact hashes, audit failure-path probes, and reproduction results. Base-model checkpoints and downloaded datasets are intentionally not committed.
The v0.1 Python package remains a research implementation. Its basic API exposes frozen backbone wrappers, adapters, typed Choice/Noul/Score decision cores, experiment runners, and audit scripts. No DecPort v0.1 software release was made.
The conclusion is bounded to three target families, one Open-Jev/Qwen source representation boundary, the archived datasets and budgets, one nonlinear core design, and the preregistered criteria. It does not prove mathematical equivalence between learned and random cores, and it does not imply that nothing transferred.
Historical evidence remains in Git so the commit SHAs cited by the article stay valid. Future large publication artifacts should preferably use release assets or an archival deposition rather than rewriting existing provenance.
The code is licensed under the Apache License 2.0.