Skip to content

feat: advance the evidence-gated research core (R01/R02/R03/R04 batch) - #2

Open
SHi-ON wants to merge 554 commits into
Ethical-Tech-CoLab:mainfrom
SHi-ON:feat/verifiable-core
Open

SHi-ON wants to merge 554 commits into
Ethical-Tech-CoLab:mainfrom
SHi-ON:feat/verifiable-core

Conversation

@SHi-ON

@SHi-ON SHi-ON commented Sep 27, 2026

Copy link
Copy Markdown

Successor to closed #1 on rewritten history (plans/ purged, see comments). R01/R02/R03/R04.1 batch.

Bind the passing v3 corrective cases and cumulative six-case service-death coverage to immutable receipts. Keep uncertainty, detector, resource, B12, behavioral, and publication claims open.
Add bounded response-delay and client-deadline controls for prospective selected-topology fault qualification while leaving production defaults unchanged. Cover the commit-without-confirmation path, durable unresolved intent, terminal Gateway quarantine, and no-retry behavior.
Lock a single-use selected-topology development packet that delays commitTurn confirmation past the Gateway deadline, checks durable uncertainty and no retry, then recreates the Gateway to test fail-closed recovery.
Preserve the failed v1 startup as portable evidence and register a fresh v2 packet under the application-qualified signer namespace. Keep the fault schedule and safety thresholds unchanged.
Preserve the failed v2 recovery wrapper and register a fresh v3 supplement that removes every Gateway-owned Unix socket before process recreation. Bind the observed first-emission plus intention ledger accounting without changing the safety claim.
Preserve the failed v3 pre-execution build and add a temporary Alpine native-build toolchain so better-sqlite3 can compile when no prebuild is available. Remove build tools from the runtime image and bind the repaired image into a fresh v4 packet.
Preserve the failed v4 recovery wrapper and align uncertain-write recovery with the selected lifecycle restart set: both model adapters, both Baby hosts, the Gateway, and the Controller.
Preserve the completed-but-failed v5 wrapper with portable provenance, and register v6 with separate live uncertainty and reconstructed incomplete-turn expectations. No scientific threshold or execution behavior changes.
Record the bounded v6 result, add portable and live receipt reconciliation, and retain explicit exclusions for B12 and behavioral research claims.
Add a schema-gated delay seam and a single-use selected-topology packet proving that a model completion after the outer normalized deadline cannot add evidence or resume the attempt.
Preserve the v1 signed-budget mismatch as a failed qualification and register a fresh v2 identity whose run budget matches the unchanged one-second adapter deadline.
Seal the bounded v2 late-callback result with portable and live reconciliation while retaining explicit B12, detector, resource, and behavioral exclusions.
Freeze the outcome-blind LV01 design, analysis, and seed/resource policy with a mutation-tested contract gate.
Allow only the frozen LV01 identifier alongside legacy E## cards across configuration and integrity records, without widening experiment identifiers.
Generate balanced, domain-separated LV01 cases through a scoped interface while preserving the original scenario split contract.
Bind LV01 configurations to their separate design, analysis, resource, predictor, and partition contract without changing E16.
Restore LV01's exact recurrent receiver into an isolated scorer and reject any architecture or candidate shape outside the frozen study contract.
Select authenticated training-prefix associations by verified sequence and normalize them only over the current four receiver candidates.
Fit the five fixed ordinary-record predictors per receiver role, select only on the held validation fold, and reject target-bearing input.
Score actual receiver actions separately from task targets and apply the fixed equal-role, component, member, and Holm reductions.
Reduce blinded pilot dispersion, select only from the frozen candidate sizes, and qualify numerical power against an independent base-R reference.
Bind the synthetic power-selector qualification to its exact source commit and preserve its no-pilot, no-selection claim boundary.
SHi-ON added 25 commits October 1, 2026 12:47
- study-design v2 evaluation gains receiverDecisionsDerivation
  (5160 composition); new digest
  5784f7b5311148a30346b38962947818244a3f982ddeb04809b012a65aaf98f6.
- Checker pins the string (v2-only; v1 frozen untouched) + live
  arithmetic composition from partitions; v2+v1 both green.
- Manuscript digest lines (2) queued for HANDYM fill; LV02/LG01 pins
  stay refresh-at-activation. Review queued file-backed.
- Blocked-underpowered (nulls, joints < 0.90, reserves 16/20/24/29/37/52
  all adequate), smallest-qualifying-N (75, reserve 16), malformed
  pilot/seed fail-closed; file 5/5 green.
- v2 selector itself verified present/correct vs plan 4.6 (candidates,
  30k reps, 0.05/4 screening, 0.05/42 MC level, union power >= 0.90,
  Wilson reserves); R cross-check needs Rscript (absent; CI-owned).
- Review queued file-backed (HANDYM unreachable); no verdict claimed.
- Removed 8 unused bindings (launch 4, preflight POLICY pair now held
  by headroom module, lv01 assert import, frozen basename import);
  releaseLease call preserved (return dropped only).
- eslint.config.js no longer ignores scripts/**/*.mjs: all 124 files
  lint clean; A0 batch eslint stage no longer hollow.
- Regression: A0 4 files + lv01-cli 47/47 green. Repo-wide still has
  2 pre-existing errors in untouched isolation.test.ts (owner lane).
- Review queued file-backed (HANDYM unreachable); no verdict claimed.
- review-package: M-Q/M-POP/M-ACCESS/M-FOUND rows for sections 1-4; file:line pins on all M-rows; study-design digest a483f25f->5784f7b5 (G1 sum-string 7588e94, HANDYM-verified PASS)

- first-paper-draft: 1680 [240x7] decomposition; range unbound hedge; partitionSubdomains transcription; canonical main/repeat notes

- estimates-and-intervals: component-range unbound hedge. HANDYM lane; static methods only, no results language; no push
- New lv01-ordinary-wiring: builds the collector ordinary input from a
  real validation-window selection (selectLv01OrdinaryPredictor) with a
  folds-digest receipt; barrel-exported. Records stay caller inputs;
  parent-evidence record loader queued.
- 4 new tests (3 pure + 1 live slot collect with fitted provider);
  31/31 across both files, tsc + eslint clean.
- Review queued file-backed; no verdict claimed.
- New lv01-parent-training: slot-seeded RunConfig on SPEC 18 defaults,
  createRun + runToCompletion on caller runtime, sealed bundle export,
  seeds-binding receipt; barrel-exported. Train-collect-audit
  composition queued next.
- 3 live tests (valid bundle, distinct-slot binding, fail-closed);
  3/3 green, tsc + eslint clean.
- Review queued file-backed; no verdict claimed.
- Journal slots link their own trained parent (stageDir-relative ref,
  containment-checked); collect accepts parentBundleDirFor resolver;
  audit resolves per-slot parents with cache, legacy single-parent
  fallback intact (parentBundleDir now optional, fail-closed).
- 3 new tests (journal unit, 2-slot per-slot collect+audit green,
  forged-ref escape rejection); 33/33, tsc + eslint clean.
- Review queued file-backed; no verdict claimed.
- New train-lv01-slot-parents script: verified packet+binding,
  per-slot training on createProductionRuntime, receipts file;
  explicit training flags (no invented study defaults).
- collect gains --per-slot-parents (exactly-one with --parent-bundle);
  audit drops required --parent-bundle (refs suffice); dispatcher
  train-parents case + optional-flag forwarding.
- 3 CLI tests incl. full script-level train-collect-audit e2e
  (v999 fixtures, verified receipt); 11/11, eslint clean.
- Review queued file-backed; no verdict claimed.
- partitionLv01OrdinaryRowsBySchedule: caller rows into
  training/validationFit/validationSelection by schedule case
  membership; test-partition, unlisted, and duplicate caseIds
  fail closed. Row extraction from raw evidence stays a
  methods-defined mapping (questionnaire queued for HANDYM).
- 1 new test (assignment + 3 rejections); file 4/4, tsc + eslint clean.
- Review queued file-backed; no verdict claimed.
- lv01WorkloadForCollection accepts prototype allocation blocks only
  on exact match to LV01_FULL_WORKLOAD_COUNTS (3000/240/240/240,
  cited to study-design v2 partitions); dual-block and resized
  allocations fail closed; small-fixture path unchanged.
- Extended workload test (accept + 2 rejections); tsc + eslint clean.
- Review queued file-backed; no verdict claimed.
- tracked mirror of plans/drafts/lv01-row-extraction-mapping.v1.md (v1 frozen; plans/ is ignored so git-invisible)

- Q1-Q5 all determined, loader unblocked. HANDYM lane; no push
- extractLv01OrdinaryRecords: parent evidence joined to schedule (Q1-Q5, fail-closed)

- typedTurns/typedLedger export-only reuse; barrel export; 8-test suite incl. composition
- composeLv01OrdinaryProvidersFromParent: extract, per-role split, partition, wire
- per-episode-turn matching + scenarioStateHash join key (compatible with v1; single-turn identical)

- loader/composition verdicts recorded file-backed. HANDYM lane; no push
- v1.1 C3a: test-membership pre-check before ordinary lookup; collision test
- withdraw false builder-assertion-absent claim; assert verified present at lv01-schedule.ts:127-131

- C3a review PASS recorded file-backed. HANDYM lane; no push
- selectLv01OrdinaryProviderForCase: case receiver binds composed role fit, no substitution
- C2 join anchor: all-partition uniqueness + test-vs-ordinary non-overlap
- inventory-32 design source traced (SPEC S9.1 S01..S32); coordinated window after d3ca4ab

- HANDYM lane; no push
- channel.deliveredTokenInventory S01..S32 + 3 window keys; v2 checker asserts + key freezes
- §5.1 table + review-package rows rebind study-design.v2 (0259…) and
  analysis-plan.v2 (c401…) committed bytes; shifted line citations
  updated; design-only, no collection
- observationSchemas sender/receiver mirror; derivation.roleStreamRoleCodes; v2 pins
- §5.1 table + review-package rows rebind study-design.v2 (be1f…) and
  seed-policy.v2 (8556…) committed bytes; shifted line citations
  updated; design-preparation only, no lock, no collection
Reuse the frozen family decision for dyad-level component vectors and cover equivalence, known t/Holm answers, degenerate inputs, and invalid data. Add a resource-admitted timing recipe for subsequent cost measurement.

Validation: 19 focused tests passed across four suites; TypeScript build, affected-file lint, script syntax, and secret scan passed. No operating-characteristics campaign or collection was executed. Working version: 0.1.528.
Replace retired si fort instructions with the native fort command, preserve files-mode secret materialization, and name the public Fort host in runnable examples.

Validation: configuration and signer suites passed within the 19-test focused run; affected-file lint, TypeScript build, and secret scan passed. Native fort help confirms the documented command and host option. Working version: 0.1.529.
@SHi-ON

SHi-ON commented Oct 4, 2026

Copy link
Copy Markdown
Author

Pushed the latest work through 847cdb2 (0.1.529), including the earlier local commits. The final two chunks add the LV01 dyad-vector analysis entry with corrected tests and a timing recipe, then switch current configuration and runbooks to the native Fort CLI. Validation passed: 19 focused tests, TypeScript build, affected-file lint, script syntax, secret scan, and whitespace checks. The working tree is clean. Research collection and the timing recipe have not been run; empirical qualification and pilot milestones remain open.

SHi-ON added 4 commits October 4, 2026 00:33
…k runbook

Fresh unauthenticated API observation finds upstream PR Ethical-Tech-CoLab#2 open at
847cdb2 with its workflow action_required, and both upstream and fork
main reporting protected=false with zero rulesets, so merge-blocking is
affirmatively unconfigured. Adds the exact admin sequence to approve the
PR run, require consolidated-suite + mode-r on main, and demonstrate
blocking. Working version: 0.1.529.
Second-person runbook execution by ALDM (not the ALD-060 implementer):
scripts/run-ald079-independent-restore.mjs follows Preparation 1-5 and
Restore 1-6 against a scratch recovery unit (real SQLite store, 8 stepped
turns), asserting every Restore gate. Record at
reports/research/ald-079-independent-restore-validation.json: restored/ok
true, turn and both policies match, all seven stream prefixes intact and
walking, §7.3 recovery event plus new checkpoint appended, one post-restore
step proves resumability. BACKLOG 258/261; status snapshots and conformance
matrix regenerated (verified green in a clean worktree). Working version:
0.1.529.
Verifies the runs-on-every-change half end to end: workflow triggers,
check:ci -> test:ci -> vitest wiring, and a 60/60 green run of the four
suites ALD-078 names (ALD-011 crash-safety with 20 SIGKILL trials,
ALD-036 gateway conformance, ALD-057 semantic-leakage battery, ALD-067/068
red-team suites). Merge-blocking stays open pending the admin sequence in
the same doc. Working version: 0.1.529.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant