Skip to content

Issue #117 Arm-C SWA allocation remediation (InferSwarm #166): lifecycle-owned full->swa mapping in the standalone R6 stage runtime - #34

Merged
ezutfen merged 1 commit into
inferswarm-researchfrom
issue-166-swa-remediation
Sep 13, 2026
Merged

ezutfen merged 1 commit into
inferswarm-researchfrom
issue-166-swa-remediation

Conversation

@ezutfen

@ezutfen ezutfen commented Sep 13, 2026

Copy link
Copy Markdown

Closes InferSwarm FlashML-org#166 (implementation remediation only; NO Arm-C requalification).

Starting heads (verified exact at session start)

The defect (FlashML-org#157-localized)

The standalone R6 stage path builds the same HybridSWAKVCache pool as the scheduler but never reaches the scheduler's alloc_swa lifecycle (CacheManager.allocate_paged): full_to_swa_index_mapping stayed at the all-zero sentinel, every SWA-layer KV store raced on swa slot 0, and every prefix read consumed slot-0 bytes (chunk-2 nondeterminism; earliest varying boundary L0_kv_slice_post_write).

The remediation

GemmaDenseStage owns the session SWA allocation lifecycle at the same lifecycle points the scheduler uses:

  • prefill()/decode() (all four role branches) call _ensure_swa_session_mapping BEFORE _prepare (prepare_metadata translates the full page-table range through the mapping — the first prefix-consuming SWA operation of each chunk);
  • incremental ownership frontier [0, allocated) matching CacheManager.allocate_paged's per-chunk granularity; each position's slot allocated exactly once; mapping persists across chunks and decode;
  • reset_session_state rebuilds the allocator AFTER the zero loop (the zero loop wipes _swa_free to all-zeros, which would hand out sentinel slot 0 on the next allocation — the Feat/sm75 cuda128 FlashML-org/FreeToken#157-recorded trap), releasing all ownership wholesale: reuse cannot inherit stale mapping;
  • exhaustion and position gaps fail CLOSED (RuntimeError) — execution never continues on the sentinel;
  • non-SWA pools untouched no-ops; pool sized max_seq_len+1 (sentinel convention) so a full-capacity session fits.

No parallel allocator (existing pool primitives reused verbatim); no case/regime/model special cases; 64-row boundary contract, geometry, planner, tokenizer, thresholds unchanged.

Validation (CPU-only; no GPU/model execution)

  • NEW tests/research/test_issue166_swa_remediation.py: 35 lifecycle tests incl. negative controls (skipped allocation, delayed allocation, slot-0 aliasing, stale-mapping-survives-reset, ownership leak, ignored exhaustion, case-gating)
  • Focused regression sweep 189/189 green (issue117_arm_c_remediation, r6_gate_contracts, r6_final_row_fp32_promotion, r6_inplace_softcap, r6_single_gpu_control, r6_single_gpu_control_gate)
  • tests/research/ 563 passed; the single failure (test_r6_localization::test_stage_chain_capture_ops_are_explicit) is proven pre-existing at base 55e8baa (stash-and-rerun); scheduler/kvcache SWA suite failures likewise byte-identical pre-existing on this CPU-only host
  • Causal applicability: the lifecycle's alloc_swa(arange(frontier, need_end)) produces the byte-identical mapping the accepted Feat/sm75 cuda128 FlashML-org/FreeToken#157 intervention produced (test_standalone_reaches_the_scheduler_allocation_semantic)

Boundary

Terminal: ISSUE117_ARM_C_SWA_REMEDIATION_READY (InferSwarm record PR). This producer is NOT promoted to qualification authority; fresh Arm-C requalification remains separately blocked pending maintainer acceptance/merge. InferSwarm-side record: Zutfen-LLC/inferswarm branch issue-166-swa-remediation @ e13cf58.

…): lifecycle-owned full->swa mapping in the standalone R6 stage runtime

The standalone stage path built the same HybridSWAKVCache pool as the
scheduler but never reached the scheduler's alloc_swa lifecycle
(CacheManager.allocate_paged): full_to_swa_index_mapping stayed at the
all-zero sentinel, so every SWA-layer KV store raced on swa slot 0 and
every prefix read consumed slot-0 bytes (FlashML-org#157 chunk-2 nondeterminism,
earliest varying boundary L0_kv_slice_post_write).

The stage runtime now owns the session SWA allocation lifecycle at the
same points the scheduler uses:
- prefill()/decode() (all four role branches) call
  _ensure_swa_session_mapping BEFORE _prepare (prepare_metadata
  translates the full page-table range through the mapping -- the first
  prefix-consuming SWA operation of the chunk);
- ownership is an incremental frontier ([0, allocated)) matching
  CacheManager.allocate_paged's per-chunk granularity; each position's
  slot is allocated exactly once, mapping persists across chunks/decode;
- reset_session_state rebuilds the allocator AFTER the zero loop (the
  zero loop wipes _swa_free to all-zeros, which would hand out sentinel
  slot 0 -- the FlashML-org#157-observed trap), releasing all ownership wholesale:
  reuse cannot inherit stale mapping;
- exhaustion and position gaps fail CLOSED (RuntimeError); execution
  never continues on the sentinel;
- non-SWA pools are untouched no-ops; pool sized max_seq_len+1 so a
  full-capacity session fits (sentinel convention, scheduler parity).

No parallel allocator: existing pool primitives reused verbatim. No
case/regime/model special cases. 64-row boundary contract, geometry,
planner, tokenizer, thresholds unchanged. 35 CPU lifecycle tests with
negative controls (skipped allocation, delayed allocation, slot-0
aliasing, stale-mapping-survives-reset, leak, ignored exhaustion,
case-gating) in tests/research/test_issue166_swa_remediation.py.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant