Skip to content

MadSpin: read mg7's decay pools as .npy again, and stop splitting them - #90

Merged
oliviermattelaer merged 1 commit into
claude/madspin-performance-optimization-4a418afrom
claude/mg7-npy-pools
Sep 2, 2026
Merged

oliviermattelaer merged 1 commit into
claude/madspin-performance-optimization-4a418afrom
claude/mg7-npy-pools

Conversation

@oliviermattelaer

Copy link
Copy Markdown
Contributor

Reinstates output_format = lhe_npy for decay_generator = mg7, which T145 dropped in favour of LHE.

The objection was wrong, and no sidecar is added

T145's reason for LHE was that a pool has to survive being handed to another process by path alone carrying the channel's partial width, which an LHE <init> block does and a .npy does not -- and that supplying it for a .npy would mean a width sidecar in the owner->waiter publish contract, "the one protocol surface where a publish/open race can live".

Both hops were checked directly:

  1. generation fork to parent. Already carries the width explicitly: _generate_decay_entry writes {'files', 'width', 'channel_widths'} into its JSON and _generate_decays hands it back to _reader_from_paths as cross. Predates this change.
  2. refill owner to waiters. Nothing has to be published. Note the width here is not unused, as was suggested: get_decay_from_file weights the choice between a pdg's channels by channels[key].cross, and after a refill the reader it reads that off is the refill slice. But a channel's partial width does not change between generations of its pool, so the reader that just ran dry already has it and _worker_refill passes it on in process. ms_refill.gen stays the only thing on disk a waiter looks at.

The .npy removes the split rather than optimising it

A numpy pool is memory-mapped and randomly addressable, so worker i of N takes rows i, i+N, ... by index arithmetic. No per-worker file to write, no other worker's events to parse past. _StridedEvents over LHE text has to next() over them, which is precisely why the pool had to be dealt out into one file per worker first.

This makes PR #86 unnecessary. #86 taught madspace's combine_to_lhe to write N LHE files from C++ so MadSpin would not have to split one; its own author measured 0.34 s / 10k and 1.11 s / 100k and judged that not worth a compiled dependency. With .npy there is no split at all, so there is nothing left for a multi-file writer to optimise -- and none is needed for .npy either.

The three call sites

blocker real? what was done
_open_refill_slice builds EventFile(path) directly yes dispatches on suffix; takes the channel's width from its caller
_count_lhe_events returns 0 on a .npy yes, and the nastiest renamed _count_pool_events, dispatches, has a test. The silent zero capped the owner's slice at zero events so it refilled on its first draw for ever -- it presents as a refill loop, not a type error
_split_pool_round_robin on npy sources moot never reached for a numpy pool; now says so instead of raising UnicodeDecodeError three frames inside the LHE parser

_refill_pool_paths tells the layouts apart by what the backend actually left in the refill directory rather than by a flag. _StridedEvents and _LimitedEvents were re-checked on an NpyDecayPool and do still work unchanged (T144's measurement holds), but native striding replaces the former.

Measured

ttbar_full, PA, seed 42, 18 cores, warm, back to back, same production sample. mg7/LHE is 554d8e981 itself, so this is the same tree either side of the format.

madevent mg7/LHE mg7/.npy npy vs LHE
10 000 35.13 s 20.45 s 19.69 s -3.7%
100 000 89.08 s 38.10 s 35.74 s -6.2%

Medians of 3 (10k) and 6 (100k). One 100k pass was slow for all three backends (madevent's me_generation went 9.6 to 15.1 s in the same pass); dropping it gives 38.10 to 35.38 s, -7.1%. Spread otherwise 0.4 s at 10k, 0.7 s at 100k.

By phase at 100k: decay_mg7_split 1.1 s to 0, decay_mg7_launch 14.2 to 12.6 s (madspace writes binary instead of formatting text), decay_loop 15.4 to 14.7 s.

Price: peak RSS 259 to 308 MiB at 100k, and a .npy pool is ~35% larger on disk (every event padded to the longest). Both transient.

Physics agrees -- BR within 0.2% of madevent's 0.917999 throughout.

Refill exercised deliberately

The benchmark triggers no refill at either size, so it was forced: with MADSPIN_OWNER_UNDERSIZE=0.9 the 10k run drove 14+ generations through _worker_refill / _owner_generate / _open_refill_slice, and every Events/ms_refill_<gen> came back holding exactly one events.npy and nothing beside it. BR 0.917343, cross section 462.667 pb.

Tests

MadSpin unit suite 570 to 579, all green. The LHE-pool test class is replaced by npy equivalents, plus new classes for the strided reader, the _count_pool_events silent zero, the refill width hop, and the splitter guard.

🤖 Generated with Claude Code

T145 moved mg7's decay pools from .npy to LHE on the grounds that a pool has to
survive being handed to another process BY PATH ALONE -- across the per-particle
generation fork, and from a channel's refill owner to its waiters -- carrying the
channel's partial width with it, which an LHE <init> block does and a .npy does
not; and that supplying it for a .npy would mean a width sidecar in the
owner->waiter publish contract, "the one protocol surface where a publish/open
race can live".

That is wrong on both hops, and no sidecar is added here.

  * The fork already marshals the width explicitly. _generate_decay_entry writes
    {'files', 'width', 'channel_widths'} into its JSON and _generate_decays hands
    it back to _reader_from_paths as ``cross``. That predates this change.

  * The refill needs nothing published. The width IS read back after a refill --
    get_decay_from_file weights the choice between a pdg's channels by
    ``channels[key].cross``, and the refill reader is what it reads it off, so it
    is not, as claimed, unused -- but a channel's partial width does not change
    from one generation of its pool to the next. The reader that just ran dry
    already has it, and _worker_refill passes it on in process. ms_refill.gen
    remains the only thing on disk the waiters look at.

WHAT THE .NPY BUYS

Not only the ~2.3 us an event of parsing it saves (measured: 6.48 us against
8.76 us building the same Event objects). A numpy pool is memory-mapped and
randomly addressable, so worker i of N takes the rows at positions i, i+N, ...
by index arithmetic. That removes the round-robin split outright rather than
making it cheaper: there is no per-worker file to write, and no other worker's
events to parse past. _StridedEvents over LHE text has to next() over them,
which is exactly why the pool had to be dealt out into one file per worker
first.

THE THREE CALL SITES

  * _open_refill_slice built an EventFile from the path. It now dispatches on
    the suffix and takes the channel's width from its caller.
  * _count_lhe_events counted "<event" as text and so returned 0 on a .npy: the
    owner's slice was capped at zero events and it refilled on its very first
    draw, for ever. It presents as a refill loop, not as a type error. Renamed
    _count_pool_events, dispatches, and has a test of its own.
  * _split_pool_round_robin is not reached at all for a numpy pool, and now says
    so rather than raising UnicodeDecodeError three frames inside the LHE
    parser.

_refill_pool_paths tells the two layouts apart by what the backend actually left
in the refill directory, not by a flag: mg7 leaves one events.npy there, madevent
one LHE file per worker. Both are the canonical layout, so _generate_refill_pool
recognises either and does no further work; only the gridpack backend, which
writes somewhere else entirely, still gets materialised. Probing is race-free --
a waiter only looks once ms_refill.gen says the generation is complete.

MEASURED, p p > t t~ with both tops fully decayed, spinmode PA, seed 42,
18 cores, warm, back to back, same production sample. mg7/LHE is 554d8e9
itself, so this is the same tree either side of the format.

               madevent   mg7/LHE   mg7/.npy   npy vs LHE
    10 000      35.13 s   20.45 s    19.69 s      -3.7%
    100 000     89.08 s   38.10 s    35.74 s      -6.2%

Medians of 3 (10k) and 6 (100k) runs. One pass of the 100k set was slow for all
three backends (a machine hiccup: madevent's me_generation went 9.6 -> 15.1 s in
the same pass); dropping that pass gives 38.10 -> 35.38 s, -7.1%. Run-to-run
spread otherwise is 0.4 s at 10k and 0.7 s at 100k.

Where the 100k saving comes from, by phase: decay_mg7_split 1.1 s -> 0 (there is
no split), decay_mg7_launch 14.2 -> 12.6 s (madspace writes binary instead of
formatting LHE text), decay_loop 15.4 -> 14.7 s.

The price is memory and disk: peak RSS 259 -> 308 MiB at 100k (the mmapped pool
pages), and a .npy pool is ~35% larger than the LHE one because every event is
padded to the longest in the file. Both are transient.

Physics agrees. Branching ratio 0.9177-0.9197 against madevent's 0.917999 at
10k and 0.917999 at 100k; cross section after decay within 0.2% throughout.

The refill path was exercised deliberately rather than assumed: with
MADSPIN_OWNER_UNDERSIZE=0.9 the 10k run drove 14+ generations through
_worker_refill / _owner_generate / _open_refill_slice, and each
Events/ms_refill_<gen> came back holding exactly one events.npy and nothing
else. BR 0.917343, cross section 462.667 pb.

MadSpin unit suite: 570 pass before, 579 after.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@oliviermattelaer
oliviermattelaer merged commit ec56a72 into claude/madspin-performance-optimization-4a418a Sep 2, 2026
507 checks passed
@oliviermattelaer
oliviermattelaer deleted the claude/mg7-npy-pools branch September 2, 2026 16:38
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant