MadSpin: generate decay events with mg7, and cut the overheads around them (2.0x at 10k, 4.4x at 100k) - #63
Conversation
Measure where a MadSpin run actually spends its time before changing any of
it. MadSpin already logged three timings, but the max-weight scan and the
mid-loop decay-pool refills -- which turn out to be 4.6% and 29.9% of a
tt~ full-decay run -- were invisible, and the LHE parser timers could not be
switched on at all.
- MadSpinInterface accumulates wall time per phase and emits it as one JSON
line ("MadSpin phase timings: {...}"), so a driver can pull the whole split
out of a log with one regex instead of tracking the wording of each
human-readable line. Refills are charged to their own bucket rather than to
the pre-generation pass.
- lhe_parser._ENABLE_LHE_TIMERS was hard-coded False and nothing set it; it
now follows MG_LHE_TIMERS in the environment.
- madspin_benchmark.py runs one production sample through MadSpin and reports
the split, reusing MadSpinFactory. The production sample is cached under
--workdir and reused, so re-timing after a change costs one MadSpin run.
Two bugs in MadSpinFactory surfaced while building this:
- The run_card "set" lines were written after the "done" that closes the
launch menu and starts the run, so they were never executed. nevents, iseed
and the beam energies silently kept their defaults -- invisible while the
tests asked for the default 10000 events, but it also means the parallel
MadSpin tests have been running with a random seed.
- A production run that crashed mid-script was not detected, so the factory
carried on with a half-configured sample. Here systematics died on lhapdf's
broken python 3.14 bindings; the step is now disabled outright (MadSpin does
not use the reweighting information) and "interrupted with error" aborts.
Baseline for p p > t t~ with both tops fully decayed, 10000 events, spinmode
PA, 18 cores: 79.5 s total, of which decay-event generation via Fortran
MadEvent is 64.0 s (80.5%) and the accept/reject loop proper is 4.9 s (6.1%).
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Diagram rejected anything but two incoming particles, so a decay such as t > b w+, w+ > all all could not be given to madspace at all. This is the first step towards generating MadSpin's decay events with the mg7/madmatrix engine instead of a Fortran MadEvent run per decaying particle. A decay needs no new phase-space machinery: it is the s-channel cascade the 2 -> n mappings already build, with the t-channel chain absent and the root virtuality fixed instead of sampled. The changes are therefore mostly about not assuming the number of incoming particles is 2: - Diagram accepts one or two incoming particles; _incoming_vertices is sized by that count rather than being a fixed array of two. - Topology::topologies skips find_t_vertices for a decay -- that search walks in from the *second* incoming particle to find the vertices between the two beams, which has no meaning here. The vertex the single incoming particle attaches to is the root of the cascade and its other lines are the root's children, so _t_integration_order stays empty and every consumer takes its no-t-channel path. - PhaseSpaceMapping emits n_out + n_in momenta, sizes its default cuts to match, maps no luminosity, and builds p_in = (M, 0, 0, 0) where the collision path builds two back-to-back beams. The random-number budget is 3n-4, the same as a leptonic fixed-s collision. - DifferentialCrossSection gains a "decay" mode: the differential rate is |M|^2 / (2 M) in GeV, not a hadronic cross section in pb. The flux is a compile-time constant so this is one multiply, no new instruction. - LHECompleter learns each subprocess's incoming count and uses it wherever it had hard-coded 2: which leading particles are initial state, what the outgoing particles' mothers are, where resonances get inserted, and how far the propagator momentum masks are shifted. Validated against results known analytically in tests/test_decay_topology.py: the 1 -> 2 massless and massive phase-space volumes to 1e-10, the 1 -> 3 massless volume to 0.02% (0.8 sigma) with a broad propagator, plus momentum conservation, the external masses, the decaying particle at rest and an exact forward/inverse round trip. 1482 tests pass, up from 1469, so the 2 -> n path is unchanged. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
OneProcessExporterMG7 hard-coded two initial legs. For a 1 -> n process it therefore wrote garbage without complaining: incoming came out as [pdg, None], and the outgoing offset of 3 sent the first final-state particle to outgoing[-1], overwriting the last one -- "t > b w+, w+ > all all" was exported as outgoing = [81, -81], with the b silently gone. Nothing caught this because the existing decay test (test_ungroup_decay_mg7) only looks at directory names, never at subprocesses.json. Derive the initial-leg count from the legs themselves and offset the outgoing ones by it. The same count now groups the flavor combinations by initial state, and the beam-swap mirror flag is forced off for a decay, which has no beam pair to swap. Leg numbering is asserted rather than assumed, so a process that does not follow the "initial state first, numbered 1..n" convention fails loudly instead of writing a wrong topology. For a collision every expression reduces to what it was before: verified by generating p p > t t~ with both versions and diffing subprocesses.json and proc_characteristics (identical). The new test fails on the old exporter. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
With the exporter able to describe a 1 -> n process, teach the launcher to integrate one. Whether a directory is a decay is read off the exported subprocesses rather than the run card, so the two cannot disagree. For a decay: - the total energy is the decaying particle's mass and there is no parton luminosity (which is what the mappings already call "leptonic"); - the renormalisation scale is fixed at that mass, so alpha_s is constant. Rather than demand an LHAPDF set -- possibly downloading one -- purely to evaluate a coupling, write a minimal .info holding the param card's alpha_s. This is the right answer for a fixed scale, not an approximation; - the flux is 1/(2M), giving a partial width in GeV (see the DifferentialCrossSection decay mode); - the phase space is multichannel: the flat mapping is built from a synthetic two-incoming diagram, which a decay has no counterpart for; - the <init> block reports the decaying particle at rest as a single beam. Two things a decay is the first process to hit: - clean_pids knew merged ids 81 and 82 but not 83 (neutrinos), so any final state containing one crashed on a param-card lookup. Unknown ids in the reserved 81..99 window now raise instead of being passed through as if they were pdgs. - A multiparticle decay definition enumerates closed channels too: "t > b w+, w+ > all all" yields t > b t b~, b W+ Z and b W+ h, none of which the top is heavy enough for. Their width is exactly zero, but the mapping has no physical point to return -- the invariant's lower bound lands above its upper bound -- so it produced NaN momenta and poisoned the whole integral. They are dropped up front. Also fixes a hard-coded offset of 2 in LHECompleter::init_propagator_data, which reads the color flow of the wrong particle when the initial state is not a beam pair (it threw "Incompatible with color singlet" on every decay). t > b w+, w+ > all all now runs end to end and writes a correct LHE: the top at rest with status -1, the W as a status-2 resonance with a single mother, and a consistent color flow. KNOWN ISSUE, not yet resolved: the partial width does not match Fortran MadEvent. For t > b w+, w+ > e+ ve, mg7 gives 0.15728 +- 0.00020 GeV against MadEvent's 0.161870 +- 0.000104, a 2.9% deficit -- well outside the 1% tolerance the mg7 cross-section tests hold collisions to. Widening bw_cutoff accounts for about 1% of it and then saturates; the rest is unexplained. This must be understood before MadSpin is pointed at this path, since MadSpin uses the partial width to normalise the branching ratio. Nothing reaches this code from MadSpin yet, and the collision path is unchanged (1482 madspace tests pass; p p > t t~ still integrates and writes a correct LHE). Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The partial width of t > b w+, w+ > e+ ve came out 3% below Fortran MadEvent (0.15534 vs 0.161870 GeV). The cause was not the matrix element: a decay directory shipped the hadron-collider cut defaults, so a 20 GeV pt cut sat on the b and a 10 GeV one on the e+. Inside a 173 GeV decay that rejects ~3% of the phase space (the run log's "samps: 63.0k, samps. after cuts: 60.9k"), and a partial width is an inclusive quantity, so the cut biases it straight down. RunCardMG7.create_default_for_process already had the right rule -- ninitial == 1 means clear the cuts -- but it was dead code: ProcCharacteristic defaults ninitial to 0 and ProcessExporterMG7 .finalize() called create_run_card *before* create_proc_characteristics filled it in. Swap the two, so the run card is derived from a populated proc_characteristic, and note the dependency where it now matters. Fixed at output time rather than in the launcher: the card the user reads is then the card that is run, and a user who does want a cut on a decay can still set one. Zeroing the cuts at run time would leave the directory advertising cuts it silently ignores. The Breit-Wigner cutoff is deliberately untouched (still 15): it bounds how far off shell the propagators are sampled, it is not a cut on the final state. t > b w+, w+ > e+ ve 0.161717 +- 0.000073 vs ME 0.161870 +- 0.000104 t > b w+, w+ > all all 1.46034 +- 0.00062 vs ME 1.458317 +- 0.003125 i.e. 1.2 and 0.6 sigma, against 2.9% and 18% before. Collisions are untouched: run_card.toml, run_card_default.toml and proc_characteristics generated for p p > t t~ are byte-identical across this change. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Generating the decay events was 80.5% of a MadSpin run: for p p > t t~ with
both tops fully decayed, 64.0 s of 79.5 s went into Fortran MadEvent runs, one
per decaying particle plus a mid-loop refill whenever a pool ran dry. The
accept/reject loop those pools feed was 4.9 s.
Point that at the mg7 (madmatrix/madspace) generator, now that it can integrate
a 1 -> n process. A new madspin-card option `decay_generator` selects the
backend; it defaults to mg7, with madevent kept as a one-line fallback and
forced for gridpack mode, which drives the decay directory through run.sh --
something the mg7 output does not provide.
The launcher runs out of process. It chdirs, installs signal handlers and holds
its own madspace context, none of which belongs in MadSpin's interpreter next
to the f2py matrix elements. Its exit code is checked and its output kept in
decay_dir/mg7_generation.log, so a failure is a clear error rather than an
empty pool. The partial width comes back from the LHE <init> block, which for a
1 -> n process carries a width in GeV.
Timing, same benchmark and seed, 18 cores:
phase mg7 madevent speedup
decay_event_generation 23.97 40.26 1.68x
me_generation 5.47 5.10 0.93x
max_weight_scan 3.52 3.66 1.04x
decay_loop 5.23 28.61 5.47x
of which refill 0.38 23.74 63.21x
output_gzip 1.73 1.68 0.97x
total 40.07 79.46 1.98x
The refill collapse is the striking one, and it is not a subtle effect: a
madevent refill costs ~12 s of fixed survey/refine/combine overhead whatever
its size (the baseline paid 23.7 s for 4708 events). mg7 also returns exactly
the number of events asked for, where madevent is asked for 0.8x and overshoots,
so the pools run dry far less often. The accept/reject loop itself is unchanged
at 4.85 s vs 4.88 s, as it must be.
Physics agrees. Partial widths for the two decay directories:
t > b w+, w+ > all all 1.4579204(12874) vs 1.458317(31247) 0.12 sigma
t~ > b~ w-, w- > all all 1.4600359(12868) vs 1.460172(30610) 0.04 sigma
and the mg7 errors are 2.4x smaller for half the wall time. Cross-section after
decay 482.600 vs 483.288 pb (-0.14%, within the width errors), unweighting
efficiency 0.3728 vs 0.3756.
Also fixes the decay-directory output format, which was pinned to madevent with
a comment saying the runner required it; it now follows whichever runner will
actually be used.
Note: the mg7 run card has no seed parameter, so decay pools are not
reproducible run to run (MadSpin's own accept/reject RNG is still seeded). That
is worth adding but is not a correctness issue here.
Three of the six MadSpin acceptance tests fail on this machine both before and
after this commit -- the production step crashes in systematics on lhapdf's
broken python 3.14 bindings, before MadSpin is even reached. Verified by running
them against the pre-change file.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Once MadSpin's decay events came from mg7, decay-event generation was still
60% of a run. Splitting it showed the integration was not the cost at all:
decay_me_generate 0.04 s MG5 amplitude generation
decay_me_output 0.96 s writing the C++
decay_mg7_launch 32.35 s the launcher subprocess
of which integrate 0.27 s survey + generate + combine
So essentially all of it was compiling the matrix elements, and the launcher
was compiling them one file at a time:
- misc.compile defaults to nb_core=1, i.e. a serial make. A subprocess has
several independent translation units, so this left the machine idle:
3.05 s serial against 0.90 s at -j18 for one t > b w+ subprocess on 18 cores.
- Even with -j, a subprocess has only ~5 objects, so make cannot fill a large
machine however high the job count. The subprocesses are independent, so
compile_subprocesses now builds them concurrently up front and splits the job
budget between them. A decay of t > b w+, w+ > all all has 4 open channels;
p p > t t~ has 2, built 2-at-a-time with -j9.
The budget honours cpu_thread_pool_size when the user has capped it, and
otherwise takes the machine.
Measured back to back on the same (busy) machine, so the ratios are meaningful
even though the absolute numbers are inflated:
decay_mg7_launch 32.35 -> 9.81 s 3.30x
total 56.47 -> 33.02 s 1.71x
Every other phase moved by 1.00-1.15x, i.e. not at all.
Also stops running systematics on a decay. Scale and PDF variations are a beam
quantity and a decay has neither, so it could only ever fail ("not supported
for pdlabel=none") -- and it is not free, since MadSpin reruns the launcher for
every pool refill. It now logs one line and skips, instead of writing a crash
log per run. Collisions are unaffected: verified that p p > t t~ still compiles
(2 at a time), integrates, runs systematics and writes a correct LHE.
On the MadSpin side, generate_events_mg7 appends to its log rather than
truncating -- a pool refill reruns the launcher in the same directory, and
overwriting threw away the log of the run that did the compile.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
mg7 can write its events either as LHE text or as a numpy structured array
carrying the same fields. MadSpin consumes one decay event per accept/reject
trial and never shows the pool to anything else, so the text round trip was
pure overhead. Ask the launcher for "lhe_npy" and build the lhe_parser.Event
objects straight from the array.
Two corrections on the way, both of which changed the answer:
- The benchmark was reporting the LHE cost as 6.88 s (17.2% of wall) by adding
up timers that nest inside one another. next_event_readline_total contains
the Event() it builds, which is also counted in event_parse_total, which
contains the particle block, which contains the per-particle parse. The
honest figure is 2.73 s (6.8%), of which ~92% is the decay pools. That is the
actual size of this prize, and the report now says so.
- The first implementation held one numpy column array per field and read
scalars out of it with .item(). That measured 22.6 us/event against the text
parser's 16.6 us -- slower than what it replaced. Converting a whole chunk
with .tolist() and then indexing plain python scalars is 10.0 us/event
instead. The conversion is done 4096 events at a time so the pool stays
memory-mapped rather than being materialised as python objects all at once.
Measured back to back on the same machine:
decay_loop 7.83 -> 7.15 s 1.09x
max_weight_scan 4.94 -> 4.28 s 1.15x
total 32.15 -> 30.73 s 1.05x
with every other phase at 1.00x. Peak RSS goes from 142 to 172 MiB, the cost of
the mapped arrays and the converted chunks.
Physics unchanged: partial widths 1.4586029(12908) and 1.4581234(13700) against
MadEvent's 1.458317(31247) and 1.460172(30610), i.e. 0.09 and 0.61 sigma. Also
checked directly that 2000 pool events conserve momentum, carry exactly one
incoming particle, resolve every mother to a real particle, and render as valid
LHE. The three MadSpin acceptance tests that pass on this machine still pass --
one of them, test_lhe_none_decay, exercises the run_bridge path, which consumes
the pools through the same .cross/next() surface.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
At 100k events, gzip was 19.7% of a MadSpin run. Almost none of that work was
necessary.
misc.gzip has two branches: files over 256 MB are handed to the external gzip
tool, which compresses at level 6, while everything smaller goes through
gzip.open, whose default is level 9. So the same data was compressed
differently depending on its size, and the common case took the slow branch.
Level 9 is a bad trade on LHE text -- on a 172 MB decayed file it costs 18.6 s
against 4.5 s at level 6 and buys 4% (38.1 MB against 39.7 MB). Nobody chose
it; it is the module default. Both branches now use level 6, exposed as
GZIP_COMPRESSLEVEL and overridable per call.
That branch also read the whole file in as a str and encoded it, so a 172 MB
LHE needed both copies resident before a byte was written. It streams now.
Separately, run_onshell gunzipped its input LHE at the start and re-gzipped the
identical content at the end -- 6 s at 100k spent reconstructing a file that
was never modified. EventFile reads a gzipped file directly (it says so in its
docstring), so the input is now left exactly as it was found. The output
filename is derived from the input with any .gz stripped, because otherwise
appending _decayed to a .gz name would ask EventFile to write a gzip stream
where the code below expects a plain file to compress.
Measured back to back at 100k on the same machine:
output_gzip 16.69 -> 3.98 s 4.20x (12.7 s off an ~85 s run)
peak RSS 781 -> 393 MiB 2.0x less
The memory halving is the streaming fix; the time is level 6 plus not touching
the input at all.
Physics unchanged: cross-section after decay 483.211 against 483.552 pb and
BR 0.957775 against 0.958450, i.e. within the partial-width MC errors, with the
unweighting efficiency flat at 0.3290 against 0.3295. The decayed output is
still a readable gzipped LHE with all 100000 events and the right banner
cross-section, and the input keeps its original compression instead of being
repacked at level 9. Four MadSpin acceptance tests pass, including the hepmc
and none-spinmode paths that read their input through the same code.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
… trial
Two lookups in the accept/reject loop recomputed run constants:
- get_onshell_evt_and_wgt asked the banner for the pole mass and the width of
every decaying particle on every trial -- 4 card walks per trial. Those values
are already in decay_dict, which run_onshell builds from the same param card
before the loop starts ('param' and 'param_card' both resolve to the SLHA
card, so they are the same numbers). The BW cut is hoisted out of the inner
loop while there.
- get_pdir maps an event to its process directory, which is a property of the
flavour tag alone, and a run sees a handful of tags. It was doing the dict
walks and the pdg2prefix tuple build 124k times per 10k events. Memoized on
the tag -- under the tag asked about, not the anti-particle tag the 1 -> n
fallback may rewrite it to.
Worth 1.03x on the decay loop at 100k (47.80 -> 46.48 s), measured back to back.
That is much less than it looked under cProfile, which put the banner lookups at
~10% of the loop; they are nearer 2%. cProfile charges per-call overhead to
call-heavy python, which is exactly what this code is, so its ranking of this
loop should not be trusted without a wall-clock check.
Direct timers around the density evaluation, which is what the loop is for:
get_density, 124089 calls per 10k events
prep 1.03 s 8.3 us/call python: momenta, pdgs, pdir
fortran 0.51 s 4.1 us/call the matrix element itself
densmat 0.35 s 2.8 us/call DensityMatrix construction
total 1.89 s of a 5.07 s decay loop
So the Fortran matrix element is 10% of the decay loop and the python around it
is the rest -- preparing one 4.1 us call costs 8.3 us. The remaining ~3.2 s sits
in Event.boost (a FourMomentum per particle per call) and the density algebra
(trace/tensor_product/scalar_multiplication on matrices small enough that numpy
call overhead dominates). Both are per-trial work on small objects, and the
trials of one production event are independent, so the real fix is to batch
them rather than to keep shaving the glue.
Physics unchanged: cross-section after decay 483.331 against 482.967 pb, BR
0.958014 against 0.957291, unweighting efficiency 0.3291 against 0.3266 -- all
within the MC spread of the pools. Four MadSpin acceptance tests pass.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Event.boost built two FourMomentum objects for every particle -- one to convert
the particle, one for the result of FourMomentum.boost -- and threw both away
after copying four floats back. MadSpin calls this once per decay per
accept/reject trial, so a 100k-event run churned through millions of them: it
was the largest single item in the decay loop after the density evaluation
itself.
The boost vector is the same for every particle in the event, so its norm and
its mass are loop invariant. Hoist those, inline the rest of
FourMomentum.boost, and work on floats. The arithmetic is left in exactly the
order FourMomentum.boost used it -- including computing the mass as
E^2 - px^2 - py^2 - pz^2 rather than E^2 - pnorm, which rounds differently --
so the result is bit-for-bit unchanged.
4.65 us -> 1.35 us per event 3.45x
Verified bit-identical, not merely close: 39976 momentum components across 2000
real decay events compared with ==, zero differences, plus the zero-momentum
branch. The spacelike-boost-vector case still raises ZeroDivisionError exactly
as before; that hazard is untouched.
End to end at 100k, back to back:
decay_loop (net) 46.88 -> 44.69 s 1.05x
max_weight_scan 5.14 -> 4.48 s 1.15x (same density path)
total 73.23 -> 68.87 s 1.06x
The 24 lhe_parser unit tests and four MadSpin acceptance tests pass, and the
cross-section after decay is 483.431 against 483.219 pb with the unweighting
efficiency at 0.3315 against 0.3306, i.e. the usual pool-to-pool spread.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Follow-up to the previous commit, now that changing the rounding is on the
table. Two more things in Event.boost were per-particle work that need not be:
- the two divisions. mass and pnorm are shared by every particle, so 1/mass and
(bE-mass)/pnorm are computed once and multiplied in, leaving only multiplies
and adds inside the loop.
- the copy of the boost momentum. The helas sign flip is applied to locals
instead of to a FourMomentum copy, which drops an allocation per call and
still leaves the caller's object untouched, as the copy did.
original (allocating) 4.693 us/event
previous commit 1.167 us/event 4.02x
this commit 0.942 us/event 4.98x
floor (attributes only) 0.868 us/event
The floor is reading and writing p.px/py/pz/E for five particles with no
arithmetic at all, so this is within 8% of what the function can cost while
Particle stores its components as attributes. __slots__ does not help: measured
at 3% on this access pattern, because CPython 3.14 already specialises instance
attribute access. Event.boost is done.
No longer bit-identical, as authorised. Over 40k momentum components of real
decay events the worst relative difference against the original is 1.2e-12,
arising where E+k*s3 cancels; typical components agree to ~1e-16. The
zero-momentum branch, the empty event, and the caller's momentum being left
unmodified are all checked.
End to end this is below the noise -- decay_loop 44.91 -> 44.28 s at 100k, 1.01x
-- because the previous commit had already taken boost to the attribute floor.
It is committed for the function-level gain, not for the run-level one. 24
lhe_parser unit tests and 4 MadSpin acceptance tests pass.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The density algebra is 15.5% of the decay loop, and it is not arithmetic: the
matrices have 4 to 16 entries, so numpy's per-call dispatch dwarfs the work.
Measured in situ on a 5.19 s loop (10k events):
DensityMatrix.__init__ 0.342 s 124325 calls 2.75 us 16 values
trace 0.238 s 171375 calls 1.39 us 4 values, 2 diagonal
tensor_product 0.116 s 57125 calls 2.03 us (4,2) (x) (4,2)
scalar_multiplication 0.108 s 57125 calls 1.88 us 16 x 16
trace was summing *two numbers* with np.sum(values[bool_mask]) at 0.90 us, of
which 0.74 us is np.sum's fixed cost on any array at all. It now indexes a
cached tuple of diagonal positions and adds them in python: 0.11 us, 8x, and
4.2x on the 16-entry production matrices.
scalar_multiplication used np.sum(a*b), which is two calls and a temporary.
np.dot is identical for complex -- it does not conjugate, unlike vdot -- and
takes 0.23 us against 0.90 us.
Both are summation-order changes only, so results move at complex64 epsilon:
over 300 random matrices of the shapes above the worst relative difference is
1.2e-07 for trace and 5.1e-07 for scalar_multiplication.
Back to back at 100k:
decay_loop 42.92 -> 42.11 s 1.02x
max_weight_scan 4.50 -> 4.00 s 1.12x
total 66.59 -> 64.17 s 1.04x
The 23 MadSpin unit tests (which cover the density mapping) and 4 acceptance
tests pass.
The other two operations are left alone deliberately: __init__ and
tensor_product would only get materially faster by storing values as python
complex rather than a complex64 array, which changes the representation every
consumer of .values depends on. That is worth about 10% of the loop and is a
real refactor, so it should be a decision of its own.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
PY_GET_DENSITY evaluates one phase-space point per call, and at these process
sizes the crossing into fortran costs more than the density it computes.
PY_GET_DENSITY_BATCH takes NBATCH points and loops over f77_density inside,
so the entry is paid once.
Measured on 64 real decay-pool points for t > b w+, w+ > all all:
single-point loop 2.835 us/point
batched, NB=64 0.585 us/point 4.8x
and bit-identical, max|diff| exactly 0 -- it is the same routine called in the
same order, only the boundary moved.
Everything is per point except POS and ALLOW_HEL, which describe the helicity
structure and are shared by construction. PDGS is per point because a decay
pool mixes flavours (t > b u d~ and t > b c s~ come from the same pool), and
ALPHAS/SCALE2 are per point because nothing guarantees the points share a
scale. NBATCH is an argument rather than inferred, so callers can use whatever
width they have -- MadSpin's max-weight scan batches at max_weight_ps_point,
which is a card option and not always its default of 400.
Nothing calls this yet; the caller side is a separate change.
A note for whoever tests this next: f77_density has first-call state, so a
comparison harness must warm it up before recording a reference, and it
returns zeros for unphysical momenta. Both together made an early version of
this check look like a 0.49 disagreement when the batch was in fact exact.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The scan evaluates Nevents_for_max_weight production events x max_weight_ps_point decay configurations and keeps the maximum. It has no accept/reject, so every trial is used: the decay sets of one production event can all be drawn up front and their densities evaluated in one fortran call instead of one call per trial. get_density_batch is get_density for a list of events sharing a helicity structure. It groups the points by external multiplicity -- one call carries a single NEXT and a decay pool can hold 1 -> 2 and 1 -> 3 channels -- and builds the momenta straight into a fortran-ordered array, which also takes the pure-python invert_momenta transpose off the per-point path. calculate_matrix_element_from_density grows a decay_densities argument that makes its loop skip both the boost and the density call; the boost mutates the decay event in place, so in that mode the caller owns it and does it once, up front. Slots that share a helicity structure share a call, so t t~ batches two slots per trial. The first trial of each production event still goes through the unbatched path: it is what populates production._ms_density_static, which says what to boost by and which helicities to ask for. Outside the pole approximation the production and every decay are reshuffled *inside* the ME call, so their momenta are not known beforehand -- there the draws stay lazy and the loop is untouched. Verified against the per-trial path trial by trial, with the Breit-Wigner sampling pinned so both calls draw the same masses: bit-identical on p p > t t~ with both tops fully decayed (1475/1475 trials, max|diff| exactly 0, all_maxwgt equal element by element), on the semileptonic sample, and on a mixed 1 -> 3 / 1 -> 2 decay set that exercises the multiplicity grouping. Batch widths 1, 2 and 7 all reproduce. Back to back at 100k events, max_weight_scan 4.08 s -> 3.57 s (1.14x) with every other phase within 2%. Short of the 2x hoped for, and the probe says why: inside the batched call the fortran is now 0.24 s of 1.27 s, against 0.69 s of per-event get_momenta/get_pdg and 0.22 s of DensityMatrix construction. The f2py entry is no longer what the scan pays for -- the remaining cost is python that batching does not touch. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|
Hi Theo, While this is mainly moving madspin backend from madevent to madspace. The rest is some additional validation for the mg7 paper to say that madspin2 is faster in mg7 compare to the madspin2 paper (soon to be released). Olivier |
theoheimel
left a comment
There was a problem hiding this comment.
The madspace changes look good to me
a1afe9d made the mg7 launcher compile subprocesses concurrently with a ThreadPoolExecutor. That is unsafe: every P* makefile builds the shared common library by recursing into the shared src/, and nothing serialises them, so N concurrent subprocess builds means N processes compiling src/build.<backend>/ Parameters.o and read_slha.o to the same paths and linking the same lib<...>_common.so. It broke check_xsec_processes (ttx1j) on CI: p p > t t~ j is the only entry in that section with more than one subprocess, so it is the only one that took the concurrent path, and it failed with a compilation error in P0_QQx_ttxg. The same commit passed a re-run minutes later, which is the giveaway. The race does not reproduce on macOS (0/6 at -j4, 0/10 at -j1 with 18 cores): src/ has two objects and the window is short. Widening it with a sleep in the src rules shows it plainly -- four concurrent writers to each of Parameters.o and read_slha.o. Two further reasons to drop rather than patch this: - PR #61 proposed the same ThreadPoolExecutor design and was closed unmerged. This commit reintroduced it without knowing. - PR #67 does it properly, at the make level: SubProcesses/makefile becomes a jobserver dispatcher, the common library is built once up front, and P* directories are told so with MADMATRIX_COMMONLIB_EXTERNAL=1. Its jobserver also shares slots dynamically, where the static build_jobs // len(pending) split here gave -j1 per make on a 4-core CI runner -- neutralising the parallel-make win (3.4x, the larger half) while keeping the racy half. Removes compile_subprocesses, build_subprocess, resolve_api_path, build_jobs and the concurrent.futures import; MadgraphSubprocess's compile loop and init_subprocesses are byte-identical to main again, so this file no longer conflicts with #67 (verified with git merge-tree). Everything else a1afe9d added is unrelated to the build and stays: the decay-phase timing split, and on this branch the mg7 decay mode, clean_pids, drop_closed_channels and skipping systematics for decays. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|
| partial widths | σ×BR | |
|---|---|---|
| madevent, run 1 | 0.3237400, 0.9708000 | 71.25510 |
| madevent, run 2 | 0.3237400, 0.9708000 | 71.25510 |
| mg7, run 1 | 0.3248777, 0.9719431 | 71.32298 |
| mg7, run 2 | 0.3237563, 0.9676568 | 71.24660 |
madevent reproduces bit-for-bit; mg7 does not.
The failures are within one sigma
madspace quotes the width errors itself — 0.282% and 0.283% on the two decays. BR is their product, so per run the relative error on BR is 0.40%, and comparing two independent runs gives an expected 1σ spread of 0.567%.
Observed: 0.411% and 0.479%, i.e. both under 1σ. The test asks for 0.100%, which is 0.18σ, so its pass probability is roughly 14%. test_short_madspin_zz passing is luck, not a distinction.
Not specific to MadSpin
madspace seeds nothing — std::random_device at seven sites across the CPU runtime, GPU runtime, event generator and MLP. Same directory, same card, p p > t t~ j:
run 1: integral: 1006(11)
run 2: integral: 987(11)
This branch did not create that gap. It made a test notice it, by swapping a seeded generator for an unseeded one inside a comparison that assumed determinism.
Why we need #51
#51 ("Implement reproducible event generation") is the fix, and this PR should wait for it. Its stated goals — full reproducibility of generation in regular and gridpack mode, independent of worker thread count — are exactly what is required here. Once decay pools are reproducible, the modes again share an identical BR by construction and the 0.1% tolerance goes back to measuring the unweighting.
Worth stressing, because it is the tempting shortcut: please do not fix this by loosening rel_tol. The assertion's own docstring calls it "the strict invariant catching real bugs". Relaxing it to ~2% (3σ) would make the tests green while leaving them unable to catch a genuine 0.5% error in the unweighting normalisation — precisely the class of bug it exists for. Loosening removes the test's power rather than restoring it.
Until #51 lands
The two tests will keep failing about six runs in seven. The interim option, if this PR needs to land first, is to flip decay_generator back to madevent by default and document mg7 as opt-in (set decay_generator mg7) — that restores determinism and keeps the tests meaningful, at the cost of the speedup being off by default. Preference is to wait for #51 rather than ship that.
Also worth adding once #51 is in: the comparison would be more honest as an n-sigma test against the combined MC error rather than a fixed relative tolerance, since a fixed tolerance is only valid while the two sides are correlated by construction. That is a follow-up, not a substitute for #51.
MadSpin's density modes (spinmode madspin/PA/onshell -- the default) were evaluating every production and decay matrix element with the *model's* default parameters, not the ones the input events were generated with. `output standalone` writes `<me_dir>/Cards/param_card.dat` from the model, because it knows nothing about the event file, and nothing ever refreshed it. `initialise_f2py_module` handed that file to the compiled density library; its only other branch required `<me_dir>/param_card.dat` -- a file `output standalone` never writes -- to exist *and* `<me_dir>/Cards/param_card.dat` -- which it always writes -- not to, so the run's own card was unreachable. Verified on disk across every MadSpin run directory on the development machine: MT 173 vs 172.76, WT 1.4915 vs 1.33, and an EFT run whose density matrices were built at ctGRe = +0.37 while the events carried ctGRe = -1. The legacy onshell_v1 read site had a third source: `path_me/Cards/ param_card.dat`, which is the *process directory's* card when MadSpin runs under MadEvent, and which may have been edited after the events were made. There is exactly one correct source, and MadSpin already materialises it: `path_me/param_card.dat`, written from `banner['slha']` on every run. It is the card the events were generated with, including the `import model <MODEL> <CARD>` override -- the one supported way to ask MadSpin for different parameters, which is checked against the banner before being accepted. The production matrix element must match the events; the decay one must match the decay events MadSpin itself generates from that same card and the partial widths it measures there. So: copy it into each ME directory at compile time (`refresh_me_param_cards`, the same idiom as the legacy `compile_fortran`), and route both read sites through one resolver (`me_param_card`). A copy rather than a symlink -- these trees are compiled in place, tarred into gridpacks and moved to other machines, and the copy left behind is the record of what the run used. Also close the cache side. A reused `ms_dir`/`use_old_dir` directory caches the decay gridpacks, the partial widths measured while building them, the maximum weights and the pickled branching ratios, and re-measures none of them. `run_from_pickle` compared the pickled banner's blocks but skipped every `decay` block -- exactly the widths that drive the branching ratios and the Breit-Wigner sampling. Stamp the directory with the card it was built with (`ms_param_card.dat`) and refuse to reuse it under a different one, naming the blocks that moved; warn when a directory predates the stamp. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
test_one_mode takes four positional arguments and is a helper: the
eleven real tests call it with a spinmode, the particles to decay and
two file names. unittest collects by the "test" prefix, so it was also
collected on its own, called with no arguments, and raised
TypeError: test_one_mode() missing 4 required positional arguments
on every run of tests/acceptance_tests/test_madspin -- the suite has
been exiting non-zero unconditionally, on every branch, which is how a
red result stops meaning anything.
Renamed to _run_one_mode. The leading underscore is what does the work;
any name keeping the "test" prefix would only move the bug. Verified
against origin/madspin_density (02ee49a), which fails identically, so
this is pre-existing and not a regression from any branch in flight.
before Ran 6 tests FAILED (errors=1)
after Ran 5 tests OK
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
MG5: make the custom_fcts lookup case-insensitive, and name the function in the error
…llection tests: stop collecting the MadSpin acceptance helper as a test
0a1007b routed every density matrix element through the run's own card, but its tests only ever laid out a tree-level ``output standalone`` directory. Loop-induced is the one spinmode family that takes a visibly different route into the density matrix -- a MadLoop tree, a second branch in ``initialise_f2py_module`` that rewrites ``MadLoopParams.dat`` and calls ``set_madloop_path``, and parameters that MadLoop reads through ``SubProcesses/MadLoop5_resources/param_card.dat`` rather than through the path handed to ``initialise``. Nothing said whether the fix reached it. It does, and for a reason worth recording: that MadLoop5_resources card is a symlink ``loop_exporters.py`` points at ``Cards/param_card.dat``, so ``refresh_me_param_cards`` overwriting the file in ``Cards/`` is what keeps MadLoop in step; and the single ``mymod.initialise`` call precedes the MadLoop branch, so both sides get the same card. Neither is stated anywhere, and either could be undone by a plausible edit. So: run the whole param_card class again over a loop-induced layout, plus four tests for the parts that only exist there -- that the MadLoop branch is the one being exercised (otherwise these silently become a second copy of the tree-level case), that the library is initialised with the run's card, that the card MadLoop resolves through its own link is the run's card too, and that ``MadLoopParams.dat`` survives the in-place rewrite (failing to parse it takes MadLoop down with a Fortran STOP, i.e. exit code 0). Measured on a ``g g > z z [noborn=QCD]`` run whose events carry MZ = 93, WZ = 2.5: before the fix both matrix-element directories sat at the model's 91.188 / 2.441404, and the production matrix element evaluated at a real event of that run moved 6.845e-04 -> 7.645e-04, +11.7%, between the two cards. 1000 events decayed in spinmode onshell differ in 942 of them. All 15 fail on 02ee49a and pass here. Unit suite 512 -> 527 green. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Raised by Sourcery on #395. An unstamped ``ms_dir``/``use_old_dir`` holding cached material was warned about and then stamped with the running card anyway. The warning was the whole point and the stamp undid it: the *next* run with that same card would meet the stamp, find it matching, and reuse decay events, partial widths and maximum weights of unknown provenance in silence. A first run vouching for a cache it explicitly said it could not vouch for -- and the upgrade boundary, where every unstamped directory lives, is the one realistic moment for it. (A *changed* card was never the exposure: it would meet that same stamp and be refused. Same card, second run, is the hole.) So: return after the warning. Left unstamped rather than refused, because the stamp is missing because the directory predates it, not because anything is known to be wrong -- the great majority of these are consistent, and refusing would break every existing ms_dir on upgrade. The directory stays usable, the warning comes back on every reuse, and no file ever claims to know what that cache was built with. What earns a stamp is a directory with no cache left in it: a fresh one, or this one once cleared. That makes the fresh/already-built distinction load-bearing rather than cosmetic, so it gets a name (``_holds_cached_material``) and the read sites' caches are enumerated rather than sampled: the previous test missed ``max_wgt_sequential``, ``pure_interference*`` and ``production_me/all_ME.pkl``, and a cache it failed to notice is now precisely a cache that gets authenticated with a card it may never have been built with. Three of the new tests fail before this change and pass after. Unit suite 527 -> 532 green. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
MadSpin: evaluate the matrix elements with the run's own param_card
…mance-optimization-4a418a # Conflicts: # MadSpin/decay.py # MadSpin/interface_madspin.py # madgraph/iolibs/export_cpp.py # madgraph/iolibs/template_files/mg7/madevent.py # madspace/src/driver/lhe_output.cpp # tests/parallel_tests/madspin_comparator.py
a310351 to
54b5a0a
Compare
…n-performance-optimization-4a418a # Conflicts: # tests/acceptance_tests/test_madspin.py
54b5a0a to
5f9c2ea
Compare
…ere nothing forks
Two independent defects in the decay_generator = mg7 guard, both reachable
today. Neither depends on what happens to the pool format afterwards.
1. `set nb_core 1` -- the escape hatch the demotion warning tells the user to
take -- crashed.
_generate_decays forks one process per DECAYING PARTICLE, not per core, so
nb_core does not switch it off. The readers cannot come back across that fork,
so it marshals them as paths and rebuilds them in the parent with
_reader_from_paths, which built an lhe_parser.EventFile unconditionally. On
mg7's numpy pool that reads the \x93 of the \x93NUMPY magic as utf-8:
UnicodeDecodeError: 'utf-8' codec can't decode byte 0x93 in position 0
Two decaying particles is p p > t t~ with both tops decayed, i.e. the
configuration the mg7 work was benchmarked on. So with the demotion escaped,
mg7 was not merely unsupported in its headline case -- it was broken, and the
warning pointed at the breakage.
_reader_from_paths now dispatches on the path suffix. A numpy pool also has no
banner, so unlike an LHE one it cannot re-read its own cross section (the
channel's partial width) on the parent side: the width is marshalled with the
paths -- it was already in that JSON as `channel_widths` -- and handed back in.
Missing width raises here rather than leaving cross = None to break the
cross-section weighted channel choice silently, much later.
2. The guard demoted spinmodes that never fork a worker.
Corrected against the code rather than the assessment that flagged it: the
assessment says the parallel drivers are reachable only from spinmode
PA/onshell and proposes gating on `spinmode in ('PA','onshell')`. That is
backwards for the default. run_madspin dispatches 'madspin' -- and 'full',
which it rewrites to 'madspin' -- to run_onshell(density_method=True), exactly
like PA and onshell, so those fork too; that gate would have LEFT mg7 selected
on the default spinmode and walked it into the very refill path the guard
exists to avoid. The spinmodes that genuinely never fork are 'none'
(run_bridge) and 'madspin_v1' (decay_all_events), which consume their pools
serially through .cross/next() and never reach the refill machinery. Only
those are exempted.
Verified by regression tests that fail on the parent commit (4 errors + 1
failure) and pass here. MadSpin unit suite: 549 pass before, 557 pass after.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
`add_param('decay_generator', 'mg7', ...)` sat AFTER the `return` of
MadSpinOptions.__setitem__, i.e. it never ran. The merge with main
(5a99859) is where it landed there: main added the `unweighting` parameter
and the __setitem__ override at exactly the point where this add_param was the
last line of default_setup, and the resolution kept it attached to the wrong
side of the new boundary.
The consequence is not cosmetic. With the parameter undeclared:
* `set decay_generator mg7` (or madevent) in a madspin card raises
InvalidCmd('Unknown options decay_generator') out of check_set, which aborts
reading the card -- so the option could not be used at all, in either
direction;
* generate_events reads it as `self.options['decay_generator']` inside a
try/except KeyError whose except branch is `output_format = 'madevent'`.
ConfigFile.__getitem__ raises exactly that KeyError, so every run took the
fallback.
So the mg7 decay generator has been off, unconditionally and silently, since
that merge -- which also means the nb_core > 1 demotion that is documented as
the only thing turning mg7 off was unreachable code. Found by running the PR's
own benchmark with `--set decay_generator=madevent`, which failed on the `set`
line rather than on anything to do with the backend.
Moved back to the end of default_setup, where f96f3b6 put it, with tests that
the parameter exists, that a card can set it, and that it still only takes the
two backends.
MadSpin unit suite: 557 pass before, 560 after.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Turns decay_generator = mg7 on for everybody, by making its pools look like
every other backend's instead of teaching the refill path a second format.
WHY LHE AND NOT .npy
A decay pool has to survive being handed to another process BY PATH ALONE.
Twice: once across the per-particle generation fork (_reader_paths ->
_reader_from_paths), and again from a channel's refill owner to every waiter,
which opens its own slice by path and has no other channel back to the
generator. What travels with the path has to include the channel's partial
width -- it is what picks a channel among a pdg's and what sets the branching
ratio. An LHE file carries it in its <init> block. A .npy has no banner, so the
width would have to become a sidecar file in the owner->waiter publish
contract: an extension of the one protocol surface where a publish/open race
can live, and where PR #378 already had to fix one.
Verified on a real refill rather than argued: at 100k events the t~ channel ran
dry, and each of the 18 slices of Events/ms_refill_1 came back carrying
1.4602897691e+00 in its own <init> block. Nothing had to be published beside
them.
The price is ~4 us an event of parsing. Measured here at 8.7% of wall for the
whole LHE parser at 100k, of which the pools are part.
WHERE THE POOL IS WRITTEN
mg7's combine_to_lhe writes one file; MadSpin deals it out round-robin into one
file per unweighting worker under Events/<run_name>/unweighted_events_<i>.lhe.
That is not a cosmetic tidy-up:
* A worker handed a single file has to STRIDE it, and _StridedEvents skips by
calling next() on the other workers' events -- so each of 18 workers parses
the whole pool to read 1/18 of it. One file per worker is what madevent
already arranges through the run_card's nb_unweight_output.
* Those paths ARE _refill_pool_paths for run_name = ms_refill_<gen>, so on a
refill _generate_refill_pool finds sources == targets and skips
_materialise_refill_pool entirely. mg7 lands in the same branch madevent
does, not the gridpack one. Confirmed on the 100k run above: 18 files, at
the canonical paths, no .part left behind, gen 1 published.
The split costs one raw-line pass (parsing off) over the pool.
WHAT DOES NOT CHANGE
The refill machinery: not one line. NpyDecayPool and _reader_from_paths' numpy
branch stay -- they are the reader for mg7's other output format and are
exercised by the tests -- but no shipped path produces a .npy pool now.
Reproducibility does NOT improve. RunCardMG7 declares no seed parameter (there
is an 'iseed' accessor reading self['generation']['seed'], but no such
add_toml_param), and madspace seeds std::mt19937 from std::random_device per
thread. So main's fixed-owner + deterministic det_seed design still protects a
property mg7 pools do not have: gen N gets the same OWNER no matter who runs
out first, but not the same events. mg7 pools are compatible with that
machinery, not reproducible under it. That is an mg7/madspace change (a seed in
RunCardMG7, plumbed to the C++ RNG), not a MadSpin one, and it equally affects
the nb_core = 1 behaviour that already shipped.
THE DEMOTION IS GONE
The nb_core > 1 -> madevent demotion existed only because the refill path could
not read a .npy. Removed, along with its warning. Gridpack (ms_dir) still
forces madevent: it drives the decay directory through run.sh, which the mg7
output does not provide.
MEASURED, p p > t t~ with both tops fully decayed, spinmode PA, seed 42,
18 cores, same production sample, back to back:
madevent mg7/LHE speedup
10 000 44.37 s 25.57 s 1.73x
100 000 94.79 s 41.36 s 2.29x
Both baselines already have main's parallel unweighting in them, which is why
they are not the 79.5 s the mg7 work was originally measured against; against
that serial baseline the same 10k run would read 3.1x. At 100k the gain is
where the original work said it would be -- decay_loop 33.20 -> 18.08 s, i.e.
the refill collapse, mg7 returning exactly the number of events asked for
instead of madevent's 0.8x undershoot.
Physics agrees. Branching ratio 0.919455 vs 0.917999 at 10k and 0.917185 vs
0.917978 at 100k; cross section after decay 463.73 vs 463.00 pb (+0.16%) and
462.73 vs 463.13 pb (-0.09%).
MadSpin unit suite: 560 pass before, 566 after.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…rkers The per-worker slices hold, between them, every event of the file mg7 wrote -- _split_pool_round_robin reads the source to the end or raises, and the byte totals match to within the 17 extra copies of the banner (241,685,010 in vs 241,829,748 out for the 254,250-event t~ pool of the 100k benchmark). Keeping the source beside them therefore just doubles what a pool costs on disk, and a refill ADDS a generation rather than replacing one, so it doubles again for every refill: 461 MiB in one channel's Events/run_01 on that run, where 241 was enough. Nothing reads it after the split. Each slice carries its own banner (that is what makes the partial width travel with the path), and the run directory keeps info.json and the generation log, so the provenance is not in this file. Removed once every slice is confirmed present, and only on the split path -- a serial run is handed mg7's file itself and must keep it. MadSpin unit suite: 566 pass before, 568 after. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
… dying
Consequence of the previous two commits, found by CI rather than reasoned
about: with decay_generator restored as a real option and the nb_core demotion
gone, mg7 is now what every MadSpin run actually uses -- and mg7 integrates
through madspace, a compiled C++/pybind extension MadGraph builds on demand
rather than shipping built.
Where it is not built, bin/generate_events tries to build it from inside the
run. When that works it costs minutes of cmake; when it does not, madevent.py
raises at import and generate_events_mg7 turns it into
InvalidCmd: the mg7 decay generator failed in .../decay_6_0 (exit 1)
which fails the whole run AFTER MadSpin has generated its own matrix elements
and has nothing to show for them. The MadSpin parallel factory workflow is
exactly such an environment -- it installs meson/ninja and runs, with no
build_madspace step, unlike acceptancetest_mg7.yml -- and all seven of its jobs
died this way. So would a user on a fresh checkout: before these commits they
silently got madevent (decay_generator was dead code), so shipping mg7 without
this is a robustness regression, not just a CI gap.
generate_events now probes for the built extension and falls back to madevent
with a warning naming the one-off fix. The probe looks for the compiled
_madspace_py*.so rather than for madspace/install/madspace, because a failed
build leaves that directory behind, empty. Memoised on the class: it asks about
the installation, which does not change during a run, and it is consulted once
per decay channel.
Note this leaves mg7 + the parallel unweighting uncovered by CI: the MadSpin
factory matrix now falls back to madevent there rather than failing. Giving
those jobs a build_madspace dependency would cover it, at the cost of a C++
build in front of seven heavy jobs -- worth deciding deliberately rather than
as a side effect of this fix.
MadSpin unit suite: 568 pass before, 570 after. Verified separately that a
machine WITH madspace built still selects mg7 and still runs the benchmark
(ttbar_full 10k, 21.88 s).
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
T145 moved mg7's decay pools from .npy to LHE on the grounds that a pool has to
survive being handed to another process BY PATH ALONE -- across the per-particle
generation fork, and from a channel's refill owner to its waiters -- carrying the
channel's partial width with it, which an LHE <init> block does and a .npy does
not; and that supplying it for a .npy would mean a width sidecar in the
owner->waiter publish contract, "the one protocol surface where a publish/open
race can live".
That is wrong on both hops, and no sidecar is added here.
* The fork already marshals the width explicitly. _generate_decay_entry writes
{'files', 'width', 'channel_widths'} into its JSON and _generate_decays hands
it back to _reader_from_paths as ``cross``. That predates this change.
* The refill needs nothing published. The width IS read back after a refill --
get_decay_from_file weights the choice between a pdg's channels by
``channels[key].cross``, and the refill reader is what it reads it off, so it
is not, as claimed, unused -- but a channel's partial width does not change
from one generation of its pool to the next. The reader that just ran dry
already has it, and _worker_refill passes it on in process. ms_refill.gen
remains the only thing on disk the waiters look at.
WHAT THE .NPY BUYS
Not only the ~2.3 us an event of parsing it saves (measured: 6.48 us against
8.76 us building the same Event objects). A numpy pool is memory-mapped and
randomly addressable, so worker i of N takes the rows at positions i, i+N, ...
by index arithmetic. That removes the round-robin split outright rather than
making it cheaper: there is no per-worker file to write, and no other worker's
events to parse past. _StridedEvents over LHE text has to next() over them,
which is exactly why the pool had to be dealt out into one file per worker
first.
THE THREE CALL SITES
* _open_refill_slice built an EventFile from the path. It now dispatches on
the suffix and takes the channel's width from its caller.
* _count_lhe_events counted "<event" as text and so returned 0 on a .npy: the
owner's slice was capped at zero events and it refilled on its very first
draw, for ever. It presents as a refill loop, not as a type error. Renamed
_count_pool_events, dispatches, and has a test of its own.
* _split_pool_round_robin is not reached at all for a numpy pool, and now says
so rather than raising UnicodeDecodeError three frames inside the LHE
parser.
_refill_pool_paths tells the two layouts apart by what the backend actually left
in the refill directory, not by a flag: mg7 leaves one events.npy there, madevent
one LHE file per worker. Both are the canonical layout, so _generate_refill_pool
recognises either and does no further work; only the gridpack backend, which
writes somewhere else entirely, still gets materialised. Probing is race-free --
a waiter only looks once ms_refill.gen says the generation is complete.
MEASURED, p p > t t~ with both tops fully decayed, spinmode PA, seed 42,
18 cores, warm, back to back, same production sample. mg7/LHE is 554d8e9
itself, so this is the same tree either side of the format.
madevent mg7/LHE mg7/.npy npy vs LHE
10 000 35.13 s 20.45 s 19.69 s -3.7%
100 000 89.08 s 38.10 s 35.74 s -6.2%
Medians of 3 (10k) and 6 (100k) runs. One pass of the 100k set was slow for all
three backends (a machine hiccup: madevent's me_generation went 9.6 -> 15.1 s in
the same pass); dropping that pass gives 38.10 -> 35.38 s, -7.1%. Run-to-run
spread otherwise is 0.4 s at 10k and 0.7 s at 100k.
Where the 100k saving comes from, by phase: decay_mg7_split 1.1 s -> 0 (there is
no split), decay_mg7_launch 14.2 -> 12.6 s (madspace writes binary instead of
formatting LHE text), decay_loop 15.4 -> 14.7 s.
The price is memory and disk: peak RSS 259 -> 308 MiB at 100k (the mmapped pool
pages), and a .npy pool is ~35% larger than the LHE one because every event is
padded to the longest in the file. Both are transient.
Physics agrees. Branching ratio 0.9177-0.9197 against madevent's 0.917999 at
10k and 0.917999 at 100k; cross section after decay within 0.2% throughout.
The refill path was exercised deliberately rather than assumed: with
MADSPIN_OWNER_UNDERSIZE=0.9 the 10k run drove 14+ generations through
_worker_refill / _owner_generate / _open_refill_slice, and each
Events/ms_refill_<gen> came back holding exactly one events.npy and nothing
else. BR 0.917343, cross section 462.667 pb.
MadSpin unit suite: 570 pass before, 579 after.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
MadSpin: read mg7's decay pools as .npy again, and stop splitting them
Speeds up MadSpin by moving decay-event generation off Fortran MadEvent and onto
the mg7 (madmatrix/madspace) engine, then removing the overheads that dominated
what was left.
p p > t t~with both tops fully decayed (t > b w+, w+ > all alland thecharge conjugate), 18 cores:
Every number below is a back-to-back measurement against the immediately
preceding commit. This machine's load drifts by up to 40% between runs, so
same-session pairs are the only trustworthy comparison, and unrelated phases
were checked for movement each time.
Why decay-event generation moved
It was 80.5% of a run: 64.0 s of 79.5 s at 10k events, against 4.9 s for the
accept/reject loop those events feed. Most of that was fixed cost. A MadEvent
pool refill costs ~12 s of survey/refine/combine whatever its size — the
baseline paid 23.7 s to produce 4708 events.
mg7 could not do this at all: madspace rejected anything but two incoming
particles, and the launcher is built around beams, PDFs and a luminosity.
cascade the 2 → n mappings already build, with the t-channel chain absent and
the root virtuality fixed rather than sampled, so this is mostly about not
assuming the incoming count is 2. Validated against analytic phase-space
volumes: 1 → 2 massless and massive exact to 1e-10, 1 → 3 massless to 0.02%.
incoming = [pdg, None], and an outgoing offset of 3 that dropped the firstfinal-state particle.
t > b w+, w+ > all allcame out asoutgoing = [81, -81], with the b silently gone.mass,
dGamma = |M|^2/(2M)in GeV rather than a cross section in pb, and aconstant alpha_s taken from the param card rather than demanding an LHAPDF set
to evaluate a coupling.
w+ > all allenumeratesw+ > t b~,b W+ Zandb W+ h; their width is exactly zero but the mappinghas no physical point to return, so it produced NaN momenta and poisoned the
whole integral.
Result: decay-event generation 1.68x, refills 63x (mg7 returns exactly the
number of events asked for, so pools rarely run dry), total 1.98x.
The rest
misc.compiledefaults tonb_core=1,and a subprocess has only ~5 objects so
make -jalone cannot fill themachine. Subprocesses now build concurrently with the job budget split between
them: 3.30x on the launcher phase.
structured array; MadSpin builds its events from that instead of parsing.
misc.gzipused the module defaultcompresslevel=9for files under256 MB while shelling out to
gzip(level 6) above it — the same datacompressed differently depending on its size, with the common case on the slow
branch. Level 6 costs 4.5 s against 18.6 s on a 172 MB LHE and buys 4%. It also
read the whole file into memory as a str and encoded it; it streams now.
Separately, MadSpin gunzipped its read-only input and repacked identical
content at the end. 4.20x on that phase, and peak RSS halved (781 → 393 MiB).
removed;
Event.boostno longer allocates two FourMomentum objects perparticle (4.98x, within 8% of the floor set by attribute access);
traceand
scalar_multiplicationno longer call numpy to add two numbers.PY_GET_DENSITY_BATCH), used by themax-weight scan.
Physics
Unchanged, and checked directly rather than inferred from the total. Partial
widths against Fortran MadEvent at 100k:
with the mg7 errors 2.4x smaller for half the wall time. The max-weight scan
batching is bit-identical trial by trial (1475/1475, max|diff| exactly 0), as is
the first boost rewrite (39976 momentum components). The later changes are
summation-order only: complex64 epsilon for the density algebra, 1.2e-12 for the
boost.
Tests
test_madspin -p U(23),test_lhe_parser -p U(24), the madspace suite (1482,up from 1469), and four MadSpin acceptance tests. Two new regression tests:
test_output_mg7_decay_subprocess_metadata(fails on the pre-fix exporter) andtest_decay_topology.pyfor the 1 → n mappings.Three MadSpin acceptance tests fail on the development machine both before and
after this branch — the production step crashes in systematics on lhapdf's
broken Python 3.14 bindings, before MadSpin is reached. Verified against the
pre-change files.
Also worth knowing
decay_generatorin the madspin card selects the backend; it defaults tomg7, withmadeventas a one-line fallback and forced for gridpack mode.MadSpin's own accept/reject RNG is still seeded. Worth adding.
MG_LHE_TIMERS=1enables the LHE parser timers, which existed but could not beswitched on.
tests/parallel_tests/madspin_benchmark.pyreproduces every number here.setlines were written after thedonethat starts the run, so nevents,iseed and the beam energies never applied; and a crashed production run was not
detected.
What is left, with measurements
The decay loop is now ~42 s of ~64 s at 100k, and it is not the matrix element:
per trial, momenta extraction is 23 us, weight assembly 22 us, DensityMatrix
construction 7 us, and the Fortran matrix element 9 us. The largest remaining
item is vectorising
get_momenta/get_pdg(~12% of the run). Porting thedensity to the madmatrix C++ backend is bounded by Amdahl at 4.7% whatever the
vector width, and madmatrix has no density support today.
🤖 Generated with Claude Code