Skip to content

Implement concurrent MadMatrix builds across subprocesses - #61

Closed
Qubitol wants to merge 4 commits into
MadGraphTeam:mainfrom
Qubitol:madmatrix-multiprocess-build
Closed

Qubitol wants to merge 4 commits into
MadGraphTeam:mainfrom
Qubitol:madmatrix-multiprocess-build

Conversation

@Qubitol

@Qubitol Qubitol commented Aug 10, 2026

Copy link
Copy Markdown
Member

With this PR, it is possible to compile concurrently the MadMatrix libraries across several subprocesses.

The number of concurrent processes is taken from the value of cpu_thread_pool_size in the run card.

Each subprocess is compiled using misc.compile with default nb_core=1.

Backend autodetection (in case of cppauto) and commonlib targets are still realised synchronously to avoid race conditions.

Logging happens in two ways on the basis of the value of verbosity from the run card:

  • pretty: re-uses the PrettyBox from MadSpace, updating it every time a new subprocess has finished compilation, and prevents ugly box updates if the number of subprocesses makes the box height larger than the viewport; if that happens, it will automatically contract into a summary box;
  • log: prints messages one after the other

@oliviermattelaer

Copy link
Copy Markdown
Contributor

I like it ... but we need to define a strategy for this.

This is unfortunate because I was going in the opposite direction ...
I did split get_amp into multiple file to speed up the compilation (especially in presence of helicity recycling where that part is huge) and also because crossing is decreasing the number of sub-processes...

@Qubitol

Qubitol commented Aug 11, 2026

Copy link
Copy Markdown
Member Author

Yeah, fair point.

Are we really going to reach a point in which CPPProcess.cc (or whatever file contains the HELAS) is going to not be the compilation bottleneck?
Because the idea behind this PR is that even when using make -j, at some point you have only one core dealing with CPPProcess.cc, so we can speed up a bit the compilation (especially GPU builds) by compiling all of these CPPProcess.cc together across different subprocesses.

I'm convinced that this would be useful even for O(10) subprocesses.

However, if there is no bottleneck anymore, then it's just better to use make -j and that's it, I suppose.

@oliviermattelaer

Copy link
Copy Markdown
Contributor

#67 is "my" counterproposal

The main point of that proposal is to do
make -j 8
in the SubProcesses folder
this one will compile src first (to avoid race condition)
and then do
$(MAKE) -C Pdir1
$(MAKE) -C Pdir2
...

In that mode the pool of worker is shared

@Qubitol

Qubitol commented Aug 14, 2026

Copy link
Copy Markdown
Member Author

Agree, #67 is a better implementation. We can close this.

@Qubitol Qubitol closed this Aug 14, 2026
oliviermattelaer pushed a commit that referenced this pull request Aug 17, 2026
a1afe9d made the mg7 launcher compile subprocesses concurrently with a
ThreadPoolExecutor. That is unsafe: every P* makefile builds the shared common
library by recursing into the shared src/, and nothing serialises them, so N
concurrent subprocess builds means N processes compiling src/build.<backend>/
Parameters.o and read_slha.o to the same paths and linking the same
lib<...>_common.so.

It broke check_xsec_processes (ttx1j) on CI: p p > t t~ j is the only entry in
that section with more than one subprocess, so it is the only one that took the
concurrent path, and it failed with a compilation error in P0_QQx_ttxg. The
same commit passed a re-run minutes later, which is the giveaway.

The race does not reproduce on macOS (0/6 at -j4, 0/10 at -j1 with 18 cores):
src/ has two objects and the window is short. Widening it with a sleep in the
src rules shows it plainly -- four concurrent writers to each of
Parameters.o and read_slha.o.

Two further reasons to drop rather than patch this:

- PR #61 proposed the same ThreadPoolExecutor design and was closed unmerged.
  This commit reintroduced it without knowing.
- PR #67 does it properly, at the make level: SubProcesses/makefile becomes a
  jobserver dispatcher, the common library is built once up front, and P*
  directories are told so with MADMATRIX_COMMONLIB_EXTERNAL=1. Its jobserver
  also shares slots dynamically, where the static build_jobs // len(pending)
  split here gave -j1 per make on a 4-core CI runner -- neutralising the
  parallel-make win (3.4x, the larger half) while keeping the racy half.

Removes compile_subprocesses, build_subprocess, resolve_api_path, build_jobs
and the concurrent.futures import; MadgraphSubprocess's compile loop and
init_subprocesses are byte-identical to main again, so this file no longer
conflicts with #67 (verified with git merge-tree).

Everything else a1afe9d added is unrelated to the build and stays: the
decay-phase timing split, and on this branch the mg7 decay mode, clean_pids,
drop_closed_channels and skipping systematics for decays.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants