Skip to content

Merge upstream mg5amcnlo 3.8.0 into main - #117

Merged
oliviermattelaer merged 87 commits into
mainfrom
merge_380
Sep 11, 2026
Merged

oliviermattelaer merged 87 commits into
mainfrom
merge_380

Conversation

@oliviermattelaer

Copy link
Copy Markdown
Contributor

Merges upstream mg5amcnlo/3.8.0 (901080da4) into main.

mg5/3.8.0 brings MG5aMC 3.7.3 + the start of 3.8.0. Nine files conflicted; the rest merged cleanly.

The main theme

mg7 (PR #115) and upstream 3.8.0 independently implemented the same feature: content-validation of the HEPTools CI cache, to stop a green warm cache run from publishing a cache with a tool silently missing. Upstream factored it into a reusable composite action (.github/actions/check_heptools + check_heptools.sh); mg7 did it inline in bash.

Upstream's version wins here: its checks are a strict content superset of ours (everything we checked — CutTools, IREGI, collier, ninja, the pythia8 header/binary — plus boost, YODA, hepmc3, the PDF sets, and lib/lib64 variants), and it is one shared file instead of a dozen inline test -f blocks. Upstream also splits boost into its own boost_cache job, so a flaky boost download no longer discards the eMELA build.

mg7's orthogonal CI policy is re-applied on top:

  • ubuntu-24.04 only (no 22.04 matrix leg)
  • --ref "${{ github.ref }}" on every gh cache delete
  • the github.ref_name == 'main' || 'test_ci' || workflow_dispatch job gating
  • ./bin/madgraph and input/mg7_configuration.txt (bin/mg5_aMC does not exist in mg7 — upstream's new boost_cache and part of emela_cache would have hard-failed)
  • the looptools-<os> entry in the reset_heptools delete list, which upstream still lacks
  • boost-<os> added to the delete/cleanup lists, since boost now has its own cache

Per-file resolutions

File Side Why
VERSION theirs ours was stale at 3.7.2
UpdateNotes.txt theirs verified strict superset — contains both mg7 BUG FIX entries verbatim, plus 4 more and the new 3.8.0 section
Template/NLO/SubProcesses/genps_fks.f theirs both sides made the identical virtgranny_red EXTERNAL fix, only continuation-line order differed
tests/acceptance_tests/test_cmd_madevent.py theirs both agree cross=15.73; the error=0.04 set on 26 Jun was loosened to 0.514 upstream two days later in 75d821554 "fix CI test"
tests/acceptance_tests/test_cmd.py theirs --no_open
madgraph/interface/madgraph_interface.py theirs + ours upstream's do_draw rework (--no_open, --merge→PDF) had already won everywhere else in the file via auto-merge; mg7's --generate_only is kept as an alias so the acceptance tests and the flag both keep working
.github/actions/restore_heptools/action.yml theirs see above
.github/workflows/acceptancetest.yml ours (both hunks) net diff vs main is zero — see below
.github/workflows/warm_cache.yml mixed 18 hunks, resolved individually

Why acceptancetest.yml takes ours

Two traps in that file:

  1. The acceptancetest_uncovered_merged: job header is unconflicted, but upstream's side of the conflict is the body of acceptancetest_heft (a job mg7 deliberately pruned, and whose header the auto-merge had already deleted). Taking upstream would have grafted the heft steps under our job's name and deleted five mg7 jobs — uncovered_merged, uncovered_madloop, improve_ps_at_rest, check_madevent_me_vs_standalone, check_mlm_rwgt.
  2. Upstream appends acceptancetest_delphes_parallel: at EOF, but that job already exists intact at line 1288 — a duplicate YAML key, which kills the whole workflow.

Two problems found outside the conflicts

Both are clean auto-merges that produce broken YAML, so they would not have shown up as conflicts:

  • looptools_cache had the check twice. Git merged our inline for f in … loop (5-space indent) and upstream's check_heptools step (7-space indent) into the same job, same step name. warm_cache.yml did not parse. Resolved to upstream's action, with its if: cache-hit != 'true' dropped so it still runs on the cache-hit path — that was the deliberate behaviour in the mg7 comment (an already-stored incomplete cache gets reported rather than silently reused).
  • heptools_cache had a duplicate if: key. Upstream's block carries its own if: always() && … while mg7 already has one on that job. Removed upstream's; the always() semantics live in ours.

Also: Restore lhapdf (fallback when pythia8 and emela failed) was dropped by the auto-merge. That is fine — upstream restores lhapdf unconditionally first, which subsumes it.

Verification

  • no conflict markers in any file touched by the merge
  • all workflow/action YAML parses: 13 jobs in warm_cache.yml, 73 in acceptancetest.yml
  • no duplicate step names or undeclared needs.<job> references introduced (the one duplicate in acceptancetest_density predates this merge on main)
  • every uses: ./.github/actions/X resolves to an existing action.yml
  • no bin/mg5_aMC / mg5_configuration.txt references left anywhere under .github/
  • Python files parse; check_heptools.sh passes bash -n

The acceptance tests themselves were not run locally — they need a full CI environment, so CI on this PR is the real check.

🤖 Generated with Claude Code

oliviermattelaer and others added 30 commits July 3, 2026 06:16
…rnal leg

{G}(4), {H}(5), {Q}(6), {W}(7) and {S}(9) each name a rank-two piece of the
*propagator* numerator of a massive vector -- the metric, the Theta projector,
the qq/q^2 tensor, the Ward-protected full propagator and the scalar piece --
complete with its 1/(q^2-M^2+iM*Gamma) pole (aloha/create_aloha.py, the "1X"
forms). None of them is a polarisation vector, and no external wavefunction
for them exists: VXXXXX only knows nhel = -1,0,+1.

Put on a genuine external leg they were nevertheless accepted, and the raw
integer went straight into the NHEL table:

    generate p p > z{G} h   ->   DATA (NHEL(I,1),I=1,4) / 1,-1, 4, 0/
                                 CALL VXXXXX(P(0,3),MDL_MZ,NHEL(3),...)

VXXXXX then computes hel=4, hel0=1-|hel|=-3 and returns a vector that is
neither normalised nor transverse. The process compiles and check_sa prints a
matrix element: no crash, no warning, a silently wrong number. Initial-state
legs behaved the same way.

A decay-chain-aware walk over the ProcessDefinition now refuses them, so that a
leg handed over to a decay chain (the propagator use, 't > w+{G} b,
w+ > ta+ vt') keeps working while an external one is refused with a message
naming the tag. There is no option to enable them on an external leg: there is
nothing to enable. The same guard is repeated in HelasWavefunction for the
direct-API path that never reaches the command interface.

'{A}' (99) is deliberately left out of the new veto: it is already refused on
an external leg by HelasWavefunction and keeps that behaviour unchanged.

Also fix the round-trip these braces never had: nice_string()/input_string()/
base_string() printed the raw integer ('{4}', '{99}', '{0,9}'), which the
parser rejects with "polarization are between -3 and 3" -- so a printed
process line could not be read back. They now print the brace letter with no
separator ('{G}', '{A}', '{0S}'), which is exactly the syntax the parser
accepts; a ',' inside a brace additionally breaks the decay-chain split.
shell_string() is left alone so that P<n>_... directory names do not change.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…-380

MG5: refuse the propagator-only polarisation braces {G},{H},{Q},{W},{S} on an external leg
…and-build-fixes

Improve HwU processing and add Matplotlib PDF output
…owner

_scan_maxwgt_parallel cut the probe events into ceil(N/nb_core)-sized
chunks and skipped the empty trailing ones, but still handed each worker
the original nb_core. That number is the pool-addressing count (the decay
pool is split into nb_core files and a worker opens paths[shard_id]) and
it is also the modulus of _channel_owner, so ids past the last forked
worker stayed nameable as the sole regenerator of a decay channel.

At the shipped default of 75 probe events that is one unforked id on 16
cores, three on 18 and seven on 32. A live worker that runs its pool out
on such a channel blocks on a process that was never started, and neither
fail-safe rescues it: _read_worker_status returns None -- not ('D',) --
for a status file nobody ever wrote, and _wait_cycle_to_self finds no
wait-for chain to follow. It waits out MADSPIN_REFILL_WAIT, an hour by
default, and takes the scan down.

Slice the probe events into exactly nb_core balanced contiguous ranges
instead, so the two counts agree by construction and every owner id
belongs to a running worker. Where the split is even -- what both scans
arrange by rounding their probe size up to a multiple of nb_core -- these
are the same slices as before, so no sample is re-diced; the uneven cases
are the ones that were dying.

Also, belt and braces: if a slice ever does come back empty, the parent
now leaves 'D' on that id. A missing status file cannot be told apart
from an owner still working, which is what makes the wait unbounded; 'D'
makes _worker_refill's existing fail-safe fire at once instead.

_run_onshell_parallel is not affected: it reassigns nb_core to the number
of shards it actually launches before forking.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Added support for merging diagrams and text files into single outputs.
Co-authored-by: sourcery-ai[bot] <58596630+sourcery-ai[bot]@users.noreply.github.com>
Both the output PDF and every input `.eps` file are converted to absolute paths before being passed to `gs`. This guarantees they all start with `/` so Ghostscript can never misinterpret them as CLI flags 

**Explicit `shell=False`** passed directly to `subprocess.check_call`. Combined with the **list-form** argument vector (`['gs', ...]` instead of a string), this eliminates any shell-metacharacter injection risk entirely (no shell is ever invoked, so `$()`, `;`, `|`, backticks, etc. are harmless).
fixed space-in-path bug while opening merged pdf files
Add directory creation for output paths in diagram commands.
Bring the 3.8.0 release branch (57 commits) into the MadSpin density
rewrite. Nine files conflicted; the resolutions were:

 - VERSION: taken from 3.8.0 (3.7.3 / 2026-09-01) -- the release branch
   owns version metadata.
 - UpdateNotes.txt: 3.8.0's file taken wholesale, and the two MadSpin
   density entries that madspin_density had filed under a placeholder
   "3.7.2 (XX/XX/XX)" section re-filed at the top of the now-current
   3.7.3 section.
 - madgraph/core/base_objects.py: kept both. 3.8.0 introduced
   polarization_to_string() so that propagator-only polarizations
   ({G},{H},{Q},{W},{S},{A}) round-trip through nice_string(); the
   density branch had restructured the same four sites to emit its
   'offshell' '*' marker. All four now use polarization_to_string()
   inside the density branch's structure.
 - madgraph/interface/madgraph_interface.py: kept both -- the density
   branch's set2_nb_core() body (cpu_count when None/0) plus 3.8.0's new
   set2_nb_core_pythia8() / set2_nb_core_delphes() setters.
 - madgraph/various/misc.py: kept the density branch's form. Both sides
   independently fixed the same 'finally: return' that swallowed
   exceptions in stdchannel_redirected(); the two fixes are equivalent.
 - madgraph/interface/common_run_interface.py: whitespace only.
 - .github/workflows/acceptancetest.yml: kept both -- the density jobs
   and 3.8.0's acceptancetest_delphes_parallel.
 - .github/workflows/warm_cache.yml: kept both new inputs (reset_pip,
   reset_delphes) and both new steps, and adopted 3.8.0's policy of not
   resetting caches on the weekly schedule for the pip step too, so the
   delete-cache job stays internally consistent.
 - tests/acceptance_tests/test_cmd_madevent.py: kept both -- 3.8.0's new
   _get_delphes_path()/test_pythia8_delphes_parallel(), and for
   test_loop_induced_ggh the density branch's measured central value
   (15.73) with 3.8.0's deliberately widened tolerance (0.514).
…e_string

The '*' (offshell) marker was appended in the "if" branch and the leg
separator in the matching "else", so an offshell leg was rendered without
the space that separates it from the next one. Two consequences:

  * mid-string, 'u u~ > z* a' came out as 'u u~ > z*a', which the process
    parser rejects loudly ("No particle z*a in model");
  * as the last leg, the "remove last space" slice at the end of
    nice_string() chopped the '*' instead of a space, so 'u u~ > a z*'
    came out as 'u u~ > a z' -- the offshell marker silently lost, and the
    string still re-parsing as a valid but different (onshell) process.

The three sibling renderers -- Process.nice_string(), Process.input_string()
and Process.base_string() -- already emit '*' followed by a space; this
aligns the fourth with them. The misbinding dates from f631152, which
inserted the offshell test between the polarization block and the "else"
that used to belong to it.

Tests pin all four decorations (none, polarization, offshell, both) across
the four renderers, plus the last-leg case.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…space

ProcessDefinition.nice_string: keep the space after an offshell leg
…able

When the max-weight scan caps its worker count at the number of probe
events, the decay pool has already been written with the uncapped count,
so _reopen_decay_pool sees len(paths) != nb_core, drops the own-file fast
path and strides the chained pool. That is correct (the stripes stay
disjoint) and cheap (0.04-0.18 s for 16-64 workers over a scan-sized
pool, one-off: refills go through _open_refill_slice, which keys on the
post-cap count), but it was silent and undocumented.

Record at the cap why the mismatch is accepted and why both alternatives
are worse -- re-splitting the pool ties its shape to whichever phase caps
first, and a second uncapped count for pool addressing reinstates the
divergence between the pool-file count and the _channel_owner modulus
that produced the phantom owner and the 3600 s hang the cap fixes.

Log the fallback at debug level in _reopen_decay_pool, naming both
counts, so the degradation is visible rather than silent. No count logic
changes.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Follow-up to the --no_open / --merge options added in the previous
commits.

- draw(): let _draw_parser consume the options before check_draw, so
  that only the output directory is left in args. The hand-rolled flag
  filtering treated any other option as a directory, so
  'display diagrams --horizontal' created a directory literally named
  '--horizontal' and then crashed with an IndexError. check_draw is
  back to its original shape and unknown options now give a clean
  error.
- Same treatment for 'display diagrams_text' through a small dedicated
  _display_text_parser.
- Restore self.exec_cmd('open ...') instead of calling the macOS-only
  'open' binary directly: misc.open_file is cross platform, honours the
  eps_viewer configuration (including the pstopdf/ps2pdf conversion on
  mac) and opens the file in the background. This also removes the
  shell escaping of the filename, which was wrong for a list-form
  subprocess call. Opening the merged PDF is moved out of the
  ghostscript try block so that an 'open' failure is not reported as a
  merge failure.
- Forbid --no_open / --merge (and a directory for diagrams_text) for
  NLO processes: born, real and loop amplitudes are drawn by separate
  calls, so --merge would have each type overwrite the previous one,
  and the NLO diagrams_text is a pager-only implementation which would
  silently ignore both the directory and the flags.
- CheckLoop.check_display: count the positional arguments only, so that
  a trailing option no longer makes 'display diagrams . --no_open' fail
  with "Can only display born or loop diagrams, not .".
- Add --no_open to the 'display diagrams' acceptance test, so that it
  no longer opens an eps viewer while running.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
oliviermattelaer and others added 3 commits September 9, 2026 22:36
CI: generation-suffixed cache keys with prefix fallback, removing the delete-then-save window
MadSpin: fix the onshell_v1 identical-particle factor, and assert the flavour composition
Picks up the ZZ fix (b12a8a4, "MadSpin: undo MG5's decay-chain identical-
particle factor in onshell_v1") that resolves the test_short_madspin_zz failure
left open on the previous merge, plus the CI cache rework.

Conflicts (3):

* tests/unit_tests/madspin/test_madspin.py: both sides appended a block at EOF
  over an empty base. Kept both -- mg7's 7 numpy-pool classes and upstream's
  TestDecayChainIdenticalFactor. The two stub fixes from 9959eef are
  superseded upstream and merged cleanly.
* .github/actions/restore_heptools_contur/action.yml: upstream, with mg7's
  input/mg7_configuration.txt. Brings the generation-suffixed key and the
  guard that stops MG5 being pointed at a fastjet-config the cache did not
  restore.
* .github/workflows/warm_cache.yml: upstream's cache scheme, mg7's policy.

On the cache rework (3513ae7): it applies to mg7 unchanged in intent. The bug
it fixes is present here verbatim -- heptools_cache deleted the combined cache
before saving a fresh one, and every consumer restored with an exact key and no
restore-keys, so for the ~30 min in between nothing matched at all. On that miss
restore_heptools_contur still wrote a fastjet-config path that does not exist,
makefile_fks_dir saw "ifdef fastjet_config", and every aMC@NLO SubProcess
compile died on fastjet/ClusterSequence.hh. Generation-suffixed keys with a
prefix restore-key remove the window entirely, and with it the delete-before-
save that was the fragile half of #115. It also fixes reset_ufo, which targeted
"ufomodel-$ImageOS" while the key written was a bare "ufomodel" -- a no-op here
too. Storage pressure is lower in mg7 than upstream, since mg7 builds only
ubuntu-24.04, and the extra sub-cache rebuild costs nothing because the cleanup
step already deleted every sub-cache on the success path.

mg7 policy re-applied on top: ubuntu-24.04 only, the ref_name job gating,
./bin/madgraph and input/mg7_configuration.txt, the meson/ninja pip install for
numpy's f2py backend, and --ref scoping on every cache delete -- threaded into
the new prefix-delete helper as "gh cache list --ref "$GITHUB_REF"", so a reset
still only touches the ref it runs on. Verified no job's if/runs-on/needs
changed against the pre-merge tree.

tests/unit_tests/madspin/test_madspin.py: 608 tests, same 1 failure + 10 errors
as the pre-merge tree on this machine (float32 rounding and a numpy/py3.14
mismatch, both local); the 8 tests the merge adds all pass.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@oliviermattelaer

Copy link
Copy Markdown
Contributor Author

CI green after the re-merge — 278 pass, 0 fail

Every failure from the first round is resolved.

First round Now How
madspin_factory (test_short_madspin_zz) pass upstream b12a8a45d — the legacy onshell_v1 decay-chain route carried the extra identical-particle factor, not the density
unittest_10, unittest_madspin_sequential pass the two stub fixes, superseded by upstream and merged cleanly
acceptancetest_delphes_parallel pass the apt Hash Sum mismatch flake, gone on re-run
IOtest_TIR, IOtest_matchbox pass not touched — no golden regenerated. They were almost certainly the cache-window bug 3513ae7bc describes: running IOTests was red on main at the same time, and upstream saw the same signature (13 acceptance jobs red, then green on a re-run of the identical commit)
acceptancetest_96 pass on re-run pre-existing concurrency race, see below

acceptancetest_96 — a pre-existing race, not this merge

test_madspin_ON_and_onshell_atNLO failed once and passed on a plain re-run of the same commit. The mechanism is worth recording because the symptom points nowhere near the cause.

The decay-ME build died with two errors, both at column 72:

aloha_functions.f:1143:72  ...1d-99))*rHalf   -> Error: Missing exponent in real number
cuts.inc:6:72              ...mue_over_ref... -> Error: Symbol 'mu' has no IMPLICIT type

Fixed-form Fortran truncating at 72 columns, i.e. madspin_me/Source/make_opts was not in effect. Three independent symptoms in the same compile agree:

  • f77 was used — make's built-in $(FC), so FC=$(DEFAULT_F_COMPILER) never fired;
  • -ffixed-line-length-132 absent, which make_opts adds unconditionally for non-ifort;
  • the library target was ../../lib/libdhelas. with an empty $(libext), and libext is only defined in make_opts.

The job log shows five concurrent compile failures against one shared path, /tmp/amc5z2fnpa9/MGProcess/madspin_me/Source. So: five processes build the same madspin_me; make_opts is populated with shutil.copy (madgraph/__init__.py:49, madgraph_interface.py:3333), which truncates the destination before writing; and replace_make_opt_f_compiler (export_v4.py:3323) swallows the resulting IOError behind a logger.info. A process compiling inside that window gets pure make defaults.

An empty-but-present make_opts fits better than a missing one: make parsed it without complaint, and FC=$(DEFAULT_F_COMPILER) lives in the body that was truncated away, so FC stayed f77.

Nothing in this merge touches that path — MadSpin/interface_madspin.py has two hunks, both pure weight arithmetic, and amcatnlo_run_interface.py's fastjet guard writes me_dir/Source/make_opts, not madspin_me/Source, with a different failure mode. Left unfixed here deliberately: it is a real bug but out of scope for a merge, and it deserves its own change.

MG5DIR/Template/LO/Source/make_opts is shared mutable state: every MG5aMC
process using an installation rewrites it in place (MadGraphCmd.__init__
restores it from .make_opts; set_fortran_compiler / set_cpp_compiler write the
detected compilers into it), and every new output directory is seeded with a
verbatim copy of it. All of those writes truncated the destination first --
open(..,'w') in update_make_opts_full, shutil.copy everywhere else -- so a
reader landing in that window copies out an empty make_opts.

Nothing complains when that happens. An empty make_opts still *parses*, so make
simply keeps its builtins: $(FC) stays f77 (the ifeq ($(origin FC),default)
block that sets FC=$(DEFAULT_F_COMPILER) lives after the
#end_of_make_opts_variables marker, i.e. in the part that was truncated away),
$(libext) is undefined, and -ffixed-line-length-132 is never added. The build
then dies far from the cause, as column-72 errors in unrelated fixed-form
Fortran:

  aloha_functions.f:1143:72  Error: Missing exponent in real number
  cuts.inc:6:72              Error: Symbol 'mu' at (1) has no IMPLICIT type

and links against ../../lib/libdhelas. with an empty extension. Seen once in
CI on a MadSpin decay-ME build (madspin_me/Source), where MadSpin forks a
worker per core for the maxwgt scan and each worker's refill runs a full
"output <decay_dir> -f", i.e. N concurrent rewriters of the one Template file
while another export is copying it out.

Worse, the damage outlives the race: update_make_opts_full read the file line
by line and wrote back what it had read, so a partial read was persisted as a
make_opts with a variables block and no body at all.

  - misc.atomic_write / misc.atomic_copy: write a temp file in the destination
    directory and os.replace() it, so a reader always sees the whole old file
    or the whole new one. Used at every make_opts write: madgraph/__init__,
    MadGraphCmd.__init__, the auto-update restore, update_make_opts_full and
    ProcessExporterFortranSA.copy_template.
  - update_make_opts_full reads the file in one go, and raises instead of
    rewriting when what it read is empty, has lost the
    #end_of_make_opts_variables marker, or defines no DEFAULT_F_COMPILER.
  - replace_make_opt_f_compiler / replace_make_opt_c_compiler raise for the
    process directory instead of logging 'Fail to set compiler. Trying to
    continue anyway.' at info level. Continuing there can only produce the
    confusing column-72 failure above. MG5DIR/Template keeps the tolerant
    path: a shared install is legitimately read-only.

Drive-by: the madgraph/__init__ block used shutil without importing it, so on
a pre-2016 install its os.remove() took effect and the copy that was meant to
follow raised NameError into the bare except. It no longer removes anything.

tests/unit_tests/various/test_make_opts.py covers both halves: a reader racing
a writer never observes a partial make_opts (this fails against shutil.copy /
open(..,'w')), and a make_opts that has lost its body is refused rather than
rewritten.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
oliviermattelaer and others added 7 commits September 10, 2026 09:56
…vacuously

Both from the review on #411.

update_make_opts_full accepted a file that still had its variables and the
#end_of_make_opts_variables marker but had lost everything after it, and wrote
that straight back -- the same permanent corruption the validation was added to
prevent, one truncation offset away:

  in:  DEFAULT_F_COMPILER=gfortran\n#end_of_make_opts_variables\n
  out: DEFAULT_F_COMPILER=gfortran-14\nDEFAULT_F2PY_COMPILER=f2py\n#end_of_...

Require the body to still carry the two definitions whose absence produced the
reported failure: FC=$(DEFAULT_F_COMPILER) (else make compiles with its builtin
f77) and libext= (else the libraries link as '../../lib/libdhelas.'). Both are
unconditionally present in the only two make_opts sources in the tree,
Template/LO/Source/.make_opts and Template/NLO/Source/make_opts.inc.

The two concurrency tests started the writer before the reader and stopped on
the writer, so the writer could complete every iteration before the reader was
first scheduled and the test would pass having read nothing. They now share a
race() helper that meets both threads at a barrier and keeps writing until the
reader has managed MIN_READS reads, and fails outright if it never got there.
Measured 0/20 vacuous runs here before the change, so this was latent rather
than flaky -- but a single-core runner would not have been so lucky.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Three consecutive warm_cache failures on 2026-09-09 left the repo at 10.26GB
against GitHub's 10GB per-repo cache limit, and the eviction that followed took
out heptools-ubuntu24 — which is why every ubuntu-24.04 job that needs lhapdf
has been red since, on every branch:

  Exception: lhapdf executable (/home/runner/.cache/HEPtools/bin/lhapdf-config)
             is not found on your system

The intermediate caches (lhapdf/boost/pythia8/emela/contur/looptools, ~3.5GB
per run across both images) are deleted at the end of a SUCCESSFUL run by
heptools_cache, and deliberately kept on a failed one so that re-running the
one broken job does not rebuild the whole chain. Nothing ever collected them
afterwards, so each failure stranded another 3.5GB — and the loop closes on
itself: over quota, LRU evicts, run 34402367805 lost its own looptools
sub-cache 27 minutes after saving it, which failed that run, which stranded
another 3.5GB.

  20:56:15  Cache saved with key: looptools-ubuntu24-34402367805
  21:23:33  Cache not found for input keys: looptools-ubuntu24-34402367805

  - delete-cache now sweeps intermediate caches left behind by earlier runs,
    bounding the leak to one generation while keeping the re-run affordance:
    this run's own id is skipped, and so is any run still in progress, so a
    warm_cache running concurrently on another branch keeps its sub-caches.
  - the six sub-cache restores in heptools_cache get fail-on-cache-miss. The
    job gate already establishes the entry must exist (the sub-cache job
    reported success, and it saves unconditionally on success), so a miss means
    eviction and should say so. Without it the restore quietly does nothing and
    the failure surfaces two steps later as

      HEPTools cache incomplete: missing file .../CutTools/lib/libcts.a

    which sends you looking at CutTools instead of at the quota. That is
    exactly how this outage presented.

The 20 stranded entries (8.98GB) have been deleted by hand to unblock CI; this
is what keeps them from coming back.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
CI: stop failed warm_cache runs leaking the repo past its cache quota
Make every Source/make_opts write atomic, and refuse a body-less one
warm_cache's push trigger was filtered by path but not by branch, so any branch
that happened to touch warm_cache.yml or one of the restore_*/check_heptools
actions started a full warm run of its own. Every such run rebuilds ~3.5GB of
intermediate caches across both images against a 10GB per-repo quota, and two
overlapping runs are enough to exhaust it and start evicting the shared
heptools caches -- which is what took out heptools-ubuntu24 and turned every
ubuntu-24.04 lhapdf job red.

It is not hypothetical: merging the previous fix put a warm run on 3.8.0 while
the one its own feature-branch push had started was still queued.

Restrict the push trigger to 3.x and test_ci. Warming from anywhere else stays
available through the workflow_dispatch button, where it is a deliberate act
rather than a side effect of editing a workflow file.

The schedule trigger needs no equivalent filter: GitHub only ever runs cron
workflows on the default branch, which is 3.x. There is no pull_request,
workflow_call or workflow_run trigger, here or in any other workflow, so these
three are the whole surface.

test_ci does not exist yet; naming it costs nothing and the filter starts
applying the day it is created.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…lter

CI: warm the caches automatically only on 3.x and test_ci
HelasWavefunction.check_majorana_and_flip_flow ends with an is_loop-only
sweep that repoints, inside the current diagram, every mother that still
refers to the wavefunction just superseded by a reused copy. It selected
those mothers with mother_wf.get('number') == new_wf_number.

That key is not stable. The insertion branches earlier in the same
function renumber every wavefunction after the insertion point by +/-1 to
keep diagram_wavefunctions contiguous, and the reuse branch immediately
above pops the copy and decrements everything after it. By the time the
sweep runs, the slot that held new_wf_number+1 holds new_wf_number and
belongs to an unrelated wavefunction, whose mother is then silently
overwritten.

Match on object identity instead. Reported as
https://bugs.launchpad.net/mg5amcnlo/+bug/2159043, where a color-sextet
diquark model has enough flip-and-renumber churn per loop diagram to land
a collision on a live wavefunction: a g t t~ vertex loses its gluon
mother to a top line from the |dF|=2 vertex, and sort_by_pdg_codes then
raises "ValueError: 21 is not in list". The same collision has been
firing harmlessly in MSSM output since the sweep was added in 961564b.

Validated over 42 generation configurations (loop_sm, LoopSMEWTest,
loop_smgrav, loop_MSSM, plus the reporter's model), both
loop_optimized_output modes, comparing a full HELAS fingerprint before
and after: 41 byte-identical, and the reporter's process goes from crash
to clean. Only gluino loop lines reach the sweep at all; where they do,
g g > go go g changes 41 substitution decisions and u u~ > go go g
changes 85, with no change to the generated matrix element.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@oliviermattelaer oliviermattelaer added this to the Alpha release milestone Sep 10, 2026
oliviermattelaer and others added 8 commits September 10, 2026 13:58
HelasWavefunction.__eq__ deliberately ignores the pdg code, so a particle
and its antiparticle compare equal whenever nothing else distinguishes
them: same spin, mass, width, colour, state and mothers. For a coloured
fermion the colour representation separates the two, but for a colour
singlet -- a lepton -- it does not.

check_majorana_and_flip_flow reuses an already-built wavefunction with

    new_wf = wavefunctions[wavefunctions.index(new_wf)]

which therefore can return the wrong sign for a loop wavefunction sitting
on a lepton line. The wavefunction that ends up in the diagram then has a
particle identity inconsistent with the interaction it was built from, and
sort_by_pdg_codes later fails with "ValueError: ... not in list".

Require the pdg code to match when reusing a loop wavefunction. This was
reported in appendix A of arXiv:2108.11404 (scalar leptoquark pair
production at NLO), where it is worked around in a patched copy of the
code; the paper notes the bug was acknowledged upstream, but it was never
fixed. The workaround given there guards only pdg < 0, and nothing makes
the other sign safe, so the check is applied to every loop wavefunction --
instrumenting both variants over the process list below shows they select
the same wavefunction everywhere, so the wider guard costs nothing.

Reproducer, with the model from the paper (LQnlo_5FNS_v5_UFO) and its
is_perturbating addition so that lepton loops are not vetoed:

    generate d d~ > LQ1d LQ1d~ [virt=QCD]

crashes for LQ1d, LQ1dd, LQ3d and LQ3dd from d d~ and b b~, and for LQ3u
from u u~; the LQ2 family is unaffected. All nine build with this commit.

This is independent of the mother-substitution bug fixed in the previous
commit: with only that one applied the nine still crash, and with only
this one the sextet-diquark reproducer of LP#2159043 still crashes.

The reuse branch is only reached when a fermion flow is flipped, which no
SM, EW or graviton-model loop process does. Instrumented over 42
configurations (loop_sm, LoopSMEWTest, loop_smgrav, loop_MSSM and the
sextet-diquark model, both loop_optimized_output modes), the guarded
branch is entered 0 times for all of those and for the squark processes,
and up to 948 times for gluino ones -- but never selects a different
wavefunction there, so no generated code changes outside the leptoquark
case.

Tests: tests/test_manager.py -p U loop (24/24), the MadLoop and FKS IOTest
golden groups, and test_ML5MSSMQCD for g g > go go, g g > go go g and
g g > n1 n1 against the stored references.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
tests/input_files/loop_MSSM still carries the python2 compatibility
shims: object_library.py imports six and uses six.iteritems, and
write_param_card.py imports six.moves.range.

MG5aMC loads a UFO model with a sanitised sys.path that excludes user
site-packages, so an interpreter that has six only there still fails:

    models.UFOError: No module named 'six'

and the parallel MadLoop tests that use this model (test_ML5MSSMQCD)
cannot run at all. Worse, the in-place python2 to python3 conversion that
MG5aMC then attempts leaves models/loop_MSSM/__init__.py with

    import function_library
    import object_library

above its 'from __future__' line, which is a syntax error on the next
run, so the failure persists until the copied model is deleted by hand.

Replace six.iteritems with dict.items and drop both imports. This is the
same change MadGraph7 already applied to its copy of the fixture in
d8ec200 and 7e0a11f; the file is now identical there and here.

Checked with 'python3 -s' (user site-packages hidden, so six is not
importable): the model fails to import before this commit and imports
cleanly after, and test_short_mssm_vs_stored_HCR_gg_gogo_QCD passes.

Note DM_pion, fourfermion_UFO, full_sm_UFO and loop_smgrav still import
six; they are left alone here.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Drop the six dependency from the loop_MSSM test fixture
…ision

Fix two independent loop-wavefunction bugs in check_majorana_and_flip_flow (LP#2159043, arXiv:2108.11404 app. A)
Auto-merged with no conflicts. Brings:

* the make_opts truncation fix (#411): every Source/make_opts write is now
  atomic and a body-less one is refused instead of silently leaving the build
  on make's f77 defaults. That is the race behind the one-off acceptancetest_96
  failure on this PR (column-72 Fortran truncation), so that job should stop
  flaking.
* two loop-ME bug fixes (#414): a loop wavefunction was reused across a pdg
  sign collision, and the mother-substitution sweep matched on 'number', which
  the surrounding insertions reshuffle -- now matched on identity.
* two warm_cache quota fixes (#412, #413).

One mg7 adaptation, in the new push filter. Upstream restricts automatic cache
warming to its default branch, "3.x"; mg7's default branch is main, so as
merged the filter would have matched nothing here and warm_cache would never
have run automatically at all. Changed to main + test_ci.

Deliberately NOT given mg7's --ref scoping: the new "Collect intermediate
caches left behind by earlier runs" sweep. Everywhere else mg7 scopes cache
deletes to the ref they run on, so a reset never touches another branch's
entries. Here the point is the opposite -- the 10GB quota is per repo, not per
ref, so a ref-scoped sweep would leave leaks from workflow_dispatch runs on
other branches uncollected and defeat the fix. Its concurrency safety comes
from the two guards it already has: it skips this run's own id, and any run
that is not 'completed'.

Verified: all workflow/action YAML parses; no job's if/runs-on/needs changed;
no ubuntu-22.04, bin/mg5_aMC or mg5_configuration.txt anywhere under .github/;
the --ref scoping in delete_caches.sh survived. tests/unit_tests/various/
test_make_opts.py 13/13 and test_helas_objects.py 65/65 pass locally.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Keeps the PR mergeable: main has moved 12 commits ahead (README/sphinx docs,
the 0.2.0 alpha release workflow, the HERWIG6 gfortran-16 SPLIT rename), and
GitHub had the PR at CONFLICTING/DIRTY.

One conflict, VERSION, resolved to main's side:

    version = 0.2.0
    date = 2026-09-07

Note this reverses the call made on the first 3.8.0 merge, which took
upstream's "3.7.3". That was right at the time -- main was still on 3.7.2 and
mg7 had no version of its own. It is wrong now: 1a401be deliberately moved
MadGraph7 to its own numbering for the alpha, and bin/create_release.py checks
VERSION against its --version argument while madspace/pyproject.toml derives
its version from the MadGraph release version. Taking upstream's 3.7.3 here
would silently break the release workflow.

Everything else auto-merged. Re-checked after the merge: the warm_cache push
filter is still main + test_ci, and all workflow/action YAML parses.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Brings main's 89 new commits: the MadGraph7 rename (#64 follow-up), PR #60
(Del Duca-Dixon-Maltoni colour basis), the CVMFS PDF-set fallback, the banner
logo work and the CI duplicate-run gate. 1137 files, and git reported no
conflict at all.

That last part is the trap. The rename landed on main while this branch was
carrying upstream files that still say MadGraph5_aMC@NLO, and the two touch
different lines, so git merges them silently: the merge reintroduced 12 lines
of pre-rename vocabulary that nothing flagged. Renamed here, following exactly
what f5a6cf3 did at the same places:

  madgraph/various/histograms.py             3  (matplotlib + HTML renderers,
                                                 new in 3.7.3, and the plot
                                                 watermark)
  tests/unit_tests/various/test_make_opts.py 4  (file header + one comment)
  madgraph/interface/common_run_interface.py 2  (comments)
  madgraph/iolibs/export_v4.py               1  (comment)
  madgraph/interface/madgraph_interface.py   1  (comment)
  madgraph/__init__.py                       1  (comment)

Deliberately NOT renamed, because they are identifiers rather than prose:
 - MadGraph5Error, the exception class, which the rename kept (main still has
   it in base_objects.py and misc.py). A plain "MadGraph5" grep matches it and
   over-reports by ~10x;
 - the "@MG5aMC" MadAnalysis5 card escape tags in common_run_interface.py,
   which are a literal card protocol string;
 - "$HEP/MG5aMC_PY8_interface/MG5aMC_PY8_interface" in check_heptools.sh,
   which is a real path on disk that HEPToolInstaller creates.

The matplotlib plot watermark is set to 'MadGraph7_aMC@NLO' rather than plain
'MadGraph7', so it renders identically to the gnuplot watermark main kept at
histograms.py:2728 ('MadGraph7\_aMC\@nlo' -- those are gnuplot escapes). The
two renderers draw the same plot and should not disagree on their own label.

IOTest goldens: the merge changes 6 of them, all content (upstream's 3.7.3
histogram work and the matrix templates), none vocabulary -- verified that no
golden loses a MadGraph7 or regains a MadGraph5_aMC. Nothing regenerated;
testIO_gnuplot_histo_output, testIO_DJR_histograms and the two
export_matrix_element_v4_madevent goldens all pass as they stand.

VERSION stays main's 0.2.0.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Both renderers watermarked the plot with the pre-rename name. The rename
(f5a6cf3) had rewritten the gnuplot one to 'MadGraph7\_aMC\@nlo' -- gnuplot
escapes, so it drew as "MadGraph7_aMC@NLO" -- and the matplotlib renderer that
3.7.3 adds was set to match it on the previous commit. Both now read plain
"MadGraph7".

The gnuplot label is emitted verbatim into the .gnuplot output, so four stored
IOTest goldens carry it. Updated by applying the SAME substitution to them
rather than regenerating: the diff is 40 lines, all of them that one label and
nothing else. Then verified against the real generator --
testIO_gnuplot_histo_output, testIO_DJR_histograms and the 25 unit tests in
test_histograms.py all pass, so the goldens do match what the code now emits.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@oliviermattelaer
oliviermattelaer merged commit 815cacb into main Sep 11, 2026
513 checks passed
@oliviermattelaer
oliviermattelaer deleted the merge_380 branch September 11, 2026 06:55
oliviermattelaer added a commit that referenced this pull request Sep 11, 2026
main has moved 246 commits, including PR #117 which merged upstream
mg5amcnlo 3.8.0. Clean merge, no conflicts.

Checked rather than assumed: 3.8.0 brings its own frame-denominator fix
(mg5amcnlo#407) into MadSpin, so _onshell_production_norm now exists on
main. There is exactly one definition of it after the merge, and this
branch does not add a second -- the light-cone work touches the spinor
routines and the massless-parton projection, not that denominator.

642 unit tests OK across test_lhe_parser, test_aloha and test_madspin.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
oliviermattelaer added a commit that referenced this pull request Sep 11, 2026
…branch

Brings in two things:

* main, 246 commits including PR #117's merge of upstream mg5amcnlo
  3.8.0. Clean, no conflicts.
* 32e8362 "update IOtest", the regenerated FD-gauge golden. That
  commit landed on the light-cone branch after this one forked, which
  is the whole reason IOtest_fd_gauge was red here and green there --
  the drift was always the light-cone spinor change, and this branch
  was carrying the pre-regeneration golden. Verified: the golden here
  is now byte-identical to the light-cone branch's.

651 unit tests OK across test_lhe_parser, test_aloha and test_madspin.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
oliviermattelaer added a commit that referenced this pull request Sep 11, 2026
main has moved 348 commits, including PR #117's merge of upstream
mg5amcnlo 3.8.0.

One conflict, in madgraph_interface's polarisation help text, where both
sides edited the same block for different reasons:

* this branch replaced the "users need ... 'nhel=1' ..." line, because
  nhel=1 is only a variance choice (Monte-Carlo over helicities is an
  unbiased estimator of the same cross-section), not a requirement, and
  added the '{X}'-is-a-set lines that document why '{++}' and '{+T}'
  are refused;
* main added '{G}','{H}','{Q}','{W}','{S}', the 3.8.0 propagator-piece
  selectors.

Resolved as a union. Taking main's side would have silently reverted
the nhel=1 correction; taking this branch's would have dropped the new
selectors from the help.

3.8.0 also brings its own copy of the frame-denominator fix
(mg5amcnlo#407). After the merge there is exactly one definition of
_onshell_production_norm -- the merge did not stack two.

674 unit tests OK across test_fks_base, test_fks_helas_objects,
test_fks_common, test_cmd and test_madspin.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
roiser pushed a commit to roiser/MadGraph7 that referenced this pull request Sep 11, 2026
`p p > t t~ DIM6=0` in dim6top_LO_UFO aborted the mg7/madspace survey with

  RuntimeError: non-finite integral in channel '2.0' after 1000 samples
  (mean=nan, variance=nan)

while madevent integrated the same process fine. dim6top carries an FCNC
coupling order that DIM6=0 does not constrain, so MG5 keeps the four
flavour-off-diagonal channels (u c~, c u~, d s~, s d~ > t t~). They have 7
diagrams each but an amplitude of exactly zero, because their Wilson
coefficients are zero in the default param card. A channel like that should
contribute nothing; instead it poisoned the whole run.

The chain, all of it 0/0 on a channel with no amplitude anywhere:

  umami.cc built the per-diagram amp2 madspace samples from as
  numerator/denominator, the denominator being the sum of the numerators.

  kernel_collect_channel_weights then normalised the channel weights by the sum
  of those amp2.

  survey() divided by the subprocess mean to get its relative spread, which is
  a plain ZeroDivisionError once the nan above is gone and the channel
  correctly integrates to zero.

The matrix-element multichannel weights in process_sigmaKin_function.inc and
process_function_definitions.inc had the same division; they are guarded too,
for the consumers that ask for mulChannelWeight (mg7 does not, so only the
three above were load-bearing for this reproducer).

NONE OF THE C++ GUARDS MAY BE WRITTEN AS A TEST FOR ZERO. The generated code
is compiled with -ffast-math, and its -ffinite-math-only lets the compiler
assume the quotient is finite and delete such a guard as dead -- an explicit
`denominator == 0. ? 0. : ...` in umami.cc was silently removed, and only
reappeared when a printf next to it perturbed the optimiser. This is the same
trap as MadGraphTeam#117 and #516, which is why CrossSectionKernels.o is already built
-fno-fast-math. The denominators here are sums of |amp|^2 and so never
negative, so they are floored by adding the smallest normal instead: 0/tiny is
0, and every denominator a real amplitude produces is unchanged bit for bit.
That keeps the guard robust without taking -ffast-math off a hot file.

madspace's own kernel is not compiled that way, so it substitutes the norm and
falls back to the uniform distribution -- the channel weights have to remain a
normalised distribution, and with no amplitude to prefer one, none is.

Verified with dim6top_LO_UFO, ctG = 1, 13 TeV, fixed mu = 91.188,
NNPDF23_lo_as_0130_qed:

  p p > t t~ DIM6=0   mg7 736.223 +- 1.686 pb   madevent 736.2 +- 0.61 pb

and no regression: plain `sm` `p p > t t~` gives 718.8675807362097 pb, bit
identical to the same run before this change.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants