Repository navigation
Warm cache: verify looptools content before combining and cleaning - #115
Merged
Merged
Conversation
The looptools_cache job had no verification step (unlike pythia8_cache and contur_cache, which both check their artifacts). Since ./bin/madgraph exits 0 even when "install collier" / "install ninja" fails, the job went green with a half-populated cache, heptools_cache gated the combine only on needs.*.result == 'success', folded the incomplete content into heptools-$KEY and then deleted every sub-cache - destroying the only good copies. * heptools_cache: new "Validate restored heptools content" step inspects the restored files and exports complete=true/false plus the list of missing tools. The delete-previous / save / cleanup steps are now gated on that flag instead of on the job results. When something is missing nothing is saved, no intermediate cache is deleted, and a final step writes a job summary and fails the job, so re-running only the offending job is enough (the others cache-hit). * looptools_cache: new "Verify looptools landed in the cache" step, mirroring the existing pythia8/contur ones. actions/cache skips its post-save when the job fails, so an incomplete looptools-<os> is never stored. It also runs on the cache-hit path, so an already-stored bad cache is reported instead of being reused forever. * delete-cache: looptools-<os> was missing from the reset_heptools delete list, so an incomplete looptools cache could never be reset. * delete-cache: `inputs.reset_* == 'true'` is always false - workflow_dispatch boolean inputs are booleans and GitHub casts to number on a loose comparison (true -> 1, 'true' -> NaN), so all four reset_* deletes were inert. Test the boolean itself. * restore_heptools: the consumer-side validation only tested `-d .../ninja`, which a failed install still leaves behind. Check libninja.a, libcollier.a, libcts.a and libiregi.a instead, so a bad combined cache falls back to the sub-caches rather than silently breaking the MadLoop link. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
oliviermattelaer
added a commit
that referenced
this pull request
Sep 9, 2026
Picks up the ZZ fix (b12a8a4, "MadSpin: undo MG5's decay-chain identical- particle factor in onshell_v1") that resolves the test_short_madspin_zz failure left open on the previous merge, plus the CI cache rework. Conflicts (3): * tests/unit_tests/madspin/test_madspin.py: both sides appended a block at EOF over an empty base. Kept both -- mg7's 7 numpy-pool classes and upstream's TestDecayChainIdenticalFactor. The two stub fixes from 9959eef are superseded upstream and merged cleanly. * .github/actions/restore_heptools_contur/action.yml: upstream, with mg7's input/mg7_configuration.txt. Brings the generation-suffixed key and the guard that stops MG5 being pointed at a fastjet-config the cache did not restore. * .github/workflows/warm_cache.yml: upstream's cache scheme, mg7's policy. On the cache rework (3513ae7): it applies to mg7 unchanged in intent. The bug it fixes is present here verbatim -- heptools_cache deleted the combined cache before saving a fresh one, and every consumer restored with an exact key and no restore-keys, so for the ~30 min in between nothing matched at all. On that miss restore_heptools_contur still wrote a fastjet-config path that does not exist, makefile_fks_dir saw "ifdef fastjet_config", and every aMC@NLO SubProcess compile died on fastjet/ClusterSequence.hh. Generation-suffixed keys with a prefix restore-key remove the window entirely, and with it the delete-before- save that was the fragile half of #115. It also fixes reset_ufo, which targeted "ufomodel-$ImageOS" while the key written was a bare "ufomodel" -- a no-op here too. Storage pressure is lower in mg7 than upstream, since mg7 builds only ubuntu-24.04, and the extra sub-cache rebuild costs nothing because the cleanup step already deleted every sub-cache on the success path. mg7 policy re-applied on top: ubuntu-24.04 only, the ref_name job gating, ./bin/madgraph and input/mg7_configuration.txt, the meson/ninja pip install for numpy's f2py backend, and --ref scoping on every cache delete -- threaded into the new prefix-delete helper as "gh cache list --ref "$GITHUB_REF"", so a reset still only touches the ref it runs on. Verified no job's if/runs-on/needs changed against the pre-merge tree. tests/unit_tests/madspin/test_madspin.py: 608 tests, same 1 failure + 10 errors as the pre-merge tree on this machine (float32 rounding and a numpy/py3.14 mismatch, both local); the 8 tests the merge adds all pass. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
roiser
pushed a commit
to roiser/MadGraph7
that referenced
this pull request
Sep 11, 2026
Conflict resolutions (9 files): * VERSION, UpdateNotes.txt: upstream. The 3.7.3 section upstream ships is a strict superset of ours (it already contains both mg7 BUG FIX entries), and it adds the new 3.8.0 section. * genps_fks.f: upstream. Both sides declared virtgranny_red EXTERNAL; only the continuation-line order differed, so take upstream's and stop diverging. * test_cmd_madevent.py: upstream (error=0.514). The 0.04 we had was tightened on 26 Jun and loosened again upstream two days later in "fix CI test". * test_cmd.py, madgraph_interface.py: upstream's --no_open/--merge do_draw rework, with mg7's --generate_only kept as an alias of --no_open. * CI cache validation: mg7 (PR MadGraphTeam#115) and upstream 3.8.0 implemented the same feature independently. Take upstream's reusable .github/actions/check_heptools action -- its content checks are a strict superset of our inline ones -- and re-apply the mg7 policy on top: ubuntu-24.04 only, ref-scoped gh cache delete, the ref_name job gating, ./bin/madgraph and input/mg7_configuration.txt, and the looptools-<os> entry in the reset_heptools list (still missing upstream). Upstream's new boost_cache job is kept and mg7-ified; emela_cache now depends on it. * acceptancetest.yml: ours for both hunks. Upstream's side would have grafted the (deliberately pruned) heft job over five mg7 jobs, and appended a second acceptancetest_delphes_parallel -- a duplicate YAML key. Also fixed, outside the conflicts: git auto-merged both sides' looptools content check into looptools_cache as two same-named steps at different indentation, which does not parse as YAML; and upstream's heptools_cache block carries an "if:" key that duplicates the one mg7 already has on that job. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Problem
The last cache warming combined everything and deleted the intermediate caches, but looptools was not complete.
looptools_cachehas no verification step — unlikepythia8_cacheandcontur_cache, which both check their artifacts landed../bin/madgraph cmdexits 0 even wheninstall collier/install ninjafails, so the job went green with a half-populated cache.heptools_cachethen gated the combine only onneeds.*.result == 'success', folded the incomplete content intoheptools-$KEY, and deleted every sub-cache — destroying the only good copies.Changes
heptools_cache— validate the content, not the job resultsNew
Validate restored heptools contentstep inspects the actual restored files after thecache/restoresteps and exportscomplete=true/falseplus amissing=list. The three dangerous steps (delete-previous, save, cleanup) are now gated onsteps.validate.outputs.complete == 'true'.When something is missing: nothing is saved, no intermediate cache is deleted, and a final step writes a job summary naming the missing tool and fails the job (previously it went green while silently skipping the combine). Re-running just the failing job is then enough — the others cache-hit and the combine/cleanup completes.
The looptools half checks
CutTools/lib/libcts.a,CutTools/lib/mpmodule.mod,IREGI/src/libiregi.a,ninja/lib/libninja.aandcollier/libcollier.a— exactly the pathsrestore_heptools_looptoolsconfigures and thatactivate_dependenceinmadgraph/various/misc.pylooks for.looptools_cache— fail at the sourceNew
Verify looptools landed in the cachestep, mirroring the existing pythia8/contur ones.actions/cacheskips its post-save when the job fails, so an incompletelooptools-<os>is never stored. It runs on the cache-hit path too, so an already stored bad cache is reported instead of being silently reused forever.delete-cache— two fixeslooptools-<os>was missing from thereset_heptoolsdelete list, so an incomplete looptools cache could never be reset.if: inputs.reset_* == 'true'is always false:workflow_dispatchboolean inputs are booleans, and GitHub casts to number on a loose comparison (true→ 1,'true'→ NaN). All fourreset_*deletes were inert. Now testing the boolean itself.restore_heptools— tighten the consumer-side checkIts
validate cache contentonly tested-d .../ninja, which a failed install still leaves behind. Now checkslibninja.a,libcollier.a,libcts.aandlibiregi.a, so a bad combined cache falls back to the sub-caches instead of silently breaking the MadLoop link.Testing
Both new shell blocks and the summary/heredoc step were extracted from the YAML and run locally against a fake
HEPtoolstree, in the complete case (exit 0,complete=true) and the incomplete case (correct per-tool warnings, de-duplicatedmissinglist, exit 1). The workflow and action YAML both parse.To recover the current state
Re-run Warm all caches via Run workflow with
reset_heptoolsticked — that now actually fires, and now also dropslooptools-<os>.🤖 Generated with Claude Code