Skip unused Gaussian derivatives in xFit model evaluation - #65
Conversation
|
Important Review skippedAuto reviews are limited based on label configuration. 🏷️ Required labels (at least one) (1)
Please check the settings in the CodeRabbit UI or the ⚙️ Run configurationConfiguration used: Repository: NVIDIA/cuPhoton/.coderabbit.yaml Review profile: CHILL Plan: Enterprise Run ID: You can disable this status message by setting the Use the checkbox below for a quick retry:
Comment |
777538f to
f0a16d3
Compare
f3a58cc to
16e1729
Compare
melo-gonzo
left a comment
There was a problem hiding this comment.
LGTM. Local base/head comparisons preserve model values, Jacobians, and complete fit results exactly. The reported warm-batch gain and slower overall MPI launch remain clearly distinguished.
f0a16d3 to
dcfc778
Compare
16e1729 to
ed0efe9
Compare
Gaussian residual evaluation only needs each lobe value. Return that value before building derivative arrays, while retaining the existing arithmetic and full Jacobian path. Signed-off-by: Trent Nelson <trentn@nvidia.com>
ed0efe9 to
76224c4
Compare
Gaussian residual evaluation computes derivative arrays for both lobes and discards them. Return each lobe value before those calculations, preserving the model arithmetic and the existing Jacobian path.
This optimization can land independently of the FITS pipeline changes in #64. Its patch is unchanged by extraction from that stack. Validation of the extracted #63/#65 candidate: 2,653 CPU tests passed, 184 skipped, and all lint/type hooks passed. The following component and GPU results predate this extraction; no new GPU campaign was run for it. 151 CPU xFit tests pass. Ten matched GPU replay launches preserve exact scientific outputs and evaluation/status vectors, with warm fit medians improving by 1.056–1.089× across three frozen cases. Representative and slow cases include reversed-order independent repeats.
The committed wheel also passes exact full-output comparisons on 16 registered image pairs with eight GB200 GPUs, using both MPI and Dragon and forced KvikIO compatibility reads. Each baseline/candidate comparison preserves worker-to-item/GPU placement, candidates and solver settings. All 256 fits retain their evaluation/status vectors, including the same 14 fits reaching the 1,800-evaluation cap.
Each full-pipeline launch has one warmup and three measured rounds. Times below are baseline → candidate, in seconds:
Measured batch median speed ratios are 1.076× and 1.079×, respectively. The complete MPI launch is slower with its long warmup; the cause is unestablished. Complete-launch times include all rounds, startup, coordinator audits and shutdown. These are single launches per source/runtime; MPI and Dragon use separate matched comparisons. The earlier fit replay timings are a separate measurement.