Skip to content

Build permuted_dense inverse arrays lazily - #108

Open
Transurgeon wants to merge 3 commits into
mainfrom
lazy-inv-arrays
Open

Transurgeon wants to merge 3 commits into
mainfrom
lazy-inv-arrays

Conversation

@Transurgeon

@Transurgeon Transurgeon commented Aug 2, 2026 •

Copy link
Copy Markdown
Member

Motivation

Every permuted_dense (PD) used to build two dense inverse permutations at construction: col_inv (global n ints) and row_inv (global m ints). On the dense left_matmul kron path, the Jacobian is a stacked_pd of p blocks, and every block's column space is the full variable space. That made the index metadata O(p·n_vars + p²·m0), compared with O(nnz) for the values. On CVXPY's OptimalAdvertising benchmark this put the engine's peak at about 6 GB, while the actual nonzeros need about 5 MB.

Change

The inverse arrays are now built lazily, under one rule:

  • No constructor builds col_inv / row_inv; they stay NULL.unchanged, and the fills assert the array is present.
  • row_gather_pd_alloc is the one exception. Its fill copies rows through bound_iwork, so at alloc it only needs a membership test, which it answers with a binary search (sorted_pos) over row_perm instead of building a global-size table.

Ensure calls were added in BA_pd_csc_alloc, BTA_pd_csc_alloc, BTA_csc_pd_alloc, the coalesce output blocks, BTA_pd_spd_alloc, the kron scratch / BT, and old-code.

An earlier revision added a second "compact" PD mode instead. That needed NULL fallbacks in about six places, and only the kron-csc outputs were made compact. The single lazy rule replaces it with a smaller diff (src+include: +88/−23) and also covers the kron-pd/spd outputs and the ATA raw blocks.

Results

main this PR
OptimalAdvertising engine peak (g_peak_bytes) 5973 MB 628 MB
OptimalAdvertising process peak RSS delta 6065 MB 698 MB
OptimalAdvertising get_problem_data (median of 3) 1.06 s 0.62 s
row+col-sum Jacobian init, 128×128 (C test) 1324 B/var 292 B/var
sum(exp(A @ X)) Hessian init, 96×96 94.9 MB 62.8 MB
  • OptimalAdvertising P/A/c/b are bit-identical to main.
  • CVaRBenchmark and SemidefiniteProgramming peaks are unchanged; they come from a different cause.

Tests

  • New test_peak_memory_kron_jacobian, which asserts < 400 B/var. It is guarded by SP_TRACK_MEMORY.
  • test_permuted_dense_lazy_inv and test_permuted_dense_times_csc_lazy_inv.
  • The full suite passes in Debug, with -DSP_TRACK_MEMORY=ON, and under ASan+UBSan.
  • The existing transpose(A @ X) problem-Jacobian test covers the row_gather path: it segfaults if row_gather reads row_inv directly.

🤖 Generated with Claude Code

…blocks

Every permuted_dense allocated col_inv (global n ints) and row_inv
(global m ints) at construction. On the dense left_matmul kron path this
metadata dominates: p independent blocks each carry a full n_vars-sized
col_inv, O(p*n_vars + p^2*m0) ints per node against exact-nnz values
(5.9 GB engine peak on the OptimalAdvertising benchmark where true nnz
needs ~5 MB).

This introduces a compact PD variant whose storage stays proportional to
the block rather than the global shape, scoped to the kron-csc family:

- new_permuted_dense_compact leaves both inverse arrays NULL; the
  default constructor is unchanged, so every other PD keeps today's
  layout and O(1) code paths.
- BA_pd_csc_alloc builds its output compact, and compactness propagates
  through copy_sparsity/transpose/index chains of compact sources.
- Init-path membership tests (idxs_hits_set callers, index_pd_alloc)
  gate on a NULL inv and fall back to binary-search scans of the
  sorted perms (sorted_pos / sorted_hits in utils).
- Eval-path consumers ensure the operand-side array at alloc time only
  when the product is non-empty (permuted_dense_ensure_col_inv /
  _row_inv), keeping every fill kernel untouched (asserted in
  BA_pd_csc_fill_values); index_pd_fill_values of a compact source
  reads a source-row map precomputed into kernel_iwork at alloc.
- The mutable kron scratch keeps its arrays (its kernels write col_inv).

New peak-memory regression test builds the row-sum + col-sum Jacobian
shape at 128x128: init peak drops from 1324 to 292 bytes/var (21.7 MB
to 4.8 MB); the test asserts < 400 bytes/var.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@Transurgeon Transurgeon changed the title Lazy inverse-permutation arrays for kron-path permuted_dense blocks Compact permuted_dense (no inverse-permutation arrays) for kron-path blocks Aug 2, 2026
@Transurgeon
Transurgeon marked this pull request as draft August 3, 2026 00:39
Transurgeon and others added 2 commits October 5, 2026 21:25
# Conflicts:
#	src/utils/permuted_dense.c
#	tests/all_tests.c
#	tests/utils/test_permuted_dense.h
Replace the two-mode design (full vs compact PDs, with compactness
propagating through copies and NULL-guarded fallbacks at every reader)
with one rule: no constructor builds col_inv / row_inv; an alloc whose
paired fill reads one calls permuted_dense_ensure_col_inv / _row_inv on
that operand, and the fills assert it is present. row_gather is the one
reader that only needs membership at alloc (its fill copies through
bound_iwork), so it binary-searches row_perm via sorted_pos instead.

Deletes new_permuted_dense_compact, pd_is_compact, sorted_hits, and the
unreachable sorted_pos branches in coalesce_spd_scatter and
BTA_pd_spd_fill_values. Same Jacobian-init saving on the kron-csc path
(1324 -> 292 B/var at 128x128), and it also covers PDs no fill ever
reads inverse arrays of (kron-pd/spd outputs, ATA raw blocks): an
exp(A @ X) Hessian init drops 94.9 MB -> 62.8 MB (PR-as-written: 85.3).

CVXPY OptimalAdvertising (m=250, n=1000) through the DIFFENGINE canon
backend: engine peak 5973 MB -> 628 MB, get_problem_data 1.06 s -> 0.62 s
(median of 3), and P/A/c/b bit-identical to main. CVaRBenchmark and
SemidefiniteProgramming peaks are unchanged (a different cause).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@Transurgeon
Transurgeon marked this pull request as ready for review October 6, 2026 01:27
@Transurgeon Transurgeon changed the title Compact permuted_dense (no inverse-permutation arrays) for kron-path blocks Build permuted_dense inverse arrays lazily Oct 6, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant