Follow-up from #273.
The GPU kernel is validated against the CPU path by a fixed list of shapes in matmul_b1_demo. Until recently every one of them had all three dimensions a multiple of 64, so the ragged tail and the partially-occupied last limb of a row went entirely untested; #273 added a single (65, 65, 65) case to cover that. One hand-picked tuple is thin evidence for a three-dimensional space with limb-boundary interactions on every axis — this is what proptest is for.
Target the raw API, not the dispatch path
Generate against fp_cuda::matmul_b1_raw rather than fp's <&Matrix as Mul>::mul:
matmul_b1_raw takes row-major limb slices with no size threshold, so shapes in the 64–512 range are fair game. A proptest run is cheap there and shrinking converges quickly.
- The dispatch path only reaches the GPU above
FP_CUDA_THRESHOLD (default 2048), which would force every generated case to be ≥2048³. That makes both the run and the shrink impractical.
Reference is fp::blas on the same inputs, compared bit-exactly — the same oracle the demo already uses.
Open question: what does CI assert?
Every GPU test in the tree is behind #[cfg(feature = "gpu")], and the runners have no CUDA. A proptest that silently passes when GpuContext::new fails is worse than no test, because it reports green on a machine that never ran the kernel. Needs a deliberate decision — skip loudly, or require the feature to imply a device.
proptest 1.7 is already a workspace dependency and fp ships a proptest feature, so there is no new dependency cost.
Shrinking cost
Each shrink step is another GPU launch, so the generator should bias toward small dimensions and shrink toward them.
Follow-up from #273.
The GPU kernel is validated against the CPU path by a fixed list of shapes in
matmul_b1_demo. Until recently every one of them had all three dimensions a multiple of 64, so the ragged tail and the partially-occupied last limb of a row went entirely untested; #273 added a single(65, 65, 65)case to cover that. One hand-picked tuple is thin evidence for a three-dimensional space with limb-boundary interactions on every axis — this is what proptest is for.Target the raw API, not the dispatch path
Generate against
fp_cuda::matmul_b1_rawrather thanfp's<&Matrix as Mul>::mul:matmul_b1_rawtakes row-major limb slices with no size threshold, so shapes in the 64–512 range are fair game. A proptest run is cheap there and shrinking converges quickly.FP_CUDA_THRESHOLD(default 2048), which would force every generated case to be ≥2048³. That makes both the run and the shrink impractical.Reference is
fp::blason the same inputs, compared bit-exactly — the same oracle the demo already uses.Open question: what does CI assert?
Every GPU test in the tree is behind
#[cfg(feature = "gpu")], and the runners have no CUDA. A proptest that silently passes whenGpuContext::newfails is worse than no test, because it reports green on a machine that never ran the kernel. Needs a deliberate decision — skip loudly, or require the feature to imply a device.proptest1.7 is already a workspace dependency andfpships aproptestfeature, so there is no new dependency cost.Shrinking cost
Each shrink step is another GPU launch, so the generator should bias toward small dimensions and shrink toward them.