Skip to content

feat(turbo): add PQ int4/int8 quantizers and PQFast FastScan support - #664

Open
zzlin237 wants to merge 1 commit into
alibaba:mainfrom
zzlin237:refactor/turbo_pq_int4
Open

feat(turbo): add PQ int4/int8 quantizers and PQFast FastScan support#664
zzlin237 wants to merge 1 commit into
alibaba:mainfrom
zzlin237:refactor/turbo_pq_int4

Conversation

@zzlin237

@zzlin237 zzlin237 commented Aug 7, 2026

Copy link
Copy Markdown
Collaborator
  • Add PqInt4Quantizer / PqInt8Quantizer with ADC/SDC/batch distance kernels for avx2, avx512, neon and scalar, plus precomputed residual distance table support for IVF search
  • Add PqFastQuantizer (QuantizeType::kPQFast): 4-bit PQ with FastScan packed-block scan kernels (32 vectors per block, affine-quantized LUT)
  • Wire PQ build/search params through index_param and index_param_builders
  • Add unit tests for int4/int8/fast quantizers and FastScan kernels
    The specific performance metrics are as follows:

sift

  • centroid_count:1024
  • num_chunk:32,64
  • use_residual: true
  • nprobe: 8, 16, 32
  • metric:l2
num_chunk nprobe fast recall int4 recall fast 精排后 int4 精排后 fast QPS int4 QPS 加速比
32 8 56.88% 56.92% 81.65% 80.88% 15868 15509 1.02×
32 16 58.89% 59.39% 88.41% 88.19% 14984 14443 1.04×
32 32 59.88% 60.21% 91.83% 91.58% 13510 11834 1.14×
64 8 68.46% 68.86% 82.93% 82.63% 15342 15092 1.02×
64 16 72.45% 73.07% 91.92% 91.75% 14351 13154 1.09×
64 32 74.36% 75.00% 96.95% 96.97% 12607 9441 1.34×

gist

  • centroid_count:1024
  • num_chunk:240,480
  • use_residual: true
  • nprobe: 8, 16, 32
  • metric:l2
num_chunk nprobe fast recall int4 recall fast 精排后 int4 精排后 fast QPS int4 QPS 加速比
240 8 42.47% 42.41% 56.90% 56.77% 10939 7395 1.48×
240 16 47.61% 47.60% 70.28% 70.40% 8427 4741 1.78×
240 32 50.26% 50.36% 79.91% 80.07% 5758 2805 2.05×
480 8 52.22% 53.45% 57.26% 58.04% 8591 4308 2.00×
480 16 63.08% 64.37% 73.14% 73.37% 5900 2497 2.36×
480 32 69.74% 71.37% 85.78% 86.24% 3707 1371 2.70×

cohere

  • centroid_count:1024
  • num_chunk:192,384
  • use_residual: true
  • nprobe: 8, 16, 32
  • metric:cosine
num_chunk nprobe fast recall int4 recall fast 精排后 int4 精排后 fast QPS int4 QPS 加速比
192 8 53.76% 53.89% 74.58% 74.63% 9709 4616 2.10×
192 16 56.89% 57.03% 81.66% 81.85% 7014 2826 2.48×
192 32 58.39% 58.60% 85.67% 85.86% 4610 1708 2.70×
384 8 72.22% 72.18% 79.11% 79.06% 7548 2573 2.93×
384 16 77.76% 77.77% 88.20% 88.17% 5047 1508 3.35×
384 32 80.70% 80.81% 93.95% 94.01% 3188 877 3.64×

- Add PqInt4Quantizer / PqInt8Quantizer with ADC/SDC/batch distance
  kernels for avx2, avx512, neon and scalar, plus precomputed residual
  distance table support for IVF search
- Add PqFastQuantizer (QuantizeType::kPQFast): 4-bit PQ with FastScan
  packed-block scan kernels (32 vectors per block, affine-quantized LUT)
- Wire PQ build/search params through index_param and index_param_builders
- Add unit tests for int4/int8/fast quantizers and FastScan kernels
Copilot AI lite review requested due to automatic review settings August 7, 2026 11:14

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants