Skip to content

Commit 414c85f

Browse files
committed
Release v0.2.4: Stateless Searcher pipeline, contiguous batch search, and unified Python search API
- Refactor SearcherBase & SearcherImpl to be fully stateless and thread-safe with per-query eps and rerank_factor parameters - Introduce SearchResultBatch flat contiguous container and zero-allocation span-based batch search - Separate result unpacking and distance copying into branchless sequential loops in populate_graph_results and drain_heap_into - Modernize candidate reranking API (deglib::search::rerank) with explicit return_distances and unsorted flags - Unify Python Searcher.search to handle single-query vectors (1D, 1xDim, Dimx1) and multithreaded 2D batches seamlessly - Optimize FP16 scalar quantization via 64-element chunked SIMD streaming and encapsulate internal quantizer transforms
1 parent dcca712 commit 414c85f

17 files changed

Lines changed: 1017 additions & 797 deletions

File tree

‎RELEASE_NOTES.md‎

Lines changed: 31 additions & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -1,5 +1,35 @@
11
# Release Notes
22

3+
## deglib v0.2.4
4+
5+
### Overview
6+
deglib v0.2.4 refactors the query execution pipeline to be fully stateless and thread-safe, eliminates query-time allocations, streamlines candidate reranking, and optimizes FP16 quantization kernels. It updates `Searcher` to accept search parameters (`eps`, `rerank_factor`) directly per query, adds contiguous flat batch search (`SearchResultBatch`), separates inner loop branches for result extraction, and consolidates the Python search interface.
7+
8+
---
9+
10+
### ⚠️ Breaking Changes & Migration Guide
11+
12+
* **Stateless `Searcher` API:** Removed mutable internal search parameters (`search_eps_`, `rerank_factor_`) and their setters/getters (`set_query_arguments`, `set_search_eps`, `set_rerank_factor`). `eps` and `rerank_factor` are now query-level parameters passed directly to `search()` and `search_batch()`.
13+
* **Rerank API Consolidation (`deglib::search::rerank`):** Removed disparate internal helpers (`rerank_into`, etc.) and exposed overloaded `deglib::search::rerank` functions with explicit `return_distances` and `unsorted` boolean flags.
14+
* **Unified Python `Searcher.search`:** Single entry point handling 1D vectors, row-vectors `(1, dim)`, column-vectors `(dim, 1)`, and 2D batch matrices `(N, dim)` with `threads`. Removed redundant `search_batch` from the Python class.
15+
16+
---
17+
18+
### 🚀 Key Features & Improvements
19+
20+
#### Stateless & Thread-Safe Search Execution (`deglib::search`)
21+
* **Reentrant Querying:** Searcher instances are fully thread-safe and reentrant, allowing multiple threads to query the same instance with varying `eps` and `rerank_factor` trade-offs without data races or locks.
22+
* **Branch-Free Result Extraction:** Separated neighbor index unpacking and optional distance copying into distinct, branchless loops (`populate_graph_results` and `drain_heap_into`) to eliminate per-element conditional branches on the hot path.
23+
* **Contiguous Batch Search (`SearchResultBatch`):** Flat 1D vector layout for indices and distances (`n_queries * k`) with zero-overhead `get_indices(q)` and `get_distances(q)` row slicing. Supports zero-allocation execution with user-provided `std::span` buffers.
24+
* **Small Buffer Optimization (SBO):** Uses stack buffers for transformed query bytes and candidate indices for vectors $\le 512$ bytes, avoiding heap allocations during search.
25+
26+
#### Quantization & Kernel Optimizations (`deglib::optimization::quantization`)
27+
* **Chunked SIMD FP16 Quantization:** Processes FP16 (`uint16_t`) vectors in fixed 64-element chunks via `fp16_to_floats`, ensuring complete auto-vectorization (AVX2/F16C/AVX-512) without dimension limits.
28+
* **In-Place EVP Query Quantization:** Direct query quantization into preallocated search buffers without intermediate vector allocations.
29+
* **API Clean-Up:** Encapsulated internal quantizer row transformations (`transform`, `transform_row`, `transform_back`) as private. Removed obsolete methods (`dequantize`, `fit_quantize`, `inv_scale`).
30+
31+
---
32+
333
## deglib v0.2.3
434

535
### Overview
@@ -13,7 +43,7 @@ deglib v0.2.3 introduces a zero-overhead C++20 templated **`Searcher`** pipeline
1343
* **Zero-Overhead End-to-End Querying:** Integrates query quantization (FP32/FP16 $\rightarrow$ INT8/UINT8/EVP), graph traversal on quantized indices, and SIMD candidate refinement against FP16/FP32 base features in a single C++ call.
1444
* **Static Template Specialization (`SearcherImpl<QuantT, RefinerT>`):**
1545
* Compile-time `if constexpr` branch elimination avoiding virtual method dispatch in the hot search loop.
16-
* Full support for all quantizers (`NoQuantizer`, `ScalarInt8Quantizer`, `ScalarInt8PerDimQuantizer`, `ScalarUint8Quantizer`, `ScalarUint8PerDimQuantizer`, `EVPQuantizer`).
46+
* Full support for all quantizers (`NoQuantizer`, `ScalarQuantizerInt8`, `ScalarQuantizerInt8PerDim`, `ScalarQuantizerUint8`, `ScalarQuantizerUint8PerDim`, `EVPQuantizer`).
1747
* Support for unquantized and quantized candidate refiners (`NoRefiner`, `ExactRefiner<uint16_t>`, `ExactRefiner<float>`).
1848
* **Modern C++20 Interface:**
1949
* `std::span<const T>` and `SearchResult` helpers for safe, expressive, and allocation-free querying.

‎cpp/deglib/include/deglib/optimization/quantization/evp_quantize.h‎

Lines changed: 80 additions & 8 deletions
Original file line numberDiff line numberDiff line change
@@ -17,17 +17,17 @@ namespace deglib::quantization::evp {
1717
// ============================================================================
1818

1919
/**
20-
* Quantizes a single fp32 vector to EVP bytes.
20+
* Quantizes a single fp32 vector directly into a pre-allocated EVP byte buffer.
2121
*
2222
* Layout: [ones (dim/8 bytes)][negative_ones (dim/8 bytes)]
2323
*
2424
* @param embedding Pointer to dim float values (row-major)
2525
* @param dim Dimension (must be divisible by 8)
2626
* @param non_zeros Number of top-K elements by absolute value
27-
* @return std::vector<std::byte> with 2 * dim/8 bytes
27+
* @param out_bytes Pointer to buffer of at least 2 * dim/8 bytes
2828
* @throws std::invalid_argument if dim % 8 != 0 or non_zeros >= dim
2929
*/
30-
inline std::vector<std::byte> quantize_single(const float* embedding, uint32_t dim, uint32_t non_zeros) {
30+
inline void quantize_single_into(const float* embedding, uint32_t dim, uint32_t non_zeros, std::byte* out_bytes) {
3131
if (dim % 8 != 0) {
3232
throw std::invalid_argument("quantize_single: dim must be divisible by 8, got " + std::to_string(dim));
3333
}
@@ -36,8 +36,6 @@ inline std::vector<std::byte> quantize_single(const float* embedding, uint32_t d
3636
}
3737

3838
const size_t mask_bytes = dim / 8;
39-
std::vector<std::byte> result(2 * mask_bytes);
40-
4139
std::vector<std::pair<float, uint32_t>> abs_vals(dim);
4240
for (uint32_t i = 0; i < dim; ++i) {
4341
abs_vals[i] = {std::abs(embedding[i]), i};
@@ -51,8 +49,8 @@ inline std::vector<std::byte> quantize_single(const float* embedding, uint32_t d
5149
is_top[abs_vals[j].second] = 1;
5250
}
5351

54-
std::byte* ones_dst = result.data();
55-
std::byte* negs_dst = result.data() + mask_bytes;
52+
std::byte* ones_dst = out_bytes;
53+
std::byte* negs_dst = out_bytes + mask_bytes;
5654

5755
for (uint32_t byte_idx = 0; byte_idx < mask_bytes; ++byte_idx) {
5856
int byte_val = 0;
@@ -77,7 +75,23 @@ inline std::vector<std::byte> quantize_single(const float* embedding, uint32_t d
7775
}
7876
negs_dst[byte_idx] = static_cast<std::byte>(byte_val);
7977
}
78+
}
8079

80+
/**
81+
* Quantizes a single fp32 vector to EVP bytes.
82+
*
83+
* Layout: [ones (dim/8 bytes)][negative_ones (dim/8 bytes)]
84+
*
85+
* @param embedding Pointer to dim float values (row-major)
86+
* @param dim Dimension (must be divisible by 8)
87+
* @param non_zeros Number of top-K elements by absolute value
88+
* @return std::vector<std::byte> with 2 * dim/8 bytes
89+
* @throws std::invalid_argument if dim % 8 != 0 or non_zeros >= dim
90+
*/
91+
inline std::vector<std::byte> quantize_single(const float* embedding, uint32_t dim, uint32_t non_zeros) {
92+
const size_t mask_bytes = dim / 8;
93+
std::vector<std::byte> result(2 * mask_bytes);
94+
quantize_single_into(embedding, dim, non_zeros, result.data());
8195
return result;
8296
}
8397

@@ -378,11 +392,69 @@ inline std::vector<std::byte> quantize_batch(const std::vector<std::vector<std::
378392
return result;
379393
}
380394

395+
/**
396+
* Quantizes a single FP16 (uint16_t) vector directly into a pre-allocated EVP byte buffer.
397+
*
398+
* Layout: [ones (dim/8 bytes)][negative_ones (dim/8 bytes)]
399+
*/
400+
inline void quantize_single_into(const uint16_t* embedding, uint32_t dim, uint32_t non_zeros, std::byte* out_bytes) {
401+
if (dim % 8 != 0) {
402+
throw std::invalid_argument("quantize_single: dim must be divisible by 8, got " + std::to_string(dim));
403+
}
404+
if (non_zeros >= dim) {
405+
throw std::invalid_argument("quantize_single: non_zeros must be < dim");
406+
}
407+
408+
const size_t mask_bytes = dim / 8;
409+
std::vector<std::pair<uint16_t, uint32_t>> abs_vals(dim);
410+
for (uint32_t i = 0; i < dim; ++i) {
411+
abs_vals[i] = {static_cast<uint16_t>(embedding[i] & 0x7FFFu), i};
412+
}
413+
std::nth_element(abs_vals.begin(), abs_vals.begin() + (dim - non_zeros), abs_vals.end(), [](const auto& a, const auto& b) {
414+
return a.first != b.first ? a.first < b.first : a.second < b.second;
415+
});
416+
417+
std::vector<uint8_t> is_top(dim, 0);
418+
for (uint32_t j = dim - non_zeros; j < dim; ++j) {
419+
is_top[abs_vals[j].second] = 1;
420+
}
421+
422+
std::byte* ones_dst = out_bytes;
423+
std::byte* negs_dst = out_bytes + mask_bytes;
424+
425+
for (uint32_t byte_idx = 0; byte_idx < mask_bytes; ++byte_idx) {
426+
int bv = 0;
427+
for (int bit = 0; bit < 8; ++bit) {
428+
const uint32_t bit_idx = byte_idx * 8 + bit;
429+
if (bit_idx >= dim) break;
430+
if (is_top[bit_idx] && (embedding[bit_idx] & 0x8000u) == 0 && (embedding[bit_idx] & 0x7FFFu) > 0) {
431+
bv |= (1 << bit);
432+
}
433+
}
434+
ones_dst[byte_idx] = static_cast<std::byte>(bv);
435+
}
436+
437+
for (uint32_t byte_idx = 0; byte_idx < mask_bytes; ++byte_idx) {
438+
int bv = 0;
439+
for (int bit = 0; bit < 8; ++bit) {
440+
const uint32_t bit_idx = byte_idx * 8 + bit;
441+
if (bit_idx >= dim) break;
442+
if (is_top[bit_idx] && (embedding[bit_idx] & 0x8000u) != 0 && (embedding[bit_idx] & 0x7FFFu) > 0) {
443+
bv |= (1 << bit);
444+
}
445+
}
446+
negs_dst[byte_idx] = static_cast<std::byte>(bv);
447+
}
448+
}
449+
381450
/**
382451
* Quantizes a single FP16 (uint16_t) vector to EVP bytes.
383452
*/
384453
inline std::vector<std::byte> quantize_single(const uint16_t* embedding, uint32_t dim, uint32_t non_zeros) {
385-
return quantize_batch(embedding, 1, dim, non_zeros, 1);
454+
const size_t mask_bytes = dim / 8;
455+
std::vector<std::byte> result(2 * mask_bytes);
456+
quantize_single_into(embedding, dim, non_zeros, result.data());
457+
return result;
386458
}
387459

388460
} // namespace deglib::quantization::evp

0 commit comments

Comments
 (0)