perf: reduce sketch hashing overhead - #265
Merged
Merged
Conversation
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Design Notes
The optimization keeps the public API and hash algorithms unchanged. It focuses on the existing hot path from a sketch update through
Hashand the internal hasher, making that path easier for the compiler to specialize while keeping full-word processing separate from partial tails.The benchmarks exercise public sketch APIs rather than isolated hash helpers. Integer inputs benefit most from exposing the complete fixed-size hashing path to the optimizer. Raw-byte inputs additionally benefit from exposing the single contiguous write and using fixed-width reads for complete words. Count-Min performs four seeded MurmurHash3 passes per update in this benchmark, while Bloom performs two XXHash64 passes, so reductions in per-hash overhead accumulate in those sketches.
Benchmarks
Local Divan benchmark medians for 10,000 updates:
u64u64u64u64u64u64u64The baseline uses the same benchmark code with
datasketches/src/hashrestored frommain; the optimized run uses this PR. Results are machine-dependent and document local before-and-after evidence rather than a portable performance guarantee.Validation
cargo x checkcargo x testcargo x lintcargo bench --package benchmarks --bench benchmarks -- bloom::update countmin::update cpc::update frequencies::update hll::update theta::update tuple::update --min-time 1 --sample-count 20