Pinned Loading
-
iron-batch
iron-batch PublicLLM serving primitives in Rust: paged KV allocation, continuous batching, streaming HTTP.
Rust
-
Topolens
Topolens PublicReads NUMA/PCIe topology from sysfs and recommends CPU-to-accelerator pinning
Rust
-
bit-floor
bit-floor Public1-bit and ternary LLM inference engine in Rust and CUDA. Weights stay quantized — no dequantization during inference.
Rust
-
sift-core
sift-core PublicHybrid search from scratch in Rust: LSM storage, BM25, flat vector index, RRF fusion. Zero deps.
Rust
Something went wrong, please refresh the page to try again.
If the problem persists, check the GitHub status page or contact support.
If the problem persists, check the GitHub status page or contact support.

