High-efficiency floating-point neural network inference operators for mobile, server, and Web
-
Updated
Aug 22, 2026 - C
High-efficiency floating-point neural network inference operators for mobile, server, and Web
Efficient Deep Learning Systems course materials
BladeDISC is an end-to-end DynamIc Shape Compiler project for machine learning workloads.
The Tensor Algebra SuperOptimizer for Deep Learning
Krasis is a Hybrid LLM runtime which focuses on efficient running of larger models on consumer grade VRAM limited hardware
Everything you need to know about LLM inference
Learn LLM Inference Engineering step by step - from KV cache, PagedAttention, and continuous batching to vLLM, SGLang, and GPUs.
[MLSys 2021] IOS: Inter-Operator Scheduler for CNN Acceleration
Batch normalization fusion for PyTorch. This is an archived repository, which is not maintained.
LLM infrastructure cost reduction via NUMA-aware weight banking: 147 t/s (8.8x stock llama.cpp) on refurbished enterprise POWER8. Self-hosted inference, no cloud APIs. Part of the Proof of Physical AI stack.
Optimize layers structure of Keras model to reduce computation time
Accelerating Long Context LLM Inference with Accuracy-Preserving Context Optimization in SGLang, vLLM, llama.cpp, OpenClaw, RAG, and Agentic AI.
the agi compiler: records llm agent behavior, proves what repeats, and compiles it into verified, sandboxed wasm binaries that run for microdollars. nothing figured out twice, paper: https://arxiv.org/abs/2607.04542
Ultrafast Qwen3-TTS: sub-50 ms time-to-first-audio at 10 requests per second.
[CVPR 2025] DivPrune: Diversity-based Visual Token Pruning for Large Multimodal Models
A set of tool which would make your life easier with Tensorrt and Onnxruntime. This Repo is designed for YoloV3
Run 70B+ LLMs on a single 4GB GPU — no quantization required.
Official Repo for SparseLLM: Global Pruning of LLMs (NeurIPS 2024)
SpiderBrain v3 is a multi-platform skill/framework to reduce token usage and AI hallucinations across Claude, Cursor, and other AI tools.
A physics-grounded, cost-aware optimization loop for vLLM
Add a description, image, and links to the inference-optimization topic page so that developers can more easily learn about it.
To associate your repository with the inference-optimization topic, visit your repo's landing page and select "manage topics."