vkop 是一个基于 Vulkan 实现的迷你AI推理引擎, 仅在GPU上运行.
首先需要安装项目的依赖项shaderc或者vulkan sdk:
wget https://sdk.lunarg.com/sdk/download/latest/linux/vulkan-sdk.tar.gz
tar xvf vulkan-sdk.tar.gz
source path/to/VulkanSDK/setup-env.sh
export PATH=$VULKAN_SDK/x86_64_bin:$PATH对于模型转换
export CMAKE_POLICY_VERSION_MINIMUM=3.5
pip install onnx onnx-simplifier onnxsim onnxruntime
对于压测模型下载
pip install torchvision
对于测试依赖libtorch, cmake过程自动下载解压
设置 Vulkan ICD 加载器,以 NVIDIA 为例:
export VK_ICD_FILENAMES=/usr/share/vulkan/icd.d/nvidia_icd.json确保使用正确的 Vulkan 版本:
source path/to/VulkanSDK/setup-env.shcmake .. -DENABLE_TESTS=ON -DUSE_VALIDATION_LAYERS=ON -DENABLE_ASAN=OFF -DUSE_DEBUG_LAYERS=OFF -DUSE_FP16=OFF -DUSE_MEASURE_TIME=OFF -DPython3_EXECUTABLE=$(which python3)
如果是交叉编译,需要设置交叉编译环境变量,借鉴参考toolchain.cmake
cmake .. -DCMAKE_TOOLCHAIN_FILE=../toolchain.cmake -DENABLE_TESTS=OFF
必须使用 Homebrew LLVM 的 clang++(/opt/homebrew/opt/llvm/bin/clang++):
- Apple clang 17 编译
core/function.cpp/core/runtime.cpp时前端在TransformCXXFoldExpr处无限递归段错误,不可用; - GCC + Apple libc++ 桥接 ABI 不兼容(std::string/流运行时输出乱码),不可用。
安装依赖并配置:
brew install cmake shaderc vulkan-loader vulkan-headers molten-vk \
glfw utf8proc pkgconf googletest libomp
cmake .. -DENABLE_TESTS=ON \
-DCMAKE_CXX_COMPILER=/opt/homebrew/opt/llvm/bin/clang++ \
-DCMAKE_C_COMPILER=/opt/homebrew/opt/llvm/bin/clang \
-DCMAKE_PREFIX_PATH="/opt/homebrew;/opt/homebrew/opt/llvm" \
-DPython3_EXECUTABLE=$(which python3)运行时 VulkanLib.cpp 通过 dlopen 加载 Vulkan loader,必须设置环境变量:
export DYLD_FALLBACK_LIBRARY_PATH=/opt/homebrew/lib
export VK_ICD_FILENAMES=/opt/homebrew/etc/vulkan/icd.d/MoltenVK_icd.json代码中已处理的 clang/macOS 兼容性点:
core/Tensor.hppfp16 内联汇编的h0/s0寄存器别名是 GCC 专有扩展, 已通过!defined(__clang__)守卫,clang 走可移植路径;- macOS 上
size_t(unsigned long)≠uint64_t(unsigned long long),Tensor标量构造函数需显式接受std::size_t,否则Tensor<int64_t>(v.size())会误配到Tensor(bool)构造出空张量; VK_EXT_host_image_copy在 Vulkan 1.4 晋升为核心功能后,MoltenVK 的 loader 对 EXT 后缀入口点 dispatch 为 NULL,vulkan/VulkanImage.cpp已改为优先解析无后缀核心名。
python3 -m onnx2vkop.cli -i resnet18-v2-7.onnx模型文件使用 FlatBuffers 格式(file identifier VKOP,version 1)。旧版
struct.pack 格式不再生成;已有的旧 .vkopbin 需用本命令重新转换。
- 支持量化:fp16, int8 对称量化
- 支持指定batch size
- 支持针对3D,4D NCHW to RGBA转换到模型
- 支持tensor合并,以便节约内存
usage: cli.py [-h] [-q QUANT] -i INPUT [-u] [-b BATCH] [-r]
options:
-h, --help show this help message and exit
-q, --quant QUANT Override input_model
-i, --input INPUT input_model file
-u, --unify convert initializers to a single memory block
-b, --batch BATCH batch size for inference
-r, --rgba nchw to rgba conversion for initializers
./benchmark/vkbench ../resnet18-v2-7.vkopbin dog.jpeg
支持将postproc 手动注册到gpu 处理,比如softmax,topk减少CPU与GPU间的内存吞吐
- KV cache 已 GPU 化:decode 每轮 present→past 为单条命令缓冲内的
device→device 拷贝,无 CPU 往返(原
KV_INPLACE_PLAN已完成并删除)。 - GPU shape-meta:张量带
shape_ssbo_侧信道(产出方填充),binary 广播 shader 的 broadcast==2 SSBO 路径 +dispatch_from_shape间接派发已落地, 由VKOP_GPU_SHAPE开关控制(默认关,走 CPU dims 回退)。稳态 readback 主要通过 Reshape/Expand 的自动学习缓存(LEARNING→CONFIRMING→STABLE)和 host-authoritative int64/int32 initializer 跳过回读来消除(原PHASE2_4_PLAN已完成/被替代并删除)。
一条 prompt 直接出 PNG,运行期不读任何 python/ORT 中间产物:C++ 分词 + 原始 ChatML
模板 + drop_idx 截取 + padding 屏蔽,以及 DiT 的 joint rope 频率表和 FlowMatch σ 调度
都在驱动内现算(--ref DIR 只用于数值对拍,可选)。
cd image/exporter
DYLD_FALLBACK_LIBRARY_PATH=/opt/homebrew/lib \
VK_ICD_FILENAMES=/opt/homebrew/etc/vulkan/icd.d/MoltenVK_icd.json \
../../build/image_gen dit_prefill_static.vkopbin dit_decode_static.vkopbin \
"A red fox sitting on a wooden bench in a sunlit park" 40 7
四份图按阶段顺序装载、用完立刻释放(单份 ~14 GB,任意两份同时驻留就超 36 GB 统一内存):
文本塔 text_encoder.vkopbin + 1.24 GB 输入查表 text_encoder_embeds.bin(mmap,一条
prompt 只碰其中几十行)、DiT prefill/decode、VAE vae_decoder_512.vkopbin;分词器复用
LLM 的 llm/tokenizer/qwen3_vl.bin。产物路径可用 TEXT_ENCODER_VKOPBIN /
TEXT_ENCODER_EMBEDS / TOKENIZER_BIN / VAE_VKOPBIN 覆盖。512×512 / 40 步实测 145 秒
(decode 2.6 s/步,文本塔前向 0.26 s,VAE 10.7 s)。导出、转换与逐位对拍口径见
image/exporter/BASELINE.md。
vkop is a mini AI inference engine based on Vulkan, with runtime logic under 1000 lines of code.
First, install the required dependencies, such as shaderc or Vulkan SDK:
wget https://sdk.lunarg.com/sdk/download/latest/linux/vulkan-sdk.tar.gz
tar xvf vulkan-sdk.tar.gz
source path/to/VulkanSDK/setup-env.sh
export PATH=$VULKAN_SDK/x86_64_bin:$PATHFor model conversion:
export CMAKE_POLICY_VERSION_MINIMUM=3.5
pip install onnx onnx-simplifier onnxsim onnxruntimeFor benchmarking models:
pip install torchvision
For testing dependencies, libtorch is downloaded and extracted automatically during the cmake process.
Set up the Vulkan ICD loader, using NVIDIA as an example:
export VK_ICD_FILENAMES=/usr/share/vulkan/icd.d/nvidia_icd.jsonEnsure the correct Vulkan version is used:
source path/to/VulkanSDK/setup-env.shcmake .. -DENABLE_TESTS=ON -DUSE_VALIDATION_LAYERS=OFF -DENABLE_ASAN=OFF -DUSE_DEBUG_LAYERS=OFF -DUSE_FP16=OFF -DUSE_MEASURE_TIME=OFFIf you are cross-compiling, set up the cross-compilation environment variables, based on toolchain.cmake:
cmake .. -DCMAKE_TOOLCHAIN_FILE=../toolchain.cmake -DENABLE_TESTS=OFF
Homebrew LLVM clang++ (/opt/homebrew/opt/llvm/bin/clang++) is required:
- Apple clang 17 segfaults in its frontend (infinite recursion in
TransformCXXFoldExpr) when compilingcore/function.cpp/core/runtime.cpp; - GCC + Apple libc++ has a broken ABI (garbled std::string/stream output at runtime).
Install dependencies and configure:
brew install cmake shaderc vulkan-loader vulkan-headers molten-vk \
glfw utf8proc pkgconf googletest libomp
cmake .. -DENABLE_TESTS=ON \
-DCMAKE_CXX_COMPILER=/opt/homebrew/opt/llvm/bin/clang++ \
-DCMAKE_C_COMPILER=/opt/homebrew/opt/llvm/bin/clang \
-DCMAKE_PREFIX_PATH="/opt/homebrew;/opt/homebrew/opt/llvm" \
-DPython3_EXECUTABLE=$(which python3)VulkanLib.cpp loads the Vulkan loader via dlopen at runtime, so these
environment variables must be set:
export DYLD_FALLBACK_LIBRARY_PATH=/opt/homebrew/lib
export VK_ICD_FILENAMES=/opt/homebrew/etc/vulkan/icd.d/MoltenVK_icd.jsonmacOS/clang compatibility points already handled in the code:
- The
h0/s0register aliases in the fp16 inline asm ofcore/Tensor.hppare GCC-only extensions, guarded by!defined(__clang__); clang takes the portable path; - On macOS
size_t(unsigned long) is NOTuint64_t(unsigned long long), so theTensorscalar ctor must explicitly acceptstd::size_t, otherwiseTensor<int64_t>(v.size())silently resolves toTensor(bool)and yields an empty tensor; - Once
VK_EXT_host_image_copywas promoted to Vulkan 1.4 core, the MoltenVK loader leaves the EXT-suffixed entry points with a NULL dispatch;vulkan/VulkanImage.cppresolves the unsuffixed core names first.
python3 -m onnx2vkop.cli -i resnet18-v2-7.onnxModel files use the FlatBuffers format (file identifier VKOP, version 1). The
legacy struct.pack format is no longer produced; existing old .vkopbin
files must be reconverted with this command.
- Supports quantization: fp16, int8 symmetric quantization
- Supports specifying batch size
- Supports 3D/4D NCHW to RGBA model conversion
- Supports tensor merging to save memory
usage: cli.py [-h] [-q QUANT] -i INPUT [-u] [-b BATCH] [-r]
options:
-h, --help show this help message and exit
-q, --quant QUANT Override input_model
-i, --input INPUT input_model file
-u, --unify convert initializers to a single memory block
-b, --batch BATCH batch size for inference
-r, --rgba nchw to rgba conversion for initializers
./benchmark/vkbench ../resnet18-v2-7.vkopbin dog.jpegSupports manually registering post-processing operations like softmax and top-k on the GPU to reduce memory throughput between CPU and GPU.
- KV cache is GPU-resident: each decode round copies present→past
device→device inside one command buffer, with no CPU round-trip (the
former
KV_INPLACE_PLANis done and removed). - GPU shape-meta: tensors carry a
shape_ssbo_side-channel (populated by the producing op); the broadcast==2 SSBO path in binary shaders plusdispatch_from_shapeindirect dispatch are wired up behind theVKOP_GPU_SHAPEenv flag (off by default, CPU dims fallback). Steady-state readbacks are instead eliminated via the Reshape/Expand auto-learning cache (LEARNING→CONFIRMING→STABLE) and host-authoritative int64/int32 initializers skipping readback (the formerPHASE2_4_PLANis done or superseded and removed).
One prompt in, one PNG out, and the driver reads no python/ORT intermediate
artifacts at runtime: tokenization, the raw ChatML template, drop_idx
slicing and padding masking happen in C++, and so do the DiT joint-rope
frequency table and the FlowMatch σ schedule (--ref DIR exists only for
numeric alignment and is optional).
cd image/exporter
DYLD_FALLBACK_LIBRARY_PATH=/opt/homebrew/lib \
VK_ICD_FILENAMES=/opt/homebrew/etc/vulkan/icd.d/MoltenVK_icd.json \
../../build/image_gen dit_prefill_static.vkopbin dit_decode_static.vkopbin \
"A red fox sitting on a wooden bench in a sunlit park" 40 7
The four graphs load one stage at a time and are released right after use (each
is ~14 GB; any two together exceed 36 GB of unified memory): the text tower
text_encoder.vkopbin plus its 1.24 GB input table text_encoder_embeds.bin
(mmapped — a prompt touches a few dozen of its rows), DiT prefill/decode, and
the VAE vae_decoder_512.vkopbin. Tokenization reuses the LLM's
llm/tokenizer/qwen3_vl.bin. Override paths with TEXT_ENCODER_VKOPBIN,
TEXT_ENCODER_EMBEDS, TOKENIZER_BIN, VAE_VKOPBIN. Measured 512×512 at 40
steps: 145 s (2.6 s/step decode, 0.26 s text-tower forward, 10.7 s VAE). Export,
conversion and the bit-wise alignment protocol are in
image/exporter/BASELINE.md.