Instructions to use ProprietaryLegal/sm70-fp8-block32-kernels with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Kernels
How to use ProprietaryLegal/sm70-fp8-block32-kernels with Kernels:
# !pip install kernels from kernels import get_kernel # a version (or an explicit revision) is required; see the "Files and versions" tab for the available ones kernel = get_kernel("ProprietaryLegal/sm70-fp8-block32-kernels", version=1) - Notebooks
- Google Colab
- Kaggle
SM70 FP8 block-32 kernels (PLI Labs)
These are weight-only FP8 kernels for NVIDIA Volta (V100, SM70). They serve checkpoints whose weights are E4M3 with one power-of-two (UE8M0) scale per 32 × 32 block, the dense format of DeepSeek-V4.1-Flash (weight_block_size: [32, 32], scale_fmt: ue8m0). Volta has no FP8 tensor cores, so the weights stay FP8 in memory and are expanded to FP16 inside the kernels. Weight memory falls from 2 bytes per parameter (FP16 dequantization) to 1.0625 bytes. For DeepSeek-V4.1-Flash's per-layer dense set at TP4, decode-size products (1–16 rows) run 9–36% faster than FP16 weights with cuBLAS. Individual projections vary (see Results).
This repository is a distribution of a patch series for 1Cat-vLLM, the vLLM fork for Volta. It contains source, tests and measurements, not model weights.
Contents
| path | what |
|---|---|
patches/0001-…patch |
TurboMind SM70 W8A16: a Config_E4M3 GroupSizeV=32 tile family (M 8..128); fp8_sm70_prepare, fp8_gemm_sm70_out and fp8_sm70_dequantize_out accept group_size=32 |
patches/0002-…patch |
vllm/model_executor/layers/quantization/utils/sm70_fp8_block32.py: packing, Triton GEMV, TurboMind and dequant paths, measured dispatch, custom op; tests and benchmark |
patches/0003-…patch |
fp8.py: [32, 32] block-FP8 linears on SM70 use these kernels |
src/ |
the Python/Triton module, its tests and the benchmark, as standalone files |
results/ |
benchmark JSON and tables, test outputs, the GPU clock/power trace |
LICENSE |
Apache-2.0, as 1Cat-vLLM / vLLM |
How it works
Each weight is stored once, in TurboMind's SM70 B-operand layout: 32-row panels [N/32][K/8][32][8] with byte order 0,2,1,3 inside each 8-byte chunk, plus one FP16 scale per 32 K-values of each row. Three kernels read that one copy, chosen per call inside the opaque custom op torch.ops.vllm.sm70_fp8_block32_linear:
| rows M (FP16 output) | kernel |
|---|---|
| ≤ 2 (≤ 4 when N ≤ 1152) | Triton GEMV. It decodes E4M3 by integer ops ((b & 0x7F) << 7 | (b & 0x80) << 8 is the value × 2^-8 as FP16) and accumulates in FP32. It can also write FP32 output, up to M = 8. |
| up to 16 / 32 / 64 / 128 (by N, K) | TurboMind group-32 HMMA (FP32 accumulation) |
| above | FP16 rebuild of the weight + cuBLAS |
Exactness. E4M3 values are multiples of 2^-9 with at most 4 significant bits and |v| ≤ 448. For scales 2^e with e ∈ [-14, 7], every v × 2^e is an exact FP16 value, so all three kernels multiply exactly the weights of the load-time FP16 dequantization. They differ from it only in FP32 accumulation order.
The packer refuses, loudly:
- scales outside that window, since 448 × 2^8 would overflow FP16;
- scales that are not powers of two;
- E4M3 NaN codes.
In vLLM, a layer whose scales fail the window check is dequantized to FP16 at load instead.
Results
V100-SXM2-32GB, SM clock capped at 1477 MHz, 180 W board power limit, driver 580.159.03, CUDA 12.8.93, torch 2.10.0+cu128. Timings are CUDA-graph replays with weight rotation, so the 6 MB L2 never serves a repeat. The full tables are in results/bench_tables.md, and every kernel at every size is in results/bench_all_paths.txt.
DeepSeek-V4.1-Flash, per layer per GPU at TP4. The set is the FP16-output dense projections (wq_a+wkv, wq_b, indexer wq_b, 2 × wo_a), which is 63.5 MiB of FP16 weights versus 33.7 MiB as FP8 block-32.
| M | FP16 weights + cuBLAS (µs) | FP8 block-32 (µs) | change |
|---|---|---|---|
| 1 | 115.2 | 74.2 | −36% |
| 2 | 129.3 | 93.8 | −27% |
| 4 | 132.1 | 103.4 | −22% |
| 8 | 135.4 | 112.4 | −17% |
| 16 | 137.7 | 125.3 | −9% |
| 32 | 185.2 | 209.7 | +13% |
| 64 | 233.7 | 287.1 | +23% |
| 128 | 430.5 | 414.8 | −4% |
| 512 | 687.4 | 782.6 | +14% |
| 4096 | 4994.4 | 4335.4 | −13% |
From M ≈ 32 upward, holding only the FP8 copy costs time against FP16 weights at several shapes (up to +75% on the narrow wo_a group at M 32–64). Choose FP16 weights when speed at mid-size batches matters more than memory. The board's 180 W limit throttles the SM clock to about 950–1100 MHz in compute-bound rows (results/gpu_clock_power_trace.txt).
Tests (results/new_tests_pytest.txt): 38/38 pass. They cover:
- bitwise one-hot decode over all finite codes and both window ends;
- products against the FP64 reference at 11 shapes × 16 M values;
- refusals;
- the unchanged group-128 ops;
- the vLLM route.
The 22 existing 1Cat SM70 FP8 suites give 435 passed and 4 failed, identical with and without the patches (results/onecat_sm70_fp8_suites_*.txt).
Build and use
No prebuilt binary is published. The kernels live in 1Cat-vLLM's vllm/_C extension (about 340 MB for SM70), which only loads with the exact Python tree it was built from. A copy detached from that tree would not be self-contained.
Build from source:
git clone https://github.com/1CatAI/1Cat-vLLM && cd 1Cat-vLLM
git checkout b5e926550 # base the patches were made on
git am /path/to/patches/*.patch # or: git fetch https://github.com/ProprietaryLegal/1Cat-vLLM feat/sm70-fp8-block32-weight-only
# source build for SM70 as 1Cat-vLLM's README describes (CUDA 12.8, Python 3.12):
export TORCH_CUDA_ARCH_LIST=7.0 CMAKE_CUDA_ARCHITECTURES=70 MAX_JOBS=8
pip install -r requirements/build/cuda.txt
pip install -e . --no-build-isolation
pytest tests/kernels/quantization/test_sm70_fp8_block32.py tests/quantization/test_sm70_fp8_block32_route.py
python benchmarks/kernels/benchmark_sm70_fp8_block32.py --out results/
Toolchain used for these results: CUDA 12.8.93, torch 2.10.0+cu128, Python 3.12, x86_64.
Serving a [32, 32] block-FP8 checkpoint on V100 then needs no flag. Under the default kernel_config.sm70_fp8 policy, eligible layers are packed automatically. VLLM_SM70_FP8_TURBOMIND=0 (kernel_config.sm70_fp8.enabled=False) keeps the FP16 load-time dequantization.
Direct use of the module:
from vllm.model_executor.layers.quantization.utils import sm70_fp8_block32 as b32
w = b32.prepare_fp8_block32(weight_e4m3, scale_ue8m0) # [N, K], [N/32, K/32]
y = b32.fp8_block32_linear(x_fp16, w) # FP16 out; out_dtype=torch.float32 also supported
Limitations
- Mid-size batches (M 32–64, and 256–512 on some shapes) are slower than FP16 weights with cuBLAS, as described above.
- There is no FP32-output TurboMind epilogue yet. Row-parallel outputs kept in FP32 use the GEMV (M ≤ 8) or the rebuild path.
- Grouped (
is_bmm) layers keep 1Cat's existing grouped route. - The fused gated-SiLU and exact-8K pre-scaled TurboMind variants remain group-128 only.
Credits
The kernels were developed for the PLI Labs DeepSeek-V4.1-Flash port to V100. They build on 1Cat-vLLM's SM70 TurboMind integration (lmdeploy TurboMind) and on vLLM. The upstream pull request to 1CatAI/1Cat-vLLM carries the same patch series.