SM70 FP8 block-32 kernels (PLI Labs)

These are weight-only FP8 kernels for NVIDIA Volta (V100, SM70). They serve checkpoints whose weights are E4M3 with one power-of-two (UE8M0) scale per 32 × 32 block, the dense format of DeepSeek-V4.1-Flash (weight_block_size: [32, 32], scale_fmt: ue8m0). Volta has no FP8 tensor cores, so the weights stay FP8 in memory and are expanded to FP16 inside the kernels. Weight memory falls from 2 bytes per parameter (FP16 dequantization) to 1.0625 bytes. For DeepSeek-V4.1-Flash's per-layer dense set at TP4, decode-size products (1–16 rows) run 9–36% faster than FP16 weights with cuBLAS. Individual projections vary (see Results).

This repository is a distribution of a patch series for 1Cat-vLLM, the vLLM fork for Volta. It contains source, tests and measurements, not model weights.

Contents

path what
patches/0001-…patch TurboMind SM70 W8A16: a Config_E4M3 GroupSizeV=32 tile family (M 8..128); fp8_sm70_prepare, fp8_gemm_sm70_out and fp8_sm70_dequantize_out accept group_size=32
patches/0002-…patch vllm/model_executor/layers/quantization/utils/sm70_fp8_block32.py: packing, Triton GEMV, TurboMind and dequant paths, measured dispatch, custom op; tests and benchmark
patches/0003-…patch fp8.py: [32, 32] block-FP8 linears on SM70 use these kernels
src/ the Python/Triton module, its tests and the benchmark, as standalone files
results/ benchmark JSON and tables, test outputs, the GPU clock/power trace
LICENSE Apache-2.0, as 1Cat-vLLM / vLLM

How it works

Each weight is stored once, in TurboMind's SM70 B-operand layout: 32-row panels [N/32][K/8][32][8] with byte order 0,2,1,3 inside each 8-byte chunk, plus one FP16 scale per 32 K-values of each row. Three kernels read that one copy, chosen per call inside the opaque custom op torch.ops.vllm.sm70_fp8_block32_linear:

rows M (FP16 output) kernel
≤ 2 (≤ 4 when N ≤ 1152) Triton GEMV. It decodes E4M3 by integer ops ((b & 0x7F) << 7 | (b & 0x80) << 8 is the value × 2^-8 as FP16) and accumulates in FP32. It can also write FP32 output, up to M = 8.
up to 16 / 32 / 64 / 128 (by N, K) TurboMind group-32 HMMA (FP32 accumulation)
above FP16 rebuild of the weight + cuBLAS

Exactness. E4M3 values are multiples of 2^-9 with at most 4 significant bits and |v| ≤ 448. For scales 2^e with e ∈ [-14, 7], every v × 2^e is an exact FP16 value, so all three kernels multiply exactly the weights of the load-time FP16 dequantization. They differ from it only in FP32 accumulation order.

The packer refuses, loudly:

  • scales outside that window, since 448 × 2^8 would overflow FP16;
  • scales that are not powers of two;
  • E4M3 NaN codes.

In vLLM, a layer whose scales fail the window check is dequantized to FP16 at load instead.

Results

V100-SXM2-32GB, SM clock capped at 1477 MHz, 180 W board power limit, driver 580.159.03, CUDA 12.8.93, torch 2.10.0+cu128. Timings are CUDA-graph replays with weight rotation, so the 6 MB L2 never serves a repeat. The full tables are in results/bench_tables.md, and every kernel at every size is in results/bench_all_paths.txt.

DeepSeek-V4.1-Flash, per layer per GPU at TP4. The set is the FP16-output dense projections (wq_a+wkv, wq_b, indexer wq_b, 2 × wo_a), which is 63.5 MiB of FP16 weights versus 33.7 MiB as FP8 block-32.

M FP16 weights + cuBLAS (µs) FP8 block-32 (µs) change
1 115.2 74.2 −36%
2 129.3 93.8 −27%
4 132.1 103.4 −22%
8 135.4 112.4 −17%
16 137.7 125.3 −9%
32 185.2 209.7 +13%
64 233.7 287.1 +23%
128 430.5 414.8 −4%
512 687.4 782.6 +14%
4096 4994.4 4335.4 −13%

From M ≈ 32 upward, holding only the FP8 copy costs time against FP16 weights at several shapes (up to +75% on the narrow wo_a group at M 32–64). Choose FP16 weights when speed at mid-size batches matters more than memory. The board's 180 W limit throttles the SM clock to about 950–1100 MHz in compute-bound rows (results/gpu_clock_power_trace.txt).

Tests (results/new_tests_pytest.txt): 38/38 pass. They cover:

  • bitwise one-hot decode over all finite codes and both window ends;
  • products against the FP64 reference at 11 shapes × 16 M values;
  • refusals;
  • the unchanged group-128 ops;
  • the vLLM route.

The 22 existing 1Cat SM70 FP8 suites give 435 passed and 4 failed, identical with and without the patches (results/onecat_sm70_fp8_suites_*.txt).

Build and use

No prebuilt binary is published. The kernels live in 1Cat-vLLM's vllm/_C extension (about 340 MB for SM70), which only loads with the exact Python tree it was built from. A copy detached from that tree would not be self-contained.

Build from source:

git clone https://github.com/1CatAI/1Cat-vLLM && cd 1Cat-vLLM
git checkout b5e926550          # base the patches were made on
git am /path/to/patches/*.patch # or: git fetch https://github.com/ProprietaryLegal/1Cat-vLLM feat/sm70-fp8-block32-weight-only
# source build for SM70 as 1Cat-vLLM's README describes (CUDA 12.8, Python 3.12):
export TORCH_CUDA_ARCH_LIST=7.0 CMAKE_CUDA_ARCHITECTURES=70 MAX_JOBS=8
pip install -r requirements/build/cuda.txt
pip install -e . --no-build-isolation
pytest tests/kernels/quantization/test_sm70_fp8_block32.py tests/quantization/test_sm70_fp8_block32_route.py
python benchmarks/kernels/benchmark_sm70_fp8_block32.py --out results/

Toolchain used for these results: CUDA 12.8.93, torch 2.10.0+cu128, Python 3.12, x86_64.

Serving a [32, 32] block-FP8 checkpoint on V100 then needs no flag. Under the default kernel_config.sm70_fp8 policy, eligible layers are packed automatically. VLLM_SM70_FP8_TURBOMIND=0 (kernel_config.sm70_fp8.enabled=False) keeps the FP16 load-time dequantization.

Direct use of the module:

from vllm.model_executor.layers.quantization.utils import sm70_fp8_block32 as b32
w = b32.prepare_fp8_block32(weight_e4m3, scale_ue8m0)   # [N, K], [N/32, K/32]
y = b32.fp8_block32_linear(x_fp16, w)                     # FP16 out; out_dtype=torch.float32 also supported

Limitations

  • Mid-size batches (M 32–64, and 256–512 on some shapes) are slower than FP16 weights with cuBLAS, as described above.
  • There is no FP32-output TurboMind epilogue yet. Row-parallel outputs kept in FP32 use the GEMV (M ≤ 8) or the rebuild path.
  • Grouped (is_bmm) layers keep 1Cat's existing grouped route.
  • The fused gated-SiLU and exact-8K pre-scaled TurboMind variants remain group-128 only.

Credits

The kernels were developed for the PLI Labs DeepSeek-V4.1-Flash port to V100. They build on 1Cat-vLLM's SM70 TurboMind integration (lmdeploy TurboMind) and on vLLM. The upstream pull request to 1CatAI/1Cat-vLLM carries the same patch series.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support