anm2211 commited on
Commit
8c5fe01
·
verified ·
1 Parent(s): 25664a5

Update README.md

Browse files
Files changed (1) hide show
  1. README.md +0 -31
README.md CHANGED
@@ -94,37 +94,6 @@ humming-kernels`). See **Serving with vLLM** below.
94
 
95
  ## Serving with vLLM
96
 
97
- > **Temporary vLLM compatibility note:** the upstream vLLM Humming MoE
98
- > implementation currently has a bug that prevents this checkpoint from being
99
- > served correctly. Until the fix is available upstream, use one of the
100
- > following workarounds:
101
- >
102
- > 1. Build vLLM from our fork:
103
- >
104
- > ```bash
105
- > git clone https://github.com/adotdad/vllm.git
106
- > cd vllm
107
- > pip install -e .
108
- > ```
109
- >
110
- > 2. Or patch your existing vLLM installation by replacing:
111
- >
112
- > ```text
113
- > vllm/model_executor/layers/quantization/humming.py
114
- > ```
115
- >
116
- > with the version from our fork:
117
- >
118
- > ```text
119
- > https://github.com/adotdad/vllm
120
- > ```
121
-
122
- Install the Humming kernels (required for vLLM to load this checkpoint):
123
-
124
- ```bash
125
- pip install humming-kernels
126
- ```
127
-
128
  Hopper (sm_90) or Ampere (sm ≥ 80) GPUs required for serving.
129
 
130
  ```bash
 
94
 
95
  ## Serving with vLLM
96
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
97
  Hopper (sm_90) or Ampere (sm ≥ 80) GPUs required for serving.
98
 
99
  ```bash