fatal: run with vllm 0.3 0.30.0+cu129

#6
by rt5566 - opened

(APIServer pid=220981) pydantic_core._pydantic_core.ValidationError: 1 validation error for ModelConfig
(APIServer pid=220981) Value error, Unknown quantization method: exl3. Must be one of ['awq', 'auto_awq', 'fp8', 'fbgemm_fp8', 'fp_quant', 'modelopt', 'modelopt_fp4', 'modelopt_mxfp8', 'modelopt_mixed', 'auto_gptq', 'gptq', 'gptq_marlin', 'awq_marlin', 'humming', 'compressed-tensors', 'experts_int8', 'quark', 'moe_wna16', 'torchao', 'inc', 'mxfp4', 'gpt_oss_mxfp4', 'deepseek_v4_fp8', 'online', 'fp8_per_tensor', 'fp8_per_block', 'fp8_per_channel', 'int8_per_channel_weight_only', 'nvfp4_per_token', 'mxfp8']. [type=value_error, input_value=ArgsKwargs((), {'model': ...nderer_num_workers': 1}), input_type=ArgsKwargs]
(APIServer pid=220981) For further information visit https://errors.pydantic.dev/2.13/v/value_error

Unknown quantization method: exl3 means vLLM started without the OrcaSAQ2 plugin loaded. This checkpoint's config.json has "quant_method": "exl3" (3.21 bpw average), which stock vLLM doesn't know. The plugin registers it through the vllm.general_plugins entry point, so it has to be installed into the same Python env that runs vllm serve:

pip install git+https://github.com/Continuum-AI-Corp/OrcaSAQ2-kernel

It also needs an exllamav3 wheel built for that env's torch/CUDA (they're release assets, not on PyPI; see the kernel repo README), and TP>1 isn't implemented, so run with a single GPU. On memory: the four safetensors shards come to 12.27 GB (11.43 GiB). The repo's presets/16gb.env (15.5 GiB cap, fp8_e4m3 KV, block size 32, MAXLEN=32768) is the tested starting point for a 16 GB card. Their measured KV pool is 59,753 tokens without MTP, or 35,617 with MTP.

Sign up or log in to comment