What happened
NVD published CVE-2026-105922 (CVSS 4.3) on 2026-10-06 for a denial-of-service flaw in vLLM's penalty handler with a public exploit released: manipulating the penalty-handling path in model_executor/layers/utils.py crashes the server. A companion OOB read in the mamba mixer (CVE-2026-105775) was also published for vLLM up to 0.31.0.
Why it matters
vLLM is the most widely deployed open-source LLM serving engine; remote DoS against the serving endpoint takes the whole inference cluster offline with a single crafted request — cheap to weaponize, no credentials required if the endpoint is network-reachable, and the public exploit shortens time-to-attack.
Attack vector
A specially crafted request reaches the Penalty Handler path where get_token_bin_counts_and_mask performs an improper resource release / invalid manipulation, causing the vLLM inference process to crash — an availability attack against the model-serving endpoint.
Affected systems
vLLM up to 0.31.0 (vllm/model_executor/layers/utils.py get_token_bin_counts_and_mask)
Mitigation
Monitor NVD for a fixed vLLM release (none published yet as of writing); the issue was reported via issue tracker without a vendor response. Apply network filtering and request-rate limiting in front of vLLM endpoints.