Vulnerability  ·  2026-10-07

vLLM EngineCore DoS / cache-poisoning cluster — authenticated caller can crash the shared LLM serving engine

VulnerabilityMedium impactGlobalCVE-2026-105754
A coordinated batch of vLLM advisories published Oct 5 (fixed in 0.28.0/0.30.0): the multimodal-mirror LRU cache commits a media hash before engine admission so a rejected request leaves caches desynchronized and a later request hits an assertion (105753); the disaggregated /inference/v1/generate endpoint accepts caller-supplied tensors/hashes/placeholders that can terminate the shared EngineCore and poison or retrieve cross-request encoder-cache state (105754); a malformed cache_salt can raise an uncaught ValueError in LMCache-MP lookups (105756); structured-output failures reach the EngineCore fatal path (105757); unauthenticated arbitrary HTTP method tokens grow Prometheus label sets until memory is exhausted (105759); and /score,/rerank derive query_key from the caller-controlled X-Request-Id header enabling query-embedding overwrite (105755).
vLLM is among the most widely deployed open-source LLM serving backends. These are reached by ordinary API callers (front/back-end mode or OpenAI-compatible endpoints), and each lets a single low-privilege or unauthenticated request kill the shared EngineCore — a one-request shared-availability outage — or corrupt cache state affecting other tenants sharing the serving node.
Send crafted multimodal/constrained-generation requests, malformed cache_salt values, or repeated unique HTTP method tokens to the serving endpoints to trigger assertions/engine termination or unbounded metric growth.
vLLM < 0.28.0 (CVE-2026-105753 mirror-cache assertion); < 0.30.0 (CVE-2026-105754 /inference/v1/generate EngineCore termination + encoder-cache poisoning; CVE-2026-105756 LMCache-MP cache_salt crash; CVE-2026-105757 structured-output fatal path; CVE-2026-105759 Prometheus label memory exhaustion; CVE-2026-105755 /score, /rerank X-Request-Id cache overwrite)
Upgrade to vLLM 0.28.0 (105753) and 0.30.0 (105754/105756/105757/105759/105755); constrain request-level media/cache arguments and rate-limit /metrics where possible.
NVD CVE-2026-105754GitHub advisory GHSA-ph72-cqr5-qpp7NVD CVE-2026-105753NVD CVE-2026-105759
See this in the live feed Explore related AI security and governance findings — updated every morning.
Open the feed →