Database/AI/ML frameworks & serving
vLLM: mirrored multimodal cache desync lets a rejected request assert-kill the engine on reuse
Impact
The default mirrored multimodal LRU cache commits a media hash in the frontend sender cache during rendering, before the engine admits the request. If that request is then rejected, the engine-side receiver cache never gets the payload, and the two caches disagree. A later request reusing the same media hash causes the sender to send nothing and the receiver to hit an assertion ("Expected a cached item"), taking down the shared serving process. On a GPU node this is one process serving many tenants: everyone's in-flight generations die and the replica must restart, reloading weights into HBM before it can serve again. Confidentiality and integrity are not affected.
Who can reach it
Any authenticated client that can submit multimodal requests over the network (CVSS AV:N/PR:L). The trigger is a request that gets rejected after rendering, followed by reuse of the same media - reachable without special privileges and plausibly hit by accident.
What to do
Upgrade vLLM to 0.28.0 or later and restart the serving replicas. Expect the usual weight-reload warm-up per replica; roll replicas one at a time to keep the endpoint up. No GPU driver, firmware or node-level change is required.
References
Related entries
- vLLM: forged multimodal features on the disaggregated generate path kill EngineCore and poison encoder cacheCVE-2026-105754 · vLLM /inference/v1/generate (disaggregated scale-out: caller-supplied tensors, mm_hashes, mm_placeholders)Medium
- vLLM: unvalidated cache_salt raises an uncaught ValueError and terminates EngineCoreCVE-2026-105756 · vLLM OpenAI-compatible API (cache_salt validation with the LMCache-MP connector)Medium
- vLLM: structured-output request failures escape request scope and terminate the shared engineCVE-2026-105757 · vLLM structured-output path (grammar compilation, ngram_gpu speculative decoding, Rust frontend validation)Medium
- vLLM: unbounded frame count in video/jpeg base64 data URLs crashes the server with OOMCVE-2026-34755 · vLLM OpenAI-compatible API server (video/jpeg base64 multimodal path)Medium
- vLLM: no upper bound on the n parameter lets a single request OOM the API serverCVE-2026-34756 · vLLM OpenAI-compatible API server (ChatCompletionRequest/CompletionRequest n parameter)Medium
- vLLM (revision pinning): Revision pinning does not apply to all model artifactsCVE-2026-47155 · vLLM (revision pinning)Medium
This entry is curated: imported from vendor advisories with machine assistance, not yet individually verified. Confirm against your vendor's advisory before acting, and report anything wrong.