GPU VulnDB

Database/AI/ML frameworks & serving

vLLM: forged multimodal features on the disaggregated generate path kill EngineCore and poison encoder cache

CVSS 6.5CVE-2026-105754AI/ML frameworks & servingcurated

Impact

On the disaggregated scale-out path, /inference/v1/generate accepted caller-supplied tensors (features.kwargs_data), cache identifiers (features.mm_hashes), placeholder ranges (features.mm_placeholders) and wire-selected multimodal field processors without rebinding them to the active model's renderer contract. Forged grid geometry, wrong field types or non-positive placeholder lengths terminate the shared EngineCore, dropping every in-flight request on that server and the GPU memory state behind them. Worse, an attacker who knows or can induce a victim's content hash can use forged cache hashes to poison or retrieve cross-request encoder-cache state - a tenant-boundary break inside one serving process, not just a crash. Dropped sparse placeholder masks can also change replayed transport semantics.

Who can reach it

A client that can reach the /inference/v1/generate endpoint with low-privilege credentials (CVSS AV:N/PR:L) on a deployment running the disaggregated scale-out path. The cross-request cache read additionally requires knowing or inducing the victim's media content hash.

What to do

Upgrade vLLM to 0.30.0 and restart the affected serving replicas; model weights and GPU nodes are untouched, so this is a rolling restart of the inference deployment, not a node drain. If you cannot upgrade immediately, do not expose the disaggregated /inference/v1/generate endpoint to untrusted callers - keep it on an internal network between prefill and decode workers only.

References

Related entries

All AI/ML frameworks & serving entries

This entry is curated: imported from vendor advisories with machine assistance, not yet individually verified. Confirm against your vendor's advisory before acting, and report anything wrong.