GPU VulnDB

Database/AI/ML frameworks & serving

vLLM: mirrored multimodal cache desync lets a rejected request assert-kill the engine on reuse

CVSS 6.5CVE-2026-105753AI/ML frameworks & servingcurated

Impact

The default mirrored multimodal LRU cache commits a media hash in the frontend sender cache during rendering, before the engine admits the request. If that request is then rejected, the engine-side receiver cache never gets the payload, and the two caches disagree. A later request reusing the same media hash causes the sender to send nothing and the receiver to hit an assertion ("Expected a cached item"), taking down the shared serving process. On a GPU node this is one process serving many tenants: everyone's in-flight generations die and the replica must restart, reloading weights into HBM before it can serve again. Confidentiality and integrity are not affected.

Who can reach it

Any authenticated client that can submit multimodal requests over the network (CVSS AV:N/PR:L). The trigger is a request that gets rejected after rendering, followed by reuse of the same media - reachable without special privileges and plausibly hit by accident.

What to do

Upgrade vLLM to 0.28.0 or later and restart the serving replicas. Expect the usual weight-reload warm-up per replica; roll replicas one at a time to keep the endpoint up. No GPU driver, firmware or node-level change is required.

References

Related entries

All AI/ML frameworks & serving entries

This entry is curated: imported from vendor advisories with machine assistance, not yet individually verified. Confirm against your vendor's advisory before acting, and report anything wrong.