GPU VulnDB

Database/AI/ML frameworks & serving

vLLM: video decoder limit bypass via sampler subclass shadowing exhausts unaccounted GPU memory

CVSS 6.3CVE-2026-100649AI/ML frameworks & servingcurated

Impact

Decoder limits are tracked per sampler subclass, so a caller who selects different sampler subclasses in video requests gets independent counters and can allocate more hardware video decoders than the operator configured. The extra allocations consume GPU memory that vLLM does not account for, which is the memory the model weights and KV cache depend on. On a shared GPU node this pushes the serving process toward out-of-memory failures that the scheduler cannot anticipate, and the practical fix is restarting the affected server. NVD records availability impact only.

Who can reach it

Anyone who can reach the vLLM API and submit video requests; the advisory describes the attacker as unauthenticated. Attack complexity is rated high and the attack requires specific conditions, so it is not a trivial one-shot.

What to do

Upgrade vLLM to 0.29.0 or later and restart the serving processes - a rolling restart per replica, no node reboot. Meanwhile disable video input support or put the endpoint behind a gateway that authenticates and rate-limits video requests.

References

Related entries

All AI/ML frameworks & serving entries

This entry is curated: imported from vendor advisories with machine assistance, not yet individually verified. Confirm against your vendor's advisory before acting, and report anything wrong.