GPU VulnDB

Database/AI/ML frameworks & serving

vLLM: audio clip size limit not enforced, letting unauthenticated clients exhaust node memory and CPU

CVSS 6.9CVE-2026-100648AI/ML frameworks & servingcurated

Impact

The configured maximum audio clip size is not applied in the multimodal chat path, so a client can submit an oversized audio file and force the server to decode it. Decoding runs on the inference node and consumes host memory and CPU outside the GPU memory accounting, which starves the model server that shares the node. On a GPU fleet that means a single unauthenticated caller can degrade or stall an expensive accelerator node serving many users, and recovery usually means restarting the server process rather than a cheap throttle. NVD records availability impact only, no confidentiality or integrity effect.

Who can reach it

Anyone who can reach the vLLM chat endpoint over the network. The advisory states the clients are unauthenticated, so any exposure of the OpenAI-compatible API to tenants or to the internet is enough.

What to do

Upgrade vLLM to 0.29.0 or later and restart the serving processes; a rolling restart per replica is enough, no node drain or reboot. Until then, keep the multimodal chat endpoints behind an authenticating gateway that enforces its own request body size limit, or disable audio inputs.

References

Related entries

All AI/ML frameworks & serving entries

This entry is curated: imported from vendor advisories with machine assistance, not yet individually verified. Confirm against your vendor's advisory before acting, and report anything wrong.