GPU VulnDB

Database/AI/ML frameworks & serving

vLLM: remote media is fully materialized before size and per-prompt limits are enforced

CVSS 7.1CVE-2026-100650AI/ML frameworks & servingcurated

Impact

vLLM through 0.29.0 fetches and fully materializes remote or inline media before applying the compressed-audio size cap and the per-modality item limits. Across four ingress paths the server reads the whole HTTP body, base64-decodes the inline payload, or spawns a fetch/decode task per media part, and only then rejects the request. A caller can force the API server or batch runner to allocate memory and burn outbound bandwidth proportional to a size they choose, exhausting the host before inference even starts, which takes the model replica down on that GPU node. The chat and batch surfaces require an API key when one is configured; the Rust frontend /tokenize route is unauthenticated by design. No code execution or data disclosure.

Who can reach it

Any client that can reach the API - authenticated for the chat and batch paths when a key is set, unauthenticated for the Rust frontend POST /tokenize. The remote-fetch path also lets the attacker point the server at a host they control to size the response.

What to do

Upgrade to vLLM 0.29.0 or later (fix commit 752a3a5) and restart each serving and batch-runner process. Meanwhile, bound request body size at the ingress, restrict or block outbound egress from serving pods so remote media URLs cannot be fetched arbitrarily, and set a memory limit on the pod so one replica cannot take the node with it.

References

Related entries

All AI/ML frameworks & serving entries

This entry is curated: imported from vendor advisories with machine assistance, not yet individually verified. Confirm against your vendor's advisory before acting, and report anything wrong.