Database/AI/ML frameworks & serving
vLLM: remote media is fully materialized before size and per-prompt limits are enforced
Impact
vLLM through 0.29.0 fetches and fully materializes remote or inline media before applying the compressed-audio size cap and the per-modality item limits. Across four ingress paths the server reads the whole HTTP body, base64-decodes the inline payload, or spawns a fetch/decode task per media part, and only then rejects the request. A caller can force the API server or batch runner to allocate memory and burn outbound bandwidth proportional to a size they choose, exhausting the host before inference even starts, which takes the model replica down on that GPU node. The chat and batch surfaces require an API key when one is configured; the Rust frontend /tokenize route is unauthenticated by design. No code execution or data disclosure.
Who can reach it
Any client that can reach the API - authenticated for the chat and batch paths when a key is set, unauthenticated for the Rust frontend POST /tokenize. The remote-fetch path also lets the attacker point the server at a host they control to size the response.
What to do
Upgrade to vLLM 0.29.0 or later (fix commit 752a3a5) and restart each serving and batch-runner process. Meanwhile, bound request body size at the ingress, restrict or block outbound egress from serving pods so remote media URLs cannot be fetched arbitrarily, and set a memory limit on the pod so one replica cannot take the node with it.
References
Related entries
- vLLM: overlong token_ids on the disaggregated serving endpoint crash the workerCVE-2026-100651 · vLLM (disaggregated serving endpoint /inference/v1/generate, decoder prompt-length validation)High
- vLLM: out-of-range stop_token_ids trigger a CUDA device assertion and wedge EngineCoreCVE-2026-100654 · vLLM (OpenAI-compatible /v1/completions and /v1/chat/completions, stop_token_ids validation)High
- vLLM (`MediaConnector`): SSRF, recurrence of CVE-2025-6242CVE-2026-24779 · vLLM (`MediaConnector`)High
- vLLM (`load_from_url_async`): Bypass of the CVE-2026-24779 SSRF fixCVE-2026-25960 · vLLM (`load_from_url_async`)High
- picklescan: `scan_pytorch` bypass via forged magic numbersCVE-2026-53875 · picklescanHigh
- torchvision (GIF decoder): Out-of-bounds heap read in `read_from_tensor` GIF decodeCVE-2026-65918 · torchvision (GIF decoder)High
This entry is curated: imported from vendor advisories with machine assistance, not yet individually verified. Confirm against your vendor's advisory before acting, and report anything wrong.