Database/AI/ML frameworks & serving
vLLM: audio clip size limit not enforced, letting unauthenticated clients exhaust node memory and CPU
Impact
The configured maximum audio clip size is not applied in the multimodal chat path, so a client can submit an oversized audio file and force the server to decode it. Decoding runs on the inference node and consumes host memory and CPU outside the GPU memory accounting, which starves the model server that shares the node. On a GPU fleet that means a single unauthenticated caller can degrade or stall an expensive accelerator node serving many users, and recovery usually means restarting the server process rather than a cheap throttle. NVD records availability impact only, no confidentiality or integrity effect.
Who can reach it
Anyone who can reach the vLLM chat endpoint over the network. The advisory states the clients are unauthenticated, so any exposure of the OpenAI-compatible API to tenants or to the internet is enough.
What to do
Upgrade vLLM to 0.29.0 or later and restart the serving processes; a rolling restart per replica is enough, no node drain or reboot. Until then, keep the multimodal chat endpoints behind an authenticating gateway that enforces its own request body size limit, or disable audio inputs.
References
Related entries
- Hugging Face Accelerate: unsanitized shard paths in a checkpoint index give arbitrary file read and hangsCVE-2026-69112 · Hugging Face Accelerate (load_checkpoint_in_model / load_checkpoint_and_dispatch weight_map handling)Medium
- llama.cpp ggml RPC server: unvalidated tensor op and op_params in deserialize_tensorCVE-2026-78147 · llama.cpp ggml RPC server (deserialize_tensor op / op_params validation)Medium
- llama.cpp ggml RPC server: null pointer dereference in graph_compute kills the GPU workerCVE-2026-78148 · llama.cpp ggml RPC server (rpc_server::graph_compute)Medium
- BentoML: SSRF filter misses 100.64.0.0/10, so serving pods fetch from internal CGNAT hostsCVE-2026-78205 · BentoML make_safe_connect (SSRF address filter, RFC 6598 range)Medium
- vLLM: DeepStream backend misclassification skips pixel limits and lets unauthenticated video exhaust GPU decodeCVE-2026-78684 · vLLM (DeepStream GPU decode path, pixel-limit enforcement)Medium
- llama.cpp RPC server: crafted tensor dimensions hit a reachable assertion and abort the processCVE-2026-86317 · llama.cpp RPC server (rpc_server::deserialize_tensor in ggml/src/ggml-rpc/ggml-rpc.cpp)Medium
This entry is curated: imported from vendor advisories with machine assistance, not yet individually verified. Confirm against your vendor's advisory before acting, and report anything wrong.