Database/AI/ML frameworks & serving
vLLM: unbounded cache_salt stalls the EngineCore scheduler thread, denying service to all requests
Impact
An unauthenticated caller can post a multi-hundred-megabyte cache_salt value on the OpenAI- or Anthropic-compatible endpoints. The value is pickled and SHA-256 hashed on the single EngineCore scheduler thread, so one request stalls scheduling for every other request on that server. On a GPU node this takes the whole model replica offline, not one session, and the GPUs sit idle while the scheduler churns CPU. No code execution or data disclosure.
Who can reach it
Anyone who can reach the inference HTTP endpoint. No authentication needed on this path, so any tenant or, where the endpoint is exposed beyond the cluster, any network caller.
What to do
Upgrade to vLLM 0.29.0 and restart each serving process; model weights reload, so expect a cold start per replica. Until then, cap request body size at the ingress or gateway in front of vLLM and keep the endpoint off untrusted networks.
References
Related entries
- vLLM: audio clip size limit not enforced, letting unauthenticated clients exhaust node memory and CPUCVE-2026-100648 · vLLM (multimodal chat audio decoding, VLLM_MAX_AUDIO_CLIP_FILESIZE_MB)Medium
- Hugging Face Accelerate: unsanitized shard paths in a checkpoint index give arbitrary file read and hangsCVE-2026-69112 · Hugging Face Accelerate (load_checkpoint_in_model / load_checkpoint_and_dispatch weight_map handling)Medium
- llama.cpp ggml RPC server: unvalidated tensor op and op_params in deserialize_tensorCVE-2026-78147 · llama.cpp ggml RPC server (deserialize_tensor op / op_params validation)Medium
- llama.cpp ggml RPC server: null pointer dereference in graph_compute kills the GPU workerCVE-2026-78148 · llama.cpp ggml RPC server (rpc_server::graph_compute)Medium
- BentoML: SSRF filter misses 100.64.0.0/10, so serving pods fetch from internal CGNAT hostsCVE-2026-78205 · BentoML make_safe_connect (SSRF address filter, RFC 6598 range)Medium
- vLLM: DeepStream backend misclassification skips pixel limits and lets unauthenticated video exhaust GPU decodeCVE-2026-78684 · vLLM (DeepStream GPU decode path, pixel-limit enforcement)Medium
This entry is curated: imported from vendor advisories with machine assistance, not yet individually verified. Confirm against your vendor's advisory before acting, and report anything wrong.