GPU VulnDB

Database/AI/ML frameworks & serving

vLLM: unbounded cache_salt stalls the EngineCore scheduler thread, denying service to all requests

CVSS 6.9CVE-2026-100647AI/ML frameworks & servingcurated

Impact

An unauthenticated caller can post a multi-hundred-megabyte cache_salt value on the OpenAI- or Anthropic-compatible endpoints. The value is pickled and SHA-256 hashed on the single EngineCore scheduler thread, so one request stalls scheduling for every other request on that server. On a GPU node this takes the whole model replica offline, not one session, and the GPUs sit idle while the scheduler churns CPU. No code execution or data disclosure.

Who can reach it

Anyone who can reach the inference HTTP endpoint. No authentication needed on this path, so any tenant or, where the endpoint is exposed beyond the cluster, any network caller.

What to do

Upgrade to vLLM 0.29.0 and restart each serving process; model weights reload, so expect a cold start per replica. Until then, cap request body size at the ingress or gateway in front of vLLM and keep the endpoint off untrusted networks.

References

Related entries

All AI/ML frameworks & serving entries

This entry is curated: imported from vendor advisories with machine assistance, not yet individually verified. Confirm against your vendor's advisory before acting, and report anything wrong.