GPU VulnDB

Database/AI/ML frameworks & serving

vLLM: rejected prefill requests leak Mooncake transfer placeholders, stalling valid requests

CVSS 6.9CVE-2026-94625AI/ML frameworks & servingcurated

Impact

In vLLM through 0.29.0 a rejected prefill request in the MooncakeConnector leaves an ownerless transfer placeholder that is never reclaimed, so the sender task pool fills up. An attacker only needs to keep sending requests that get rejected to exhaust the pool, after which legitimate requests are delayed by up to 480 seconds. The operational sting is that health checks keep returning success throughout, so a disaggregated-prefill serving tier degrades to unusable latency while orchestration sees a healthy pod and neither restarts it nor drains traffic. Availability and latency only - no disclosure or code execution.

Who can reach it

Anyone who can send inference requests to a vLLM deployment using the Mooncake KV-transfer connector. No authentication required in the default configuration. Deployments not using MooncakeConnector are unaffected.

What to do

Pick up the fix from vllm-project/vllm PR 51236 (not present in 0.29.0) and restart the serving processes; the record names no released fixed version, so track the upstream release. Until then, put authentication and rate limiting in front of the endpoint and monitor sender pool occupancy rather than relying on the HTTP health check, which does not reflect this state.

References

Related entries

All AI/ML frameworks & serving entries

This entry is curated: imported from vendor advisories with machine assistance, not yet individually verified. Confirm against your vendor's advisory before acting, and report anything wrong.