GPU VulnDB

Database/AI/ML frameworks & serving

vLLM: Mooncake transfer-ID collision leaks GPU KV cache blocks until restart

CVSS 8.7CVE-2026-94627AI/ML frameworks & servingcurated

Impact

When concurrent child requests of one multi-prompt completion share a single Mooncake transfer ID, KV cache block ownership is mishandled and blocks are orphaned instead of freed. The leak is in GPU memory, so it accumulates request after request until the cache cannot serve legitimate work; only a process restart reclaims it. On shared inference capacity this is a slow, self-inflicted-looking degradation - latency and admission failures rise long before anything crashes - which makes it easy to misdiagnose as capacity pressure rather than an attack or a bug.

Who can reach it

Any client that can submit multi-prompt completion requests to a disaggregated vLLM deployment using the Mooncake connector. No authentication required per the record.

What to do

Apply vllm-project/vllm PR #49796; no fixed release is named (reported through 0.29.0). Deployment is an image update plus a restart of the affected serving processes, which is also the only way to reclaim already-leaked blocks. Monitor GPU KV cache utilisation as a leak indicator in the meantime.

References

Related entries

All AI/ML frameworks & serving entries

This entry is curated: imported from vendor advisories with machine assistance, not yet individually verified. Confirm against your vendor's advisory before acting, and report anything wrong.