GPU VulnDB

Database/AI/ML frameworks & serving

vLLM: unvalidated tp_size in kv_transfer_params drives the decode worker into OOM-kill

CVSS 8.7CVE-2026-94626AI/ML frameworks & servingcurated

Impact

The OpenAI-compatible completion endpoints accept a tp_size value inside kv_transfer_params without checking it against the deployment's actual tensor-parallel size. An arbitrary value makes vLLM allocate unbounded host memory until the kernel OOM-killer takes the decode worker process. On a GPU node this evicts a process holding model weights, and on a node packed with several serving replicas the OOM killer's choice of victim may not even be the attacked process, so the blast radius can exceed the targeted tenant.

Who can reach it

Any client that can call the completion endpoints and include kv_transfer_params; no authentication required per the record. Applies to prefill/decode disaggregated deployments.

What to do

Apply vllm-project/vllm PR #51137; the record names no fixed release (reported through 0.29.0). Rollout is an image update and a restart of the serving replicas, drained one at a time. Interim: have the gateway reject or overwrite tp_size in client-supplied kv_transfer_params, and set per-container memory limits so an OOM stays inside one replica.

References

Related entries

All AI/ML frameworks & serving entries

This entry is curated: imported from vendor advisories with machine assistance, not yet individually verified. Confirm against your vendor's advisory before acting, and report anything wrong.