GPU VulnDB

Database/AI/ML frameworks & serving

vLLM: attacker-chosen P2P offload peers exhaust ZeroMQ contexts and crash EngineCore

CVSS 8.7CVE-2026-94624AI/ML frameworks & servingcurated

Impact

With OffloadingConnector configured for a peer-to-peer secondary tier, a client can put arbitrary remote host and port values in kv_transfer_params. Each unreachable peer keeps a ZeroMQ socket alive until the context quota runs out, at which point an uncaught ZMQError crashes EngineCore and all inference on that instance stops. The same parameter also lets a caller point KV offload traffic at a host of their choosing, which is worth noting for anyone who treats the offload path as internal-only. Recovery needs an operator restart; the node's GPUs are idle in the meantime.

Who can reach it

Any client that can reach the API and set kv_transfer_params, with no authentication per the record. Limited to deployments that enable OffloadingConnector with TieringOffloadingSpec and a P2P secondary tier.

What to do

Take the fix in vllm-project/vllm PR #51504; no fixed release is named (reported through 0.29.0). Deploying it is an image update and an EngineCore restart per replica. Until then, strip client-supplied peer host/port at the gateway or allowlist offload peers, and disable the P2P secondary tier if you are not relying on it.

References

Related entries

All AI/ML frameworks & serving entries

This entry is curated: imported from vendor advisories with machine assistance, not yet individually verified. Confirm against your vendor's advisory before acting, and report anything wrong.