GPU VulnDB

Database/AI/ML frameworks & serving

Mooncake transfer engine: zero-length handshake frame crashes the hosting inference process

CVSS 8.7CVE-2026-104433AI/ML frameworks & servingcurated

Impact

Mooncake is the KV-cache transfer engine behind disaggregated serving setups such as SGLang. Its handshake listener binds on all interfaces, and an out-of-bounds read in readString() in include/common.h lets an eight-byte zero-length frame terminate the entire hosting process - not just the transfer thread. On a GPU node that means the inference server dies and its loaded weights and KV cache go with it, so recovery costs a full model reload across the GPUs it occupied. Because the port is reachable from anywhere that can route to the node, any tenant or workload sharing the cluster network can knock serving replicas over repeatedly. Availability impact only; the record describes no read of tenant data or code execution.

Who can reach it

Anyone with network reach to the Mooncake handshake port, which listens on all interfaces. No authentication is required.

What to do

Upgrade Mooncake to 0.3.12 or later (fix in commit c142b405) and restart the inference servers that embed it - a rolling restart of serving replicas, no node drain or reboot. Until then, bind or firewall the handshake port so it is reachable only from the peer nodes that need it.

References

Related entries

All AI/ML frameworks & serving entries

This entry is curated: imported from vendor advisories with machine assistance, not yet individually verified. Confirm against your vendor's advisory before acting, and report anything wrong.