Database/AI/ML frameworks & serving
Mooncake transfer engine: zero-length handshake frame crashes the hosting inference process
Impact
Mooncake is the KV-cache transfer engine behind disaggregated serving setups such as SGLang. Its handshake listener binds on all interfaces, and an out-of-bounds read in readString() in include/common.h lets an eight-byte zero-length frame terminate the entire hosting process - not just the transfer thread. On a GPU node that means the inference server dies and its loaded weights and KV cache go with it, so recovery costs a full model reload across the GPUs it occupied. Because the port is reachable from anywhere that can route to the node, any tenant or workload sharing the cluster network can knock serving replicas over repeatedly. Availability impact only; the record describes no read of tenant data or code execution.
Who can reach it
Anyone with network reach to the Mooncake handshake port, which listens on all interfaces. No authentication is required.
What to do
Upgrade Mooncake to 0.3.12 or later (fix in commit c142b405) and restart the inference servers that embed it - a rolling restart of serving replicas, no node drain or reboot. Until then, bind or firewall the handshake port so it is reachable only from the peer nodes that need it.
References
Related entries
- KubeAI (Ollama engine controller): Injection in `ollamaStartupProbeScript()`CVE-2026-34940 · KubeAI (Ollama engine controller)High
- Xinference: model launch API executes attacker-supplied Python because trust_remote_code is always onCVE-2026-76841 · Xinference (Xorbits Inference) model loaders - trust_remote_codeHigh
- Ollama: model pull follows cross-host redirects, giving SSRF to internal and metadata endpointsCVE-2026-85180 · Ollama (tensor-layer blob download, cross-host redirect handling)High
- Axolotl: multipack patch loads Hugging Face base models with trust_remote_code=True, giving RCE on the training nodeCVE-2026-86169 · Axolotl (multipack patch path, trust_remote_code guard)High
- vLLM: rejected requests leak decode-worker metadata until the worker exhausts memoryCVE-2026-93436 · vLLM NIXL KV connector (decode-side metadata cleanup for rejected requests)High
- vLLM: negative token IDs in embeddings requests poison the CUDA context and wedge the engineCVE-2026-93592 · vLLM (/v1/embeddings and /pooling token ID validation)High
This entry is curated: imported from vendor advisories with machine assistance, not yet individually verified. Confirm against your vendor's advisory before acting, and report anything wrong.