Database/AI/ML frameworks & serving
LightLLM: unbounded key-value writes on the NCCL control channel exhaust worker memory
Impact
When LightLLM runs with --pd_trans_mode nccl, the KV-transfer worker exposes an RPC method exposed_set_value that stores caller-supplied key-value pairs with no size or count limit. An unauthenticated caller can grow that store until the worker process dies of memory exhaustion, which takes the node's serving role down with it. On a GPU node this is expensive: the process holds model weights and KV cache, so a restart means reloading weights and losing warm cache, and in a multi-node disaggregated setup the peer waiting on that transfer fails too.
Who can reach it
Anyone with network reach to the LightLLM KV-transfer control channel. No authentication is described on that channel, so exposure depends entirely on whether the port is confined to the internal fabric.
What to do
No fixed version is published in the record; the issue is tracked as ModelTC/LightLLM#1595 and reported through 1.2.0. Until a fix lands, bind the KV-transfer control channel to the internal fabric only, never to a tenant-reachable network, and cap worker memory so the OOM is contained. Recovery is restarting the affected worker process.
References
Related entries
- KubeAI (Ollama engine controller): Injection in `ollamaStartupProbeScript()`CVE-2026-34940 · KubeAI (Ollama engine controller)High
- Xinference: model launch API executes attacker-supplied Python because trust_remote_code is always onCVE-2026-76841 · Xinference (Xorbits Inference) model loaders - trust_remote_codeHigh
- Ollama: model pull follows cross-host redirects, giving SSRF to internal and metadata endpointsCVE-2026-85180 · Ollama (tensor-layer blob download, cross-host redirect handling)High
- Axolotl: multipack patch loads Hugging Face base models with trust_remote_code=True, giving RCE on the training nodeCVE-2026-86169 · Axolotl (multipack patch path, trust_remote_code guard)High
- vLLM: rejected requests leak decode-worker metadata until the worker exhausts memoryCVE-2026-93436 · vLLM NIXL KV connector (decode-side metadata cleanup for rejected requests)High
- vLLM: negative token IDs in embeddings requests poison the CUDA context and wedge the engineCVE-2026-93592 · vLLM (/v1/embeddings and /pooling token ID validation)High
This entry is curated: imported from vendor advisories with machine assistance, not yet individually verified. Confirm against your vendor's advisory before acting, and report anything wrong.