GPU VulnDB

Database/AI/ML frameworks & serving

vLLM: unvalidated bad_words token indices corrupt logits of other in-flight requests

CVSS 2.3CVE-2026-93989AI/ML frameworks & servingcurated

Impact

vLLM through 0.29.0 does not check that bad_words token indices fall inside the model's generation output width before they reach the GPU sampling kernel, so an out-of-bounds index writes into logits memory belonging to other requests batched alongside it. On a shared inference endpoint that means one caller can make a different caller's completion return wrong tokens - a cross-request integrity failure inside a single continuous-batching engine, not just a crash of the caller's own request. For an operator this is the kind of bug that surfaces as unreproducible bad output on a busy node and is very hard to attribute. Impact is limited to output integrity; the record claims no disclosure of the other request's content and no code execution.

Who can reach it

Any client authenticated to the vLLM OpenAI-compatible API that is allowed to set sampling parameters, in any deployment where requests from different callers share one engine. No privileged access to the node is needed; the attacker only needs to submit a request with a crafted bad_words value.

What to do

Upgrade vLLM past the fix in vllm-project/vllm PR 48824 and restart each serving process - a rolling restart of the model replicas, no node reboot. Until then, the practical mitigation is to reject or strip bad_words at the gateway in front of vLLM, or to stop batching requests from different tenants on one engine.

References

Related entries

All AI/ML frameworks & serving entries

This entry is curated: imported from vendor advisories with machine assistance, not yet individually verified. Confirm against your vendor's advisory before acting, and report anything wrong.