Database/AI/ML frameworks & serving
vLLM: unvalidated bad_words token indices corrupt logits of other in-flight requests
Impact
vLLM through 0.29.0 does not check that bad_words token indices fall inside the model's generation output width before they reach the GPU sampling kernel, so an out-of-bounds index writes into logits memory belonging to other requests batched alongside it. On a shared inference endpoint that means one caller can make a different caller's completion return wrong tokens - a cross-request integrity failure inside a single continuous-batching engine, not just a crash of the caller's own request. For an operator this is the kind of bug that surfaces as unreproducible bad output on a busy node and is very hard to attribute. Impact is limited to output integrity; the record claims no disclosure of the other request's content and no code execution.
Who can reach it
Any client authenticated to the vLLM OpenAI-compatible API that is allowed to set sampling parameters, in any deployment where requests from different callers share one engine. No privileged access to the node is needed; the attacker only needs to submit a request with a crafted bad_words value.
What to do
Upgrade vLLM past the fix in vllm-project/vllm PR 48824 and restart each serving process - a rolling restart of the model replicas, no node reboot. Until then, the practical mitigation is to reject or strip bad_words at the gateway in front of vLLM, or to stop batching requests from different tenants on one engine.
References
Related entries
- Langflow: authenticated user reaches eval() through component input options and runs code on the hostCVE-2026-101861 · Langflow schema.py (eval() on component input option values)Low
- mistral.rs: out-of-bounds read parsing GGUF token id metadata crashes the inference serverCVE-2026-75090 · mistral.rs GGUF tokenizer (convert_gguf_to_hf_tokenizer)Low
- Ollama: integer overflow in the GGUF v1 string reader when parsing a crafted model fileCVE-2026-86289 · Ollama GGUF decoder (readGGUFV1String in fs/ggml/gguf.go)Low
- vLLM: attacker-supplied chat_template burns server resources on the GPU nodeCVE-2026-90878 · vLLM OpenAI-compatible server (/v1/chat/completions Jinja chat_template rendering)Low
- vLLM: malformed tiktoken vocab file crashes the tokenizer backend, denying service on the GPU nodeCVE-2026-90713 · vLLM (Rust tiktoken vocab file handler, TiktokenTokenizer::new)Low
- LangChain4j agentic: unsafe Jackson default typing in AgenticScope deserialization allows arbitrary class instantiationCVE-2026-97869 · LangChain4j agentic module (AgenticScopeSerializer JSON deserialization)Low
This entry is curated: imported from vendor advisories with machine assistance, not yet individually verified. Confirm against your vendor's advisory before acting, and report anything wrong.