Database/AI/ML frameworks & serving
vLLM: engine crash via the penalty handler when prompt embeds are combined with penalties
Impact
A client request that combines prompt embeddings with penalty sampling parameters crashes the vLLM engine in get_token_bin_counts_and_mask. The published reproducer kills the engine process, which on a GPU node takes down the whole served model, not just the offending request - every other tenant on that endpoint loses service until the daemon restarts and reloads weights onto the GPUs. Because the trigger is ordinary API parameters rather than anything privileged, one tenant can repeatedly deny the node to the rest. A public reproducer script exists and the project had not responded at the time of publication.
Who can reach it
Any authenticated client that can submit a completions request with prompt embeds and penalty parameters. No special privilege beyond normal API access to the endpoint.
What to do
No fixed version is published in this record; the report is an upstream GitHub issue with no vendor response yet. Mitigation available today is input validation at the gateway in front of vLLM - reject or sanitise penalty parameters on requests carrying prompt embeds - plus a supervisor that restarts the engine. When an upstream fix lands, the action is an upgrade and a serving-daemon restart.
References
Related entries
- mistral.rs: out-of-bounds read parsing GGUF token id metadata crashes the inference serverCVE-2026-75090 · mistral.rs GGUF tokenizer (convert_gguf_to_hf_tokenizer)Low
- Ollama: integer overflow in the GGUF v1 string reader when parsing a crafted model fileCVE-2026-86289 · Ollama GGUF decoder (readGGUFV1String in fs/ggml/gguf.go)Low
- vLLM: attacker-supplied chat_template burns server resources on the GPU nodeCVE-2026-90878 · vLLM OpenAI-compatible server (/v1/chat/completions Jinja chat_template rendering)Low
- vLLM: malformed tiktoken vocab file crashes the tokenizer backend, denying service on the GPU nodeCVE-2026-90713 · vLLM (Rust tiktoken vocab file handler, TiktokenTokenizer::new)Low
- LangChain4j agentic: unsafe Jackson default typing in AgenticScope deserialization allows arbitrary class instantiationCVE-2026-97869 · LangChain4j agentic module (AgenticScopeSerializer JSON deserialization)Low
- Langchain-Chatchat: arbitrary file write outside the upload and knowledge-base directoriesCVE-2026-51882 · Langchain-Chatchat (file upload and knowledge-base endpoints, path traversal)Unscored
This entry is curated: imported from vendor advisories with machine assistance, not yet individually verified. Confirm against your vendor's advisory before acting, and report anything wrong.