Database/AI/ML frameworks & serving
vLLM: out-of-bounds read in Mamba2 mixer reachable from a completions request
Impact
An authenticated client of a vLLM endpoint can craft a completions request that drives an out-of-bounds read in conv_ssm_forward in the Mamba2 mixer layer. The reported effect is availability loss on the serving process, which on a GPU node means the worker holding the model weights goes down and the GPU sits idle until the server is restarted and weights are reloaded - minutes of lost capacity per incident on a large model, and a tenant on a shared endpoint can repeat it. The record does not establish that memory contents are returned to the caller, so treat this as a crash/stability issue rather than a disclosure one. Exploit details are public and the project had not responded at the time of publication.
Who can reach it
Any client that can send a completions request to a vLLM server running a Mamba2-based model. Authentication is required at the privilege level the CVSS vector assumes (PR:L), so this is a tenant or any holder of an API key, not an unauthenticated internet attacker - unless the endpoint is exposed without auth.
What to do
No fixed version is published in this record; the report is an upstream GitHub issue with no vendor response yet. Until a release lands, restrict who can reach the endpoint and consider not serving Mamba2-architecture models on shared endpoints. When a fix ships, the action is a vLLM upgrade and a restart of the serving daemon, which drops in-flight requests and requires a weight reload.
References
Related entries
- vLLM: engine crash via the penalty handler when prompt embeds are combined with penaltiesCVE-2026-105922 · vLLM (penalty handler, get_token_bin_counts_and_mask in model_executor/layers/utils.py)Low
- mistral.rs: out-of-bounds read parsing GGUF token id metadata crashes the inference serverCVE-2026-75090 · mistral.rs GGUF tokenizer (convert_gguf_to_hf_tokenizer)Low
- Ollama: integer overflow in the GGUF v1 string reader when parsing a crafted model fileCVE-2026-86289 · Ollama GGUF decoder (readGGUFV1String in fs/ggml/gguf.go)Low
- vLLM: attacker-supplied chat_template burns server resources on the GPU nodeCVE-2026-90878 · vLLM OpenAI-compatible server (/v1/chat/completions Jinja chat_template rendering)Low
- vLLM: malformed tiktoken vocab file crashes the tokenizer backend, denying service on the GPU nodeCVE-2026-90713 · vLLM (Rust tiktoken vocab file handler, TiktokenTokenizer::new)Low
- LangChain4j agentic: unsafe Jackson default typing in AgenticScope deserialization allows arbitrary class instantiationCVE-2026-97869 · LangChain4j agentic module (AgenticScopeSerializer JSON deserialization)Low
This entry is curated: imported from vendor advisories with machine assistance, not yet individually verified. Confirm against your vendor's advisory before acting, and report anything wrong.