GPU VulnDB

Database/AI/ML frameworks & serving

vLLM: out-of-bounds read in Mamba2 mixer reachable from a completions request

CVSS 2.1CVE-2026-105775AI/ML frameworks & servingcurated

Impact

An authenticated client of a vLLM endpoint can craft a completions request that drives an out-of-bounds read in conv_ssm_forward in the Mamba2 mixer layer. The reported effect is availability loss on the serving process, which on a GPU node means the worker holding the model weights goes down and the GPU sits idle until the server is restarted and weights are reloaded - minutes of lost capacity per incident on a large model, and a tenant on a shared endpoint can repeat it. The record does not establish that memory contents are returned to the caller, so treat this as a crash/stability issue rather than a disclosure one. Exploit details are public and the project had not responded at the time of publication.

Who can reach it

Any client that can send a completions request to a vLLM server running a Mamba2-based model. Authentication is required at the privilege level the CVSS vector assumes (PR:L), so this is a tenant or any holder of an API key, not an unauthenticated internet attacker - unless the endpoint is exposed without auth.

What to do

No fixed version is published in this record; the report is an upstream GitHub issue with no vendor response yet. Until a release lands, restrict who can reach the endpoint and consider not serving Mamba2-architecture models on shared endpoints. When a fix ships, the action is a vLLM upgrade and a restart of the serving daemon, which drops in-flight requests and requires a weight reload.

References

Related entries

All AI/ML frameworks & serving entries

This entry is curated: imported from vendor advisories with machine assistance, not yet individually verified. Confirm against your vendor's advisory before acting, and report anything wrong.