Database/AI/ML frameworks & serving
vLLM: unbounded prompt token ids write out of bounds in the penalty bincount Triton kernel
Impact
The Triton _bincount_kernel used to build the repetition-penalty prompt-presence bitset indexes with prompt token ids without checking them against vocabulary size. A request carrying a token id equal to the vocabulary size - reachable through the multimodal audio path, per the advisory - writes past the end of the bitset in GPU memory. The corrupted region is sampler state shared with the other requests in the batch, so the penalty behaviour of concurrent requests on the same GPU changes. On a multi-tenant inference server that means one tenant's input alters another tenant's decoding. The record describes out-of-bounds writes into adjacent sampler state only; it does not claim a path to code execution or to reading another request's tokens. This is a distinct code path from CVE-2026-93840 with its own fix.
Who can reach it
Anyone who can send a request - specifically a multimodal audio request - to the vLLM server. No authentication beyond the deployment's own API gating; batching with the victim request is the default behaviour.
What to do
No fixed release is named in the record: the advisory says vLLM through 0.29.0 is affected and points to PR #49081. Track that PR, apply it when it lands in a release, and restart each vLLM replica (drain, restart, reload weights). Until then, validate prompt token ids against vocabulary size before they reach vLLM, or disable the audio multimodal path if you do not need it.
References
Related entries
- OpenLLM: Local file inclusion via the web applicationCVE-2024-8982 · OpenLLMMedium
- Weights & Biases OpenUI: Unauthenticated endpoints allow file upload and downloadCVE-2024-10649 · Weights & Biases OpenUIMedium
- Dask distributed (+ Jupyter proxy): Exposure when Dask, JupyterLab and jupyter-server-proxy are combinedCVE-2026-23528 · Dask distributed (+ Jupyter proxy)Medium
- tract: unchecked size multiplication when reading an NNEF tensor gives a heap over-read on model loadCVE-2026-55093 · tract-nnef (read_tensor in nnef/src/tensors.rs)Medium
- tract: ONNX external_data path is not sanitised, so loading a model reads arbitrary local filesCVE-2026-55832 · tract-onnx (external_data path handling, get_external_resources / MmapDataResolver)Medium
- BentoML 1.3.9 (open redirect in the serving UI): A crafted URL against the BentoML server bounces the visitor to anNCVD-2025-017-bentoml-1-3-9-open-redirect-in-t · BentoML 1.3.9 (open redirect in the serving UI)Medium
This entry is curated: imported from vendor advisories with machine assistance, not yet individually verified. Confirm against your vendor's advisory before acting, and report anything wrong.