GPU VulnDB

Database/AI/ML frameworks & serving

vLLM: unbounded prompt token ids write out of bounds in the penalty bincount Triton kernel

CVSS 6.3CVE-2026-93841AI/ML frameworks & servingcurated

Impact

The Triton _bincount_kernel used to build the repetition-penalty prompt-presence bitset indexes with prompt token ids without checking them against vocabulary size. A request carrying a token id equal to the vocabulary size - reachable through the multimodal audio path, per the advisory - writes past the end of the bitset in GPU memory. The corrupted region is sampler state shared with the other requests in the batch, so the penalty behaviour of concurrent requests on the same GPU changes. On a multi-tenant inference server that means one tenant's input alters another tenant's decoding. The record describes out-of-bounds writes into adjacent sampler state only; it does not claim a path to code execution or to reading another request's tokens. This is a distinct code path from CVE-2026-93840 with its own fix.

Who can reach it

Anyone who can send a request - specifically a multimodal audio request - to the vLLM server. No authentication beyond the deployment's own API gating; batching with the victim request is the default behaviour.

What to do

No fixed release is named in the record: the advisory says vLLM through 0.29.0 is affected and points to PR #49081. Track that PR, apply it when it lands in a release, and restart each vLLM replica (drain, restart, reload weights). Until then, validate prompt token ids against vocabulary size before they reach vLLM, or disable the audio multimodal path if you do not need it.

References

Related entries

All AI/ML frameworks & serving entries

This entry is curated: imported from vendor advisories with machine assistance, not yet individually verified. Confirm against your vendor's advisory before acting, and report anything wrong.