GPU VulnDB

Database/AI/ML frameworks & serving

vLLM: allowed_token_ids validated against tokenizer length, corrupting shared GPU logit-bias state

CVSS 6.3CVE-2026-93840AI/ML frameworks & servingcurated

Impact

vLLM validated allowed_token_ids against the tokenizer vocabulary size rather than the width of the model's output logits, so a caller can pass token ids that sit past the end of the output vocabulary and still clear validation. Those ids land in LogitBiasState, which is shared across the requests batched together on the GPU, and the resulting corruption lets other in-flight requests sample tokens outside their own allowlists. On a shared inference endpoint this is a cross-request integrity problem: one tenant's request payload silently changes the constrained decoding of another tenant's request on the same GPU. It does not leak the other request's content, and the record shows no memory-safety or code-execution consequence.

Who can reach it

Anyone who can submit a completion request with allowed_token_ids to the vLLM server. No authentication is needed beyond whatever the deployment puts in front of the API; the attacker only needs their request to be batched with the victim's, which is the normal case on a busy server.

What to do

Upgrade to vLLM 0.29.0 or later (fix in commit 5b0e5b69, PR #49080) and restart the server process. That is a daemon restart per replica - model weights must be reloaded, so drain each replica before cycling it, but no node reboot or driver change is involved. If you cannot upgrade immediately, reject or clamp allowed_token_ids at the gateway against the model's real output vocabulary size.

References

Related entries

All AI/ML frameworks & serving entries

This entry is curated: imported from vendor advisories with machine assistance, not yet individually verified. Confirm against your vendor's advisory before acting, and report anything wrong.