Database/AI/ML frameworks & serving
vLLM: allowed_token_ids validated against tokenizer length, corrupting shared GPU logit-bias state
Impact
vLLM validated allowed_token_ids against the tokenizer vocabulary size rather than the width of the model's output logits, so a caller can pass token ids that sit past the end of the output vocabulary and still clear validation. Those ids land in LogitBiasState, which is shared across the requests batched together on the GPU, and the resulting corruption lets other in-flight requests sample tokens outside their own allowlists. On a shared inference endpoint this is a cross-request integrity problem: one tenant's request payload silently changes the constrained decoding of another tenant's request on the same GPU. It does not leak the other request's content, and the record shows no memory-safety or code-execution consequence.
Who can reach it
Anyone who can submit a completion request with allowed_token_ids to the vLLM server. No authentication is needed beyond whatever the deployment puts in front of the API; the attacker only needs their request to be batched with the victim's, which is the normal case on a busy server.
What to do
Upgrade to vLLM 0.29.0 or later (fix in commit 5b0e5b69, PR #49080) and restart the server process. That is a daemon restart per replica - model weights must be reloaded, so drain each replica before cycling it, but no node reboot or driver change is involved. If you cannot upgrade immediately, reject or clamp allowed_token_ids at the gateway against the model's real output vocabulary size.
References
Related entries
- vLLM: unbounded prompt token ids write out of bounds in the penalty bincount Triton kernelCVE-2026-93841 · vLLM (Triton _bincount_kernel, penalty prompt-presence bitset)Medium
- OpenLLM: Local file inclusion via the web applicationCVE-2024-8982 · OpenLLMMedium
- Weights & Biases OpenUI: Unauthenticated endpoints allow file upload and downloadCVE-2024-10649 · Weights & Biases OpenUIMedium
- Dask distributed (+ Jupyter proxy): Exposure when Dask, JupyterLab and jupyter-server-proxy are combinedCVE-2026-23528 · Dask distributed (+ Jupyter proxy)Medium
- tract: unchecked size multiplication when reading an NNEF tensor gives a heap over-read on model loadCVE-2026-55093 · tract-nnef (read_tensor in nnef/src/tensors.rs)Medium
- tract: ONNX external_data path is not sanitised, so loading a model reads arbitrary local filesCVE-2026-55832 · tract-onnx (external_data path handling, get_external_resources / MmapDataResolver)Medium
This entry is curated: imported from vendor advisories with machine assistance, not yet individually verified. Confirm against your vendor's advisory before acting, and report anything wrong.