GPU VulnDB

Database/AI/ML frameworks & serving

vLLM: out-of-vocabulary stop_token_ids kill EngineCore and take the model server down until restart

CVSS 8.2CVE-2026-100652AI/ML frameworks & servingcurated

Impact

The Rust HTTP and gRPC frontends in 0.22.0 through 0.23.0 do not bound-check stop_token_ids against the vocabulary before they reach MinTokensLogitsProcessor. A request with min_tokens greater than zero and an out-of-vocabulary stop token id triggers a CUDA tensor indexing failure that leaves EngineCore in a fatal state - the server does not recover on its own and has to be restarted. Any client that can send an inference request can do this, so a single malformed request drops a model replica that may hold tens of gigabytes of weights across several GPUs. Recovery means a cold start including model load, which on large models is minutes of lost capacity per hit, and the request can simply be repeated. No data disclosure or code execution is claimed.

Who can reach it

Anyone who can send an inference request to the vLLM endpoint. No authentication is required unless the operator has put one in front of it.

What to do

Upgrade vLLM past 0.23.0 to the version named in the project advisory and restart the serving processes - a per-replica daemon restart with model reload, which can be done as a rolling replacement behind the router. If you cannot upgrade now, reject or clamp stop_token_ids at the gateway so ids outside the model's vocabulary never reach vLLM, and make sure the supervisor restarts a wedged EngineCore automatically.

References

Related entries

All AI/ML frameworks & serving entries

This entry is curated: imported from vendor advisories with machine assistance, not yet individually verified. Confirm against your vendor's advisory before acting, and report anything wrong.