GPU VulnDB

Database/AI/ML frameworks & serving

vLLM: out-of-range stop_token_ids trigger a CUDA device assertion and wedge EngineCore

CVSS 7.1CVE-2026-100654AI/ML frameworks & servingcurated

Impact

stop_token_ids are checked only for being integers, not for being inside the model vocabulary. With min_tokens > 0 an out-of-range id is used as a logits index and hits a CUDA device-side assertion. One malformed request puts EngineCore into a fatal state: every subsequent request fails until the service is restarted. The GPUs stay allocated to a dead replica, so a single authenticated tenant can park a GPU node's worth of capacity with one HTTP call, repeatedly.

Who can reach it

Any authenticated API user of the inference endpoint - in a multi-tenant serving cluster, any tenant with a key for that model.

What to do

Upgrade to vLLM 0.29.0 and restart each serving process. If you cannot upgrade now, reject or clamp stop_token_ids at a gateway in front of vLLM, and make sure the process supervisor restarts the server on the fatal-state condition so the outage is short rather than open-ended.

References

Related entries

All AI/ML frameworks & serving entries

This entry is curated: imported from vendor advisories with machine assistance, not yet individually verified. Confirm against your vendor's advisory before acting, and report anything wrong.