Database/AI/ML frameworks & serving
vLLM: out-of-range stop_token_ids trigger a CUDA device assertion and wedge EngineCore
Impact
stop_token_ids are checked only for being integers, not for being inside the model vocabulary. With min_tokens > 0 an out-of-range id is used as a logits index and hits a CUDA device-side assertion. One malformed request puts EngineCore into a fatal state: every subsequent request fails until the service is restarted. The GPUs stay allocated to a dead replica, so a single authenticated tenant can park a GPU node's worth of capacity with one HTTP call, repeatedly.
Who can reach it
Any authenticated API user of the inference endpoint - in a multi-tenant serving cluster, any tenant with a key for that model.
What to do
Upgrade to vLLM 0.29.0 and restart each serving process. If you cannot upgrade now, reject or clamp stop_token_ids at a gateway in front of vLLM, and make sure the process supervisor restarts the server on the fatal-state condition so the outage is short rather than open-ended.
References
Related entries
- vLLM (`MediaConnector`): SSRF, recurrence of CVE-2025-6242CVE-2026-24779 · vLLM (`MediaConnector`)High
- vLLM (`load_from_url_async`): Bypass of the CVE-2026-24779 SSRF fixCVE-2026-25960 · vLLM (`load_from_url_async`)High
- picklescan: `scan_pytorch` bypass via forged magic numbersCVE-2026-53875 · picklescanHigh
- torchvision (GIF decoder): Out-of-bounds heap read in `read_from_tensor` GIF decodeCVE-2026-65918 · torchvision (GIF decoder)High
- MLflow: model version creation reads another user's run artifacts without READ permissionCVE-2026-69148 · MLflow tracking server (CreateModelVersion source run/model validation)High
- MLflow AI Gateway: unvalidated api_base in gateway secrets turns the proxy endpoint into an authenticated SSRFCVE-2026-71211 · MLflow AI Gateway (gateway secret auth_config.api_base, raw_proxy endpoint)High
This entry is curated: imported from vendor advisories with machine assistance, not yet individually verified. Confirm against your vendor's advisory before acting, and report anything wrong.