Database/AI/ML frameworks & serving
vLLM: negative token IDs in embeddings requests poison the CUDA context and wedge the engine
Impact
A single unauthenticated request carrying a negative token ID to /v1/embeddings or /pooling trips a CUDA device-side assertion. Once that assertion fires the GPU context is poisoned, so every subsequent request on that engine fails until the process is restarted - one malformed request takes the whole model replica out of service, not just its own request. On a GPU fleet this converts a trivially cheap request into a persistent denial of service against an expensive accelerator, and it can be repeated immediately after each restart. Where the inference endpoint is reachable by tenants or by an internal gateway without auth, any caller can keep a GPU idle indefinitely.
Who can reach it
Anyone who can send HTTP requests to the vLLM API server. No authentication is required - vLLM's OpenAI-compatible server does not authenticate by default, so exposure equals whatever network reach the endpoint has (tenant pods, an internal gateway, or the public internet if fronted directly).
What to do
Upgrade vLLM to 0.28.0 or later and restart each serving process; no node reboot or driver change is needed, but every model replica has to be cycled, so drain traffic per replica to avoid a serving gap. Until then, put authentication and an input-validation proxy in front of the embeddings and pooling endpoints, or disable those endpoints if unused, and make sure the supervisor restarts a wedged engine automatically.
References
Related entries
- Gradio (`/queue/join`): SSRFCVE-2024-4325 · Gradio (`/queue/join`)High
- ONNX: Security-control bypass through 1.20.1CVE-2026-28500 · ONNXHigh
- ONNX (`ExternalDataInfo`): Security control bypass in external-data path handlingCVE-2026-34445 · ONNX (`ExternalDataInfo`)High
- JupyterLab: saved HTML cell output can run arbitrary JupyterLab commands on one user clickCVE-2026-42557 · JupyterLab (HTML sanitizer / CommandLinker command dispatch)High
- LocalAI (`/models/apply`): Unauthenticated SSRF fetching arbitrary internal URLsCVE-2026-59707 · LocalAI (`/models/apply`)High
- Text Generation Inference (TGI): SSRF in the OpenAI-compatible multimodal chat endpointCVE-2026-63086 · Text Generation Inference (TGI)High
This entry is curated: imported from vendor advisories with machine assistance, not yet individually verified. Confirm against your vendor's advisory before acting, and report anything wrong.