Database/AI/ML frameworks & serving
vLLM: catastrophic regex backtracking in several request paths lets a caller burn API-server CPU
Impact
Several regular expressions reachable from request data - in vllm/lora/utils.py, the phi4mini tool parser, and the OpenAI-compatible chat serving path - backtrack catastrophically on crafted nested or repeated input. A single authenticated caller can pin the Python frontend process at full CPU, which stalls request admission for every other tenant sharing that endpoint. On a GPU node the practical cost is idle accelerators: the model stays resident and the GPUs stay allocated while the frontend is unable to schedule work. Availability only; no disclosure or code execution is claimed.
Who can reach it
Any caller that can submit requests to an affected vLLM HTTP endpoint. The CVSS vector requires low-privilege authentication (PR:L), so a tenant with a valid API key, or anyone at all on a deployment that fronts vLLM without auth.
What to do
Upgrade to vLLM 0.9.0 or later and restart the server processes; versions >= 0.6.3 and < 0.9.0 are affected. Rolling the API/engine processes behind a load balancer avoids draining the node, since nothing below the serving layer changes - expect model reload time per replica. Until then, cap request body size and per-request CPU time at the gateway and disable the phi4mini tool parser if it is not in use.
References
Related entries
- vLLM: unauthenticated /tokenize caller sets video max_frames and fps, exhausting frontend memoryCVE-2026-105758 · vLLM Qwen2VLVideoBackend / Qwen3VLVideoBackend (request-level media_io_kwargs video limits)Medium
- vLLM: GLMGA video backend builds an attacker-sized frame-index list, starving the shared media loaderCVE-2026-105760 · vLLM GLMGA video backend (pre-decode frame-index list from request media_io_kwargs)Medium
- BentoML OpenLLM 0.6.30 (async_run_command in src/openllm/common.py): A model repository directory name flows unescapedCVE-2026-15035 · BentoML OpenLLM 0.6.30 (async_run_command in src/openllm/common.py)Medium
- JupyterHub: unauthenticated logins write unbounded usernames to the log, exhausting storageCVE-2026-54338 · JupyterHub form-based login authenticators (failed-login logging)Medium
- vLLM: malformed JSON to the OpenAI-compatible endpoints returns server paths and versionsCVE-2026-73555 · vLLM OpenAI-compatible API server (validation_exception_handler, sanitize_message)Medium
- vLLM: attacker-supplied structured-output regex pins a CPU core and stalls the engine pathCVE-2026-73556 · vLLM structured outputs, lm-format-enforcer backend (structured_outputs.regex)Medium
This entry is curated: imported from vendor advisories with machine assistance, not yet individually verified. Confirm against your vendor's advisory before acting, and report anything wrong.