GPU VulnDB

Database/AI/ML frameworks & serving

vLLM: catastrophic regex backtracking in several request paths lets a caller burn API-server CPU

CVSS 5.3CVE-2025-71379AI/ML frameworks & servingcurated

Impact

Several regular expressions reachable from request data - in vllm/lora/utils.py, the phi4mini tool parser, and the OpenAI-compatible chat serving path - backtrack catastrophically on crafted nested or repeated input. A single authenticated caller can pin the Python frontend process at full CPU, which stalls request admission for every other tenant sharing that endpoint. On a GPU node the practical cost is idle accelerators: the model stays resident and the GPUs stay allocated while the frontend is unable to schedule work. Availability only; no disclosure or code execution is claimed.

Who can reach it

Any caller that can submit requests to an affected vLLM HTTP endpoint. The CVSS vector requires low-privilege authentication (PR:L), so a tenant with a valid API key, or anyone at all on a deployment that fronts vLLM without auth.

What to do

Upgrade to vLLM 0.9.0 or later and restart the server processes; versions >= 0.6.3 and < 0.9.0 are affected. Rolling the API/engine processes behind a load balancer avoids draining the node, since nothing below the serving layer changes - expect model reload time per replica. Until then, cap request body size and per-request CPU time at the gateway and disable the phi4mini tool parser if it is not in use.

References

Related entries

All AI/ML frameworks & serving entries

This entry is curated: imported from vendor advisories with machine assistance, not yet individually verified. Confirm against your vendor's advisory before acting, and report anything wrong.