Database/AI/ML frameworks & serving
vLLM: no upper bound on the n parameter lets a single request OOM the API server
Impact
The Pydantic models for chat and text completion accept an arbitrarily large n, and vLLM materialises that many request-object copies on the heap before the request ever reaches the scheduler. One HTTP request blocks the asyncio event loop and then crashes the process out of memory. The whole endpoint stops serving, not just the offending request, and recovery means reloading model weights onto the GPU — so a single packet costs the fleet minutes of idle accelerator. The advisory describes the attacker as unauthenticated; the CVSS vector assigned to the record says PR:L, so treat any client able to reach the API as capable of triggering it.
Who can reach it
Any client that can reach the vLLM HTTP API and submit a completion request. The vendor description states no authentication is required; the scored vector assumes a low-privilege credential. Either way this is reachable by any tenant sharing the inference endpoint.
What to do
Upgrade to vLLM 0.19.0 or later and restart the serving processes; roll replica by replica to avoid a full outage, accepting a weight reload per worker. Red Hat errata cover AI Inference Server, RHEL AI 3 and OpenShift AI. Until patched, reject or clamp the n field at the gateway in front of vLLM. No node drain or reboot.
References
Related entries
- vLLM (revision pinning): Revision pinning does not apply to all model artifactsCVE-2026-47155 · vLLM (revision pinning)Medium
- Starlette: malformed Host header makes request.url.path diverge from the routed pathCVE-2026-48710 · Starlette (Host header validation when reconstructing request.url)Medium
- vLLM - sampling parameter validation: Temperature validation uses strict comparison operators, so boundary values slipCVE-2026-54235 · vLLM - sampling parameter validationMedium
- vLLM: audio input in chat completions skips the decode-duration limit, letting a small clip OOM the workerCVE-2026-57173 · vLLM (input_audio path in /v1/chat/completions, AudioMediaIO decode duration guard)Medium
- MLflow: missing permission check on runs/log-inputs lets any user forge another run's lineageCVE-2026-69146 · MLflow tracking server auth plugin (mlflow/server/auth, runs/log-inputs endpoint)Medium
- vLLM: request-selected video decoder backend allocates GPU memory outside the KV-cache budgetCVE-2026-69147 · vLLM (MediaConnector video backend selection / GPU memory budgeting)Medium
This entry is curated: imported from vendor advisories with machine assistance, not yet individually verified. Confirm against your vendor's advisory before acting, and report anything wrong.