Database/AI/ML frameworks & serving
vLLM: arbitrary HTTP method tokens create unbounded Prometheus label sets and exhaust the service
Impact
The Rust frontend's track_http_metrics middleware records the raw HTTP method token as a Prometheus label on requests that reach registered routes. Sending unique arbitrary method tokens to an unguarded route such as /tokenize makes Prometheus Family::get_or_create permanently allocate a new counter and histogram label set each time. Memory grows without bound and the /metrics response balloons, degrading or killing the inference server and the scrape path that monitors it. On a GPU fleet the second effect is the nastier one: a bloated /metrics can take down the shared Prometheus that the autoscaler and on-call dashboards depend on, so the blast radius extends past the one replica. No confidentiality or integrity impact.
Who can reach it
Unauthenticated network attacker who can reach any registered route on the vLLM HTTP frontend, including routes not behind auth such as /tokenize. CVSS rates attack complexity high (AV:N/AC:H/PR:N).
What to do
Upgrade vLLM to 0.30.0 and restart the serving replicas. Until then, keep the vLLM frontend off untrusted networks and put a proxy in front that rejects non-standard HTTP methods; also cap scrape response size or series per target on the Prometheus side so a bloated endpoint cannot take monitoring down with it.
References
Related entries
- Ray (dashboard DELETE endpoints): Browser-origin protection covers POST/PUT but not DELETECVE-2026-27482 · Ray (dashboard DELETE endpoints)Medium
- LocalAI (`/models/apply`): SSRF and partial local file inclusionCVE-2024-6095 · LocalAI (`/models/apply`)Medium
- NVIDIA NemoClaw: insufficiently protected credentials allow information disclosure and data tamperingCVE-2026-65087 · NVIDIA NemoClaw (credential storage)Medium
- Linux perf/x86/amd/uncore - memory leak in the events array: Per-CPU northbridge and last-level-cache uncore contextsCVE-2022-49784 · Linux perf/x86/amd/uncore - memory leak in the events arrayMedium
- PyTorch (flatbuffer loader): Out-of-bounds read parsing flatbuffer modelCVE-2024-31584 · PyTorch (flatbuffer loader)Medium
- vLLM: crafted request to the Gemma4 unified parser crashes the inference serverCVE-2026-103241 · vLLM (Gemma4UnifiedParser, rust/src/parser/src/unified/gemma4.rs)Medium
This entry is curated: imported from vendor advisories with machine assistance, not yet individually verified. Confirm against your vendor's advisory before acting, and report anything wrong.