GPU VulnDB

Database/AI/ML frameworks & serving

vLLM: arbitrary HTTP method tokens create unbounded Prometheus label sets and exhaust the service

CVSS 5.9CVE-2026-105759AI/ML frameworks & servingcurated

Impact

The Rust frontend's track_http_metrics middleware records the raw HTTP method token as a Prometheus label on requests that reach registered routes. Sending unique arbitrary method tokens to an unguarded route such as /tokenize makes Prometheus Family::get_or_create permanently allocate a new counter and histogram label set each time. Memory grows without bound and the /metrics response balloons, degrading or killing the inference server and the scrape path that monitors it. On a GPU fleet the second effect is the nastier one: a bloated /metrics can take down the shared Prometheus that the autoscaler and on-call dashboards depend on, so the blast radius extends past the one replica. No confidentiality or integrity impact.

Who can reach it

Unauthenticated network attacker who can reach any registered route on the vLLM HTTP frontend, including routes not behind auth such as /tokenize. CVSS rates attack complexity high (AV:N/AC:H/PR:N).

What to do

Upgrade vLLM to 0.30.0 and restart the serving replicas. Until then, keep the vLLM frontend off untrusted networks and put a proxy in front that rejects non-standard HTTP methods; also cap scrape response size or series per target on the Prometheus side so a bloated endpoint cannot take monitoring down with it.

References

Related entries

All AI/ML frameworks & serving entries

This entry is curated: imported from vendor advisories with machine assistance, not yet individually verified. Confirm against your vendor's advisory before acting, and report anything wrong.