GPU VulnDB

Database/AI/ML frameworks & serving

vLLM: attacker-chosen X-Request-Id overwrites another tenant's cached query embedding in /score

CVSS 4.2CVE-2026-105755AI/ML frameworks & servingcurated

Impact

Flash late-interaction scoring derives each worker's query_key from the caller-supplied X-Request-Id header. A concurrent request that reuses a victim's identifier overwrites the cached query embedding, so the victim's documents get scored against the attacker's query and come back with results the victim did not ask for. Shared use counters can also produce a late-interaction cache-miss error. On a shared reranking endpoint this is silent cross-tenant corruption of results rather than a crash, which is harder to notice than an outage: downstream retrieval pipelines act on scores that belong to someone else's query.

Who can reach it

An authenticated tenant able to send requests to /score or /rerank and set the X-Request-Id header, who must land a request concurrently with the victim's (CVSS marks attack complexity high).

What to do

Upgrade to vLLM 0.30.0 and restart the serving processes; all prior versions are affected. Frontend change only - roll replicas, no node drain. As an interim control, have the gateway overwrite rather than pass through client-supplied X-Request-Id.

References

Related entries

All AI/ML frameworks & serving entries

This entry is curated: imported from vendor advisories with machine assistance, not yet individually verified. Confirm against your vendor's advisory before acting, and report anything wrong.