GPU VulnDB

Database/AI/ML frameworks & serving

vLLM: Harmony tool continuations lose cache_salt, leaking prefix-cache hits across tenants

CVSS 3.1CVE-2026-105752AI/ML frameworks & servingcurated

Impact

When a Harmony tool continuation is submitted through POST /v1/responses, vLLM rebuilds the next-turn engine input without carrying the caller's cache_salt, so the continuation prefix lands in the global unsalted cache namespace even though the tenant asked for salting. With prefix caching on - the default - a tenant who can reconstruct a victim's low-entropy post-tool history can replay that continuation and read cached_tokens_per_turn to learn whether the prefix had already been processed. That defeats the isolation cache_salt exists to provide: on a multi-tenant inference endpoint it is a side channel confirming what another tenant ran, not a disclosure of the content itself.

Who can reach it

An authenticated tenant on a shared vLLM deployment with prefix caching enabled, who can guess or reconstruct the victim's post-tool conversation prefix. Attack complexity is rated high.

What to do

Upgrade to vLLM 0.30.0 and restart the serving processes; all prior versions are affected. No node-level work. If upgrading has to wait, the blunt mitigation is disabling prefix caching for tenants that rely on cache_salt, at a real throughput cost, or giving those tenants their own replica.

References

Related entries

All AI/ML frameworks & serving entries

This entry is curated: imported from vendor advisories with machine assistance, not yet individually verified. Confirm against your vendor's advisory before acting, and report anything wrong.