GPU VulnDB

Database/AI/ML frameworks & serving

vLLM: multi-prompt completion request trips a NIXL prefix-caching assertion and kills the decode worker

CVSS 8.7CVE-2026-94623AI/ML frameworks & servingcurated

Impact

The NIXL connector's prefix-caching path does not validate block counts across a completion request that carries several prompts of differing lengths. A single such request fails an assertion in the decode worker and terminates it, leaving that worker unavailable until restarted. This is an ordinary, well-formed API shape - batching multiple prompts in one completion call - so it can be tripped accidentally as well as deliberately, and on a shared disaggregated tier it removes serving capacity that other tenants depend on. The fix is separate from the other NIXL denial-of-service issues in this batch.

Who can reach it

Any client able to submit a completion request with multiple prompts to a vLLM deployment running prefill/decode disaggregation with the NIXL connector. No authentication required per the record.

What to do

Apply the change in vllm-project/vllm PR #51505; the record names no fixed release (reported through 0.29.0). Rollout is an image update plus a restart of decode workers, done per replica with traffic drained so the weight reload is not user-visible. As a stopgap, reject multi-prompt completion requests at the gateway on NIXL disaggregated routes.

References

Related entries

All AI/ML frameworks & serving entries

This entry is curated: imported from vendor advisories with machine assistance, not yet individually verified. Confirm against your vendor's advisory before acting, and report anything wrong.