GPU VulnDB

Database/AI/ML frameworks & serving

vLLM: overlong token_ids on the disaggregated serving endpoint crash the worker

CVSS 7.1CVE-2026-100651AI/ML frameworks & servingcurated

Impact

On the disaggregated serving endpoint, a multimodal request builds an EngineInput straight from caller-supplied token_ids without checking them against max_model_len. For processors that skip the prompt-length check (Nemotron Parse, Whisper, FireRedLID are named), the overlong prompt reaches the worker's copy into a fixed max_model_len-wide array and the worker dies. Disaggregated prefill/decode deployments spread one model across several GPU nodes, so losing a worker takes down more than the request that caused it. Affects only those model configurations.

Who can reach it

Any client that can reach /inference/v1/generate on an affected model configuration. This endpoint is normally an internal prefill/decode link rather than a tenant-facing one, so the realistic attacker is someone already inside the serving network or a tenant whose pod can route to it.

What to do

Upgrade to vLLM 0.29.0 and restart the disaggregated serving processes. Until then, keep /inference/v1/generate on a network path reachable only by the paired prefill and decode workers, not by tenant pods.

References

Related entries

All AI/ML frameworks & serving entries

This entry is curated: imported from vendor advisories with machine assistance, not yet individually verified. Confirm against your vendor's advisory before acting, and report anything wrong.