GPU VulnDB

Database/AI/ML frameworks & serving

vLLM: engine crash via the penalty handler when prompt embeds are combined with penalties

CVSS 2.1CVE-2026-105922AI/ML frameworks & servingcurated

Impact

A client request that combines prompt embeddings with penalty sampling parameters crashes the vLLM engine in get_token_bin_counts_and_mask. The published reproducer kills the engine process, which on a GPU node takes down the whole served model, not just the offending request - every other tenant on that endpoint loses service until the daemon restarts and reloads weights onto the GPUs. Because the trigger is ordinary API parameters rather than anything privileged, one tenant can repeatedly deny the node to the rest. A public reproducer script exists and the project had not responded at the time of publication.

Who can reach it

Any authenticated client that can submit a completions request with prompt embeds and penalty parameters. No special privilege beyond normal API access to the endpoint.

What to do

No fixed version is published in this record; the report is an upstream GitHub issue with no vendor response yet. Mitigation available today is input validation at the gateway in front of vLLM - reject or sanitise penalty parameters on requests carrying prompt embeds - plus a supervisor that restarts the engine. When an upstream fix lands, the action is an upgrade and a serving-daemon restart.

References

Related entries

All AI/ML frameworks & serving entries

This entry is curated: imported from vendor advisories with machine assistance, not yet individually verified. Confirm against your vendor's advisory before acting, and report anything wrong.