GPU VulnDB

Database/AI/ML frameworks & serving

vLLM: structured-output request failures escape request scope and terminate the shared engine

CVSS 6.5CVE-2026-105757AI/ML frameworks & servingcurated

Impact

Failures in constrained-generation handling reach EngineCore's fatal-error path instead of failing just the request. Three routes are named: a per-request backend mismatch re-raises a grammar compilation exception; padding produced by the ngram_gpu speculative-decoding mode passes a negative token into guidance validation; and the Rust frontend admits empty structured-output values that the Python frontend rejects. Any of these lets an ordinary-looking JSON-schema or grammar request kill the serving process, dropping all concurrent tenants on that replica and forcing a weight reload. Structured output is heavily used by agent and tool-calling workloads, so this is reachable in normal traffic, not only by a deliberate attacker.

Who can reach it

Any authenticated client that can send a structured-output (guided decoding / JSON schema / grammar) request over the network, CVSS AV:N/PR:L. No special configuration is needed for the frontend-validation route; the speculative-decoding route requires ngram_gpu to be enabled.

What to do

Upgrade vLLM to 0.30.0 and restart the serving replicas, rolling them to keep the endpoint available. As an interim measure, disable ngram_gpu speculative decoding and validate structured-output fields at the gateway; neither closes the grammar-backend route, so the upgrade is the real fix.

References

Related entries

All AI/ML frameworks & serving entries

This entry is curated: imported from vendor advisories with machine assistance, not yet individually verified. Confirm against your vendor's advisory before acting, and report anything wrong.