GPU VulnDB

Database/AI/ML frameworks & serving

SGLang: duplicate bootstrap_room values crash or hang the disaggregated scheduler

CVSS 8.7CVE-2026-102634AI/ML frameworks & servingcurated

Impact

In prefill/decode disaggregated mode with the Mooncake KV transfer backend, /generate does not check that bootstrap_room is unique. Concurrent requests carrying the same bootstrap_room crash the scheduler process or leave other users' requests hanging until the transfer timeout expires. On a shared inference tier this is a cross-tenant availability problem: one caller's requests stall or kill a scheduler that many tenants route through, and the GPUs behind it sit idle while the process is restarted. Any unauthenticated client that can reach the endpoint can repeat it.

Who can reach it

Anyone who can send HTTP requests to the SGLang /generate endpoint. No authentication required, per the record. Only deployments running prefill/decode disaggregation with the Mooncake backend are affected.

What to do

No fixed version is named in the record - the flaw is reported through 0.5.20. Track the sgl-project/sglang repository for the fix. Meanwhile keep the serving endpoint behind a gateway that authenticates callers and strips or regenerates bootstrap_room, and be ready to restart scheduler processes; recovery is a daemon restart, not a node reboot.

References

Related entries

All AI/ML frameworks & serving entries

This entry is curated: imported from vendor advisories with machine assistance, not yet individually verified. Confirm against your vendor's advisory before acting, and report anything wrong.