Database/Control plane, storage & DevOps

Slurm (NULL pointer dereference in RPC handling): A crafted message crashes the Slurm daemon. On slurmctld that stalls
Impact
A crafted message crashes the Slurm daemon. On slurmctld that stalls every scheduling decision on the cluster - no new job starts, no allocations, and GPUs sit idle until the controller is back. This is the sixth of the December 2023 batch and is the only one of that batch not already in the database.
Who can reach it
Network reach to a Slurm daemon. Same exposure surface as the rest of the 2023-12 batch, so anything that can send RPCs to slurmctld or slurmd.
What to do
Upgrade to Slurm 22.05.11, 23.02.7 or 23.11.1 and restart the daemons. slurmctld restarts preserve running jobs, so this is a low-drama upgrade - do it in the same window as the rest of the 2023-12 fixes if you have not already.
References
This entry is curated: imported from vendor advisories with machine assistance, not yet individually verified. Confirm against your vendor's advisory before acting, and report anything wrong.