GPU VulnDB

Database/Container, Kubernetes & orchestration

Envoy: HTTP/1.1 requests or responses with many 1-byte chunks exhaust proxy memory

CVSS 7.5CVE-2020-8659Container, Kubernetes & orchestrationcurated

Impact

An unauthenticated client can drive an Envoy instance into excessive memory consumption by sending HTTP/1.1 traffic composed of very many tiny (1-byte) chunks, which Envoy buffers far less efficiently than the wire size suggests. On a GPU cluster this matters because Envoy is typically the shared ingress and the service-mesh sidecar in front of inference endpoints: an OOM in the gateway takes every tenant's model endpoint offline at once, and an OOM in a sidecar kills the pod it fronts even though the GPU workload itself is healthy. Availability only - the record claims no confidentiality or integrity loss (CVSS C:N/I:N/A:H). Pods that are restarted lose any warmed model weights and re-pay the load time on the next request.

Who can reach it

Anyone who can reach an Envoy HTTP/1.1 listener - no authentication needed. For a public inference gateway that is the internet; for a mesh sidecar it is any workload allowed to open a connection to that pod. A malicious or compromised upstream can also trigger it via response bodies.

What to do

Upgrade Envoy to 1.13.1 or later (1.12.3, 1.11.2 and 1.10.0 patch releases also carry the fix per the upstream version history); OpenShift Service Mesh users take RHSA-2020:0734, Debian users the LTS update. Rollout is a proxy restart: gateways roll behind a load balancer with connection draining, mesh sidecars require restarting (or hot-restarting) each injected pod, which for GPU pods means a reschedule and a model reload. No node drain or reboot is involved.

References

Related entries

All Container, Kubernetes & orchestration entries

This entry is curated: imported from vendor advisories with machine assistance, not yet individually verified. Confirm against your vendor's advisory before acting, and report anything wrong.