GPU VulnDB

Database/Container, Kubernetes & orchestration

Envoy: generating internal responses to pipelined HTTP/1.1 requests consumes excessive memory

CVSS 7.5CVE-2020-8661Container, Kubernetes & orchestrationcurated

Impact

When Envoy answers requests itself - local replies such as 404s, 413s or filter-generated errors - rather than proxying them upstream, a pipelined stream of such requests on one HTTP/1.1 connection makes it accumulate response data in memory without bound. This is a separate code path from the small-chunk buffering issue in CVE-2020-8659 and needs no upstream service to cooperate: the attacker only has to make Envoy reject things quickly, which any listener does by default. In a GPU cluster the practical effect is that one connection can OOM the shared ingress in front of every tenant's inference endpoint, or a mesh sidecar in front of a single GPU pod. Availability impact only per the record.

Who can reach it

Any unauthenticated client that can open a TCP connection to an Envoy HTTP/1.1 listener. Requests that Envoy rejects before routing still count, so authentication filters in front of the backend do not prevent reaching the vulnerable path.

What to do

Upgrade Envoy to 1.13.1 or later (the upstream 1.13.1 version history lists this with the parallel fixes for the older 1.12/1.11/1.10 branches); OpenShift Service Mesh users take RHSA-2020:0734. Cost is a proxy restart only - roll gateway replicas behind the load balancer, and restart injected sidecars pod by pod, accepting that each GPU pod restarted has to reload its model.

References

Related entries

All Container, Kubernetes & orchestration entries

This entry is curated: imported from vendor advisories with machine assistance, not yet individually verified. Confirm against your vendor's advisory before acting, and report anything wrong.