
Upwind Security MDR@UpwindMDR
🚨HIGH - vLLM Decode Worker Memory Exhaustion via max_tokens=0 (CVE-2026-93436) In vLLM disaggregated prefill/decode deployments, rejected inference requests don’t fully clean up decode-side metadata. An attacker can spam requests with max_tokens=0 to force unbounded memory growth on the decode worker, exhausting RAM until restart (DoS). Non-disaggregated deployments aren’t affected. 👉Affected: vllm <= 0.29.0
0001035
304 followersView on X
