
Recent K8s + AI updates worth catching up on vLLM v0.19.1 → Patch release fixing CVE-2026-0994 (protobuf deserialization). If you’re running vLLM in production, this is a required upgrade. Most major features : async scheduling by default, Gemma 4 support, KV cache optimizations) landed in v0.19.0 / v0.18 this is stabilization + security. llm-d → Cloud Native Computing Foundation Sandbox (March 2026) Donated by IBM Research, Red Hat, and Google Cloud, with backing from NVIDIA, AMD, Hugging Face, Intel, etc.The important shift is treating distributed inference as a first-class Kubernetes workload. Concepts like prefill/decode disaggregation and prefix-cache-aware routing (EPP) are pushing infra closer to model-aware scheduling. Karpenter v1.11.x → NodePool limits, topology-aware scheduling, and provider-level hooks are now part of the baseline. If your mental model of Karpenter is pre-1.0, you’re already behind. Kubernetes AI Conformance v1.35 → Not another “drop,” but a directional signal. In-place pod resizing (scale inference without restarts) Workload-aware scheduling (reduce deadlocks in distributed training) Takeaway: This isn’t hype cycles. The Kubernetes ecosystem is actively evolving to natively support AI workloads from inference runtimes to scheduling semantics. If you’re building in this space, these aren’t optional reads.
Post summary
The post announces a patch release for CVE‑2026‑0994, noting the vulnerability involves protobuf deserialization and advises upgrading to vLLM v0.19.1.




