ข้ามไปยังเนื้อหา

Pod Disruption กับ Cluster Autoscaling

PodDisruptionBudget จำกัดว่า replica ของ workload ถูก evict พร้อมกันได้กี่ตัวระหว่าง voluntary disruption — ไม่มีอำนาจเหนือ involuntary disruption เลย

การเสีย Pod ไม่ได้เป็น event แบบเดียวกันเสมอไป และความต่างนี้สำคัญต่อสิ่งที่คุณวางแผนรับมือได้จริง

  • Voluntary disruption คือสิ่งที่ operator หรือ automation ตั้งใจสั่งเอง จึงควบคุม เลื่อน หรือจำกัดอัตราได้ เช่น kubectl drain ตอน upgrade node, kubectl cordon ตามด้วย rolling replacement หรือ Cluster Autoscaler ที่ลบ Node ที่ใช้งานต่ำ
  • Involuntary disruption คือสิ่งที่ไม่มีใครสั่ง เช่น hardware ล่ม kernel panic หรือ Node หายไปเฉย ๆ ไม่มี budget หรือ setting ไหนป้องกันเรื่องนี้ได้ — พอเกิดขึ้นแล้ว Node ก็หายไปเลย

PodDisruptionBudget มีผลแค่กับกลุ่มแรกเท่านั้น

PodDisruptionBudget (policy/v1) ประกาศ minAvailable หรือ maxUnavailable สำหรับกลุ่ม Pod ที่เลือกด้วย label และทุกเส้นทาง voluntary eviction — kubectl drain, Eviction API, Cluster Autoscaler ที่ scale ลด Node — ต้องเคารพ budget นี้

apiVersion: policy/v1
kind: PodDisruptionBudget
metadata:
name: web-app-pdb
spec:
minAvailable: 2
selector:
matchLabels:
app: web-app
Terminal window
# kubectl drain blocks an eviction that would push availability below the PDB
kubectl drain node-3 --ignore-daemonsets --delete-emptydir-data
kubectl get pdb web-app-pdb

ด้วย minAvailable: 2 บน Deployment ที่รัน web-app 3 replica การ drain จะ evict Pod ได้ทีละตัวเท่านั้น — คือรอให้ replacement ขึ้น Ready บน Node อื่นก่อนถึงจะ evict ตัวถัดไปได้ ถ้า hardware ล่มพา Node ที่มีสองในสาม replica นั้นหายไปพร้อมกัน PDB จะทำอะไรไม่ได้เลย เพราะไม่ได้ถูกสร้างมาเพื่อกันเรื่องนั้นตั้งแต่แรก PDB ปกป้องกรณีที่ operation เลือกจะเอาออกมากเกินไปพร้อมกัน ไม่ใช่กรณีที่โลกทำแบบนั้นใส่คุณเอง

HPA กับ VPA ทำงานระดับ Pod ส่วน Cluster Autoscaler ทำงานสูงกว่าหนึ่งชั้น คือระดับ Node

  • เพิ่ม Node เมื่อ Pod ค้างอยู่ที่ Pending เพราะไม่มี Node ไหนมี capacity ว่างพอจะ schedule Pod นั้นได้
  • ลบ Node เมื่อ Node นั้นถูกใช้งานต่ำมาก และ Pod ทุกตัวบนนั้น reschedule ไปที่อื่นได้ — แต่ทำโดย drain Node นั้นก่อน โดยเคารพ PodDisruptionBudget ตลอดทาง เหมือนกับตอนทำ kubectl drain ด้วยมือทุกประการ
Terminal window
# Pods stuck here because nothing fits them is the classic trigger for Cluster Autoscaler to add a Node
kubectl get pods --field-selector=status.phase=Pending

ทางเลือกใหม่ที่ควรรู้จักคือ Karpenter ซึ่งข้ามโมเดล fixed node-group-template ไปเลย โดย provision Node ที่ขนาดพอดีตรง ๆ ตอบสนอง Pod ที่ schedule ไม่ได้ แทนที่จะ scale group ที่กำหนดไว้ล่วงหน้าขึ้นลง

flowchart LR
  drain["kubectl drain /\nCluster Autoscaler scale-down"] --> pdb{"PodDisruptionBudget\nminAvailable satisfied?"}
  pdb -->|yes| evict["Evict one Pod at a time"]
  pdb -->|no, would violate| wait["Wait before evicting more"]
  evict --> empty["Node fully drained"]
  empty --> remove["Node removed"]
A drain or Cluster Autoscaler scale-down respecting a PodDisruptionBudget before a Node is removed
ความต่างหลักระหว่าง voluntary กับ involuntary disruption คืออะไร
การตั้ง minAvailable ของ PodDisruptionBudget ปกป้องอะไรจริง ๆ
อะไรที่มักกระตุ้นให้ Cluster Autoscaler เพิ่ม Node ใหม่
Karpenter ต่างจากแนวทาง Cluster Autoscaler แบบดั้งเดิมอย่างไร