ข้ามไปยังเนื้อหา

Monitoring & Consumer Lag

Kafka เปิดเผย metric หลายร้อยตัวผ่าน JMX แต่สัญญาณระดับ application ที่สำคัญที่สุดคือ consumer lag ช่องว่างระหว่าง log-end-offset ของ partition กับ committed offset ของ consumer group ซึ่งบอกว่า consumer ของคุณตามทันหรือไม่

producer append record เรื่อย ๆ เลื่อน log-end-offset ของแต่ละ partition (ตำแหน่งของ record ใหม่ล่าสุด) consumer group commit ว่าประมวลผลไปถึงไหน lag คือผลต่าง

lag = log-end-offset − committed offset

lag ใกล้ศูนย์แปลว่า consumer ตามทัน lag ที่โตขึ้นเรื่อย ๆ แปลว่า consumer ตามไม่ทัน record เข้ามาเร็วกว่าที่ประมวลผลได้ และ latency แบบ end-to-end กำลังไต่ขึ้น อ่านค่า lag ตรง ๆ ด้วย consumer-groups tool

Terminal window
# Describe a consumer group: shows CURRENT-OFFSET, LOG-END-OFFSET, and LAG per partition
kafka-consumer-groups.sh --bootstrap-server localhost:9092 \
--describe --group order-processor

output มีคอลัมน์ LAG ต่อ partition จับตาที่ แนวโน้ม ของตัวเอง ไม่ใช่ค่าครั้งเดียว แบนหรือลดลงคือสุขภาพดี ไต่ขึ้นเรื่อย ๆ คือสัญญาณเตือนแรกสุดว่า consumer โหลดเกิน ค้าง หรือ crash

flowchart LR
  prod["Producer appends records"] --> leo["Log-end-offset (newest)"]
  cons["Consumer group commits"] --> cur["Committed offset (processed)"]
  leo -->|"difference"| lag["LAG = log-end-offset - committed offset"]
  cur --> lag
consumer lag คือช่องว่างระหว่าง offset ใหม่สุดกับที่ประมวลผลแล้ว

broker ของ Kafka publish metric ผ่าน JMX ซึ่ง tool อย่าง Prometheus (ผ่าน JMX exporter), Grafana หรือ APM ของคุณ scrape ไป มีบางตัวที่ห้ามต่อรอง

  • Under-replicated partitions (UnderReplicatedPartitions) partition ที่ ISR เล็กกว่า replication factor ควรเป็น ศูนย์ อะไรที่เกินนั้นแปลว่า replica ตามไม่ทันหรือ broker ล่ม และ durability กำลังเสี่ยง
  • Offline partitions (OfflinePartitionsCount) partition ที่ไม่มี leader จึงใช้งานไม่ได้ ควรเป็น ศูนย์ เสมอ
  • Active controller (ActiveControllerCount) ต้องมี หนึ่ง เดียวทั้ง cluster ศูนย์หรือมากกว่าหนึ่งคือสัญญาณปัญหาของ controller
  • Request latency เวลาของ produce กับ fetch request (TotalTimeMs แยกเป็นเฟส queue, local, remote) latency ที่ไต่ขึ้นชี้ไปที่แรงกดดันของ disk, network หรือ replication ก่อนที่ client จะรู้สึก
# Enable JMX on a broker by exporting the port before start (example)
# then a JMX exporter or your APM scrapes these MBeans
JMX_PORT=9999

เปลี่ยน metric สำคัญให้เป็น alert ไม่ใช่แค่ dashboard

  • consumer lag ไต่เกิน threshold (หรือเกิน SLA เป็นวินาที) ของ group สำคัญใด ๆ
  • under-replicated partitions มากกว่าศูนย์นานเกินชั่วครู่
  • offline partitions มากกว่าศูนย์ page ทันที
  • request latency p99 ทะลุ budget ของ produce หรือ fetch
consumer lag นิยามยังไง
คำสั่งไหนอ่าน consumer lag ต่อ partition
metric under-replicated partitions ปกติควรแสดงค่าอะไร
lag ของ consumer group ไต่ขึ้นเรื่อย ๆ มาหนึ่งชั่วโมง บอกอะไร