Monitoring & Consumer Lag
ไอเดียในหนึ่งประโยค
หัวข้อที่มีชื่อว่า “ไอเดียในหนึ่งประโยค”Kafka เปิดเผย metric หลายร้อยตัวผ่าน JMX แต่สัญญาณระดับ application ที่สำคัญที่สุดคือ consumer lag ช่องว่างระหว่าง log-end-offset ของ partition กับ committed offset ของ consumer group ซึ่งบอกว่า consumer ของคุณตามทันหรือไม่
consumer lag: ตัวเลขเดียวที่ต้องจับตา
หัวข้อที่มีชื่อว่า “consumer lag: ตัวเลขเดียวที่ต้องจับตา”producer append record เรื่อย ๆ เลื่อน log-end-offset ของแต่ละ partition (ตำแหน่งของ record ใหม่ล่าสุด) consumer group commit ว่าประมวลผลไปถึงไหน lag คือผลต่าง
lag = log-end-offset − committed offset
lag ใกล้ศูนย์แปลว่า consumer ตามทัน lag ที่โตขึ้นเรื่อย ๆ แปลว่า consumer ตามไม่ทัน record เข้ามาเร็วกว่าที่ประมวลผลได้ และ latency แบบ end-to-end กำลังไต่ขึ้น อ่านค่า lag ตรง ๆ ด้วย consumer-groups tool
# Describe a consumer group: shows CURRENT-OFFSET, LOG-END-OFFSET, and LAG per partitionkafka-consumer-groups.sh --bootstrap-server localhost:9092 \ --describe --group order-processoroutput มีคอลัมน์ LAG ต่อ partition จับตาที่ แนวโน้ม ของตัวเอง ไม่ใช่ค่าครั้งเดียว แบนหรือลดลงคือสุขภาพดี ไต่ขึ้นเรื่อย ๆ คือสัญญาณเตือนแรกสุดว่า consumer โหลดเกิน ค้าง หรือ crash
flowchart LR prod["Producer appends records"] --> leo["Log-end-offset (newest)"] cons["Consumer group commits"] --> cur["Committed offset (processed)"] leo -->|"difference"| lag["LAG = log-end-offset - committed offset"] cur --> lag
metric ฝั่ง broker ผ่าน JMX
หัวข้อที่มีชื่อว่า “metric ฝั่ง broker ผ่าน JMX”broker ของ Kafka publish metric ผ่าน JMX ซึ่ง tool อย่าง Prometheus (ผ่าน JMX exporter), Grafana หรือ APM ของคุณ scrape ไป มีบางตัวที่ห้ามต่อรอง
- Under-replicated partitions (
UnderReplicatedPartitions) partition ที่ ISR เล็กกว่า replication factor ควรเป็น ศูนย์ อะไรที่เกินนั้นแปลว่า replica ตามไม่ทันหรือ broker ล่ม และ durability กำลังเสี่ยง - Offline partitions (
OfflinePartitionsCount) partition ที่ไม่มี leader จึงใช้งานไม่ได้ ควรเป็น ศูนย์ เสมอ - Active controller (
ActiveControllerCount) ต้องมี หนึ่ง เดียวทั้ง cluster ศูนย์หรือมากกว่าหนึ่งคือสัญญาณปัญหาของ controller - Request latency เวลาของ produce กับ fetch request (
TotalTimeMsแยกเป็นเฟส queue, local, remote) latency ที่ไต่ขึ้นชี้ไปที่แรงกดดันของ disk, network หรือ replication ก่อนที่ client จะรู้สึก
# Enable JMX on a broker by exporting the port before start (example)# then a JMX exporter or your APM scrapes these MBeansJMX_PORT=9999ควร alert อะไรบ้าง
หัวข้อที่มีชื่อว่า “ควร alert อะไรบ้าง”เปลี่ยน metric สำคัญให้เป็น alert ไม่ใช่แค่ dashboard
- consumer lag ไต่เกิน threshold (หรือเกิน SLA เป็นวินาที) ของ group สำคัญใด ๆ
- under-replicated partitions มากกว่าศูนย์นานเกินชั่วครู่
- offline partitions มากกว่าศูนย์ page ทันที
- request latency p99 ทะลุ budget ของ produce หรือ fetch