Composite Monitor กับ Notification Workflow
ไอเดียหลักในหนึ่งประโยค
หัวข้อที่มีชื่อว่า “ไอเดียหลักในหนึ่งประโยค”composite monitor รวม monitor ที่มีอยู่แล้วตั้งแต่สองตัวขึ้นไปเข้าด้วยกันด้วย boolean logic เพื่อให้ alert ได้จาก ชุดผสม ของเงื่อนไข แทนที่จะ alert ทุกครั้งที่ monitor ตัวใดตัวหนึ่ง fire เดี่ยว ๆ และเมื่อถึงจุดที่ต้อง notify ใครสักคนจริง ๆ @-mention กับ escalation_message คือตัวควบคุมว่าใครจะได้ยินและจะขยายวงเมื่อไหร่
Composite monitor: alert จากชุดผสม ไม่ใช่สัญญาณเดี่ยว
หัวข้อที่มีชื่อว่า “Composite monitor: alert จากชุดผสม ไม่ใช่สัญญาณเดี่ยว”composite monitor อ้างอิง ID ของ monitor ที่มีอยู่แล้ว แล้วรวม trigger state ของตัวเองด้วย boolean operator — AND (&&), OR (||), NOT (!) ผลลัพธ์ทางปฏิบัติคือลด noise ลง แทนที่จะ page ทุกครั้งที่มี monitor ไหน fire จะ page เฉพาะตอนที่ ชุดผสมของ state ตรงตามที่กำหนดเท่านั้น
{ "name": "payments-api error rate (composite)", "type": "composite", "query": "111111 && !222222", "message": "Error rate is elevated on payments-api and the upstream-dependency monitor is healthy, so this looks like our own issue."}ในที่นี้ 111111 คือ ID ของ error-rate monitor และ 222222 คือ ID ของ upstream-dependency monitor สูตรนี้จะ fire ก็ต่อเมื่อ 111111 กำลัง alert และ 222222 ไม่ได้ alert ถ้าทั้งสองตัว alert พร้อมกัน composite จะเงียบ — เพราะเรื่องที่น่าจะเกิดคือปัญหาจริงอยู่ที่ upstream และ monitor ของ upstream เองก็ page คนที่ดูแล dependency นั้นอยู่แล้ว
Route notification ด้วย @-mention
หัวข้อที่มีชื่อว่า “Route notification ด้วย @-mention”ฟิลด์ message ของ monitor ใส่ @-mention ตรงในเนื้อข้อความได้เลย — @username สำหรับคนคนหนึ่ง, @slack-channel-name สำหรับ Slack channel ที่เชื่อมไว้แล้ว, @pagerduty-service-name สำหรับ PagerDuty service ที่เชื่อมไว้แล้ว แต่ละตัวต้องตั้งเป็น integration ใน Datadog ก่อน handle ใน message คือตัวที่ route notification ไปเมื่อ connection นั้นมีอยู่แล้ว
{ "message": "@here Error rate is elevated on payments-api.\n@slack-payments-oncall @pagerduty-payments-critical"}escalation_message: message อีกแบบสำหรับ re-notification
หัวข้อที่มีชื่อว่า “escalation_message: message อีกแบบสำหรับ re-notification”ถ้า monitor ยังค้าง trigger อยู่ renotify_interval (หน่วยเป็นนาที) จะทำให้ monitor notify ซ้ำเป็นระยะแทนที่จะ notify แค่ครั้งเดียว escalation_message เป็นฟิลด์แยกที่ให้ re-notification รอบนั้นใช้ข้อความ ต่างจากเดิม ได้ ส่วนใหญ่ใช้เพื่อดึง on-call escalation วงกว้างขึ้นเมื่อปัญหายังไม่ถูกแก้นานขึ้นเรื่อย ๆ
{ "options": { "renotify_interval": 15 }, "message": "@slack-payments-oncall Error rate is elevated on payments-api.", "escalation_message": "@pagerduty-payments-escalation Still broken after 15 minutes, escalating to the secondary on-call."}page แรกไปที่ Slack channel หลัก ถ้าผ่านไปสิบห้านาทีแล้วยังไม่มีใคร resolve notification ถัดไป จะออกไปด้วยข้อความของ escalation_message แทน ดึง secondary on-call ของ PagerDuty เข้ามา
Downtime: mute โดยไม่ปิด evaluation
หัวข้อที่มีชื่อว่า “Downtime: mute โดยไม่ปิด evaluation”Downtime (mute window) ปิดเสียง notification ของ monitor สำหรับช่วง maintenance ที่รู้ล่วงหน้า โดยไม่แตะต้อง evaluation เลย monitor ยังรัน query ของตัวเองต่อไปและอัปเดต internal state ตลอดเวลา มีแค่ notification ขาออกเท่านั้นที่ถูกกัน จุดนี้สำคัญเพราะ status history ยังถูกต้องอยู่ และ notification จะกลับมาทำงานเองทันทีที่ Downtime window จบ ไม่ต้องมีขั้นตอน re-enable ด้วยมือ
flowchart LR
A[Monitor A alerting] --> C{Composite formula}
B[Monitor B alerting] --> C
C -->|A AND NOT B is true| N[Notify via @-mentions]
C -->|formula is false| Q[Stay quiet]
N --> R{Still triggered after renotify_interval}
R -->|yes| E[Send escalation_message]
R -->|no| Done[Resolved]