Checklist ก่อนขึ้น Production
ไอเดียหลักในหนึ่งประโยค
หัวข้อที่มีชื่อว่า “ไอเดียหลักในหนึ่งประโยค”การขึ้น production กับ Datadog คือ checklist ไม่ใช่สวิตช์เดียว และทุก item บนนั้นย้อนกลับไปหา decision ที่คอร์สนี้สอนวิธีตัดสินใจไปแล้ว — ตั้งแต่ tagging ของ Foundations, ผ่าน Infrastructure & Metrics, Log Management, APM & Distributed Tracing, Monitors/Alerting/SLOs ไปจนถึง Correlating Signals
Tagging และ ownership มาก่อน เพราะทุกอย่างที่เหลือขึ้นกับสองอย่างนี้
หัวข้อที่มีชื่อว่า “Tagging และ ownership มาก่อน เพราะทุกอย่างที่เหลือขึ้นกับสองอย่างนี้”- Unified service tagging สอดคล้องกันทุกที่
env,serviceและversionถูก apply แบบเดียวกันทั้งใน Agent configuration, APM tracer setup และทุกlogs:source — นี่คือ tagging model จากโมดูล Foundations ที่ทำให้ log line, trace และ metric ผูกกลับไปหา deployable unit เดียวกันได้ ถ้ามาแก้หลัง go-live ทุก dashboard กับ monitor ที่สร้างไปในระหว่างนั้นจะพลาด telemetry บางส่วนไปแบบเงียบ ๆ - Service Catalog entry มีอยู่จริงและถูกต้อง ทุก service ที่จะ page ใครสักคนต้องมี Service Catalog entry ที่มี ownership จริง — ทีมไหน, escalation path ไหน, link ไป runbook ไหน นี่คือสิ่งที่โมดูล Correlating Signals ตั้งไว้ และเป็นตัวตัดสินว่า alert จะไปถึง on-call engineer หรือไปถึง mailing list เก่า ๆ ที่ไม่มีใครอ่าน
Cost ของ log และ metric ตัดสินใจอย่างตั้งใจ ไม่ใช่ตาม default
หัวข้อที่มีชื่อว่า “Cost ของ log และ metric ตัดสินใจอย่างตั้งใจ ไม่ใช่ตาม default”- Log index ถูก size ไว้อย่างตั้งใจ retention กับ daily quota ต่อ index สะท้อนคุณค่าจริงของ log stream นั้น ไม่ใช่ “index ทุกอย่างแล้วค่อยไปคิดทีหลัง” — คือประเด็นทั้งหมดของโมดูล Log Management
- exclusion filter มีอยู่สำหรับ log ที่ volume สูงแต่คุณค่าต่ำ ก่อนที่ volume จะขยายใน production ไม่ใช่มาทำทีหลังหลังบิลแรกที่น่าตกใจ
- log-based metric ถูกตั้งไว้สำหรับอะไรก็ตามที่วางแผนจะ exclude ถ้ายังต้องการ count หรือ trend จาก log stream ที่กำลังจะ filter ออก ให้ config log-based metric ที่พึ่ง
logs_generate_metricsไว้ก่อน — การ exclude raw log ไม่ควรทำให้มองไม่เห็น trend ที่เคยสนใจแบบเงียบ ๆ - custom metric tag ถูก review โดยคำนึงถึง cardinality tag cardinality สูงอย่าง
pod_nameยัง submit ได้อยู่ Metrics without Limits คุมว่าตัวไหนถูก index และ query ได้จริง ๆ และการเลือกนั้นควรตั้งใจก่อน go-live ไม่ใช่มาค้นพบจาก cost spike
APM sampling กับ retention review แยกกันเป็นสอง layer
หัวข้อที่มีชื่อว่า “APM sampling กับ retention review แยกกันเป็นสอง layer”- ingestion sampling ตั้งไว้อย่างตั้งใจ นี่คือ layer ที่ตัดสินว่า trace volume ทั้งหมดเท่าไหร่ที่มาถึง Datadog เลย ควรเป็น rate ที่ตั้งใจ ไม่ใช่ default ที่เกิดขึ้นมาเอง
- retention filter review แยกจาก sampling ยืนยันว่า trace ไหนถูกเก็บระยะยาวไม่ว่า sampling จะเป็นยังไง — เช่น ทุกตัวที่มี error หรือทุกตัวที่เกิน latency threshold — และไม่มีใครในทีมสับสนระหว่าง “sample น้อยลง” กับ “retain น้อยลง” เพราะเป็นปุ่มคนละอันจากโมดูล APM & Distributed Tracing และ go-live คือจังหวะเช็คทั้งสองอัน ไม่ใช่แค่อันเดียว
Alerting ผูกกับสิ่งที่สำคัญจริง ๆ
หัวข้อที่มีชื่อว่า “Alerting ผูกกับสิ่งที่สำคัญจริง ๆ”- SLO define ไว้สำหรับ service กลุ่มเล็ก ๆ ที่ business พึ่งพาจริง ๆ ไม่ใช่ vanity metric ที่ copy ใส่ทุก service เพราะเพิ่มง่าย
- burn-rate alert ผูกกับ SLO เหล่านั้น เพื่อให้ทีมได้รับแจ้งเตือนตอนที่ error budget ยังเหลือให้ react ทัน ไม่ใช่หลังจากหมดไปแล้ว นี่คือ mechanism หลักของโมดูล Monitors, Alerting และ SLOs และจะทำงานได้ก็ต่อเมื่อผูกไว้จริงก่อน go-live
- notification target ของทุก monitor ถูกเช็คกับ owner ใน Service Catalog จากขั้นตอน tagging ด้านบน เพื่อให้ page ไปถึงทีมที่ action ได้จริง
จัดการทั้งหมดเป็น code review เหมือน code
หัวข้อที่มีชื่อว่า “จัดการทั้งหมดเป็น code review เหมือน code”- monitor, dashboard และ SLO เขียนเป็น Terraform resource —
datadog_monitor,datadog_dashboard_v2,datadog_service_level_objective— แทนที่จะคลิกประกอบเอง ตามที่คุยไปก่อนหน้าในโมดูลนี้ terraform planรันพร้อม query validation เปิดอยู่ เป็นส่วนหนึ่งของ pull request review ปกติ ก่อนที่ monitor หรือ dashboard change จะไปถึง production — bar review เดียวกับ infrastructure change อื่น ๆ- RBAC ถูกเช็ค ยืนยันว่ามีแค่ role ที่ตั้งใจถือ permission อย่าง
logs_modify_indexesกับlogs_write_exclusion_filtersส่วนคนอื่นมีแค่ read/query access ปกติ - cost alert monitor ของ Cloud Cost Management มีอยู่จริง สำหรับ service หรือ account ที่ rollout นี้กระทบ เพื่อให้ cost spike ได้รับ alerting treatment แบบเดียวกับ latency spike ไม่ใช่มาเซอร์ไพรส์ตอนสิ้นเดือน
Go-Live Checklist1. Unified tagging (env/service/version) verified across Agent, APM, logs2. Service Catalog ownership accurate for every paging service3. Log indexes and exclusion filters sized deliberately; log-based metrics set up first4. Custom metric tag cardinality reviewed via Metrics without Limits5. APM ingestion sampling and retention filters reviewed as two separate layers6. SLOs and burn-rate alerts wired to business-critical services only7. Monitor notification targets checked against Service Catalog owners8. Monitors/dashboards/SLOs as Terraform code, reviewed via PR with query validation on9. RBAC checked: index/exclusion/archive permissions restricted, read access broad10. Cloud Cost Management cost alert in place for the affected services/accountsflowchart LR A[Tagging and ownership] --> B[Log and metric cost decisions] B --> C[APM sampling and retention review] C --> D[SLOs and burn-rate alerting] D --> E[Infrastructure as code review] E --> F[Go-live] A -.Service Catalog owner.-> D B -.cost alert monitor.-> E