ข้ามไปยังเนื้อหา

Indexes, Exclusion Filters, and Retention

log ถูก drop หรือ filter ได้ที่สามขั้นตอนที่ต่างกันจริง ๆ — บน host ก่อนที่จะออกไปเลย, ที่ quota รายวันของ index หรือด้วย Exclusion Filter ของ index เอง — แต่ละจุดมีผลเรื่อง cost และการกู้คืนที่ต่างกัน จึงห้ามเอามาปนกัน

ก่อนที่ log จะออกจาก host เลย ตัว Agent เองก็ drop log ทิ้งได้ผ่าน entry ใน log_processing_rules ที่ type เป็น exclude_at_match โดย match regex pattern กับ raw log line

logs:
- type: file
path: /my/app/file.log
service: payments-api
source: java
log_processing_rules:
- type: exclude_at_match
name: exclude_healthcheck_lines
pattern: 'GET /healthz'

นี่คือจุด filter ที่ถูกที่สุดเท่าที่จะทำได้ เพราะบรรทัดที่ถูก drop ไม่กิน network, intake หรือ storage เลย — แต่ก็เป็นการตัดสินใจที่ final เช่นกัน log ที่ถูก exclude ตรงนี้หายไปตลอดกาล ไม่ถูก archive, ไม่ไปถึง Processing Pipeline และกู้คืนทีหลังไม่ได้ ใช้จุดนี้เฉพาะกับบรรทัดที่ไม่มีค่าจริง ๆ (noise ของ health-check, debug spam ที่ไม่ต้องการในรูปแบบไหนเลย) เพราะไม่มี undo

log ที่รอดจาก Agent แล้วผ่าน intake กับ Processing Pipeline มาถึง Log Index หนึ่งตัวหรือมากกว่า index คือจุดที่ full-price searchable retention เกิดขึ้นจริง และแต่ละ index ตั้งค่าตัวเลขสองตัวที่สำคัญต่อ cost และ behavior:

{
"name": "main",
"daily_limit": 100,
"retention_days": 30,
"exclusion_filters": [
{ "name": "exclude-debug", "query": "level:debug" }
]
}
  • retention_days — log ที่ถูกเขียนเข้า index นี้ค้นหาได้ใน Logs Explorer แบบ full price นานกี่วัน
  • daily_limit — quota GB/วันของ index นี้ พอถึง quota ของวันนั้นแล้ว index จะหยุดรับ log เพิ่ม หรือ (แล้วแต่ config) เริ่ม sample ดังนั้น daily_limit ที่ตั้งผิดขนาดสามารถจำกัดสิ่งที่ค้นหาได้แบบเงียบ ๆ โดยไม่เกี่ยวกับส่วนอื่นใน config เลย

organization ปกติจะมี index มากกว่าหนึ่งตัว (เช่น main สำหรับ production service, index retention สั้นกว่าสำหรับ log ที่ noisy หรือมูลค่าต่ำ) และ log จะไปลง index ไหนก็ถูกควบคุมด้วย filter criteria ระดับ index เองอีกที — แต่ประเด็นสำคัญของบทนี้คือ retention_days กับ daily_limit เป็น property ของ index ไม่ใช่ของ Agent หรือ pipeline

Exclusion Filter อยู่บน index ตัวใดตัวหนึ่ง และใช้ query แบบ Lucene-style — syntax เดียวกับที่พิมพ์ในช่องค้นหาของ Logs Explorer — เพื่อ sample ออกหรือ drop log ที่ match ไม่ให้ถูกเขียนเข้า index นั้น จาก config ด้านบน { "name": "exclude-debug", "query": "level:debug" } หมายความว่า log ที่ match level:debug จะไม่ถูกเขียนเข้า index main

นี่คือคันโยกหลักที่ใช้ประจำวันเพื่อคุมปริมาณที่ index, ค้นหาได้ และคิดเงินเต็มราคา และถูกออกแบบให้ยืดหยุ่นและ final น้อยกว่าการ drop ฝั่ง Agent โดยตั้งใจ:

  • log ที่ถูก exclude จาก index หนึ่งด้วย Exclusion Filter ของตัวเอง ยังถูก route ไป index อื่น ได้ (ที่มี retention หรือ cost ต่างกัน) ถ้า index routing ส่งไปทางนั้น
  • log ที่ถูก exclude จากทุก index ยังไปถึง Log Archive ได้ ถ้าตั้งค่า archiving ไว้ — บทถัดไปจะพูดเรื่องนี้
  • เพราะการ exclude เกิดที่ index ไม่ใช่บน host จึงเปลี่ยนได้จากส่วนกลางใน Datadog โดยไม่ต้องแตะ config ไฟล์ของ Agent สักตัวเดียว

เส้นทางเต็ม ๆ ที่ log เดินทางผ่าน ตั้งแต่ต้นจนจบ คือ Agent เก็บ log → rule exclude_at_match ฝั่ง Agent (ถ้ามี) อาจ drop log บางตัวก่อนที่จะออกจาก host เลย → log ที่รอดมาถึง intakeProcessing Pipeline parse และ enrich log ทุกตัวที่ match filter ของตัวเอง → log มาถึง Log Index ซึ่ง Exclusion Filter ของตัวเองเป็นตัวตัดสินว่าอะไรจะถูก index จริง ค้นหาได้ และคิดเงินเต็มราคา

flowchart LR
  A[Agent collects log] --> B{exclude_at_match\nmatches?}
  B -->|yes, dropped forever| Z[Gone: no archive, no index]
  B -->|no| C[Intake]
  C --> D[Processing Pipeline\nparse + enrich]
  D --> E{Index Exclusion Filter\ne.g. level:debug}
  E -->|excluded| F[Not written to this index\nmay still reach archive/other index]
  E -->|kept| G["Indexed: searchable\nfor retention_days, billed at full price"]
Where a log can be dropped or filtered
log ที่ถูก drop โดย rule `exclude_at_match` ฝั่ง Agent จะเป็นยังไงต่อ
`daily_limit` ของ index ควบคุมอะไร
log ที่ match กับ Exclusion Filter ของ index (เช่น `level:debug`) อธิบายได้ดีที่สุดว่า
ลำดับที่ถูกต้องที่ log เดินทางผ่านตั้งแต่ต้นจนจบคือแบบไหน