ข้ามไปยังเนื้อหา

Key กับ Partitioning

Key ของ record เป็นตัวตัดสินว่าไป partition ไหน — default partitioner จะ hash key ให้ key เดียวกันลง partition เดียวกันเสมอ ที่เป็นวิธีได้ ordering ต่อ entity (ต่อ user, ต่อ order, ต่อ account).

message ทุกตัวที่ส่งมี key แบบ optional ได้ partitioner ของ producer จะแปลง key เป็นเลข partition:

// Same key "user-42" → same partition, every time → ordered per user
await producer.send({
topic: 'orders',
messages: [
{ key: 'user-42', value: 'placed order #918' },
{ key: 'user-42', value: 'cancelled order #918' },
],
})

สำหรับ key ที่ ไม่ null default partitioner คำนวณประมาณ hash(key) % numberOfPartitions การ hash แบบ deterministic ทำให้ทุก record ของ user-42 ไป partition เดียวกัน event ของ user นั้นจึงอยู่ใน log เดียวที่ ordered ทั้งหมด ส่วน key ต่างกันก็กระจายไปคนละ partition ช่วยบาลานซ์โหลด

ถ้า key เป็น null แปลว่าคุณไม่สนว่า record ลง partition ไหน สนแค่ให้กระจายเท่ากัน Kafka ใช้ sticky partitioner: ส่ง batch ของ record null-key ไป partition เดียว จน batch เต็มหรือ linger.ms หมด แล้วค่อยเปลี่ยน partition การ stick กับ partition เดียวต่อ batch ทำให้ได้ batch ใหญ่และมีประสิทธิภาพกว่า (request น้อยลงแต่ใหญ่ขึ้น) และยังกระจายเท่ากันเมื่อเวลาผ่านไป

flowchart TB
  k1["key = user-42"] -->|hash % N| p1["partition 1 (เสมอ)"]
  k2["key = user-99"] -->|hash % N| p2["partition 2 (เสมอ)"]
  k3["key = null"] -->|sticky batch| p3["partition เดียวต่อ batch แล้วหมุนไป"]
record ที่มี key hash ไป partition คงที่ ส่วน null key batch แบบ sticky

เพราะ mapping คือ hash(key) % numberOfPartitions การ เปลี่ยนจำนวน partition ทำให้ key ลง partition เปลี่ยน เพิ่ม partition ให้ topic ที่ใช้งานอยู่ แล้ว user-42 อาจ hash ไป partition ใหม่ — event ในอนาคตของตัวเองไปคนละที่กับของเดิม ทำลาย ordering ต่อ key ที่คุณพึ่งพาอยู่

ผลในทางปฏิบัติ:

  • เตรียม partition ให้พอ ตั้งแต่แรก สำหรับ topic ที่ใช้ key และ ordering สำคัญ
  • ถ้าจำเป็นต้องขยาย ต้องเข้าใจว่า ordering ต่อ key ของเดิมถูกรักษาไว้แค่ภายในแต่ละ partition เก่า ไม่ข้ามเส้นแบ่ง

เมื่ออยาก join สอง stream ด้วย key เดียวกัน (เช่น orders กับ payments ที่ key ด้วย orderId ทั้งคู่) ให้ทั้งสอง topic มี จำนวน partition เท่ากัน และ partitioner เดียวกัน จากนั้น orderId=918 จะลง partition 3 ใน ทั้งสอง topic และ consumer หรือ stream task ตัวเดียวเห็นทั้งสองฝั่งแบบ local — ไม่ต้อง shuffle ข้าม partition นี่คือ co-partitioning และ Kafka Streams บังคับต้องมีสำหรับ join ตาม key.

อะไรเป็นตัวกำหนดว่า record ที่มี key ไป partition ไหน?
ทำไมใช้ key อย่าง user id ถึงได้ ordering ต่อ user?
sticky partitioner ทำอะไรกับ record ที่ null key?
ทำไมการเพิ่ม partition ให้ topic ที่ใช้ key ถึงเสี่ยง?