ข้ามไปยังเนื้อหา

Remote Backends and Locking

remote backend เก็บ state ไว้ที่จุดกลางอย่าง S3 แทนที่จะอยู่บน laptop ของคนใดคนหนึ่ง ส่วน locking กันไม่ให้สองคนเขียนทับ state ที่ share กันนั้นพร้อมกัน

โดย default terraform.tfstate จะอยู่บนเครื่องที่รัน apply ซึ่งโอเคสำหรับทดลองคนเดียว แต่พังทันทีเมื่อมีคนที่สองเข้ามาร่วม

  • เพื่อนร่วมทีม A รัน apply แล้วสร้าง VPC local state file บน laptop ของเขารู้เรื่องนี้แล้ว
  • เพื่อนร่วมทีม B บน laptop คนละเครื่อง ไม่เคยเห็น state file นั้นเลย เขารัน plan เทียบกับ local state ที่เก่าหรือว่างเปล่าของตัวเอง แล้วอาจพยายามสร้าง VPC ซ้ำ หรือได้ plan ที่ไม่สมเหตุสมผลเมื่อเทียบกับของจริงที่มีอยู่แล้ว
  • CI/CD pipeline เจอปัญหาเดียวกันแต่หนักกว่า pipeline runner คือ environment ใหม่ที่ใช้แล้วทิ้ง ไม่มี local state file ติดตัวมาเลย จึงต้องการ state ฉบับกลางที่อ่านและเขียนได้ทุกครั้งที่รัน ไม่ใช่ copy ที่บังเอิญไปอยู่บน laptop ของใครสักคน

ทางแก้คือย้าย state ออกจากเครื่องใครเครื่องคนหนึ่ง ไปไว้บน backend กลางที่ทั้งทีมและทุก CI run ชี้ไปที่เดียวกัน

วิธีที่แนะนำในปัจจุบันสำหรับเก็บ Terraform state บน AWS คือ S3 bucket ที่ config เป็น backend พร้อมเปิด native S3 state locking

terraform {
backend "s3" {
bucket = "my-company-terraform-state"
key = "networking/terraform.tfstate"
region = "us-east-1"
use_lockfile = true
}
}

bucket กับ key รวมกันระบุว่า object ไหนใน S3 ที่เก็บ state ของ configuration นี้ — key คือ path ภายใน bucket ดังนั้น configuration ต่างกัน (หรือ Terragrunt unit ต่างกัน) มักใช้ key คนละอันใน bucket เดียวกัน use_lockfile = true คือบรรทัดสำคัญ Terraform ใช้ conditional-write lock file ที่เก็บอยู่คู่กับ state object ตรงใน S3 เลย เพื่อจัดการ locking ไม่ต้องมี infrastructure สำหรับ locking แยกต่างหาก แค่ S3 bucket ก็พอ

คุณจะยังเจอ pattern เก่าใน codebase ที่มีอยู่แล้ว ซึ่งใช้ DynamoDB table แยกสำหรับ locking

# Legacy locking pattern — dynamodb_table is deprecated and scheduled for removal.
# Prefer use_lockfile = true (shown above) in new configurations.
terraform {
backend "s3" {
bucket = "my-company-terraform-state"
key = "networking/terraform.tfstate"
region = "us-east-1"
dynamodb_table = "terraform-state-lock"
}
}

pattern นั้นยังทำงานได้ และไม่ต้องตกใจถ้าคุณรับ codebase ที่ใช้แบบนี้มา แต่ dynamodb_table ตอนนี้เป็น argument ที่ deprecated แล้วบน S3 backend และมีแผนจะลบออกในเวอร์ชันถัดไปของ Terraform configuration ใหม่ควรใช้ use_lockfile = true แทน เพราะทำหน้าที่เดียวกันโดยไม่ต้องสร้าง AWS resource ตัวที่สองมาดูแลและจ่ายเงินเพิ่ม

ลองนึกภาพว่าไม่มี locking เลย สองคน หรือคนกับ CI run รัน terraform apply เทียบกับ state บน S3 เดียวกันในเวลาที่ใกล้กันมาก ทั้งคู่อ่าน state ปัจจุบัน ทั้งคู่คำนวณ plan และทั้งคู่เริ่มเขียนผลกลับไปที่ state object เดียวกัน ขึ้นอยู่กับ timing การเขียนครั้งที่สองอาจเขียนทับครั้งแรก — ทำให้การเปลี่ยนแปลงที่ record ไว้จาก run แรกหายไปเงียบ ๆ ทั้งที่ infrastructure ของ run แรกถูกเปลี่ยนจริงไปแล้ว หรือ state file อาจ corrupt ระหว่างเขียน

lock เปลี่ยน race นี้ให้กลายเป็นคิว apply ตัวแรกที่เริ่มจะได้ lock ไป apply ตัวที่สองจะรอให้ lock ถูกปล่อย หรือ fail ทันทีพร้อม error ที่ชัดเจน ขึ้นอยู่กับ flag ที่ใช้ ไม่ว่าทางไหน run ตัวที่สองก็ไม่มีทางเขียนทับ state ของ run แรกในระหว่างที่ run แรกยังทำงานอยู่

Terminal window
# A run that finds the state already locked fails fast with a clear message,
# rather than corrupting the state file
terraform apply
# Error: Error acquiring the state lock
flowchart LR
  eng1["Engineer A: terraform apply"] -->|acquires lock| s3["S3 state + lock file"]
  eng2["CI run: terraform apply"] -->|blocked, waits or fails| s3
  s3 -->|lock released after apply| eng2
Two applies race for the same S3-backed state; the lock lets one through and makes the other wait
ทำไม local terraform.tfstate ถึงพังลงทันทีที่มีเพื่อนร่วมทีมคนที่สองเริ่มรัน Terraform
use_lockfile = true บน S3 backend ให้อะไร และทำไมตอนนี้ถึงเป็นที่นิยมกว่า DynamoDB table
state locking ป้องกันอะไรจริง ๆ