ข้ามไปยังเนื้อหา

Goroutines

Goroutine คือฟังก์ชันที่ทำงานพร้อมกันกับ goroutines อื่น ๆ ในพื้นที่ address เดียวกัน คุณสร้างได้ด้วย keyword เดียว:

go f(args)

นั่นคือ syntax ทั้งหมด go จัดตาราง f ให้รัน concurrent แล้ว goroutine ที่เรียกก็ดำเนินต่อทันทีโดยไม่รอให้ f return

OS thread โดยทั่วไปเริ่มต้นด้วย stack ขนาดคงที่ 1–8 MB แต่ goroutine เริ่มต้นด้วย stack ขนาดประมาณ 2–8 KB ที่ขยายและหดตัวได้อัตโนมัติตามต้องการ Go runtime จัดสรร goroutines ลง OS threads โดยใช้ M:N scheduler (M goroutines บน N threads) Context switch ระหว่าง goroutines เกิดขึ้นใน user-space และถูกกว่า kernel thread switch หลายเท่า

ในระบบ production การมี goroutines หลายหมื่นตัวรันพร้อมกันเป็นเรื่องปกติ

Scheduler ทำงานบนสามแนวคิด:

สัญลักษณ์ความหมาย
GGoroutine — หน่วยงาน concurrent
MMachine — OS thread
PProcessor — scheduling context (มี GOMAXPROCS ตัว)

แต่ละ P มี run queue ของ Gs M ต้องถือ P ถึงจะรัน G ได้ เมื่อ G บล็อกบน I/O M ที่รันอยู่จะปล่อย P เพื่อให้ M อื่นรับช่วงต่อ

sync.WaitGroup เป็นวิธีมาตรฐานในการรอ goroutines จำนวนคงที่ให้เสร็จ:

  • wg.Add(n) — เพิ่ม counter ขึ้น n (เสมอก่อน go statement)
  • wg.Done() — ลด counter ลง 1 (เรียกด้วย defer ใน goroutine)
  • wg.Wait() — บล็อกจนกว่า counter จะถึงศูนย์
package main
import (
"fmt"
"sync"
)
func worker(id int, results []string, wg *sync.WaitGroup) {
defer wg.Done()
sum := 0
for j := 1; j <= 5; j++ {
sum += j
}
results[id-1] = fmt.Sprintf("worker %d: sum 1..5 = %d", id, sum)
}
func main() {
var wg sync.WaitGroup
results := make([]string, 3)
for i := 1; i <= 3; i++ {
wg.Add(1)
go worker(i, results, &wg)
}
wg.Wait()
for _, r := range results {
fmt.Println(r)
}
fmt.Println("all workers done")
}
GMP Model:
G G G G G G G G ← goroutines (อาจมีหลักแสน)
╲ │ ╱ ╲ │ ╱
[P1] [P2] ← P (= GOMAXPROCS, default = #CPU)
│ │
[M1] [M2] ← M (OS thread)
│ │
CPU0 CPU1
เมื่อ G block บน I/O:
G ──block──→ waiting queue
M release P → P ถูก M อื่นรับ → G อื่นรัน
Work stealing:
P1 run queue ว่าง → steal G จาก P2 run queue
สิ่งที่ได้ประโยชน์ต้นทุน
goroutine (2-8 KB stack)สร้างได้หลักแสนตัวscheduler overhead ต่อ goroutine
dynamic stack growthไม่เสีย memory ล่วงหน้าstack copy cost เมื่อขยาย
M:N schedulerไม่ block OS thread บน I/Oscheduler complexity, preemption cost
work stealingCPU utilization สูงcache miss เมื่อ goroutine ย้าย P
  • goroutine = OS thread — goroutine stack เริ่มต้น 2-8 KB, OS thread 1-8 MB
  • สร้าง goroutine ฟรีทั้งหมด — มี ~4 KB allocation + scheduler bookkeeping ต่อ goroutine
  • GOMAXPROCS มากขึ้น = เร็วขึ้นเสมอ — I/O-bound workload ได้ประโยชน์น้อยกว่า CPU-bound
  • goroutine leak ไม่เป็นปัญหา — goroutine ที่ไม่มีทางออกคือ memory leak

💡 ตัวอย่างจากของจริง

Kubernetes controller manager มี goroutine ต่อ resource type (Pod, Service, Deployment ฯลฯ)

etcd ใช้ goroutine ต่อ peer connection สำหรับ raft consensus algorithm

Prometheus เปิด goroutine ต่อ scrape job — scrape หลักพัน target พร้อมกัน

ขนาด stack เริ่มต้นโดยประมาณของ goroutine ใหม่คือเท่าไหร่?
ต้องเรียก wg.Add(1) ที่ไหนสัมพันธ์กับ go statement?
เกิดอะไรขึ้นถ้า main return ขณะที่ goroutines ยังรันอยู่?
GOMAXPROCS ควบคุมอะไร?