Go Diagnosis
8/23/26Less than 1 minute
Go Diagnosis
Go 服务排查先用指标确认瓶颈类型和时间窗口,再选择 profile、trace、日志或 goroutine dump。不要先采集所有数据,再寻找问题。
Symptom to Evidence
| Symptom | First evidence | Common cause |
|---|---|---|
| CPU high | CPU profile | 热循环、序列化、正则、GC assist |
| Memory growth | heap / alloc profile | retained object、缓存无界、分配过多 |
| Goroutine growth | goroutine profile | channel、锁、I/O 或取消泄漏 |
| Latency spikes | trace + block/mutex profile | 排队、锁竞争、GC、下游等待 |
| Throughput plateau | saturation metrics | CPU、连接池、并发上限或单点串行 |
Investigation Loop
- 确认用户影响、版本、实例和时间范围。
- 比较正常与异常窗口,而不是只看异常快照。
- 采集最能验证当前假设的 profile。
- 从 top function 进入调用图和源码,区分 on-CPU 与 waiting。
- 修复后用相同 workload 比较 latency、throughput、CPU、memory 和 allocations。
Leak Checklist
- goroutine 是否在等待永远不会关闭的 channel?
- context cancel 是否在所有路径调用?
- ticker、timer、response body 和文件是否释放?
- map、queue 或 cache 是否有明确上限?
- 慢消费者出现时,生产者怎样背压或丢弃?
