可觀測性:指標、日誌、追蹤
系統出事時要能回答「哪裡慢、為什麼」:指標看趨勢、日誌看細節、分散式追蹤看一個請求經過了哪些服務。再用 SLO(Service Level Objective,服務水準目標)和錯誤預算決定該先修問題還是上新功能。
看哪一種訊號
讓哪個變慢
×4
這個請求
74 ms
平常
74 ms
p50(中位數)
74 ms
p99
115 ms
在關鍵路徑上不在上面:有餘裕
這個請求花 74 ms。關鍵路徑是 gateway → auth → recs → ml-model → render:四個服務平行跑,頁面只等最慢的那一個。把 user-db 加速一倍什麼都不會變;加速 ml-model 才有用。
假設:一個商品頁請求經過閘道——先驗證身分,再平行呼叫 user、catalog、recs、inventory(各有自己的後端),最後 render。每個 span 的時間以對數常態分布起伏約 ±35%;2,000 個請求代表一分鐘的流量,超過 150 ms 就是閘道逾時、算錯誤。預算:30 天 99.9%、平常有 0.02% 的錯誤、事故在第 10 天;警報門檻取自 Google 的 SRE(Site Reliability Engineering)實務手冊。
亮起來的是這一步執行的程式碼
type Span = { traceId: string; spanId: number; parentId: number | null; name: string; start: number; end: number }; class Tracer { spans: Span[] = []; logs: string[] = []; private nextId = 1; constructor(private traceId: string) {} start(name: string, parent: Span | null, now: number): Span { const span = { traceId: this.traceId, spanId: this.nextId++, parentId: parent ? parent.spanId : null, name, start: now, end: now }; this.spans.push(span); return span; } finish(span: Span, now: number): void { span.end = now; } // One JSON object per line, carrying the trace ID so logs join up with traces. log(span: Span, level: string, msg: string, now: number): void { this.logs.push(JSON.stringify({ t: now, trace_id: span.traceId, span_id: span.spanId, level, msg })); }} // Walk back from the end: the child that finished last before the cursor// is what the parent was waiting on; before it started, the next one was.function criticalPath(spans: Span[]): string[] { const children = new Map<number, Span[]>(); for (const span of spans) { if (span.parentId === null) continue; if (!children.has(span.parentId)) children.set(span.parentId, []); children.get(span.parentId)!.push(span); } const path: string[] = []; const walk = (span: Span) => { path.push(span.name); let cursor = span.end; const kids = (children.get(span.spanId) ?? []).slice().sort((a, b) => b.end - a.end); for (const child of kids) { if (child.end <= cursor) { walk(child); cursor = child.start; } } }; walk(spans.find((span) => span.parentId === null)!); return path;} // Count a value into the first bucket whose upper bound holds it; the// last count is for everything above the largest bound.function observe(bounds: number[], counts: number[], value: number): void { let i = 0; while (i < bounds.length && value > bounds[i]) i++; counts[i]++;} // The q-quantile from bucket counts alone, interpolating inside the bucket.function histogramQuantile(bounds: number[], counts: number[], q: number): number { const rank = q * counts.reduce((a, b) => a + b, 0); let cumulative = 0; for (let i = 0; i < counts.length; i++) { if (counts[i] > 0 && cumulative + counts[i] >= rank) { if (i === bounds.length) return bounds[bounds.length - 1]; const lower = i === 0 ? 0 : bounds[i - 1]; return lower + ((bounds[i] - lower) * (rank - cumulative)) / counts[i]; } cumulative += counts[i]; } return bounds[bounds.length - 1];} // An incident of `minutes` at a steady error rate, against an SLO over a window.function errorBudget(slo: number, windowDays: number, errorRate: number, minutes: number) { const budget = 1 - slo; const windowMinutes = windowDays * 24 * 60; const consumed = (errorRate * minutes) / (budget * windowMinutes); const burnRate = errorRate / budget; const over = (window: number) => (burnRate * Math.min(minutes, window)) / window; const page = over(60) >= 14.4 && over(5) >= 14.4; const ticket = over(360) >= 6 && over(30) >= 6; return { consumed, burnRate, page, ticket };}模型假設與範圍
- 這是可重現的教學模型;延遲、容量、故障率與工作負載是設定或樣本,不能直接當作正式系統的效能承諾。
- 固定事件樣本展示 metrics、logs、traces 與取樣;沒有 telemetry collector,也未估算真實儲存基數、成本與告警噪音。
什麼時候用
- 指標回答「有沒有問題」:RED(請求率、錯誤率、處理時間)便宜、可以長期保存、適合發警報。
- 追蹤回答「慢在哪裡」:一個請求跨好幾個服務時,只有追蹤看得到誰在等誰、哪一段在關鍵路徑上。
- 日誌回答「到底發生什麼事」:結構化、帶 trace_id,才能從指標的異常跳到追蹤、再跳到那個請求的日誌。
- SLO(Service Level Objective,服務水準目標)建立在 SLI(Service Level Indicator,服務水準指標)上,例如「成功請求的比例」;對外承諾、違約要賠的是 SLA(Service Level Agreement,服務水準協議),通常比內部 SLO 寬鬆。
和其他主題的關係
- 延伸閱讀
- 容錯模式:逾時、重試、斷路器
時間與空間複雜度(Big O)
| 操作 | 平均 | 最差 |
|---|---|---|
| 記錄一個 span 或一行日誌 | O(1) | O(1) |
| 關鍵路徑 n 個 span;排序每個 span 的子節點 | O(n log n) | O(n log n) |
| 直方圖記一個值 B 個桶;用二分搜尋是 O(log B) | O(B) | O(B) |
| 從直方圖讀百分位 精確百分位要排序全部 N 個值:O(N log N) | O(B) | O(B) |
| 錯誤預算與燃燒速率 | O(1) | O(1) |
空間:O(n + B),一條追蹤的 n 個 span,加上直方圖的 B 個計數器;精確百分位則要 O(N)
和其他做法比
| p99(存全部 2,000 個值,16 KB) | 9 個桶(72 B) | 51 個 10 ms 桶(408 B) | |
|---|---|---|---|
| 正常 | 115 ms | 140 ms | 117 ms |
| user-db ×4 | 143 ms | 149 ms | 142 ms |
| catalog-db ×3 | 196 ms | 198 ms | 196 ms |
| cache ×30 | 167 ms | 180 ms | 167 ms |
同樣 2,000 個請求,p99 三種算法。存全部的值最準,但記憶體隨請求數成長、也不能跨伺服器相加;直方圖固定大小、可以相加,誤差取決於桶有多寬:粗桶在「正常」差了 24 ms,10 ms 的桶最多差幾毫秒。記憶體以每個值或計數 8 bytes 計。
真實世界裡的它
- OpenTelemetry 是追蹤、指標、日誌的開放標準;Jaeger、Zipkin 顯示追蹤,W3C(World Wide Web Consortium)的
traceparent標頭把 trace ID 傳給下游。 - Prometheus 的 histogram 和
histogram_quantile就是上面「從桶算百分位」的做法;Uber 用關鍵路徑分析找出值得優化的服務。 - Google 的 SRE(Site Reliability Engineering)手冊定義了錯誤預算與多視窗燃燒速率警報:預算用完就暫停上新功能,先修可靠度。
取捨與陷阱
- 對平均值發警報:平均 80 ms 時,1% 的使用者可能等了 2 秒。看 p99,或者看超過期限的比例。
- 把百分位相加或平均:十台伺服器各自的 p99 平均起來不是整體的 p99;要先把直方圖的桶相加再算。
- 把使用者 ID、URL 之類放進指標的標籤,每個值都變成一條新的時間序列,監控系統會被撐爆(高基數);這些放在日誌和追蹤裡。
- 優化不在關鍵路徑上的服務,請求一毫秒都不會變快。