跳到主要內容

系統設計

把資料結構放大到好幾台機器

主題 · 可觀測性:指標、日誌、追蹤

可觀測性:指標、日誌、追蹤

系統出事時要能回答「哪裡慢、為什麼」:指標看趨勢、日誌看細節、分散式追蹤看一個請求經過了哪些服務。再用 SLO(Service Level Objective,服務水準目標)和錯誤預算決定該先修問題還是上新功能。

看哪一種訊號
讓哪個變慢
×4
這個請求
74 ms
平常
74 ms
p50(中位數)
74 ms
p99
115 ms
0 ms25 ms50 ms75 msgateway:0.0 ms → 74 ms(關鍵路徑)gateway74 msauth:5.0 ms → 13 ms(關鍵路徑)auth8.0 msuser:13 ms → 31 msuser18 msuser-db:19 ms → 31 msuser-db12 mscatalog:13 ms → 43 mscatalog30 mscatalog-db:18 ms → 43 mscatalog-db25 msrecs:13 ms → 56 ms(關鍵路徑)recs43 msml-model:21 ms → 56 ms(關鍵路徑)ml-model35 msinventory:13 ms → 19 msinventory6.0 mscache:17 ms → 19 mscache2.0 msrender:56 ms → 71 ms(關鍵路徑)render15 ms
在關鍵路徑上不在上面:有餘裕

這個請求花 74 ms。關鍵路徑是 gateway → auth → recs → ml-model → render:四個服務平行跑,頁面只等最慢的那一個。把 user-db 加速一倍什麼都不會變;加速 ml-model 才有用。

假設:一個商品頁請求經過閘道——先驗證身分,再平行呼叫 user、catalog、recs、inventory(各有自己的後端),最後 render。每個 span 的時間以對數常態分布起伏約 ±35%;2,000 個請求代表一分鐘的流量,超過 150 ms 就是閘道逾時、算錯誤。預算:30 天 99.9%、平常有 0.02% 的錯誤、事故在第 10 天;警報門檻取自 Google 的 SRE(Site Reliability Engineering)實務手冊。

亮起來的是這一步執行的程式碼
type Span = { traceId: string; spanId: number; parentId: number | null; name: string; start: number; end: number };
class Tracer {
spans: Span[] = [];
logs: string[] = [];
private nextId = 1;
constructor(private traceId: string) {}
start(name: string, parent: Span | null, now: number): Span {
const span = { traceId: this.traceId, spanId: this.nextId++, parentId: parent ? parent.spanId : null,
name, start: now, end: now };
this.spans.push(span);
return span;
}
finish(span: Span, now: number): void {
span.end = now;
}
// One JSON object per line, carrying the trace ID so logs join up with traces.
log(span: Span, level: string, msg: string, now: number): void {
this.logs.push(JSON.stringify({ t: now, trace_id: span.traceId, span_id: span.spanId, level, msg }));
}
}
// Walk back from the end: the child that finished last before the cursor
// is what the parent was waiting on; before it started, the next one was.
function criticalPath(spans: Span[]): string[] {
const children = new Map<number, Span[]>();
for (const span of spans) {
if (span.parentId === null) continue;
if (!children.has(span.parentId)) children.set(span.parentId, []);
children.get(span.parentId)!.push(span);
}
const path: string[] = [];
const walk = (span: Span) => {
path.push(span.name);
let cursor = span.end;
const kids = (children.get(span.spanId) ?? []).slice().sort((a, b) => b.end - a.end);
for (const child of kids) {
if (child.end <= cursor) {
walk(child);
cursor = child.start;
}
}
};
walk(spans.find((span) => span.parentId === null)!);
return path;
}
// Count a value into the first bucket whose upper bound holds it; the
// last count is for everything above the largest bound.
function observe(bounds: number[], counts: number[], value: number): void {
let i = 0;
while (i < bounds.length && value > bounds[i]) i++;
counts[i]++;
}
// The q-quantile from bucket counts alone, interpolating inside the bucket.
function histogramQuantile(bounds: number[], counts: number[], q: number): number {
const rank = q * counts.reduce((a, b) => a + b, 0);
let cumulative = 0;
for (let i = 0; i < counts.length; i++) {
if (counts[i] > 0 && cumulative + counts[i] >= rank) {
if (i === bounds.length) return bounds[bounds.length - 1];
const lower = i === 0 ? 0 : bounds[i - 1];
return lower + ((bounds[i] - lower) * (rank - cumulative)) / counts[i];
}
cumulative += counts[i];
}
return bounds[bounds.length - 1];
}
// An incident of `minutes` at a steady error rate, against an SLO over a window.
function errorBudget(slo: number, windowDays: number, errorRate: number, minutes: number) {
const budget = 1 - slo;
const windowMinutes = windowDays * 24 * 60;
const consumed = (errorRate * minutes) / (budget * windowMinutes);
const burnRate = errorRate / budget;
const over = (window: number) => (burnRate * Math.min(minutes, window)) / window;
const page = over(60) >= 14.4 && over(5) >= 14.4;
const ticket = over(360) >= 6 && over(30) >= 6;
return { consumed, burnRate, page, ticket };
}

模型假設與範圍

  • 這是可重現的教學模型;延遲、容量、故障率與工作負載是設定或樣本,不能直接當作正式系統的效能承諾。
  • 固定事件樣本展示 metrics、logs、traces 與取樣;沒有 telemetry collector,也未估算真實儲存基數、成本與告警噪音。

什麼時候用

  • 指標回答「有沒有問題」:RED(請求率、錯誤率、處理時間)便宜、可以長期保存、適合發警報。
  • 追蹤回答「慢在哪裡」:一個請求跨好幾個服務時,只有追蹤看得到誰在等誰、哪一段在關鍵路徑上。
  • 日誌回答「到底發生什麼事」:結構化、帶 trace_id,才能從指標的異常跳到追蹤、再跳到那個請求的日誌。
  • SLO(Service Level Objective,服務水準目標)建立在 SLI(Service Level Indicator,服務水準指標)上,例如「成功請求的比例」;對外承諾、違約要賠的是 SLA(Service Level Agreement,服務水準協議),通常比內部 SLO 寬鬆。

和其他主題的關係

時間與空間複雜度(Big O)

操作平均最差
記錄一個 span 或一行日誌O(1)O(1)
關鍵路徑
n 個 span;排序每個 span 的子節點
O(n log n)O(n log n)
直方圖記一個值
B 個桶;用二分搜尋是 O(log B)
O(B)O(B)
從直方圖讀百分位
精確百分位要排序全部 N 個值:O(N log N)
O(B)O(B)
錯誤預算與燃燒速率O(1)O(1)

空間:O(n + B),一條追蹤的 n 個 span,加上直方圖的 B 個計數器;精確百分位則要 O(N)

和其他做法比

p99(存全部 2,000 個值,16 KB)9 個桶(72 B)51 個 10 ms 桶(408 B)
正常115 ms140 ms117 ms
user-db ×4143 ms149 ms142 ms
catalog-db ×3196 ms198 ms196 ms
cache ×30167 ms180 ms167 ms

同樣 2,000 個請求,p99 三種算法。存全部的值最準,但記憶體隨請求數成長、也不能跨伺服器相加;直方圖固定大小、可以相加,誤差取決於桶有多寬:粗桶在「正常」差了 24 ms,10 ms 的桶最多差幾毫秒。記憶體以每個值或計數 8 bytes 計。

真實世界裡的它

  • OpenTelemetry 是追蹤、指標、日誌的開放標準;Jaeger、Zipkin 顯示追蹤,W3C(World Wide Web Consortium)的 traceparent 標頭把 trace ID 傳給下游。
  • Prometheus 的 histogram 和 histogram_quantile 就是上面「從桶算百分位」的做法;Uber 用關鍵路徑分析找出值得優化的服務。
  • Google 的 SRE(Site Reliability Engineering)手冊定義了錯誤預算與多視窗燃燒速率警報:預算用完就暫停上新功能,先修可靠度。

取捨與陷阱

  • 對平均值發警報:平均 80 ms 時,1% 的使用者可能等了 2 秒。看 p99,或者看超過期限的比例。
  • 把百分位相加或平均:十台伺服器各自的 p99 平均起來不是整體的 p99;要先把直方圖的桶相加再算。
  • 把使用者 ID、URL 之類放進指標的標籤,每個值都變成一條新的時間序列,監控系統會被撐爆(高基數);這些放在日誌和追蹤裡。
  • 優化不在關鍵路徑上的服務,請求一毫秒都不會變快。