跳到主要內容

系統設計

把資料結構放大到好幾台機器

主題 · API 閘道

API 閘道

API(Application Programming Interface,應用程式介面)前面的閘道,是所有外部請求的單一入口:驗證身分、限流、把請求路由到後面的服務,讓每個服務不必各自處理一遍。

情境
1/5

請求:GET /orders/42,帶著 Authorization: alice.<signature>

這裡以 HMAC-SHA-256 簽署 UTF-8 使用者名稱。簽章金鑰是公開的教學用資料;真正的簽發端會保密。Bearer token 代表持有憑證,因此被竊取的有效 token 仍可能被使用。

用戶端API 閘道驗證限流路由轉送/重試使用者服務訂單服務商品服務

驗證 HMAC-SHA-256 簽章:這個 token 是可信任的簽發端為 alice 簽發的,將它當作 alice 的憑證。

回應
200
總時間
47 ms
打到後端幾次
1
亮起來的是這一步執行的程式碼
import { createHmac, timingSafeEqual } from "node:crypto";
// Public teaching fixture; real issuers keep the key private.
const SECRET = "gw-secret";
const GATEWAY_MS = 2, TIMEOUT_MS = 1000, MAX_ATTEMPTS = 2;
const CAPACITY = 5, REFILL_PER_SEC = 2; // token bucket per user
const SERVICES = ["users", "orders", "products"];
type Backend = (service: string, attempt: number) => number; // latency, ms
type Reply = { status: number; ms: number; attempts: number };
function signToken(user: string): string {
return user + "." + createHmac("sha256", SECRET).update(user, "utf8").digest("hex");
}
class Gateway {
private buckets = new Map<string, { tokens: number; at: number }>();
constructor(private backend: Backend) {}
handle(token: string, path: string, now: number): Reply {
const user = this.verify(token);
if (user === null) return { status: 401, ms: GATEWAY_MS, attempts: 0 };
if (!this.allow(user, now)) return { status: 429, ms: GATEWAY_MS, attempts: 0 };
const service = path.split("/")[1];
if (!SERVICES.includes(service)) return { status: 404, ms: GATEWAY_MS, attempts: 0 };
return this.forward(service);
}
// Backend for frontend: one call from the app, three behind it.
home(token: string, now: number): Reply {
const user = this.verify(token);
if (user === null) return { status: 401, ms: GATEWAY_MS, attempts: 0 };
if (!this.allow(user, now)) return { status: 429, ms: GATEWAY_MS, attempts: 0 };
const replies = SERVICES.map((s) => this.forward(s));
// Sent concurrently, so the slowest sets the time.
const ms = GATEWAY_MS + Math.max(...replies.map((r) => r.ms - GATEWAY_MS));
const failed = replies.some((r) => r.status !== 200);
const attempts = replies.reduce((sum, r) => sum + r.attempts, 0);
return { status: failed ? 504 : 200, ms, attempts };
}
private verify(token: string): string | null {
const dot = token.lastIndexOf(".");
if (dot <= 0) return null;
const user = token.slice(0, dot);
const supplied = token.slice(dot + 1);
if (supplied.length !== 64 || !/^[0-9a-f]{64}$/.test(supplied)) return null;
const expected = createHmac("sha256", SECRET).update(user, "utf8").digest();
return timingSafeEqual(Buffer.from(supplied, "hex"), expected) ? user : null;
}
private allow(user: string, now: number): boolean {
const b = this.buckets.get(user) ?? { tokens: CAPACITY, at: now };
b.tokens = Math.min(CAPACITY, b.tokens + (now - b.at) * REFILL_PER_SEC);
b.at = now;
this.buckets.set(user, b);
if (b.tokens < 1) return false;
b.tokens -= 1;
return true;
}
private forward(service: string): Reply {
let ms = GATEWAY_MS;
for (let attempt = 1; attempt <= MAX_ATTEMPTS; attempt++) {
const took = this.backend(service, attempt);
if (took <= TIMEOUT_MS) return { status: 200, ms: ms + took, attempts: attempt };
ms += TIMEOUT_MS;
}
return { status: 504, ms, attempts: MAX_ATTEMPTS };
}
}

模型假設與範圍

  • 這是可重現的教學模型;延遲、容量、故障率與工作負載是設定或樣本,不能直接當作正式系統的效能承諾。
  • 驗證、路由、限流與一次重試採範例協定;延遲是假設。未實作完整 OAuth/JWT、金鑰輪替或 production gateway。

什麼時候用

  • 後面有好幾個服務,而驗證身分、限流、記錄、CORS(Cross-Origin Resource Sharing,跨來源資源共用)這些每個服務都要做的事,不想在每個服務裡各寫一份。
  • 用戶端只該知道一個網址:服務拆開、合併、搬家時,外面的 App 不必跟著改。
  • 手機 App 要一次拿到好幾個服務的資料:用 BFF(Backend for Frontend)把多趟往返變成一趟。

和其他主題的關係

由這些組成
限流器負載平衡

出現在這些架構裡

時間與空間複雜度(Big O)

操作平均最差
驗證 token
簽章長度固定;不必查資料庫
O(1)O(1)
限流(每個使用者一個 token bucket)O(1)O(1)
找路由:查表O(1)O(1)
找路由:依序比對前綴
R 是路由數
O(R)O(R)
轉送與重試
最多 k 次嘗試;每次逾時都把等待時間加上去
O(1)O(k)

空間:O(U),每個有在使用的使用者一個 bucket

Big O 實測:n 變大時步數怎麼長

數的是:找到一個請求的路由要比較幾次(n 是路由數)

Big On = 10n = 100n = 1,000n = 10,000成長倍數:實測(理論)
依序比對每個路徑前綴O(n)5.550.55015,001×909 (×1,000)
用第一段路徑查表O(1)1111×1.0 (×1.0)

閘道每個請求都要找路由,服務一多,逐條比對的成本就跟著長;用路徑的第一段(或字首樹)查表則幾乎不變。

和其他做法比

打到後端的請求驗證/限流的程式要維護幾份平均延遲p99 延遲
沒有閘道:每個服務自己檢查10,000 (100%)3104 ms3000 ms
有閘道7,436 (74%)164 ms1033 ms

同一串 10,000 個請求:4% 帶著無效的 token,一個使用者送了全部的 25%;後端平均約 40 ms,2% 的呼叫會卡 3 秒。閘道在門口擋掉 2,717 個請求,服務根本不用處理;它多花 2 ms,但卡住的呼叫在 1000 ms 就放棄重試,所以尾端延遲反而短很多。打到後端的數字包含重試。兩種配置的平均與 p99 都只統計通過驗證及限流的 7,283 個原始請求,不含 401/429;有閘道時也包含最終的 504,每個原始請求只算一個樣本。p99 是第 99 百分位(percentile):把這些樣本的延遲由快到慢排好,取 99% 位置的值。

載入首頁要多久
App 依序呼叫三個服務455 ms
App 呼叫一次 BFF,BFF 同時呼叫三個202 ms

假設手機網路每趟往返 80 ms,三個服務分別要 35、120、60 ms。手機上最貴的是往返次數,不是伺服器的處理時間;BFF 把三趟變成一趟,資料中心裡的三個呼叫又能同時進行。

真實世界裡的它

  • Kong、Envoy、NGINX、AWS API Gateway、Apigee:都是做這件事的現成產品。
  • Netflix 的 Zuul 是早期知名的例子:每個裝置類型有自己的 BFF。
  • Kubernetes 的 Ingress 和服務網格(Istio)把一部分閘道的工作搬進叢集裡。

取捨與陷阱

  • 閘道是單點:它掛了,全部服務都連不到。要多台、放在負載平衡後面,而且本身不能存狀態。
  • 只有不會改變資料的請求才能自動重試;重試「扣款」或「下單」要搭配冪等鍵,否則可能做兩次。
  • 把商業邏輯塞進閘道,它就會變成一個誰都不敢改的大怪物。閘道只做跨服務共通的事。