# Usage Stats + GPU Combinations **Date:** 2026-08-14 **Status:** Approved (design) ## Summary Two extensions to the existing llama-watch actions: 1. The In-Flight Monitor action gains a **Usage stats** display mode showing token/request totals and the generation-speed P95 from llama-swap's `/api/metrics/stats` endpoint. 2. The GPU Graph action gains **GPU combinations**: a key can aggregate an arbitrary subset of GPUs (e.g. only the two Blackwells). ## Data Source (verified against llama-swap source) - `GET /api/metrics/stats` (no query param) → global aggregates. - `GET /api/metrics/stats?model=` → per-model aggregates. - Response fields used: `total_requests`, `total_input_tokens`, `total_output_tokens`, `gen_histogram.p95`. - `p95` is a **tokens/sec generation-speed percentile**, not latency. - Totals cover llama-swap's in-memory activity retention (default ~1000 requests), not lifetime counters. - The `/api/events` SSE `activity` event (payload `{"id": N}`) fires per completed request and is used as a "re-poll stats" trigger. It carries no data itself. - The Prometheus `/metrics` endpoint exposes no token/request/latency data and is not used for this feature. ## 1. Usage stats data layer - **`src/lib/stats.ts`** (new): - `interface UsageStats { totalRequests: number; totalInputTokens: number; totalOutputTokens: number; genP95: number }`. - `parseStats(json: unknown): UsageStats | null` — defensive; returns `null` on malformed input; missing fields default to `0`. - `fetchStats(cfg: LlamaSwapConfig, modelId: string): Promise` — fetches `/api/metrics/stats` (with `?model=` when `modelId` is not `"all"`), returns parsed stats or `null` on HTTP error/parse failure. - **`src/lib/stats-cache.ts`** (new): - `StatsCache` holds `Set` of registered model keys (`"all"` or a model id) and a `Map` of last-known values. - One shared 5s interval polls every registered key via `fetchStats`. - `register(key)` / `unregister(key)` start/stop the interval as needed (no ref counting — each key registers/unregisters once). - `get(key): UsageStats | undefined`. - `refresh()` re-polls all registered keys. - Change listener invoked after each poll completes (success or failure). - Fetch failure keeps the last-known value. - **`src/lib/event-feed.ts`**: `decodeEvent` learns the `activity` event — `{ type: "activity", id: number }`. `FeedEvent` (in `inflight-tracker.ts`) gains the `activity` variant. - **`src/lib/runtime.ts`**: runtime owns a `StatsCache` (created alongside feed/poller, reset on config change). Feed `activity` events trigger `statsCache.refresh()` throttled to at most once per 2s. Runtime exposes: - `getStats(modelId: string): UsageStats | undefined` - `watchStats(modelId: string): () => void` (register; returns unsubscribe). - Stats changes flow through the existing `subscribe`/`emit` mechanism so actions re-render on the normal tick. ## 2. In-Flight action modes - New settings (per key, in `inflight-monitor.ts`): - `display: "count" | "usage"` (default `count`). - `primaryStat: "requests" | "input_tokens" | "output_tokens" | "gen_p95"` (default `requests`). - Model setting semantics: - Model select (PI datasource) gains an **"All models"** option, value `all`. - Count mode: `all` → total in-flight across all models; label `ALL MODELS`. - Usage mode: `all` → global stats; otherwise per-model stats. - **Count mode**: unchanged rendering (count number, spark, `IDLE`/ `LOADING`/`OFF`/`OFFLINE`). The 1s spark sampler runs only in count mode. - **Usage mode** (`renderUsage` in `render.ts`, dark `#10131a` bg): - top: model short name or `ALL MODELS` (7px). - center (24px bold): the `primaryStat` value, compact-formatted — `< 1000` integer, else `x.xk`, else `x.xM` (`genP95` shows integer t/s). - three small rows (7px): the remaining three stats, labeled `REQ`, `IN`, `OUT`, `P95`. - offline → `OFFLINE` overlay; stats unknown → `--`. - Press opens the UI (unchanged). ## 3. GPU combinations - New setting `gpuCombo?: string` — comma-separated GPU ids, e.g. `"0,2"`. When non-empty it overrides `gpuId`. - Pure helper in `metrics-poller.ts` (exported, tested): `combineSeries(histories: number[][], kind): { value?: number; history: number[] }` — element-wise across the per-GPU rings, aligned by index, applying the existing aggregate rules: `util_percent`/`memory_util_percent`/`fan` → average, `temperature` → max, `power` → sum. A trailing `undefined` value from an empty ring leaves that element undefined. - `gpu-graph.ts` render: for a combo, build per-GPU histories via `poller.getHistory(id, metric)`, combine, use the combined value/history. Unknown ids are filtered out; if none remain → `--`. - Label: `GPU 0+2` (ids joined with `+`). ## 4. Property inspectors - `ui/inflight.html`: add "Display" select (`count`/`usage`) and "Primary stat" select. Both always visible; `primaryStat` has no effect in count mode. - `src/lib/datasources.ts`: the `models` datasource prepends `{ label: "All models", value: "all" }` to its items (so the model select offers both the aggregate and individual models). - `ui/gpu.html`: add "GPU combination (comma-separated ids)" textfield, optional, placeholder `e.g. 0,2`. ## Error Handling / Edge Cases - Stats fetch error → keep last-known values; no `OFFLINE` flip (offline is driven by feed status as today). - Server offline → existing `OFFLINE` overlay in both modes. - Stats not yet loaded → `--` placeholders. - Activity-triggered refresh throttled to 1 per 2s. - Combo with no valid ids → `--`. - No manifest change (both features extend existing actions). ## Testing - `parseStats`: happy path; malformed input; missing fields → `0`. - `StatsCache`: register/unregister lifecycle (interval start/stop), single poll per model, values cached, refresh on demand, failure keeps last value. Uses an injected fetch function. - `decodeEvent`: `activity` event decodes to `{ type: "activity", id }`. - `renderUsage`: big primary stat, small labeled rows, compact formatting (`52.1k`), `--` when no data, `ALL MODELS` label, offline overlay. - `combineSeries`: avg/max/sum per metric; empty/mismatched histories. - In-Flight action: `all` count sums models.