From 70426d4164d6917d6d74c2812518edc2a0c42b2c Mon Sep 17 00:00:00 2001 From: Bryce Zuccaro Date: Fri, 14 Aug 2026 10:25:41 -0600 Subject: [PATCH] Add llama-watch Stream Deck plugin design spec --- .../specs/2026-08-14-llama-watch-design.md | 166 ++++++++++++++++++ 1 file changed, 166 insertions(+) create mode 100644 docs/superpowers/specs/2026-08-14-llama-watch-design.md diff --git a/docs/superpowers/specs/2026-08-14-llama-watch-design.md b/docs/superpowers/specs/2026-08-14-llama-watch-design.md new file mode 100644 index 0000000..ffa6fca --- /dev/null +++ b/docs/superpowers/specs/2026-08-14-llama-watch-design.md @@ -0,0 +1,166 @@ +# llama-watch — Stream Deck Plugin Design + +**Date:** 2026-08-14 +**Status:** Approved (pending spec review) + +## Summary + +A Stream Deck plugin ("llama-watch") that monitors a llama-swap instance +(`http://localhost:9292`). Two key-action types: + +1. **In-flight Monitor** — shows whether a specific model currently has a + request in flight, with a color-coded state (OFF / LOADING / IDLE / ACTIVE). +2. **GPU Graph** — renders a live line chart of a selected GPU metric + (utilization %, VRAM %, temperature, power draw, or fan speed) on the key, + for a specific GPU or the "All GPUs" aggregate. + +Target device: standard Stream Deck (72x72 px keys), rendered at 144x144 for +crispness. macOS only. Local install, with eventual Elgato Marketplace +submission as a design goal. + +## Approach + +Official Elgato JS SDK (`streamdeck-jssdk`) + `streamdeck-cli` tooling, a +TypeScript Node.js plugin, and `@napi-rs/canvas` for image rendering. The +plugin runs locally (launched by the Stream Deck app) and talks directly to +the llama-swap instance. + +## Data Sources + +llama-swap exposes everything needed (verified live against +`localhost:9292`): + +- **`GET /api/events`** — SSE stream used by the web UI. Emits `inflight` + events (operations `snapshot` / `add` / `remove`, per-request detail: + model, timestamp, elapsed, bytes) and `modelStatus` events (per-model + state: `stopped`, `loading`, `ready`). Also emits `logData`, `uiConfig`, + `profileChanged`, `activity` events — ignored. +- **`GET /metrics`** — Prometheus text format. Relevant series: + - `llamaswap_gpu_util_percent{id,name,uuid}` + - `llamaswap_gpu_memory_util_percent{id,name,uuid}` + - `llamaswap_gpu_temperature_celsius{id,name,uuid}` + - `llamaswap_gpu_power_draw_watts{id,name,uuid}` + - `llamaswap_gpu_fan_speed_percent{id,name,uuid}` +- **`GET /v1/models`** — model list (with `status.value` loaded/unloaded), + used to populate the model dropdown in the property inspector. + +Upstream vLLM metrics (ports 10001/10003) are bound to localhost on the +server and are NOT reachable from the client Mac. Not used. + +## Architecture + +The plugin is a single Node process with three connections: + +1. WebSocket to the Stream Deck software (via `streamdeck-jssdk`). +2. Persistent SSE connection to `/api/events` (the `EventFeed`). +3. 5s-interval polling of `/metrics` (the `MetricsPoller`). + +### Modules + +- **`MetricsPoller`** (pure logic, testable) + - Polls `GET /metrics` every **5 seconds** (fixed). + - Parses the Prometheus text format into a lookup keyed by metric name × + GPU id. + - `getSnapshot(gpuSelector, metric)` resolves any selector × metric + combination into a number, including the "All GPUs" aggregate. + - Keeps a ring buffer of 60 samples (~5 minutes) per (gpu, metric). + - A missed/failed poll skips that sample (no gap in rendering beyond a + break in the line). +- **`EventFeed`** (pure logic, testable) + - Persistent SSE connection to `/api/events`. + - Reconnect with exponential backoff (1s → 30s max); the `inflight` + `snapshot` operation emitted on connect self-heals the tracker. + - Maintains per-model in-flight request counts and per-model state from + `modelStatus`. +- **Actions** + - `InflightMonitor` — one instance per monitored model. + - `GpuGraph` — one instance per GPU (or All GPUs) × metric. +- **Property inspectors** (`pi/*.html`) — settings UI for both actions. +- **Renderer** — draws 144x144 PNGs via `@napi-rs/canvas`. + +### Data Flow + +``` +llama-swap ──SSE /api/events──▶ EventFeed ──▶ per-model inflight counts + state + ──HTTP /metrics 5s──▶ MetricsPoller ──▶ ring buffers + aggregates + │ +Stream Deck app ◀──jssdk WS── Actions ────────┘ + ◀──144x144 PNG + title── Renderer +``` + +### Settings (shared + per action) + +- **Shared:** base URL (default `http://localhost:9292`), optional API + key (sent as a header when set; instance currently requires none). +- **InflightMonitor:** model (dropdown populated live from `/v1/models`). +- **GpuGraph:** GPU selector (dropdown from live GPU list, plus "All GPUs"), + metric (dropdown: Utilization %, VRAM %, Temperature, Power draw, Fan + speed). + +## Action 1 — In-flight Monitor + +- **States (rendered on key):** + - `stopped` / unloaded → grey background, model name, `OFF` + - `loading` → amber, `LOADING` + - `ready` + 0 requests → green, `IDLE` + - `ready` + ≥1 request → red, `ACTIVE` (subtle pulse, re-render ~2 Hz while + active) +- Model short name displayed on the key; if no model selected, placeholder + with an edit hint. +- **Press:** opens `http:///ui` in the default browser (spawns + `open` on macOS). + +## Action 2 — GPU Graph + +- **Rendering (graph dominates the 72x72 key):** + - Big line chart of the last 60 samples (5 min @ 5s) filling the key. + - Dark background, bright line, subtle filled gradient under the curve. + - Small current-value label (e.g. `67%`, `336W`, `55°C`) at top; metric + name (e.g. `UTIL`, `PWR`) at bottom. +- **Scale:** % metrics fixed 0–100; temperature auto 0–100°C; power + auto-scaled to observed max. +- **Colors:** severity-based — green → amber → red. Temperature red at + ≥80°C; util/power/fan scale with level. +- **Aggregates ("All GPUs"):** Utilization/VRAM/Fan → average; Temperature → + max; Power → sum. +- **Render-on-change:** idle/flat history does not re-render every poll. +- **Press:** opens the web UI (same as Action 1). + +## Error Handling + +- llama-swap unreachable → dark grey key with `!!` and `OFFLINE`; retry with + backoff; auto-recover when server returns. +- SSE drop → reconnect with backoff (1s → 30s); self-healing via `snapshot`. +- `/metrics` poll timeout → skip sample. +- Per-button error states — a bad setting on one key does not affect others. + +## Packaging & Assets + +- Package UUID: `com.bryce.llamawatch` (reverse-DNS, unique). +- `streamdeck-cli` scaffolds project, builds TypeScript, packages + `.streamDeckPlugin` (zip) → installs via double-click into + `~/Library/Application Support/com.elgato.StreamDeck/Plugins/`. +- Marketplace-ready from the start: + - 256×256 plugin icon + action icons (incl. pressed states) at marketplace + spec sizes. + - Name "llama-watch", category e.g. "System & Monitoring". + - Privacy note: reads GPU metrics + request status from the user's own + llama-swap server; only base URL + optional API key stored in Stream + Deck's local settings; no data leaves the machine. + - No hardcoded secrets/URLs — base URL is user-editable. + - Manifest declares macOS only initially. + +## Testing + +- Unit tests with `node:test` for the Prometheus parser, inflight tracker, + and aggregate math (avg/max/sum), using **real captured fixtures** from + the server instance (metrics body, inflight snapshot, modelStatus payloads). +- Manual verification: both actions render correctly on the physical Stream + Deck; key press opens the web UI. + +## Out of Scope (YAGNI) + +- Marketplace submission itself (goal for later; assets/metadata prepared now). +- Auto-update plumbing. +- Windows/Linux support (later add-on; macOS declared in manifest). +- Additional metrics or display styles.