7.1 KiB
llama-watch — Stream Deck Plugin Design
Date: 2026-08-14 Status: Approved (pending spec review)
Summary
A Stream Deck plugin ("llama-watch") that monitors a llama-swap instance
(http://localhost:9292). Two key-action types:
- In-flight Monitor — shows whether a specific model currently has a request in flight, with a color-coded state (OFF / LOADING / IDLE / ACTIVE).
- GPU Graph — renders a live line chart of a selected GPU metric (utilization %, VRAM %, temperature, power draw, or fan speed) on the key, for a specific GPU or the "All GPUs" aggregate.
Target device: standard Stream Deck (72x72 px keys), rendered at 144x144 for crispness. macOS only. Local install, with eventual Elgato Marketplace submission as a design goal.
Approach
Official Elgato JS SDK (streamdeck-jssdk) + streamdeck-cli tooling, a
TypeScript Node.js plugin, and @napi-rs/canvas for image rendering. The
plugin runs locally (launched by the Stream Deck app) and talks directly to
the llama-swap instance.
Data Sources
llama-swap exposes everything needed (verified live against
localhost:9292):
GET /api/events— SSE stream used by the web UI. Emitsinflightevents (operationssnapshot/add/remove, per-request detail: model, timestamp, elapsed, bytes) andmodelStatusevents (per-model state:stopped,loading,ready). Also emitslogData,uiConfig,profileChanged,activityevents — ignored.GET /metrics— Prometheus text format. Relevant series:llamaswap_gpu_util_percent{id,name,uuid}llamaswap_gpu_memory_util_percent{id,name,uuid}llamaswap_gpu_temperature_celsius{id,name,uuid}llamaswap_gpu_power_draw_watts{id,name,uuid}llamaswap_gpu_fan_speed_percent{id,name,uuid}
GET /v1/models— model list (withstatus.valueloaded/unloaded), used to populate the model dropdown in the property inspector.
Upstream vLLM metrics (ports 10001/10003) are bound to localhost on the server and are NOT reachable from the client Mac. Not used.
Architecture
The plugin is a single Node process with three connections:
- WebSocket to the Stream Deck software (via
streamdeck-jssdk). - Persistent SSE connection to
/api/events(theEventFeed). - 5s-interval polling of
/metrics(theMetricsPoller).
Modules
MetricsPoller(pure logic, testable)- Polls
GET /metricsevery 5 seconds (fixed). - Parses the Prometheus text format into a lookup keyed by metric name × GPU id.
getSnapshot(gpuSelector, metric)resolves any selector × metric combination into a number, including the "All GPUs" aggregate.- Keeps a ring buffer of 60 samples (~5 minutes) per (gpu, metric).
- A missed/failed poll skips that sample (no gap in rendering beyond a break in the line).
- Polls
EventFeed(pure logic, testable)- Persistent SSE connection to
/api/events. - Reconnect with exponential backoff (1s → 30s max); the
inflightsnapshotoperation emitted on connect self-heals the tracker. - Maintains per-model in-flight request counts and per-model state from
modelStatus.
- Persistent SSE connection to
- Actions
InflightMonitor— one instance per monitored model.GpuGraph— one instance per GPU (or All GPUs) × metric.
- Property inspectors (
pi/*.html) — settings UI for both actions. - Renderer — draws 144x144 PNGs via
@napi-rs/canvas.
Data Flow
llama-swap ──SSE /api/events──▶ EventFeed ──▶ per-model inflight counts + state
──HTTP /metrics 5s──▶ MetricsPoller ──▶ ring buffers + aggregates
│
Stream Deck app ◀──jssdk WS── Actions ────────┘
◀──144x144 PNG + title── Renderer
Settings (shared + per action)
- Shared: base URL (default
http://localhost:9292), optional API key (sent as a header when set; instance currently requires none). - InflightMonitor: model (dropdown populated live from
/v1/models). - GpuGraph: GPU selector (dropdown from live GPU list, plus "All GPUs"), metric (dropdown: Utilization %, VRAM %, Temperature, Power draw, Fan speed).
Action 1 — In-flight Monitor
- States (rendered on key):
stopped/ unloaded → grey background, model name,OFFloading→ amber,LOADINGready+ 0 requests → green,IDLEready+ ≥1 request → red,ACTIVE(subtle pulse, re-render ~2 Hz while active)
- Model short name displayed on the key; if no model selected, placeholder with an edit hint.
- Press: opens
http://<base-url>/uiin the default browser (spawnsopenon macOS).
Action 2 — GPU Graph
- Rendering (graph dominates the 72x72 key):
- Big line chart of the last 60 samples (5 min @ 5s) filling the key.
- Dark background, bright line, subtle filled gradient under the curve.
- Small current-value label (e.g.
67%,336W,55°C) at top; metric name (e.g.UTIL,PWR) at bottom.
- Scale: % metrics fixed 0–100; temperature auto 0–100°C; power auto-scaled to observed max.
- Colors: severity-based — green → amber → red. Temperature red at ≥80°C; util/power/fan scale with level.
- Aggregates ("All GPUs"): Utilization/VRAM/Fan → average; Temperature → max; Power → sum.
- Render-on-change: idle/flat history does not re-render every poll.
- Press: opens the web UI (same as Action 1).
Error Handling
- llama-swap unreachable → dark grey key with
!!andOFFLINE; retry with backoff; auto-recover when server returns. - SSE drop → reconnect with backoff (1s → 30s); self-healing via
snapshot. /metricspoll timeout → skip sample.- Per-button error states — a bad setting on one key does not affect others.
Packaging & Assets
- Package UUID:
com.bryce.llamawatch(reverse-DNS, unique). streamdeck-cliscaffolds project, builds TypeScript, packages.streamDeckPlugin(zip) → installs via double-click into~/Library/Application Support/com.elgato.StreamDeck/Plugins/.- Marketplace-ready from the start:
- 256×256 plugin icon + action icons (incl. pressed states) at marketplace spec sizes.
- Name "llama-watch", category e.g. "System & Monitoring".
- Privacy note: reads GPU metrics + request status from the user's own llama-swap server; only base URL + optional API key stored in Stream Deck's local settings; no data leaves the machine.
- No hardcoded secrets/URLs — base URL is user-editable.
- Manifest declares macOS only initially.
Testing
- Unit tests with
node:testfor the Prometheus parser, inflight tracker, and aggregate math (avg/max/sum), using real captured fixtures from the server instance (metrics body, inflight snapshot, modelStatus payloads). - Manual verification: both actions render correctly on the physical Stream Deck; key press opens the web UI.
Out of Scope (YAGNI)
- Marketplace submission itself (goal for later; assets/metadata prepared now).
- Auto-update plumbing.
- Windows/Linux support (later add-on; macOS declared in manifest).
- Additional metrics or display styles.