From ef4043dd4e13fa43952fa5174970648d2f086430 Mon Sep 17 00:00:00 2001 From: Bryce Zuccaro Date: Fri, 14 Aug 2026 13:16:07 -0600 Subject: [PATCH] docs: update README for usage stats display and GPU combinations --- README.md | 21 ++++++++++++++++++--- 1 file changed, 18 insertions(+), 3 deletions(-) diff --git a/README.md b/README.md index f0340e0..8936d3d 100644 --- a/README.md +++ b/README.md @@ -4,8 +4,15 @@ A Stream Deck plugin (macOS) that monitors a llama-swap instance. ## Actions -- **In-Flight Monitor** — per-model color state: `OFF` / `LOADING` / `IDLE` / `ACTIVE` (pulsing red). Powered by the real-time `/api/events` SSE feed. -- **GPU Graph** — live line chart of a GPU metric (utilization %, VRAM %, temperature, power draw, fan speed) for one GPU or all GPUs, sampled every 5 s. +- **In-Flight Monitor** — per-model request activity. The default view shows + the live in-flight request count with a 60-second activity spark (states: + `OFF` / `LOADING` / `IDLE` / count). Switch **Display** to **Usage stats** + to show request/token totals and the generation-speed P95 from + `/api/metrics/stats`, per model or across **All models**, with a pickable + primary stat. Powered by the real-time `/api/events` SSE feed. +- **GPU Graph** — live line chart of a GPU metric (utilization %, VRAM %, + temperature, power draw, fan speed) for one GPU, all GPUs, or an arbitrary + combination (e.g. GPU combination `0,2`), sampled every 5 s. Pressing either key opens `http:///ui` in your browser. @@ -28,7 +35,15 @@ npm test # unit tests (node:test + tsx) ## Configure -Per-key settings: base URL (default `http://localhost:9292`), optional API key, and the model / GPU / metric to watch. The model and GPU dropdowns are populated live from the instance. Each action instance is independent, so you can place several In-Flight Monitor keys (one per model) and several GPU Graph keys (one per GPU × metric) on the same profile. All keys share a single connection to the configured llama-swap instance, which is re-established automatically if you change the base URL or API key on any key. +Per-key settings: base URL (default `http://localhost:9292`), optional API key, and the model / GPU / metric to watch. The model and GPU dropdowns are populated live from the instance. Each action instance is independent, so you can place several In-Flight Monitor keys (one per model, or in either display mode) and several GPU Graph keys (one per GPU × metric, or per GPU combination) on the same profile. All keys share a single connection to the configured llama-swap instance, which is re-established automatically if you change the base URL or API key on any key. + +## Notes + +- Usage totals come from llama-swap's `/api/metrics/stats` and cover its + in-memory activity retention (default ~1000 most recent requests), not + lifetime counters. The P95 is a **tokens/sec** generation-speed percentile. +- Keys are SVG-rendered (`render.ts`), so no image assets or canvas are + needed at runtime. ## Marketplace