diff options
| author | Danilo M. <danix@danix.xyz> | 2026-10-05 12:36:52 +0200 |
|---|---|---|
| committer | Danilo M. <danix@danix.xyz> | 2026-10-05 12:36:52 +0200 |
| commit | d4fc981660ef4302853c109074f8a7c3ddf08e45 (patch) | |
| tree | d192c7aaa8298286c4b4e1e4eba0a6c4b18bcbdf /docs/superpowers/specs | |
| parent | 8af37014d58918da714c741954de92881dcb2554 (diff) | |
| download | quickshell-d4fc981660ef4302853c109074f8a7c3ddf08e45.tar.gz quickshell-d4fc981660ef4302853c109074f8a7c3ddf08e45.zip | |
docs(ai): design for the AI stack drawer module
Engines (llama.cpp, sd.cpp) read-only, apps get start/stop through their
existing controls. Records that /slots?model= autoloads an unloaded model
unless autoload=false is passed, found when a design probe reloaded one.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Diffstat (limited to 'docs/superpowers/specs')
| -rw-r--r-- | docs/superpowers/specs/2026-10-05-ai-module-design.md | 94 |
1 files changed, 94 insertions, 0 deletions
diff --git a/docs/superpowers/specs/2026-10-05-ai-module-design.md b/docs/superpowers/specs/2026-10-05-ai-module-design.md new file mode 100644 index 0000000..cef2e88 --- /dev/null +++ b/docs/superpowers/specs/2026-10-05-ai-module-design.md @@ -0,0 +1,94 @@ +# AI module design + +A drawer module showing what is using the local AI stack, with start and stop +for the projects built on it. + +## Scope + +Two kinds of thing are shown: + +- **Engines**: llama.cpp (`llama-server`, router mode, a root rc service on + port 8181) and stable-diffusion.cpp (`sd-server`, started in a terminal by + `~/bin/sd-webui`, port 7860). Both carry a web UI the user sometimes drives + directly. Rows are **read-only**: up or down, idle or generating, which model. + No unload, no stop. +- **Apps**: the projects using the stack. Each has a Start and a Stop button, + except fanfictioner, which has Stop only. + +| App | Status | Start | Stop | +| --- | --- | --- | --- | +| desktop-assistant | `assistant.sh status`, exit code | `assistant.sh start` | `assistant.sh stop` | +| local-chat | `llamachat --ping`, exit code | `llamachat --daemon` | `llamachat --quit` | +| imggen | `imggen status`, `daemon down` means stopped | `imggen start` | `imggen stop` | +| fanfictioner | `pgrep -u $USER -f` on the script path | none | kill its PID and its children by PID | + +The module uses each project's existing control surface. No project gains a +new script. imggen starts its default model (`sdxl`); there is no picker. +fanfictioner has no Start because a run needs a story file. + +Fanfictioner's children are `sd-cli` runs, and its stop sends TERM to the +children first, then the parent, by PID. Never by process name: a name match +kills unrelated instances of the same program. + +## Detection + +**llama.** `GET :8181/models` lists every preset with +`status.value` of `loaded` or `unloaded`. For the loaded one, +`GET /slots?model=<id>&autoload=false` returns slots with `is_processing`; +any slot processing means busy. + +`autoload=false` is not optional. A plain `/slots?model=<id>` on an unloaded +model **loads it**: measured during design, a probe put Gemma back into VRAM. +Checking `loaded` first is not enough on its own, since the model can unload +between the two requests. With the flag an unloaded model answers 400 +`model is not loaded`, which reads as idle. + +The engine is down when `/models` does not answer. + +**sd.** `pgrep -x sd-server` gives the PID. The model is the basename of the +value after `--diffusion-model` or `-m` in its arguments. sd-server has no +progress endpoint, but its log `~/.cache/sd-server.log` streams progress bars +while it generates, so busy means the log was written in the last 5 seconds. +`ponytail:` a heuristic. A generation that stalls longer than 5s without +writing reads as idle; acceptable for an indicator. + +## Pieces + +All under `desktop/modules/ai/`, the shape of `modules/kdeconnect/`. + +- `ai-state.sh`: one sweep, TSV on stdout, one line per thing: + + engine<TAB>llama<TAB>up|down<TAB>idle|busy<TAB>model or - + engine<TAB>sd<TAB>up|down<TAB>idle|busy<TAB>model or - + app<TAB>assistant<TAB>running|stopped<TAB>- + app<TAB>chat<TAB>running|stopped<TAB>- + app<TAB>imggen<TAB>running|stopped<TAB>model or - + app<TAB>fanfic<TAB>running|stopped<TAB>pid or - + + A probe that fails reports its row as down or stopped; the script always + exits 0 with all six lines. Every external command (`curl`, `pgrep`, the + four controls, `stat`) is called by name so the test can stub it on PATH. + Paths to `assistant.sh` and the fanfictioner script default to + `~/Programming/GIT/...` and are overridable by environment variable, as + imggen does with `IMGGEN_DIR`. +- `AiModule.qml`: `name: "ai"`, a `Process` running the script every 3s + while the tile exists, which is exactly while the drawer panel is loaded, + the same gating as kdeconnect. Parsed rows are kept in a property; a run + that prints nothing keeps the previous rows. `active` is true when any + engine has a model loaded or any app is running. Actions go through + `Quickshell.execDetached`, then trigger an immediate re-poll. +- `AiTile.qml`: one line, e.g. `2 apps · llama busy`, or `idle`. +- `AiPage.qml`: an **Engines** section with read-only rows (status dot, + model, idle or generating) and an **Apps** section with a row per project + carrying its Start or Stop button. +- `test-ai-state.sh`: stubs the external commands on PATH and asserts the + output for: everything down, llama loaded and busy, llama loaded and idle, + sd running and recently logged, each app running. +- `README.md` for the module. Registered in `desktop/shell.qml`; the module + list in `desktop/README.md` gains a line. + +## Not doing + +- Unloading the llama model or stopping sd-server from the drawer. +- Notifications on state change, VRAM figures, an imggen model picker. +- Polling while the drawer is closed. |
