aboutsummaryrefslogtreecommitdiffstats
path: root/docs/superpowers
diff options
context:
space:
mode:
Diffstat (limited to 'docs/superpowers')
-rw-r--r--docs/superpowers/specs/2026-10-05-ai-module-design.md94
1 files changed, 94 insertions, 0 deletions
diff --git a/docs/superpowers/specs/2026-10-05-ai-module-design.md b/docs/superpowers/specs/2026-10-05-ai-module-design.md
new file mode 100644
index 0000000..cef2e88
--- /dev/null
+++ b/docs/superpowers/specs/2026-10-05-ai-module-design.md
@@ -0,0 +1,94 @@
+# AI module design
+
+A drawer module showing what is using the local AI stack, with start and stop
+for the projects built on it.
+
+## Scope
+
+Two kinds of thing are shown:
+
+- **Engines**: llama.cpp (`llama-server`, router mode, a root rc service on
+ port 8181) and stable-diffusion.cpp (`sd-server`, started in a terminal by
+ `~/bin/sd-webui`, port 7860). Both carry a web UI the user sometimes drives
+ directly. Rows are **read-only**: up or down, idle or generating, which model.
+ No unload, no stop.
+- **Apps**: the projects using the stack. Each has a Start and a Stop button,
+ except fanfictioner, which has Stop only.
+
+| App | Status | Start | Stop |
+| --- | --- | --- | --- |
+| desktop-assistant | `assistant.sh status`, exit code | `assistant.sh start` | `assistant.sh stop` |
+| local-chat | `llamachat --ping`, exit code | `llamachat --daemon` | `llamachat --quit` |
+| imggen | `imggen status`, `daemon down` means stopped | `imggen start` | `imggen stop` |
+| fanfictioner | `pgrep -u $USER -f` on the script path | none | kill its PID and its children by PID |
+
+The module uses each project's existing control surface. No project gains a
+new script. imggen starts its default model (`sdxl`); there is no picker.
+fanfictioner has no Start because a run needs a story file.
+
+Fanfictioner's children are `sd-cli` runs, and its stop sends TERM to the
+children first, then the parent, by PID. Never by process name: a name match
+kills unrelated instances of the same program.
+
+## Detection
+
+**llama.** `GET :8181/models` lists every preset with
+`status.value` of `loaded` or `unloaded`. For the loaded one,
+`GET /slots?model=<id>&autoload=false` returns slots with `is_processing`;
+any slot processing means busy.
+
+`autoload=false` is not optional. A plain `/slots?model=<id>` on an unloaded
+model **loads it**: measured during design, a probe put Gemma back into VRAM.
+Checking `loaded` first is not enough on its own, since the model can unload
+between the two requests. With the flag an unloaded model answers 400
+`model is not loaded`, which reads as idle.
+
+The engine is down when `/models` does not answer.
+
+**sd.** `pgrep -x sd-server` gives the PID. The model is the basename of the
+value after `--diffusion-model` or `-m` in its arguments. sd-server has no
+progress endpoint, but its log `~/.cache/sd-server.log` streams progress bars
+while it generates, so busy means the log was written in the last 5 seconds.
+`ponytail:` a heuristic. A generation that stalls longer than 5s without
+writing reads as idle; acceptable for an indicator.
+
+## Pieces
+
+All under `desktop/modules/ai/`, the shape of `modules/kdeconnect/`.
+
+- `ai-state.sh`: one sweep, TSV on stdout, one line per thing:
+
+ engine<TAB>llama<TAB>up|down<TAB>idle|busy<TAB>model or -
+ engine<TAB>sd<TAB>up|down<TAB>idle|busy<TAB>model or -
+ app<TAB>assistant<TAB>running|stopped<TAB>-
+ app<TAB>chat<TAB>running|stopped<TAB>-
+ app<TAB>imggen<TAB>running|stopped<TAB>model or -
+ app<TAB>fanfic<TAB>running|stopped<TAB>pid or -
+
+ A probe that fails reports its row as down or stopped; the script always
+ exits 0 with all six lines. Every external command (`curl`, `pgrep`, the
+ four controls, `stat`) is called by name so the test can stub it on PATH.
+ Paths to `assistant.sh` and the fanfictioner script default to
+ `~/Programming/GIT/...` and are overridable by environment variable, as
+ imggen does with `IMGGEN_DIR`.
+- `AiModule.qml`: `name: "ai"`, a `Process` running the script every 3s
+ while the tile exists, which is exactly while the drawer panel is loaded,
+ the same gating as kdeconnect. Parsed rows are kept in a property; a run
+ that prints nothing keeps the previous rows. `active` is true when any
+ engine has a model loaded or any app is running. Actions go through
+ `Quickshell.execDetached`, then trigger an immediate re-poll.
+- `AiTile.qml`: one line, e.g. `2 apps · llama busy`, or `idle`.
+- `AiPage.qml`: an **Engines** section with read-only rows (status dot,
+ model, idle or generating) and an **Apps** section with a row per project
+ carrying its Start or Stop button.
+- `test-ai-state.sh`: stubs the external commands on PATH and asserts the
+ output for: everything down, llama loaded and busy, llama loaded and idle,
+ sd running and recently logged, each app running.
+- `README.md` for the module. Registered in `desktop/shell.qml`; the module
+ list in `desktop/README.md` gains a line.
+
+## Not doing
+
+- Unloading the llama model or stopping sd-server from the drawer.
+- Notifications on state change, VRAM figures, an imggen model picker.
+- Polling while the drawer is closed.