# AI module design A drawer module showing what is using the local AI stack, with start and stop for the projects built on it. ## Scope Two kinds of thing are shown: - **Engines**: llama.cpp (`llama-server`, router mode, a root rc service on port 8181) and stable-diffusion.cpp (`sd-server`, started in a terminal by `~/bin/sd-webui`, port 7860). Both carry a web UI the user sometimes drives directly. Rows are **read-only**: up or down, idle or generating, which model. No unload, no stop. - **Apps**: the projects using the stack. Each has a Start and a Stop button, except fanfictioner, which has Stop only. | App | Status | Start | Stop | | --- | --- | --- | --- | | desktop-assistant | `assistant.sh status`, exit code | `assistant.sh start` | `assistant.sh stop` | | local-chat | `llamachat --ping`, exit code | `llamachat --daemon` | `llamachat --quit` | | imggen | `imggen status`, `daemon down` means stopped | `imggen start` | `imggen stop` | | fanfictioner | `pgrep -u $USER -f` on `python3 .../fanfictioner` | none | kill its PID and its children by PID | The module uses each project's existing control surface. No project gains a new script. imggen starts its default model (`sdxl`); there is no picker. fanfictioner has no Start because a run needs a story file. Fanfictioner's children are `sd-cli` runs, and its stop sends TERM to the children first, then the parent, by PID. Never by process name: a name match kills unrelated instances of the same program. ## Detection **llama.** `GET :8181/models` lists every preset with `status.value` of `loaded` or `unloaded`. For the loaded one, `GET /slots?model=&autoload=false` returns slots with `is_processing`; any slot processing means busy. `autoload=false` is not optional. A plain `/slots?model=` on an unloaded model **loads it**: measured during design, a probe put Gemma back into VRAM. Checking `loaded` first is not enough on its own, since the model can unload between the two requests. With the flag an unloaded model answers 400 `model is not loaded`, which reads as idle. The engine is down when `/models` does not answer. **sd.** A running `sd-cli` comes first: fanfictioner runs them one-shot, each holds VRAM and generates for as long as it lives, so the row reads up, busy, and its model marked `(sd-cli)`. The first version watched only `sd-server` and showed the engine down during a fanfictioner run. Otherwise `pgrep -x sd-server` gives the PID. The model is the basename of the value after `--diffusion-model` or `-m` in its arguments. sd-server has no progress endpoint, but its log `~/.cache/sd-server.log` streams progress bars while it generates, so busy means the log was written in the last 5 seconds. `ponytail:` a heuristic. A generation that stalls longer than 5s without writing reads as idle; acceptable for an indicator. ## Pieces All under `desktop/modules/ai/`, the shape of `modules/kdeconnect/`. - `ai-state.sh`: one sweep, TSV on stdout, one line per thing: enginellamaup|downidle|busymodel or - enginesdup|downidle|busymodel or - appassistantrunning|stopped- appchatrunning|stopped- appimggenrunning|stoppedmodel or - appfanficrunning|stoppedpid or - A probe that fails reports its row as down or stopped; the script always exits 0 with all six lines. Every external command (`curl`, `pgrep`, the four controls, `stat`) is called by name so the test can stub it on PATH. The path to `assistant.sh` defaults to `~/Programming/GIT/...` and is overridable by environment variable, as imggen does with `IMGGEN_DIR`. - `AiModule.qml`: `name: "ai"`, a `Process` running the script every 3s while the tile exists, which is exactly while the drawer panel is loaded, the same gating as kdeconnect. Parsed rows are kept in a property; a run that prints nothing keeps the previous rows. `active` is true when any engine has a model loaded or any app is running. Actions go through `Quickshell.execDetached`, then trigger an immediate re-poll. - `AiTile.qml`: one line, e.g. `2 apps ยท llama busy`, or `idle`. - `AiPage.qml`: an **Engines** section with read-only rows (status dot, model, idle or generating) and an **Apps** section with a row per project carrying its Start or Stop button. - `test-ai-state.sh`: stubs the external commands on PATH and asserts the output for: everything down, llama loaded and busy, llama loaded and idle, sd running and recently logged, each app running. - `README.md` for the module. Registered in `desktop/shell.qml`; the module list in `desktop/README.md` gains a line. ## Not doing - Unloading the llama model or stopping sd-server from the drawer. - Notifications on state change, VRAM figures, an imggen model picker. - Polling while the drawer is closed.