aboutsummaryrefslogtreecommitdiffstats
path: root/docs/superpowers/specs/2026-10-05-ai-module-design.md
blob: 8171f98ef1aac1df8f8c544f1acca95e6ef05cd4 (plain)
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
# AI module design

A drawer module showing what is using the local AI stack, with start and stop
for the projects built on it.

## Scope

Two kinds of thing are shown:

- **Engines**: llama.cpp (`llama-server`, router mode, a root rc service on
  port 8181) and stable-diffusion.cpp (`sd-server`, started in a terminal by
  `~/bin/sd-webui`, port 7860). Both carry a web UI the user sometimes drives
  directly. Rows are **read-only**: up or down, idle or generating, which model.
  No unload, no stop.
- **Apps**: the projects using the stack. Each has a Start and a Stop button,
  except fanfictioner, which has Stop only.

| App | Status | Start | Stop |
| --- | --- | --- | --- |
| desktop-assistant | `assistant.sh status`, exit code | `assistant.sh start` | `assistant.sh stop` |
| local-chat | `llamachat --ping`, exit code | `llamachat --daemon` | `llamachat --quit` |
| imggen | `imggen status`, `daemon down` means stopped | `imggen start` | `imggen stop` |
| fanfictioner | `pgrep -u $USER -f` on `python3 .../fanfictioner` | none | kill its PID and its children by PID |

The module uses each project's existing control surface. No project gains a
new script. imggen starts its default model (`sdxl`); there is no picker.
fanfictioner has no Start because a run needs a story file.

Fanfictioner's children are `sd-cli` runs, and its stop sends TERM to the
children first, then the parent, by PID. Never by process name: a name match
kills unrelated instances of the same program.

## Detection

**llama.** `GET :8181/models` lists every preset with
`status.value` of `loaded` or `unloaded`. For the loaded one,
`GET /slots?model=<id>&autoload=false` returns slots with `is_processing`;
any slot processing means busy.

`autoload=false` is not optional. A plain `/slots?model=<id>` on an unloaded
model **loads it**: measured during design, a probe put Gemma back into VRAM.
Checking `loaded` first is not enough on its own, since the model can unload
between the two requests. With the flag an unloaded model answers 400
`model is not loaded`, which reads as idle.

The engine is down when `/models` does not answer.

**sd.** A running `sd-cli` comes first: fanfictioner runs them one-shot,
each holds VRAM and generates for as long as it lives, so the row reads up,
busy, and its model marked `(sd-cli)`. The first version watched only
`sd-server` and showed the engine down during a fanfictioner run. Otherwise
`pgrep -x sd-server` gives the PID. The model is the basename of the
value after `--diffusion-model` or `-m` in its arguments. sd-server has no
progress endpoint, but its log `~/.cache/sd-server.log` streams progress bars
while it generates, so busy means the log was written in the last 5 seconds.
`ponytail:` a heuristic. A generation that stalls longer than 5s without
writing reads as idle; acceptable for an indicator.

## Pieces

All under `desktop/modules/ai/`, the shape of `modules/kdeconnect/`.

- `ai-state.sh`: one sweep, TSV on stdout, one line per thing:

      engine<TAB>llama<TAB>up|down<TAB>idle|busy<TAB>model or -
      engine<TAB>sd<TAB>up|down<TAB>idle|busy<TAB>model or -
      app<TAB>assistant<TAB>running|stopped<TAB>-
      app<TAB>chat<TAB>running|stopped<TAB>-
      app<TAB>imggen<TAB>running|stopped<TAB>model or -
      app<TAB>fanfic<TAB>running|stopped<TAB>pid or -

  A probe that fails reports its row as down or stopped; the script always
  exits 0 with all six lines. Every external command (`curl`, `pgrep`, the
  four controls, `stat`) is called by name so the test can stub it on PATH.
  The path to `assistant.sh` defaults to
  `~/Programming/GIT/...` and is overridable by environment variable, as
  imggen does with `IMGGEN_DIR`.
- `AiModule.qml`: `name: "ai"`, a `Process` running the script every 3s
  while the tile exists, which is exactly while the drawer panel is loaded,
  the same gating as kdeconnect. Parsed rows are kept in a property; a run
  that prints nothing keeps the previous rows. `active` is true when any
  engine has a model loaded or any app is running. Actions go through
  `Quickshell.execDetached`, then trigger an immediate re-poll.
- `AiTile.qml`: one line, e.g. `2 apps · llama busy`, or `idle`.
- `AiPage.qml`: an **Engines** section with read-only rows (status dot,
  model, idle or generating) and an **Apps** section with a row per project
  carrying its Start or Stop button.
- `test-ai-state.sh`: stubs the external commands on PATH and asserts the
  output for: everything down, llama loaded and busy, llama loaded and idle,
  sd running and recently logged, each app running.
- `README.md` for the module. Registered in `desktop/shell.qml`; the module
  list in `desktop/README.md` gains a line.

## Not doing

- Unloading the llama model or stopping sd-server from the drawer.
- Notifications on state change, VRAM figures, an imggen model picker.
- Polling while the drawer is closed.