diff options
| author | Danilo M. <danix@danix.xyz> | 2026-10-05 13:20:29 +0200 |
|---|---|---|
| committer | Danilo M. <danix@danix.xyz> | 2026-10-05 13:20:29 +0200 |
| commit | 87ed521bad8774899f5080f937665b110f44ac27 (patch) | |
| tree | 304668661688509b480039974398af3b3f30f4fa | |
| parent | 02c350fe5b4cd4b8df38fe07031b8d2f864c7965 (diff) | |
| download | quickshell-87ed521bad8774899f5080f937665b110f44ac27.tar.gz quickshell-87ed521bad8774899f5080f937665b110f44ac27.zip | |
docs(agents): llama-server /slots autoloads the model
A slot query for an unloaded model loads it into VRAM, so any status
probe against the router needs autoload=false. Found when a probe during
the ai module's design reloaded Gemma.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
| -rw-r--r-- | AGENTS.md | 6 |
1 files changed, 6 insertions, 0 deletions
@@ -253,6 +253,12 @@ changing that component. The ones that generalise: `LazyLoader` toggled off and on (`assistant/Assistant.qml`). A dropped connection is different: `disconnected` clears the socket, so a later `connected = true` works, until one of those attempts fails. +- **llama-server's `/slots?model=<id>` loads the model.** In router mode a + slot query for an unloaded model autoloads it into VRAM, so a status probe + is not read-only: one during the `ai` module's design put Gemma back on the + GPU. Pass `autoload=false`, which answers 400 `model is not loaded` instead. + Checking `/models` for `loaded` first is not enough on its own, because the + model can unload between the two requests. ## Theme |
