aboutsummaryrefslogtreecommitdiffstats
diff options
context:
space:
mode:
authorDanilo M. <danix@danix.xyz>2026-10-05 13:20:29 +0200
committerDanilo M. <danix@danix.xyz>2026-10-05 13:20:29 +0200
commit87ed521bad8774899f5080f937665b110f44ac27 (patch)
tree304668661688509b480039974398af3b3f30f4fa
parent02c350fe5b4cd4b8df38fe07031b8d2f864c7965 (diff)
downloadquickshell-87ed521bad8774899f5080f937665b110f44ac27.tar.gz
quickshell-87ed521bad8774899f5080f937665b110f44ac27.zip
docs(agents): llama-server /slots autoloads the model
A slot query for an unloaded model loads it into VRAM, so any status probe against the router needs autoload=false. Found when a probe during the ai module's design reloaded Gemma. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
-rw-r--r--AGENTS.md6
1 files changed, 6 insertions, 0 deletions
diff --git a/AGENTS.md b/AGENTS.md
index 73212e0..af8e749 100644
--- a/AGENTS.md
+++ b/AGENTS.md
@@ -253,6 +253,12 @@ changing that component. The ones that generalise:
`LazyLoader` toggled off and on (`assistant/Assistant.qml`). A dropped
connection is different: `disconnected` clears the socket, so a later
`connected = true` works, until one of those attempts fails.
+- **llama-server's `/slots?model=<id>` loads the model.** In router mode a
+ slot query for an unloaded model autoloads it into VRAM, so a status probe
+ is not read-only: one during the `ai` module's design put Gemma back on the
+ GPU. Pass `autoload=false`, which answers 400 `model is not loaded` instead.
+ Checking `/models` for `loaded` first is not enough on its own, because the
+ model can unload between the two requests.
## Theme