From 87ed521bad8774899f5080f937665b110f44ac27 Mon Sep 17 00:00:00 2001 From: "Danilo M." Date: Mon, 5 Oct 2026 13:20:29 +0200 Subject: docs(agents): llama-server /slots autoloads the model A slot query for an unloaded model loads it into VRAM, so any status probe against the router needs autoload=false. Found when a probe during the ai module's design reloaded Gemma. Co-Authored-By: Claude Opus 5.5 --- AGENTS.md | 6 ++++++ 1 file changed, 6 insertions(+) (limited to 'AGENTS.md') diff --git a/AGENTS.md b/AGENTS.md index 73212e0..af8e749 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -253,6 +253,12 @@ changing that component. The ones that generalise: `LazyLoader` toggled off and on (`assistant/Assistant.qml`). A dropped connection is different: `disconnected` clears the socket, so a later `connected = true` works, until one of those attempts fails. +- **llama-server's `/slots?model=` loads the model.** In router mode a + slot query for an unloaded model autoloads it into VRAM, so a status probe + is not read-only: one during the `ai` module's design put Gemma back on the + GPU. Pass `autoload=false`, which answers 400 `model is not loaded` instead. + Checking `/models` for `loaded` first is not enough on its own, because the + model can unload between the two requests. ## Theme -- cgit v1.2.3