# fanfictioner Implementation Plan > **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking. **Goal:** Build `fanfictioner`, a single stdlib-only Python 3 script that turns `story.md` into a text-free illustrated book of page images, as described in `docs/superpowers/specs/2026-10-04-fanfictioner-design.md`. **Architecture:** One executable file `fanfictioner` at the repo root. Pure helpers (config, plan.md parsing/validation, edit grouping, plan.json checks, sd-cli argument building, meminfo parsing) are covered by `--selftest` asserts. Thin I/O layers wrap llama-server (urllib), sd-cli (subprocess + RAM watchdog), `magick` and `kitty +kitten icat`. All state is on disk; resume is derived from files. **Tech Stack:** Python 3 stdlib (argparse, json, pathlib, re, shlex, shutil, subprocess, urllib.request), sd-cli (stable-diffusion.cpp), ImageMagick 7 `magick`, kitty icat, llama-server router on `localhost:8181`. --- ## Facts the engineer needs - `sd-cli -b 3 -s SEED -o dir/X_%d.png --output-begin-idx 1` writes `X_1.png`..`X_3.png`; image *b* (0-based) uses seed `SEED + b` (verified in sd.cpp `src/pipeline/image.cpp`: `cur_seed = request.seed + b`). - Krea2 Turbo: 8 steps, CFG off (`--cfg-scale 1.0` in sd-cli), shift 1.15 is sd.cpp's built-in default for Krea2. Verified 2026-10-05. - `--backend te=cpu,diffusion=vulkan0,vae=vulkan0` must list all three, or diffusion lands on the iGPU and crashes. Never `--offload-to-cpu` on the edit profile. - llama-server: `GET /v1/models` lists models (`data[].id`), `POST /models/unload {"model": ...}` unloads, `POST /v1/chat/completions` with `response_format: {"type": "json_schema", ...}` constrains output. - `~/.config/book-reader.conf` holds shell lines like `LIBRARY="/some/dir"`. - Commits are GPG-signed by global git config. Just run `git commit`; the first one in a session prompts for the PIN. ## Deviations from the spec (deliberate, smaller) - plan.json `title`/`series`/`book` and each page's `model` are taken from the parsed plan.md by the script, not from Gemma. Folder names and model choice must not depend on LLM copying. A page's `model` is `""` unless plan.md has `Model:`; then `--base` applies. - The script computes edit groups (`edit_groups`) and passes them to Gemma; `check_json` rejects a compile whose groups differ, with one retry. - Picks are also kept per page in `cand/pNNN/base.png` and `cand/pNNN/editK.png`, so a quit mid-page resumes without regenerating the stages already picked. Still disk-derived, no state file. ## File structure - Create: `fanfictioner` (executable script, everything) - Create: `README.md` - Unchanged: `LICENSE`, spec --- ### Task 1: Skeleton, selftest runner, README **Files:** - Create: `fanfictioner` - Create: `README.md` - [ ] **Step 1: Create the script skeleton** `fanfictioner`: ```python #!/usr/bin/env python3 # fanfictioner: turn a story idea into a text-free illustrated book. # Copyright (C) 2026 Danilo M. # # This program is free software; you can redistribute it and/or modify # it under the terms of the GNU General Public License version 2 as # published by the Free Software Foundation. # # This program is distributed in the hope that it will be useful, # but WITHOUT ANY WARRANTY; without even the implied warranty of # MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE. See the # GNU General Public License for more details. # # You should have received a copy of the GNU General Public License along # with this program; if not, write to the Free Software Foundation, Inc., # 51 Franklin Street, Fifth Floor, Boston, MA 02110-1301 USA. import argparse import json import os import random import re import shlex import shutil import subprocess import sys import tempfile import time import urllib.request from pathlib import Path LLM_URL = "http://localhost:8181" LLM_MODEL = "Gemma4-12B-qat-mtp" SD_DIR = Path("/data/LLM-models/SD") CONF = Path.home() / ".config/book-reader.conf" def selftest(): print("selftest ok") def main(): ap = argparse.ArgumentParser(description="Turn story.md into a text-free illustrated book.") ap.add_argument("story", nargs="?", type=Path, help="story idea in free prose") ap.add_argument("--base", choices=("zimage", "krea2"), default="zimage", help="base model for pages without a Model: line") ap.add_argument("--selftest", action="store_true", help="run built-in checks and exit") a = ap.parse_args() if a.selftest: return selftest() if not a.story: ap.error("story file required") if __name__ == "__main__": main() ``` - [ ] **Step 2: Make it executable and run it** Run: `chmod +x fanfictioner && ./fanfictioner --selftest` Expected: `selftest ok` Run: `./fanfictioner` Expected: usage line and `error: story file required`, exit 2. - [ ] **Step 3: Write README.md** ```markdown # fanfictioner Turn a story idea into a text-free illustrated book. A local Gemma model (llama-server) plans the pages and writes the image prompts; `sd-cli` (stable-diffusion.cpp) draws them. You pick among three candidates at every image stage. Pages land in the book-reader library as `NNN.png`. ## Requirements - Python 3, stdlib only - llama-server in router mode on `localhost:8181` serving `Gemma4-12B-qat-mtp` - `sd-cli` with Vulkan, models under `/data/LLM-models/SD` - ImageMagick 7 (`magick`), kitty (`kitty +kitten icat`), `$EDITOR` - `~/.config/book-reader.conf` with a `LIBRARY="..."` line ## Usage fanfictioner story.md [--base zimage|krea2] fanfictioner --selftest Work files go to a directory named after the story (`stories/rooftop.md` -> `stories/rooftop/`): `plan.md`, `plan.json`, `refs/`, `cand/`. Quit any time with `q` or Ctrl-C; the next run resumes from what is on disk. Plan review keys: `c` confirm, `e` edit in `$EDITOR`, `g` ask Gemma to revise, `q` quit. Image review keys: `1`-`3` pick, `r` regenerate, `p` edit prompt, `s` skip edit (edit stage only), `q` quit. ## License Copyright (C) 2026 Danilo M. This program is free software; you can redistribute it and/or modify it under the terms of the GNU General Public License version 2 as published by the Free Software Foundation. See `LICENSE`. ## Development Approach This project is developed using AI-assisted tools. Code is generated with the help of AI based on human-provided specifications, design decisions, and iterative feedback. All contributions are reviewed, tested, and curated by the maintainer before being included in the codebase. AI is used as a productivity and exploration tool, while human oversight remains central to all decisions. The goal is to combine the flexibility of AI-assisted development with standard open-source practices such as transparency, review, and accountability. ``` - [ ] **Step 4: Commit** ```bash git add fanfictioner README.md git commit -m "Add fanfictioner skeleton and README" ``` --- ### Task 2: Config and resume detection **Files:** - Modify: `fanfictioner` (add helpers above `selftest`, asserts inside `selftest`) - [ ] **Step 1: Write the failing test** Replace `selftest` with: ```python def selftest(): with tempfile.TemporaryDirectory() as t: t = Path(t) (t / "c.conf").write_text('# comment\nLIBRARY="/data/my lib"\nPORT=8642\n') conf = load_conf(t / "c.conf") assert conf == {"LIBRARY": "/data/my lib", "PORT": "8642"}, conf pj = {"series": "S", "book": "B", "characters": [{"name": "Mara"}, {"name": "Jon"}], "pages": [{"n": 1}, {"n": 2}]} lib = t / "lib" assert page_path(lib, pj, 7) == lib / "S" / "B" / "007.png" (t / "refs").mkdir() (t / "refs" / "Mara.png").touch() page_path(lib, pj, 1).parent.mkdir(parents=True) page_path(lib, pj, 1).touch() refs, pages = todo(pj, t, lib) assert [c["name"] for c in refs] == ["Jon"], refs assert [p["n"] for p in pages] == [2], pages keep(t / "refs" / "Mara.png", t / "deep" / "x.png") assert (t / "deep" / "x.png").exists() print("selftest ok") ``` - [ ] **Step 2: Run test to verify it fails** Run: `./fanfictioner --selftest` Expected: FAIL with `NameError: name 'load_conf' is not defined` - [ ] **Step 3: Write minimal implementation** Add above `selftest`: ```python def load_conf(path=CONF): conf = {} for line in path.read_text().splitlines(): if "=" in line and not line.lstrip().startswith("#"): k, v = line.split("=", 1) conf[k.strip()] = shlex.split(v)[0] if v.strip() else "" return conf def page_path(lib, pj, n): return Path(lib) / pj["series"] / pj["book"] / f"{n:03d}.png" def todo(pj, wd, lib): """What is left to do, derived from disk: (characters without a ref, pages not in the library).""" refs = [c for c in pj["characters"] if not (wd / "refs" / f"{c['name']}.png").exists()] pages = [p for p in pj["pages"] if not page_path(lib, pj, p["n"]).exists()] return refs, pages def keep(src, dst): """Copy src to dst atomically, so a half-written file never counts as done.""" dst.parent.mkdir(parents=True, exist_ok=True) tmp = dst.with_name(dst.name + ".tmp") shutil.copyfile(src, tmp) os.replace(tmp, dst) return dst ``` - [ ] **Step 4: Run test to verify it passes** Run: `./fanfictioner --selftest` Expected: `selftest ok` - [ ] **Step 5: Commit** ```bash git add fanfictioner git commit -m "Add config loading and disk-derived resume detection" ``` --- ### Task 3: plan.md parsing and validation **Files:** - Modify: `fanfictioner` - [ ] **Step 1: Write the failing test** Add this constant above `selftest`: ```python SAMPLE_PLAN = """# Rooftop Series: Test Series Book: Book One Style: soft watercolor, muted colors ## Characters ### Mara Young woman, long red hair, green eyes, grey coat. ### Jon Tall man, short beard, blue jacket. ### Kit Small girl, black bob, yellow raincoat. ## Pages ### 1. Arrival They reach the roof at dusk. Characters: Mara, Jon, Kit Pose: Mara left facing right, Jon center facing camera, Kit right sitting. Framing: wide shot, eye level, landscape ### 2. Empty roof Nobody is there any more. Characters: none Framing: wide shot, portrait Model: krea2 """ ``` Add at the start of `selftest`, before the `with` block: ```python plan = parse_plan(SAMPLE_PLAN) assert (plan["title"], plan["series"], plan["book"]) == ("Rooftop", "Test Series", "Book One"), plan assert plan["style"] == "soft watercolor, muted colors" assert list(plan["characters"]) == ["Mara", "Jon", "Kit"] assert plan["characters"]["Mara"] == "Young woman, long red hair, green eyes, grey coat." p1, p2 = plan["pages"] assert (p1["n"], p1["title"], p1["characters"]) == (1, "Arrival", ["Mara", "Jon", "Kit"]) assert p1["text"] == "They reach the roof at dusk." and p1["pose"].startswith("Mara left") assert p1["model"] == "" and p2["model"] == "krea2" and p2["characters"] == [] assert validate_plan(plan) == [] bad = parse_plan(SAMPLE_PLAN.replace("Book: Book One\n", "") .replace("Characters: none", "Characters: Zed") .replace("Model: krea2", "Model: sdxl") .replace("### 2.", "### 1.")) errs = validate_plan(bad) assert "missing book" in errs, errs assert "page 1: unknown character Zed" in errs, errs assert "page 1: Pose missing" in errs, errs assert "page 1: unknown model sdxl" in errs, errs assert "page numbers must be unique and start at 1" in errs, errs assert validate_plan(parse_plan("")) == ["missing title", "missing series", "missing book", "no characters", "no pages"] ``` - [ ] **Step 2: Run test to verify it fails** Run: `./fanfictioner --selftest` Expected: FAIL with `NameError: name 'parse_plan' is not defined` - [ ] **Step 3: Write minimal implementation** Add above `SAMPLE_PLAN`: ```python def parse_plan(text): plan = {"title": "", "series": "", "book": "", "style": "", "characters": {}, "pages": []} section = cur = None for line in text.splitlines(): s = line.strip() if s.startswith("# ") and not plan["title"]: plan["title"] = s[2:].strip() elif s.startswith("## "): section, cur = s[3:].strip().lower(), None elif s.startswith("### "): head = s[4:].strip() if section == "characters": cur = plan["characters"].setdefault(head, []) elif section == "pages": m = re.match(r"(\d+)\.\s*(.*)", head) cur = {"n": int(m[1]) if m else 0, "title": m[2] if m else head, "text": [], "characters": [], "pose": "", "framing": "", "model": ""} plan["pages"].append(cur) elif section is None and (m := re.match(r"(Series|Book|Style):\s*(.*)", s)): plan[m[1].lower()] = m[2].strip() elif section == "pages" and cur is not None and \ (m := re.match(r"(Characters|Pose|Framing|Model):\s*(.*)", s)): key, val = m[1].lower(), m[2].strip() if key == "characters": val = [] if val.lower() == "none" else [c.strip() for c in val.split(",") if c.strip()] elif key == "model": val = val.split()[0].lower() if val else "" cur[key] = val elif cur is not None and s: (cur if isinstance(cur, list) else cur["text"]).append(s) plan["characters"] = {k: " ".join(v) for k, v in plan["characters"].items()} for p in plan["pages"]: p["text"] = " ".join(p["text"]) return plan def validate_plan(plan): errs = [f"missing {k}" for k in ("title", "series", "book") if not plan[k]] if not plan["characters"]: errs.append("no characters") if not plan["pages"]: errs.append("no pages") ns = [p["n"] for p in plan["pages"]] if len(set(ns)) != len(ns) or 0 in ns: errs.append("page numbers must be unique and start at 1") for p in plan["pages"]: errs += [f"page {p['n']}: unknown character {c}" for c in p["characters"] if c not in plan["characters"]] if p["characters"] and not p["pose"]: errs.append(f"page {p['n']}: Pose missing") if p["model"] and p["model"] not in ("zimage", "krea2"): errs.append(f"page {p['n']}: unknown model {p['model']}") return errs ``` - [ ] **Step 4: Run test to verify it passes** Run: `./fanfictioner --selftest` Expected: `selftest ok` - [ ] **Step 5: Commit** ```bash git add fanfictioner git commit -m "Parse and validate plan.md" ``` --- ### Task 4: Edit grouping and plan.json checks **Files:** - Modify: `fanfictioner` - [ ] **Step 1: Write the failing test** Append to `selftest`, after the plan.md asserts and before the `with` block: ```python assert edit_groups([]) == [] assert edit_groups(["Mara", "Jon", "Kit"]) == [["Mara", "Jon"], ["Kit"]] assert edit_groups(list("abcd")) == [["a", "b"], ["c", "d"]] assert edit_hint(plan) == "Page 1: [Mara, Jon] then [Kit]\nPage 2: no edits" good = {"characters": [{"name": n, "turnaround_prompt": "x"} for n in ("Mara", "Jon", "Kit")], "pages": [{"n": 1, "orientation": "landscape", "base_prompt": "x", "edits": [{"characters": ["Mara", "Jon"], "prompt": "x"}, {"characters": ["Kit"], "prompt": "x"}]}, {"n": 2, "orientation": "portrait", "base_prompt": "x", "edits": []}]} assert check_json(good, plan) == [] regrouped = json.loads(json.dumps(good)) regrouped["pages"][0]["edits"] = [{"characters": ["Mara"], "prompt": "x"}, {"characters": ["Jon", "Kit"], "prompt": "x"}] assert check_json(regrouped, plan) == [ "page 1: edit groups [['Mara'], ['Jon', 'Kit']], expected [['Mara', 'Jon'], ['Kit']]"] short = json.loads(json.dumps(good)) del short["characters"][2], short["pages"][1] assert check_json(short, plan) == ["characters differ from plan.md", "pages differ from plan.md"] ``` - [ ] **Step 2: Run test to verify it fails** Run: `./fanfictioner --selftest` Expected: FAIL with `NameError: name 'edit_groups' is not defined` - [ ] **Step 3: Write minimal implementation** Add after `validate_plan`: ```python def edit_groups(chars): """Qwen edit takes the scene plus at most 2 reference sheets, so chain groups of 2.""" return [chars[i:i + 2] for i in range(0, len(chars), 2)] def edit_hint(plan): return "\n".join(f"Page {p['n']}: " + (" then ".join("[" + ", ".join(g) + "]" for g in edit_groups(p["characters"])) or "no edits") for p in plan["pages"]) def check_json(pj, plan): errs = [] if [c["name"] for c in pj["characters"]] != list(plan["characters"]): errs.append("characters differ from plan.md") want = {p["n"]: p["characters"] for p in plan["pages"]} if sorted(p["n"] for p in pj["pages"]) != sorted(want): errs.append("pages differ from plan.md") for p in pj["pages"]: got, exp = [e["characters"] for e in p["edits"]], edit_groups(want.get(p["n"], [])) if got != exp: errs.append(f"page {p['n']}: edit groups {got}, expected {exp}") return errs ``` - [ ] **Step 4: Run test to verify it passes** Run: `./fanfictioner --selftest` Expected: `selftest ok` - [ ] **Step 5: Commit** ```bash git add fanfictioner git commit -m "Add edit grouping and plan.json consistency checks" ``` --- ### Task 5: Model profiles and sd-cli argument building **Files:** - Modify: `fanfictioner` - [ ] **Step 1: Write the failing test** Append to `selftest`, before the `with` block: ```python q = sd_args("qwen_edit", "edit it", Path("c/7_%d.png"), (832, 1216), 7, [Path("in1.png"), Path("mara.png")]) assert q[0] == "sd-cli" assert q[q.index("--diffusion-model") + 1] == str(SD_DIR / "Qwen-Image-2.1-viggle-turbo-v0.3-6step-Q8_0.gguf") assert q[q.index("--llm_vision") + 1] == str(SD_DIR / "mmproj-Qwen3VL-8B-Instruct-F16.gguf") assert q[q.index("--backend") + 1] == "te=cpu,diffusion=vulkan0,vae=vulkan0" assert "--offload-to-cpu" not in q assert q[q.index("-r") + 1] == "in1.png" and q.count("-r") == 2 for flag, val in (("-p", "edit it"), ("-W", "832"), ("-H", "1216"), ("-s", "7"), ("-b", "3"), ("-o", "c/7_%d.png"), ("--output-begin-idx", "1"), ("--steps", "6")): assert q[q.index(flag) + 1] == val, (flag, q) z = sd_args("zimage", "x", Path("o_%d.png"), SIZES["turnaround"], 1) assert "--offload-to-cpu" in z and "-r" not in z assert z[z.index("--vae") + 1] == str(SD_DIR / "vae/flux1-ae.safetensors") k = sd_args("krea2", "x", Path("o_%d.png"), SIZES["portrait"], 1) assert k[k.index("--llm") + 1] == str(SD_DIR / "Qwen3VL-4B-Instruct-Q8_0.gguf") assert k[k.index("--steps") + 1] == "8" and k[k.index("--cfg-scale") + 1] == "1.0" ``` - [ ] **Step 2: Run test to verify it fails** Run: `./fanfictioner --selftest` Expected: FAIL with `NameError: name 'sd_args' is not defined` - [ ] **Step 3: Write minimal implementation** Add below `CONF` at the top of the file: ```python # Settings verified by hand on 2026-10-04 (Krea2 re-checked against upstream on 2026-10-05). # Model paths are relative to SD_DIR. All profiles run at cfg 1.0, so no negative prompts. PROFILES = { "zimage": ["--diffusion-model", "z_image_turbo-Q8_0.gguf", "--vae", "vae/flux1-ae.safetensors", "--llm", "Qwen3-4B-Instruct-2507-Q8_0.gguf", "--cfg-scale", "1.0", "--steps", "8", "--diffusion-fa", "--offload-to-cpu"], "krea2": ["--diffusion-model", "Krea-2-Turbo-Q6_K.gguf", "--vae", "vae/wan_2.1_vae.safetensors", "--llm", "Qwen3VL-4B-Instruct-Q8_0.gguf", "--cfg-scale", "1.0", "--steps", "8", "--diffusion-fa", "--offload-to-cpu"], # te=cpu alone puts diffusion on the iGPU and crashes: always name all three backends. # --offload-to-cpu here OOM-killed the desktop: never add it. "qwen_edit": ["--diffusion-model", "Qwen-Image-2.1-viggle-turbo-v0.3-6step-Q8_0.gguf", "--vae", "vae/qwen_image_2.1_vae_bf16.safetensors", "--llm", "Qwen3VL-8B-Instruct-Q8_0.gguf", "--llm_vision", "mmproj-Qwen3VL-8B-Instruct-F16.gguf", "--cfg-scale", "1.0", "--steps", "6", "--sigmas", "1.0,0.9375,0.875,0.75,0.5,0.25,0.0", "--sampling-method", "euler", "--fa", "--backend", "te=cpu,diffusion=vulkan0,vae=vulkan0", "--mmap", "--vae-tiling"], } MODEL_FLAGS = {"--diffusion-model", "--vae", "--llm", "--llm_vision"} SIZES = {"portrait": (832, 1216), "landscape": (1216, 832), "turnaround": (1216, 832)} # Edit inputs are shrunk first: full-size refs blow past 12 GB VRAM. SHRINK = {"portrait": "576x832", "landscape": "832x576", "ref": "560x384"} ``` Add after `check_json`: ```python def sd_args(profile, prompt, out, size, seed, refs=()): args = ["sd-cli"] it = iter(PROFILES[profile]) for x in it: args.append(x) if x in MODEL_FLAGS: args.append(str(SD_DIR / next(it))) args += ["-p", prompt, "-W", str(size[0]), "-H", str(size[1]), "-s", str(seed), "-b", "3", "-o", str(out), "--output-begin-idx", "1"] for r in refs: args += ["-r", str(r)] return args ``` - [ ] **Step 4: Run test to verify it passes** Run: `./fanfictioner --selftest` Expected: `selftest ok` - [ ] **Step 5: Verify every profile model file exists on this machine** Run: `python3 -c "import runpy; f=runpy.run_path('fanfictioner'); [print(p, a, (f['SD_DIR']/b).exists()) for p,v in f['PROFILES'].items() for a,b in zip(v,v[1:]) if a in f['MODEL_FLAGS']]"` Expected: every line ends in `True`. - [ ] **Step 6: Commit** ```bash git add fanfictioner git commit -m "Add model profiles and sd-cli argument building" ``` --- ### Task 6: sd-cli runner with RAM watchdog, review loop **Files:** - Modify: `fanfictioner` - [ ] **Step 1: Write the failing test** Append inside the `with tempfile.TemporaryDirectory() as t:` block of `selftest`, at its end: ```python assert mem_available_kb("MemTotal: 30000000 kB\nMemAvailable: 1500000 kB\n") == 1500000 assert run_sd(["true"], t / "ok.log") is None assert (t / "ok.log").read_text().startswith("true") assert run_sd(["sh", "-c", "echo '[ERROR] boom'"], t / "e.log") == "sd-cli failed (exit 0)" assert run_sd(["false"], t / "f.log") == "sd-cli failed (exit 1)" ``` - [ ] **Step 2: Run test to verify it fails** Run: `./fanfictioner --selftest` Expected: FAIL with `NameError: name 'mem_available_kb' is not defined` - [ ] **Step 3: Write minimal implementation** Add below `SHRINK`: ```python MIN_RAM_KB = 2 * 1024 * 1024 # kill sd-cli below 2 GB MemAvailable, before the OOM killer hits the desktop ``` Add after `sd_args`: ```python def mem_available_kb(meminfo): return int(re.search(r"MemAvailable:\s+(\d+)", meminfo)[1]) def run_sd(args, log): """Run sd-cli with output in log; return an error string or None.""" with open(log, "w") as f: f.write(shlex.join(args) + "\n") f.flush() proc = subprocess.Popen(args, stdout=f, stderr=subprocess.STDOUT) try: while proc.poll() is None: time.sleep(2) if proc.poll() is None and \ mem_available_kb(Path("/proc/meminfo").read_text()) < MIN_RAM_KB: proc.kill() proc.wait() return "sd-cli killed by RAM watchdog (MemAvailable below 2 GB)" except KeyboardInterrupt: proc.kill() proc.wait() raise if proc.returncode or "[ERROR" in Path(log).read_text(errors="replace"): return f"sd-cli failed (exit {proc.returncode})" return None def editor(path): subprocess.run([*shlex.split(os.environ.get("EDITOR", "vi")), str(path)]) def generate(d, profile, prompt, size, refs=()): """One sd-cli call producing 3 candidates in d. Returns (files, seed, error).""" unload_llm() d.mkdir(parents=True, exist_ok=True) seed = random.randint(0, 2**31 - 4) log = d / f"{seed}.log" print(f"sd-cli {profile}, seed {seed} ...", flush=True) err = run_sd(sd_args(profile, prompt, d / f"{seed}_%d.png", size, seed, refs), log) files = [d / f"{seed}_{i}.png" for i in (1, 2, 3)] if not err and not all(f.exists() for f in files): err = "sd-cli wrote no output" if err: print(err, *log.read_text(errors="replace").splitlines()[-15:], sep="\n") return files, seed, err def show(files, seed, d): m = d / f"{seed}_montage.png" labelled = [x for i, f in enumerate(files) for x in ("-label", f"{i + 1} seed {seed + i}", str(f))] subprocess.run(["magick", "montage", "-pointsize", "32", *labelled, "-geometry", "+8+8", "-tile", "3x1", str(m)], check=True) subprocess.run(["kitty", "+kitten", "icat", str(m)]) def review(d, profile, prompt, size, refs=(), skip=False): """Generate, show, ask. Returns the picked file, or None when the user skips.""" keys = "1-3 pick, r regen, p edit prompt" + (", s skip edit" if skip else "") + ", q quit" while True: files, seed, err = generate(d, profile, prompt, size, refs) if not err: show(files, seed, d) while True: k = input(f"[{keys}] > ").strip().lower() if k in ("1", "2", "3") and not err: return files[int(k) - 1] if k == "r": break if k == "p": (d / "prompt.txt").write_text(prompt) editor(d / "prompt.txt") prompt = (d / "prompt.txt").read_text().strip() break if k == "s" and skip: return None if k == "q": sys.exit("quit, run again to resume") ``` `unload_llm` is defined in Task 7. Until then add this stub right above `generate` so the file runs: ```python def unload_llm(): pass ``` - [ ] **Step 4: Run test to verify it passes** Run: `./fanfictioner --selftest` Expected: `selftest ok` (takes about 6 s: each `run_sd` call polls once at 2 s) - [ ] **Step 5: Manual check of the review loop with real sd-cli** Run in a kitty terminal: ```bash python3 -c " import runpy; from pathlib import Path f = runpy.run_path('fanfictioner') print(f['review'](Path('/tmp/ff-review'), 'zimage', 'a red fox in snow, watercolor. No text.', (832, 1216)))" ``` Expected: about 7 minutes of sd-cli, then a 3-up montage with labels `1 seed N`, `2 seed N+1`, `3 seed N+2` in the terminal, then the key prompt. Press `2`; it prints `/tmp/ff-review/_2.png`. Then `rm -r /tmp/ff-review`. - [ ] **Step 6: Commit** ```bash git add fanfictioner git commit -m "Run sd-cli under a RAM watchdog and review candidates" ``` --- ### Task 7: llama-server client and plan stage **Files:** - Modify: `fanfictioner` - [ ] **Step 1: Write the failing test** Append to `selftest`, before the `with` block: ```python assert strip_fences("```markdown\n# T\nx\n```\n") == "# T\nx" assert strip_fences("# T\nx") == "# T\nx" ``` - [ ] **Step 2: Run test to verify it fails** Run: `./fanfictioner --selftest` Expected: FAIL with `NameError: name 'strip_fences' is not defined` - [ ] **Step 3: Write the implementation** Delete the `unload_llm` stub from Task 6. Add these constants below `MIN_RAM_KB`: ```python DRAFT_SYS = """You turn a story idea into a page plan for a text-free illustrated book. Each page is one full-page image. A scene may span several pages. Images never contain text: no speech bubbles, captions, signs or labels. Tell the story through action, expression and setting. Give every character one fixed visual description (age, hair, face, eyes, build, outfit) and never vary it between pages. If a page has a Model: line, keep it unchanged. Write the plan in exactly this Markdown format, with nothing before or after it: # Series: <series folder name> Book: <book folder name> Style: <one line describing the art style of every page> ## Characters ### <Name> <fixed visual description> ## Pages ### 1. <short title> <what happens, setting, mood> Characters: <Name>, <Name> (or: none) Pose: <body orientation and pose of each named character> Framing: <camera angle, shot size, portrait or landscape> """ COMPILE_SYS = """You convert an illustrated-book plan into image-generation prompts, as JSON. characters: one entry per character of the plan, in plan order, name exactly as written. turnaround_prompt: "Character turnaround reference sheet. The same <full visual description> shown three times side by side, full body: front view, side view, back view. Neutral standing pose, arms relaxed. Plain light-grey background, even studio lighting. <Style line>. No text, no labels." pages: one entry per page, n as in the plan. orientation: portrait or landscape, from Framing. base_prompt: the Style line, then the scene, then the full visual description of every character present, then their Pose, then the Framing, ending with "No text, no speech bubbles." edits: exactly the edit groups listed after the plan, in that order, names exactly as written. In an edit, image 1 is the scene, image 2 is the reference sheet of the first listed character, image 3 that of the second. For each character the prompt says: "In image 1, change only <name>'s head: give her/him the face, eyes and hair of the person in image <2 or 3>. Keep <name>'s exact pose from image 1: <that character's Pose>." and it ends with: "Keep bodies, clothing, other people, background, lighting and art style of image 1 unchanged." A page with no characters has an empty edits list. """ SCHEMA = { "type": "object", "required": ["characters", "pages"], "properties": { "characters": {"type": "array", "items": { "type": "object", "required": ["name", "turnaround_prompt"], "properties": {"name": {"type": "string"}, "turnaround_prompt": {"type": "string"}}}}, "pages": {"type": "array", "items": { "type": "object", "required": ["n", "orientation", "base_prompt", "edits"], "properties": { "n": {"type": "integer"}, "orientation": {"enum": ["portrait", "landscape"]}, "base_prompt": {"type": "string"}, "edits": {"type": "array", "items": { "type": "object", "required": ["characters", "prompt"], "properties": {"characters": {"type": "array", "items": {"type": "string"}, "maxItems": 2}, "prompt": {"type": "string"}}}}}}}, }, } ``` Add after `review`: ```python def llm(path, body=None): data = json.dumps(body).encode() if body is not None else None req = urllib.request.Request(LLM_URL + path, data=data, headers={"Content-Type": "application/json"}) with urllib.request.urlopen(req, timeout=900) as r: return json.load(r) def check_llm(): try: ids = [m["id"] for m in llm("/v1/models")["data"]] except OSError as e: sys.exit(f"llama-server unreachable at {LLM_URL}: {e}") if LLM_MODEL not in ids: sys.exit(f"llama-server does not offer {LLM_MODEL} (has: {', '.join(ids)})") def unload_llm(): """Free VRAM for sd-cli. Idempotent: an unloaded or unreachable server is fine.""" try: llm("/models/unload", {"model": LLM_MODEL}) except OSError: pass def chat(system, user, schema=None): body = {"model": LLM_MODEL, "messages": [{"role": "system", "content": system}, {"role": "user", "content": user}]} if schema: body["response_format"] = {"type": "json_schema", "json_schema": {"name": "plan", "schema": schema}} print("asking Gemma ...", flush=True) return llm("/v1/chat/completions", body)["choices"][0]["message"]["content"] def strip_fences(s): return re.sub(r"^```\w*\n|\n```$", "", s.strip()) def compile_plan(md, plan): """plan.md -> plan.json dict, or None when Gemma twice returns groups that do not match.""" for _ in range(2): pj = json.loads(chat(COMPILE_SYS, f"{md}\n\nEdit groups, in order:\n{edit_hint(plan)}", SCHEMA)) errs = check_json(pj, plan) if not errs: break print("compile mismatch:", *errs, sep="\n ") else: return None models = {p["n"]: p["model"] for p in plan["pages"]} pj.update(title=plan["title"], series=plan["series"], book=plan["book"]) for p in pj["pages"]: p["model"] = models[p["n"]] return pj def plan_stage(story, wd): md_path, js_path = wd / "plan.md", wd / "plan.json" if js_path.exists() and md_path.exists() and js_path.stat().st_mtime >= md_path.stat().st_mtime: return json.loads(js_path.read_text()) check_llm() wd.mkdir(exist_ok=True) if not md_path.exists(): md_path.write_text(strip_fences(chat(DRAFT_SYS, story.read_text())) + "\n") while True: md = md_path.read_text() plan = parse_plan(md) errs = validate_plan(plan) print(f"\n{md}\n--- {md_path}: {len(plan['characters'])} characters, {len(plan['pages'])} pages") for e in errs: print(" !", e) k = input("[c]onfirm [e]dit [g]emma revise [q]uit > ").strip().lower() if k == "c" and not errs: pj = compile_plan(md, plan) if pj: js_path.write_text(json.dumps(pj, indent=2) + "\n") return pj elif k == "e": editor(md_path) elif k == "g": ins = input("instruction for Gemma: ") md_path.write_text(strip_fences(chat( DRAFT_SYS, f"Current plan:\n\n{md}\n\nRevise it: {ins}\nReturn the complete new plan.")) + "\n") elif k == "q": sys.exit(0) ``` - [ ] **Step 4: Run test to verify it passes** Run: `./fanfictioner --selftest` Expected: `selftest ok` - [ ] **Step 5: Manual check against the live server** ```bash mkdir -p /tmp/ff && printf 'Two friends, Mara and Jon, climb to a city rooftop at dusk and find a lost kitten.\n' > /tmp/ff/story.md python3 -c " import runpy; from pathlib import Path f = runpy.run_path('fanfictioner') pj = f['plan_stage'](Path('/tmp/ff/story.md'), Path('/tmp/ff/story')) print(pj['series'], pj['book'], [(p['n'], p['model'], [e['characters'] for e in p['edits']]) for p in pj['pages']])" ``` Expected: plan.md printed, no `!` lines (or fix with `e`). Press `g`, give an instruction, check the plan changes. Press `c`; `/tmp/ff/story/plan.json` is written and the print shows series, book and edit groups matching the plan's `Characters:` lines. Run the same command again: it returns at once without asking (plan.json newer than plan.md). Keep `/tmp/ff` for Task 8. - [ ] **Step 6: Commit** ```bash git add fanfictioner git commit -m "Draft, review and compile the plan with Gemma" ``` --- ### Task 8: Image stages and main wiring **Files:** - Modify: `fanfictioner` - [ ] **Step 1: Implement the stages** Add after `plan_stage`: ```python def shrink(src, dst, geom): dst.parent.mkdir(parents=True, exist_ok=True) subprocess.run(["magick", str(src), "-resize", geom, str(dst)], check=True) return dst def refs_stage(chars, wd): for c in chars: print(f"\n== reference sheet: {c['name']}") pick = review(wd / "cand" / f"ref_{c['name']}", "zimage", c["turnaround_prompt"], SIZES["turnaround"]) keep(pick, wd / "refs" / f"{c['name']}.png") def page_stage(p, wd, out, base): """Base scene, then chained head edits; each pick kept in cand/pNNN so a quit resumes mid-page.""" d = wd / "cand" / f"p{p['n']:03d}" size = SIZES[p["orientation"]] scene = d / "base.png" if not scene.exists(): print(f"\n== page {p['n']}: base scene") keep(review(d / "base", p["model"] or base, p["base_prompt"], size), scene) for i, e in enumerate(p["edits"], 1): nxt = d / f"edit{i}.png" if not nxt.exists(): print(f"\n== page {p['n']}: edit {i} ({', '.join(e['characters'])})") refs = [shrink(scene, d / f"edit{i}_in.png", SHRINK[p["orientation"]])] refs += [shrink(wd / "refs" / f"{n}.png", wd / "cand" / "refs_small" / f"{n}.png", SHRINK["ref"]) for n in e["characters"]] pick = review(d / f"edit{i}", "qwen_edit", e["prompt"], size, refs, skip=True) keep(pick or scene, nxt) scene = nxt keep(scene, out) print(f"page {p['n']} -> {out}") ``` - [ ] **Step 2: Wire `main`** Replace the end of `main` (after the `ap.error` line) so the whole function reads: ```python def main(): ap = argparse.ArgumentParser(description="Turn story.md into a text-free illustrated book.") ap.add_argument("story", nargs="?", type=Path, help="story idea in free prose") ap.add_argument("--base", choices=("zimage", "krea2"), default="zimage", help="base model for pages without a Model: line") ap.add_argument("--selftest", action="store_true", help="run built-in checks and exit") a = ap.parse_args() if a.selftest: return selftest() if not a.story: ap.error("story file required") try: lib = load_conf()["LIBRARY"] wd = a.story.with_suffix("") pj = plan_stage(a.story, wd) chars, pages = todo(pj, wd, lib) refs_stage(chars, wd) for p in pages: page_stage(p, wd, page_path(lib, pj, p["n"]), a.base) print(f"\ndone: {page_path(lib, pj, 1).parent}") except KeyboardInterrupt: sys.exit("\ninterrupted, run again to resume") ``` - [ ] **Step 3: Selftest still passes** Run: `./fanfictioner --selftest` Expected: `selftest ok` - [ ] **Step 4: End-to-end run (user at the keyboard, about 1 hour of GPU time)** Use the `/tmp/ff` story from Task 7. To keep it short, first edit `/tmp/ff/story/plan.md` down to one page with two characters (this makes plan.md newer than plan.json, so it is recompiled: confirm with `c`). Run in kitty: `./fanfictioner /tmp/ff/story.md` Check, in order: 1. One reference-sheet review per character; pick one. `/tmp/ff/story/refs/<Name>.png` appears. 2. Page 1 base review; press `r` once and see new seeds; pick. `cand/p001/base.png` appears. 3. Press `q` at the edit review. Rerun the same command: it skips refs and base and goes straight to the edit (resume works). 4. Edit review shows heads changed, poses kept; pick. `$LIBRARY/<Series>/<Book>/001.png` exists. 5. Rerun: prints `done:` with nothing generated. Then delete the test book from the library: `rm -r "$LIBRARY/<Series>/<Book>"` (look at the path first) and `rm -r /tmp/ff`. - [ ] **Step 5: Commit** ```bash git add fanfictioner git commit -m "Generate reference sheets and pages into the library" ``` --- ## Self-review notes - Spec coverage: pipeline (Tasks 7, 8), files and resume (2, 8), plan.md format and validation (3), Gemma draft/revise/compile with json_schema (7), edit groups at most 2 and chaining (4, 8), profiles, sizes, shrink, no offload on edit, all-three backends (5, 8), `-b 3` with seed labels (6), review keys 1-3/r/p/s/q (6), error handling: server unreachable (7), validation (3, 7), sd-cli failure and `[ERROR` (6), RAM watchdog (6), unload before every sd-cli call (6, 7), Ctrl-C (6, 8), CLI (1, 8), selftest coverage list (2-6), license header and README (1). - Open check during Task 8: edit output size is the full page size (832x1216 / 1216x832). If the edit OOMs or the watchdog fires there, drop the edit output to the shrunk scene size and upscale is out of scope; report back instead of guessing.