diff options
| author | Danilo M. <danix@danix.xyz> | 2026-10-05 09:31:59 +0200 |
|---|---|---|
| committer | Danilo M. <danix@danix.xyz> | 2026-10-05 09:31:59 +0200 |
| commit | 33aaeb67006a46ac8da79dbc365e31bc1fa0746f (patch) | |
| tree | 6080d58e698320e7b9a092abfc010fee21c118ad | |
| parent | bb21fb90871dc631071a62ec9367c31fbda2623d (diff) | |
| download | fanfictioner-33aaeb67006a46ac8da79dbc365e31bc1fa0746f.tar.gz fanfictioner-33aaeb67006a46ac8da79dbc365e31bc1fa0746f.zip | |
Add fanfictioner implementation plan
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
| -rw-r--r-- | docs/superpowers/plans/2026-10-05-fanfictioner.md | 1054 |
1 files changed, 1054 insertions, 0 deletions
diff --git a/docs/superpowers/plans/2026-10-05-fanfictioner.md b/docs/superpowers/plans/2026-10-05-fanfictioner.md new file mode 100644 index 0000000..2caabee --- /dev/null +++ b/docs/superpowers/plans/2026-10-05-fanfictioner.md @@ -0,0 +1,1054 @@ +# fanfictioner Implementation Plan + +> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking. + +**Goal:** Build `fanfictioner`, a single stdlib-only Python 3 script that turns `story.md` into a text-free illustrated book of page images, as described in `docs/superpowers/specs/2026-10-04-fanfictioner-design.md`. + +**Architecture:** One executable file `fanfictioner` at the repo root. Pure helpers (config, plan.md parsing/validation, edit grouping, plan.json checks, sd-cli argument building, meminfo parsing) are covered by `--selftest` asserts. Thin I/O layers wrap llama-server (urllib), sd-cli (subprocess + RAM watchdog), `magick` and `kitty +kitten icat`. All state is on disk; resume is derived from files. + +**Tech Stack:** Python 3 stdlib (argparse, json, pathlib, re, shlex, shutil, subprocess, urllib.request), sd-cli (stable-diffusion.cpp), ImageMagick 7 `magick`, kitty icat, llama-server router on `localhost:8181`. + +--- + +## Facts the engineer needs + +- `sd-cli -b 3 -s SEED -o dir/X_%d.png --output-begin-idx 1` writes `X_1.png`..`X_3.png`; image *b* (0-based) uses seed `SEED + b` (verified in sd.cpp `src/pipeline/image.cpp`: `cur_seed = request.seed + b`). +- Krea2 Turbo: 8 steps, CFG off (`--cfg-scale 1.0` in sd-cli), shift 1.15 is sd.cpp's built-in default for Krea2. Verified 2026-10-05. +- `--backend te=cpu,diffusion=vulkan0,vae=vulkan0` must list all three, or diffusion lands on the iGPU and crashes. Never `--offload-to-cpu` on the edit profile. +- llama-server: `GET /v1/models` lists models (`data[].id`), `POST /models/unload {"model": ...}` unloads, `POST /v1/chat/completions` with `response_format: {"type": "json_schema", ...}` constrains output. +- `~/.config/book-reader.conf` holds shell lines like `LIBRARY="/some/dir"`. +- Commits are GPG-signed by global git config. Just run `git commit`; the first one in a session prompts for the PIN. + +## Deviations from the spec (deliberate, smaller) + +- plan.json `title`/`series`/`book` and each page's `model` are taken from the parsed plan.md by the script, not from Gemma. Folder names and model choice must not depend on LLM copying. A page's `model` is `""` unless plan.md has `Model:`; then `--base` applies. +- The script computes edit groups (`edit_groups`) and passes them to Gemma; `check_json` rejects a compile whose groups differ, with one retry. +- Picks are also kept per page in `cand/pNNN/base.png` and `cand/pNNN/editK.png`, so a quit mid-page resumes without regenerating the stages already picked. Still disk-derived, no state file. + +## File structure + +- Create: `fanfictioner` (executable script, everything) +- Create: `README.md` +- Unchanged: `LICENSE`, spec + +--- + +### Task 1: Skeleton, selftest runner, README + +**Files:** +- Create: `fanfictioner` +- Create: `README.md` + +- [ ] **Step 1: Create the script skeleton** + +`fanfictioner`: + +```python +#!/usr/bin/env python3 +# fanfictioner: turn a story idea into a text-free illustrated book. +# Copyright (C) 2026 Danilo M. <danix@danix.xyz> +# +# This program is free software; you can redistribute it and/or modify +# it under the terms of the GNU General Public License version 2 as +# published by the Free Software Foundation. +# +# This program is distributed in the hope that it will be useful, +# but WITHOUT ANY WARRANTY; without even the implied warranty of +# MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE. See the +# GNU General Public License for more details. +# +# You should have received a copy of the GNU General Public License along +# with this program; if not, write to the Free Software Foundation, Inc., +# 51 Franklin Street, Fifth Floor, Boston, MA 02110-1301 USA. + +import argparse +import json +import os +import random +import re +import shlex +import shutil +import subprocess +import sys +import tempfile +import time +import urllib.request +from pathlib import Path + +LLM_URL = "http://localhost:8181" +LLM_MODEL = "Gemma4-12B-qat-mtp" +SD_DIR = Path("/data/LLM-models/SD") +CONF = Path.home() / ".config/book-reader.conf" + + +def selftest(): + print("selftest ok") + + +def main(): + ap = argparse.ArgumentParser(description="Turn story.md into a text-free illustrated book.") + ap.add_argument("story", nargs="?", type=Path, help="story idea in free prose") + ap.add_argument("--base", choices=("zimage", "krea2"), default="zimage", + help="base model for pages without a Model: line") + ap.add_argument("--selftest", action="store_true", help="run built-in checks and exit") + a = ap.parse_args() + if a.selftest: + return selftest() + if not a.story: + ap.error("story file required") + + +if __name__ == "__main__": + main() +``` + +- [ ] **Step 2: Make it executable and run it** + +Run: `chmod +x fanfictioner && ./fanfictioner --selftest` +Expected: `selftest ok` + +Run: `./fanfictioner` +Expected: usage line and `error: story file required`, exit 2. + +- [ ] **Step 3: Write README.md** + +```markdown +# fanfictioner + +Turn a story idea into a text-free illustrated book. A local Gemma model +(llama-server) plans the pages and writes the image prompts; `sd-cli` +(stable-diffusion.cpp) draws them. You pick among three candidates at every +image stage. Pages land in the book-reader library as `NNN.png`. + +## Requirements + +- Python 3, stdlib only +- llama-server in router mode on `localhost:8181` serving `Gemma4-12B-qat-mtp` +- `sd-cli` with Vulkan, models under `/data/LLM-models/SD` +- ImageMagick 7 (`magick`), kitty (`kitty +kitten icat`), `$EDITOR` +- `~/.config/book-reader.conf` with a `LIBRARY="..."` line + +## Usage + + fanfictioner story.md [--base zimage|krea2] + fanfictioner --selftest + +Work files go to a directory named after the story (`stories/rooftop.md` -> +`stories/rooftop/`): `plan.md`, `plan.json`, `refs/`, `cand/`. Quit any time +with `q` or Ctrl-C; the next run resumes from what is on disk. + +Plan review keys: `c` confirm, `e` edit in `$EDITOR`, `g` ask Gemma to revise, +`q` quit. Image review keys: `1`-`3` pick, `r` regenerate, `p` edit prompt, +`s` skip edit (edit stage only), `q` quit. + +## License + +Copyright (C) 2026 Danilo M. <danix@danix.xyz> + +This program is free software; you can redistribute it and/or modify it under +the terms of the GNU General Public License version 2 as published by the Free +Software Foundation. See `LICENSE`. + +## Development Approach + +This project is developed using AI-assisted tools. Code is generated with the help of AI based on human-provided specifications, design decisions, and iterative feedback. + +All contributions are reviewed, tested, and curated by the maintainer before being included in the codebase. AI is used as a productivity and exploration tool, while human oversight remains central to all decisions. + +The goal is to combine the flexibility of AI-assisted development with standard open-source practices such as transparency, review, and accountability. +``` + +- [ ] **Step 4: Commit** + +```bash +git add fanfictioner README.md +git commit -m "Add fanfictioner skeleton and README" +``` + +--- + +### Task 2: Config and resume detection + +**Files:** +- Modify: `fanfictioner` (add helpers above `selftest`, asserts inside `selftest`) + +- [ ] **Step 1: Write the failing test** + +Replace `selftest` with: + +```python +def selftest(): + with tempfile.TemporaryDirectory() as t: + t = Path(t) + (t / "c.conf").write_text('# comment\nLIBRARY="/data/my lib"\nPORT=8642\n') + conf = load_conf(t / "c.conf") + assert conf == {"LIBRARY": "/data/my lib", "PORT": "8642"}, conf + + pj = {"series": "S", "book": "B", + "characters": [{"name": "Mara"}, {"name": "Jon"}], + "pages": [{"n": 1}, {"n": 2}]} + lib = t / "lib" + assert page_path(lib, pj, 7) == lib / "S" / "B" / "007.png" + (t / "refs").mkdir() + (t / "refs" / "Mara.png").touch() + page_path(lib, pj, 1).parent.mkdir(parents=True) + page_path(lib, pj, 1).touch() + refs, pages = todo(pj, t, lib) + assert [c["name"] for c in refs] == ["Jon"], refs + assert [p["n"] for p in pages] == [2], pages + + keep(t / "refs" / "Mara.png", t / "deep" / "x.png") + assert (t / "deep" / "x.png").exists() + print("selftest ok") +``` + +- [ ] **Step 2: Run test to verify it fails** + +Run: `./fanfictioner --selftest` +Expected: FAIL with `NameError: name 'load_conf' is not defined` + +- [ ] **Step 3: Write minimal implementation** + +Add above `selftest`: + +```python +def load_conf(path=CONF): + conf = {} + for line in path.read_text().splitlines(): + if "=" in line and not line.lstrip().startswith("#"): + k, v = line.split("=", 1) + conf[k.strip()] = shlex.split(v)[0] if v.strip() else "" + return conf + + +def page_path(lib, pj, n): + return Path(lib) / pj["series"] / pj["book"] / f"{n:03d}.png" + + +def todo(pj, wd, lib): + """What is left to do, derived from disk: (characters without a ref, pages not in the library).""" + refs = [c for c in pj["characters"] if not (wd / "refs" / f"{c['name']}.png").exists()] + pages = [p for p in pj["pages"] if not page_path(lib, pj, p["n"]).exists()] + return refs, pages + + +def keep(src, dst): + """Copy src to dst atomically, so a half-written file never counts as done.""" + dst.parent.mkdir(parents=True, exist_ok=True) + tmp = dst.with_name(dst.name + ".tmp") + shutil.copyfile(src, tmp) + os.replace(tmp, dst) + return dst +``` + +- [ ] **Step 4: Run test to verify it passes** + +Run: `./fanfictioner --selftest` +Expected: `selftest ok` + +- [ ] **Step 5: Commit** + +```bash +git add fanfictioner +git commit -m "Add config loading and disk-derived resume detection" +``` + +--- + +### Task 3: plan.md parsing and validation + +**Files:** +- Modify: `fanfictioner` + +- [ ] **Step 1: Write the failing test** + +Add this constant above `selftest`: + +```python +SAMPLE_PLAN = """# Rooftop +Series: Test Series +Book: Book One +Style: soft watercolor, muted colors + +## Characters + +### Mara +Young woman, long red hair, +green eyes, grey coat. + +### Jon +Tall man, short beard, blue jacket. + +### Kit +Small girl, black bob, yellow raincoat. + +## Pages + +### 1. Arrival +They reach the roof at dusk. +Characters: Mara, Jon, Kit +Pose: Mara left facing right, Jon center facing camera, Kit right sitting. +Framing: wide shot, eye level, landscape + +### 2. Empty roof +Nobody is there any more. +Characters: none +Framing: wide shot, portrait +Model: krea2 +""" +``` + +Add at the start of `selftest`, before the `with` block: + +```python + plan = parse_plan(SAMPLE_PLAN) + assert (plan["title"], plan["series"], plan["book"]) == ("Rooftop", "Test Series", "Book One"), plan + assert plan["style"] == "soft watercolor, muted colors" + assert list(plan["characters"]) == ["Mara", "Jon", "Kit"] + assert plan["characters"]["Mara"] == "Young woman, long red hair, green eyes, grey coat." + p1, p2 = plan["pages"] + assert (p1["n"], p1["title"], p1["characters"]) == (1, "Arrival", ["Mara", "Jon", "Kit"]) + assert p1["text"] == "They reach the roof at dusk." and p1["pose"].startswith("Mara left") + assert p1["model"] == "" and p2["model"] == "krea2" and p2["characters"] == [] + assert validate_plan(plan) == [] + + bad = parse_plan(SAMPLE_PLAN.replace("Book: Book One\n", "") + .replace("Characters: none", "Characters: Zed") + .replace("Model: krea2", "Model: sdxl") + .replace("### 2.", "### 1.")) + errs = validate_plan(bad) + assert "missing book" in errs, errs + assert "page 1: unknown character Zed" in errs, errs + assert "page 1: Pose missing" in errs, errs + assert "page 1: unknown model sdxl" in errs, errs + assert "page numbers must be unique and start at 1" in errs, errs + assert validate_plan(parse_plan("")) == ["missing title", "missing series", "missing book", + "no characters", "no pages"] +``` + +- [ ] **Step 2: Run test to verify it fails** + +Run: `./fanfictioner --selftest` +Expected: FAIL with `NameError: name 'parse_plan' is not defined` + +- [ ] **Step 3: Write minimal implementation** + +Add above `SAMPLE_PLAN`: + +```python +def parse_plan(text): + plan = {"title": "", "series": "", "book": "", "style": "", "characters": {}, "pages": []} + section = cur = None + for line in text.splitlines(): + s = line.strip() + if s.startswith("# ") and not plan["title"]: + plan["title"] = s[2:].strip() + elif s.startswith("## "): + section, cur = s[3:].strip().lower(), None + elif s.startswith("### "): + head = s[4:].strip() + if section == "characters": + cur = plan["characters"].setdefault(head, []) + elif section == "pages": + m = re.match(r"(\d+)\.\s*(.*)", head) + cur = {"n": int(m[1]) if m else 0, "title": m[2] if m else head, "text": [], + "characters": [], "pose": "", "framing": "", "model": ""} + plan["pages"].append(cur) + elif section is None and (m := re.match(r"(Series|Book|Style):\s*(.*)", s)): + plan[m[1].lower()] = m[2].strip() + elif section == "pages" and cur is not None and \ + (m := re.match(r"(Characters|Pose|Framing|Model):\s*(.*)", s)): + key, val = m[1].lower(), m[2].strip() + if key == "characters": + val = [] if val.lower() == "none" else [c.strip() for c in val.split(",") if c.strip()] + elif key == "model": + val = val.split()[0].lower() if val else "" + cur[key] = val + elif cur is not None and s: + (cur if isinstance(cur, list) else cur["text"]).append(s) + plan["characters"] = {k: " ".join(v) for k, v in plan["characters"].items()} + for p in plan["pages"]: + p["text"] = " ".join(p["text"]) + return plan + + +def validate_plan(plan): + errs = [f"missing {k}" for k in ("title", "series", "book") if not plan[k]] + if not plan["characters"]: + errs.append("no characters") + if not plan["pages"]: + errs.append("no pages") + ns = [p["n"] for p in plan["pages"]] + if len(set(ns)) != len(ns) or 0 in ns: + errs.append("page numbers must be unique and start at 1") + for p in plan["pages"]: + errs += [f"page {p['n']}: unknown character {c}" for c in p["characters"] + if c not in plan["characters"]] + if p["characters"] and not p["pose"]: + errs.append(f"page {p['n']}: Pose missing") + if p["model"] and p["model"] not in ("zimage", "krea2"): + errs.append(f"page {p['n']}: unknown model {p['model']}") + return errs +``` + +- [ ] **Step 4: Run test to verify it passes** + +Run: `./fanfictioner --selftest` +Expected: `selftest ok` + +- [ ] **Step 5: Commit** + +```bash +git add fanfictioner +git commit -m "Parse and validate plan.md" +``` + +--- + +### Task 4: Edit grouping and plan.json checks + +**Files:** +- Modify: `fanfictioner` + +- [ ] **Step 1: Write the failing test** + +Append to `selftest`, after the plan.md asserts and before the `with` block: + +```python + assert edit_groups([]) == [] + assert edit_groups(["Mara", "Jon", "Kit"]) == [["Mara", "Jon"], ["Kit"]] + assert edit_groups(list("abcd")) == [["a", "b"], ["c", "d"]] + assert edit_hint(plan) == "Page 1: [Mara, Jon] then [Kit]\nPage 2: no edits" + + good = {"characters": [{"name": n, "turnaround_prompt": "x"} for n in ("Mara", "Jon", "Kit")], + "pages": [{"n": 1, "orientation": "landscape", "base_prompt": "x", + "edits": [{"characters": ["Mara", "Jon"], "prompt": "x"}, + {"characters": ["Kit"], "prompt": "x"}]}, + {"n": 2, "orientation": "portrait", "base_prompt": "x", "edits": []}]} + assert check_json(good, plan) == [] + regrouped = json.loads(json.dumps(good)) + regrouped["pages"][0]["edits"] = [{"characters": ["Mara"], "prompt": "x"}, + {"characters": ["Jon", "Kit"], "prompt": "x"}] + assert check_json(regrouped, plan) == [ + "page 1: edit groups [['Mara'], ['Jon', 'Kit']], expected [['Mara', 'Jon'], ['Kit']]"] + short = json.loads(json.dumps(good)) + del short["characters"][2], short["pages"][1] + assert check_json(short, plan) == ["characters differ from plan.md", "pages differ from plan.md"] +``` + +- [ ] **Step 2: Run test to verify it fails** + +Run: `./fanfictioner --selftest` +Expected: FAIL with `NameError: name 'edit_groups' is not defined` + +- [ ] **Step 3: Write minimal implementation** + +Add after `validate_plan`: + +```python +def edit_groups(chars): + """Qwen edit takes the scene plus at most 2 reference sheets, so chain groups of 2.""" + return [chars[i:i + 2] for i in range(0, len(chars), 2)] + + +def edit_hint(plan): + return "\n".join(f"Page {p['n']}: " + (" then ".join("[" + ", ".join(g) + "]" + for g in edit_groups(p["characters"])) + or "no edits") + for p in plan["pages"]) + + +def check_json(pj, plan): + errs = [] + if [c["name"] for c in pj["characters"]] != list(plan["characters"]): + errs.append("characters differ from plan.md") + want = {p["n"]: p["characters"] for p in plan["pages"]} + if sorted(p["n"] for p in pj["pages"]) != sorted(want): + errs.append("pages differ from plan.md") + for p in pj["pages"]: + got, exp = [e["characters"] for e in p["edits"]], edit_groups(want.get(p["n"], [])) + if got != exp: + errs.append(f"page {p['n']}: edit groups {got}, expected {exp}") + return errs +``` + +- [ ] **Step 4: Run test to verify it passes** + +Run: `./fanfictioner --selftest` +Expected: `selftest ok` + +- [ ] **Step 5: Commit** + +```bash +git add fanfictioner +git commit -m "Add edit grouping and plan.json consistency checks" +``` + +--- + +### Task 5: Model profiles and sd-cli argument building + +**Files:** +- Modify: `fanfictioner` + +- [ ] **Step 1: Write the failing test** + +Append to `selftest`, before the `with` block: + +```python + q = sd_args("qwen_edit", "edit it", Path("c/7_%d.png"), (832, 1216), 7, + [Path("in1.png"), Path("mara.png")]) + assert q[0] == "sd-cli" + assert q[q.index("--diffusion-model") + 1] == str(SD_DIR / "Qwen-Image-2.1-viggle-turbo-v0.3-6step-Q8_0.gguf") + assert q[q.index("--llm_vision") + 1] == str(SD_DIR / "mmproj-Qwen3VL-8B-Instruct-F16.gguf") + assert q[q.index("--backend") + 1] == "te=cpu,diffusion=vulkan0,vae=vulkan0" + assert "--offload-to-cpu" not in q + assert q[q.index("-r") + 1] == "in1.png" and q.count("-r") == 2 + for flag, val in (("-p", "edit it"), ("-W", "832"), ("-H", "1216"), ("-s", "7"), ("-b", "3"), + ("-o", "c/7_%d.png"), ("--output-begin-idx", "1"), ("--steps", "6")): + assert q[q.index(flag) + 1] == val, (flag, q) + z = sd_args("zimage", "x", Path("o_%d.png"), SIZES["turnaround"], 1) + assert "--offload-to-cpu" in z and "-r" not in z + assert z[z.index("--vae") + 1] == str(SD_DIR / "vae/flux1-ae.safetensors") + k = sd_args("krea2", "x", Path("o_%d.png"), SIZES["portrait"], 1) + assert k[k.index("--llm") + 1] == str(SD_DIR / "Qwen3VL-4B-Instruct-Q8_0.gguf") + assert k[k.index("--steps") + 1] == "8" and k[k.index("--cfg-scale") + 1] == "1.0" +``` + +- [ ] **Step 2: Run test to verify it fails** + +Run: `./fanfictioner --selftest` +Expected: FAIL with `NameError: name 'sd_args' is not defined` + +- [ ] **Step 3: Write minimal implementation** + +Add below `CONF` at the top of the file: + +```python +# Settings verified by hand on 2026-10-04 (Krea2 re-checked against upstream on 2026-10-05). +# Model paths are relative to SD_DIR. All profiles run at cfg 1.0, so no negative prompts. +PROFILES = { + "zimage": ["--diffusion-model", "z_image_turbo-Q8_0.gguf", + "--vae", "vae/flux1-ae.safetensors", + "--llm", "Qwen3-4B-Instruct-2507-Q8_0.gguf", + "--cfg-scale", "1.0", "--steps", "8", "--diffusion-fa", "--offload-to-cpu"], + "krea2": ["--diffusion-model", "Krea-2-Turbo-Q6_K.gguf", + "--vae", "vae/wan_2.1_vae.safetensors", + "--llm", "Qwen3VL-4B-Instruct-Q8_0.gguf", + "--cfg-scale", "1.0", "--steps", "8", "--diffusion-fa", "--offload-to-cpu"], + # te=cpu alone puts diffusion on the iGPU and crashes: always name all three backends. + # --offload-to-cpu here OOM-killed the desktop: never add it. + "qwen_edit": ["--diffusion-model", "Qwen-Image-2.1-viggle-turbo-v0.3-6step-Q8_0.gguf", + "--vae", "vae/qwen_image_2.1_vae_bf16.safetensors", + "--llm", "Qwen3VL-8B-Instruct-Q8_0.gguf", + "--llm_vision", "mmproj-Qwen3VL-8B-Instruct-F16.gguf", + "--cfg-scale", "1.0", "--steps", "6", + "--sigmas", "1.0,0.9375,0.875,0.75,0.5,0.25,0.0", "--sampling-method", "euler", + "--fa", "--backend", "te=cpu,diffusion=vulkan0,vae=vulkan0", "--mmap", "--vae-tiling"], +} +MODEL_FLAGS = {"--diffusion-model", "--vae", "--llm", "--llm_vision"} +SIZES = {"portrait": (832, 1216), "landscape": (1216, 832), "turnaround": (1216, 832)} +# Edit inputs are shrunk first: full-size refs blow past 12 GB VRAM. +SHRINK = {"portrait": "576x832", "landscape": "832x576", "ref": "560x384"} +``` + +Add after `check_json`: + +```python +def sd_args(profile, prompt, out, size, seed, refs=()): + args = ["sd-cli"] + it = iter(PROFILES[profile]) + for x in it: + args.append(x) + if x in MODEL_FLAGS: + args.append(str(SD_DIR / next(it))) + args += ["-p", prompt, "-W", str(size[0]), "-H", str(size[1]), "-s", str(seed), + "-b", "3", "-o", str(out), "--output-begin-idx", "1"] + for r in refs: + args += ["-r", str(r)] + return args +``` + +- [ ] **Step 4: Run test to verify it passes** + +Run: `./fanfictioner --selftest` +Expected: `selftest ok` + +- [ ] **Step 5: Verify every profile model file exists on this machine** + +Run: `python3 -c "import runpy; f=runpy.run_path('fanfictioner'); [print(p, a, (f['SD_DIR']/b).exists()) for p,v in f['PROFILES'].items() for a,b in zip(v,v[1:]) if a in f['MODEL_FLAGS']]"` +Expected: every line ends in `True`. + +- [ ] **Step 6: Commit** + +```bash +git add fanfictioner +git commit -m "Add model profiles and sd-cli argument building" +``` + +--- + +### Task 6: sd-cli runner with RAM watchdog, review loop + +**Files:** +- Modify: `fanfictioner` + +- [ ] **Step 1: Write the failing test** + +Append inside the `with tempfile.TemporaryDirectory() as t:` block of `selftest`, at its end: + +```python + assert mem_available_kb("MemTotal: 30000000 kB\nMemAvailable: 1500000 kB\n") == 1500000 + assert run_sd(["true"], t / "ok.log") is None + assert (t / "ok.log").read_text().startswith("true") + assert run_sd(["sh", "-c", "echo '[ERROR] boom'"], t / "e.log") == "sd-cli failed (exit 0)" + assert run_sd(["false"], t / "f.log") == "sd-cli failed (exit 1)" +``` + +- [ ] **Step 2: Run test to verify it fails** + +Run: `./fanfictioner --selftest` +Expected: FAIL with `NameError: name 'mem_available_kb' is not defined` + +- [ ] **Step 3: Write minimal implementation** + +Add below `SHRINK`: + +```python +MIN_RAM_KB = 2 * 1024 * 1024 # kill sd-cli below 2 GB MemAvailable, before the OOM killer hits the desktop +``` + +Add after `sd_args`: + +```python +def mem_available_kb(meminfo): + return int(re.search(r"MemAvailable:\s+(\d+)", meminfo)[1]) + + +def run_sd(args, log): + """Run sd-cli with output in log; return an error string or None.""" + with open(log, "w") as f: + f.write(shlex.join(args) + "\n") + f.flush() + proc = subprocess.Popen(args, stdout=f, stderr=subprocess.STDOUT) + try: + while proc.poll() is None: + time.sleep(2) + if proc.poll() is None and \ + mem_available_kb(Path("/proc/meminfo").read_text()) < MIN_RAM_KB: + proc.kill() + proc.wait() + return "sd-cli killed by RAM watchdog (MemAvailable below 2 GB)" + except KeyboardInterrupt: + proc.kill() + proc.wait() + raise + if proc.returncode or "[ERROR" in Path(log).read_text(errors="replace"): + return f"sd-cli failed (exit {proc.returncode})" + return None + + +def editor(path): + subprocess.run([*shlex.split(os.environ.get("EDITOR", "vi")), str(path)]) + + +def generate(d, profile, prompt, size, refs=()): + """One sd-cli call producing 3 candidates in d. Returns (files, seed, error).""" + unload_llm() + d.mkdir(parents=True, exist_ok=True) + seed = random.randint(0, 2**31 - 4) + log = d / f"{seed}.log" + print(f"sd-cli {profile}, seed {seed} ...", flush=True) + err = run_sd(sd_args(profile, prompt, d / f"{seed}_%d.png", size, seed, refs), log) + files = [d / f"{seed}_{i}.png" for i in (1, 2, 3)] + if not err and not all(f.exists() for f in files): + err = "sd-cli wrote no output" + if err: + print(err, *log.read_text(errors="replace").splitlines()[-15:], sep="\n") + return files, seed, err + + +def show(files, seed, d): + m = d / f"{seed}_montage.png" + labelled = [x for i, f in enumerate(files) for x in ("-label", f"{i + 1} seed {seed + i}", str(f))] + subprocess.run(["magick", "montage", "-pointsize", "32", *labelled, + "-geometry", "+8+8", "-tile", "3x1", str(m)], check=True) + subprocess.run(["kitty", "+kitten", "icat", str(m)]) + + +def review(d, profile, prompt, size, refs=(), skip=False): + """Generate, show, ask. Returns the picked file, or None when the user skips.""" + keys = "1-3 pick, r regen, p edit prompt" + (", s skip edit" if skip else "") + ", q quit" + while True: + files, seed, err = generate(d, profile, prompt, size, refs) + if not err: + show(files, seed, d) + while True: + k = input(f"[{keys}] > ").strip().lower() + if k in ("1", "2", "3") and not err: + return files[int(k) - 1] + if k == "r": + break + if k == "p": + (d / "prompt.txt").write_text(prompt) + editor(d / "prompt.txt") + prompt = (d / "prompt.txt").read_text().strip() + break + if k == "s" and skip: + return None + if k == "q": + sys.exit("quit, run again to resume") +``` + +`unload_llm` is defined in Task 7. Until then add this stub right above `generate` so the file runs: + +```python +def unload_llm(): + pass +``` + +- [ ] **Step 4: Run test to verify it passes** + +Run: `./fanfictioner --selftest` +Expected: `selftest ok` (takes about 6 s: each `run_sd` call polls once at 2 s) + +- [ ] **Step 5: Manual check of the review loop with real sd-cli** + +Run in a kitty terminal: + +```bash +python3 -c " +import runpy; from pathlib import Path +f = runpy.run_path('fanfictioner') +print(f['review'](Path('/tmp/ff-review'), 'zimage', 'a red fox in snow, watercolor. No text.', (832, 1216)))" +``` + +Expected: about 7 minutes of sd-cli, then a 3-up montage with labels `1 seed N`, `2 seed N+1`, `3 seed N+2` in the terminal, then the key prompt. Press `2`; it prints `/tmp/ff-review/<seed>_2.png`. Then `rm -r /tmp/ff-review`. + +- [ ] **Step 6: Commit** + +```bash +git add fanfictioner +git commit -m "Run sd-cli under a RAM watchdog and review candidates" +``` + +--- + +### Task 7: llama-server client and plan stage + +**Files:** +- Modify: `fanfictioner` + +- [ ] **Step 1: Write the failing test** + +Append to `selftest`, before the `with` block: + +```python + assert strip_fences("```markdown\n# T\nx\n```\n") == "# T\nx" + assert strip_fences("# T\nx") == "# T\nx" +``` + +- [ ] **Step 2: Run test to verify it fails** + +Run: `./fanfictioner --selftest` +Expected: FAIL with `NameError: name 'strip_fences' is not defined` + +- [ ] **Step 3: Write the implementation** + +Delete the `unload_llm` stub from Task 6. Add these constants below `MIN_RAM_KB`: + +```python +DRAFT_SYS = """You turn a story idea into a page plan for a text-free illustrated book. +Each page is one full-page image. A scene may span several pages. +Images never contain text: no speech bubbles, captions, signs or labels. Tell the story +through action, expression and setting. +Give every character one fixed visual description (age, hair, face, eyes, build, outfit) +and never vary it between pages. +If a page has a Model: line, keep it unchanged. +Write the plan in exactly this Markdown format, with nothing before or after it: + +# <Title> +Series: <series folder name> +Book: <book folder name> +Style: <one line describing the art style of every page> + +## Characters + +### <Name> +<fixed visual description> + +## Pages + +### 1. <short title> +<what happens, setting, mood> +Characters: <Name>, <Name> (or: none) +Pose: <body orientation and pose of each named character> +Framing: <camera angle, shot size, portrait or landscape> +""" + +COMPILE_SYS = """You convert an illustrated-book plan into image-generation prompts, as JSON. + +characters: one entry per character of the plan, in plan order, name exactly as written. +turnaround_prompt: "Character turnaround reference sheet. The same <full visual description> +shown three times side by side, full body: front view, side view, back view. Neutral standing +pose, arms relaxed. Plain light-grey background, even studio lighting. <Style line>. +No text, no labels." + +pages: one entry per page, n as in the plan. orientation: portrait or landscape, from Framing. +base_prompt: the Style line, then the scene, then the full visual description of every +character present, then their Pose, then the Framing, ending with +"No text, no speech bubbles." + +edits: exactly the edit groups listed after the plan, in that order, names exactly as written. +In an edit, image 1 is the scene, image 2 is the reference sheet of the first listed character, +image 3 that of the second. For each character the prompt says: +"In image 1, change only <name>'s head: give her/him the face, eyes and hair of the person in +image <2 or 3>. Keep <name>'s exact pose from image 1: <that character's Pose>." +and it ends with: "Keep bodies, clothing, other people, background, lighting and art style of +image 1 unchanged." +A page with no characters has an empty edits list. +""" + +SCHEMA = { + "type": "object", "required": ["characters", "pages"], + "properties": { + "characters": {"type": "array", "items": { + "type": "object", "required": ["name", "turnaround_prompt"], + "properties": {"name": {"type": "string"}, "turnaround_prompt": {"type": "string"}}}}, + "pages": {"type": "array", "items": { + "type": "object", "required": ["n", "orientation", "base_prompt", "edits"], + "properties": { + "n": {"type": "integer"}, + "orientation": {"enum": ["portrait", "landscape"]}, + "base_prompt": {"type": "string"}, + "edits": {"type": "array", "items": { + "type": "object", "required": ["characters", "prompt"], + "properties": {"characters": {"type": "array", "items": {"type": "string"}, + "maxItems": 2}, + "prompt": {"type": "string"}}}}}}}, + }, +} +``` + +Add after `review`: + +```python +def llm(path, body=None): + data = json.dumps(body).encode() if body is not None else None + req = urllib.request.Request(LLM_URL + path, data=data, headers={"Content-Type": "application/json"}) + with urllib.request.urlopen(req, timeout=900) as r: + return json.load(r) + + +def check_llm(): + try: + ids = [m["id"] for m in llm("/v1/models")["data"]] + except OSError as e: + sys.exit(f"llama-server unreachable at {LLM_URL}: {e}") + if LLM_MODEL not in ids: + sys.exit(f"llama-server does not offer {LLM_MODEL} (has: {', '.join(ids)})") + + +def unload_llm(): + """Free VRAM for sd-cli. Idempotent: an unloaded or unreachable server is fine.""" + try: + llm("/models/unload", {"model": LLM_MODEL}) + except OSError: + pass + + +def chat(system, user, schema=None): + body = {"model": LLM_MODEL, + "messages": [{"role": "system", "content": system}, {"role": "user", "content": user}]} + if schema: + body["response_format"] = {"type": "json_schema", "json_schema": {"name": "plan", "schema": schema}} + print("asking Gemma ...", flush=True) + return llm("/v1/chat/completions", body)["choices"][0]["message"]["content"] + + +def strip_fences(s): + return re.sub(r"^```\w*\n|\n```$", "", s.strip()) + + +def compile_plan(md, plan): + """plan.md -> plan.json dict, or None when Gemma twice returns groups that do not match.""" + for _ in range(2): + pj = json.loads(chat(COMPILE_SYS, f"{md}\n\nEdit groups, in order:\n{edit_hint(plan)}", SCHEMA)) + errs = check_json(pj, plan) + if not errs: + break + print("compile mismatch:", *errs, sep="\n ") + else: + return None + models = {p["n"]: p["model"] for p in plan["pages"]} + pj.update(title=plan["title"], series=plan["series"], book=plan["book"]) + for p in pj["pages"]: + p["model"] = models[p["n"]] + return pj + + +def plan_stage(story, wd): + md_path, js_path = wd / "plan.md", wd / "plan.json" + if js_path.exists() and md_path.exists() and js_path.stat().st_mtime >= md_path.stat().st_mtime: + return json.loads(js_path.read_text()) + check_llm() + wd.mkdir(exist_ok=True) + if not md_path.exists(): + md_path.write_text(strip_fences(chat(DRAFT_SYS, story.read_text())) + "\n") + while True: + md = md_path.read_text() + plan = parse_plan(md) + errs = validate_plan(plan) + print(f"\n{md}\n--- {md_path}: {len(plan['characters'])} characters, {len(plan['pages'])} pages") + for e in errs: + print(" !", e) + k = input("[c]onfirm [e]dit [g]emma revise [q]uit > ").strip().lower() + if k == "c" and not errs: + pj = compile_plan(md, plan) + if pj: + js_path.write_text(json.dumps(pj, indent=2) + "\n") + return pj + elif k == "e": + editor(md_path) + elif k == "g": + ins = input("instruction for Gemma: ") + md_path.write_text(strip_fences(chat( + DRAFT_SYS, f"Current plan:\n\n{md}\n\nRevise it: {ins}\nReturn the complete new plan.")) + "\n") + elif k == "q": + sys.exit(0) +``` + +- [ ] **Step 4: Run test to verify it passes** + +Run: `./fanfictioner --selftest` +Expected: `selftest ok` + +- [ ] **Step 5: Manual check against the live server** + +```bash +mkdir -p /tmp/ff && printf 'Two friends, Mara and Jon, climb to a city rooftop at dusk and find a lost kitten.\n' > /tmp/ff/story.md +python3 -c " +import runpy; from pathlib import Path +f = runpy.run_path('fanfictioner') +pj = f['plan_stage'](Path('/tmp/ff/story.md'), Path('/tmp/ff/story')) +print(pj['series'], pj['book'], [(p['n'], p['model'], [e['characters'] for e in p['edits']]) for p in pj['pages']])" +``` + +Expected: plan.md printed, no `!` lines (or fix with `e`). Press `g`, give an instruction, check the plan changes. Press `c`; `/tmp/ff/story/plan.json` is written and the print shows series, book and edit groups matching the plan's `Characters:` lines. Run the same command again: it returns at once without asking (plan.json newer than plan.md). Keep `/tmp/ff` for Task 8. + +- [ ] **Step 6: Commit** + +```bash +git add fanfictioner +git commit -m "Draft, review and compile the plan with Gemma" +``` + +--- + +### Task 8: Image stages and main wiring + +**Files:** +- Modify: `fanfictioner` + +- [ ] **Step 1: Implement the stages** + +Add after `plan_stage`: + +```python +def shrink(src, dst, geom): + dst.parent.mkdir(parents=True, exist_ok=True) + subprocess.run(["magick", str(src), "-resize", geom, str(dst)], check=True) + return dst + + +def refs_stage(chars, wd): + for c in chars: + print(f"\n== reference sheet: {c['name']}") + pick = review(wd / "cand" / f"ref_{c['name']}", "zimage", c["turnaround_prompt"], SIZES["turnaround"]) + keep(pick, wd / "refs" / f"{c['name']}.png") + + +def page_stage(p, wd, out, base): + """Base scene, then chained head edits; each pick kept in cand/pNNN so a quit resumes mid-page.""" + d = wd / "cand" / f"p{p['n']:03d}" + size = SIZES[p["orientation"]] + scene = d / "base.png" + if not scene.exists(): + print(f"\n== page {p['n']}: base scene") + keep(review(d / "base", p["model"] or base, p["base_prompt"], size), scene) + for i, e in enumerate(p["edits"], 1): + nxt = d / f"edit{i}.png" + if not nxt.exists(): + print(f"\n== page {p['n']}: edit {i} ({', '.join(e['characters'])})") + refs = [shrink(scene, d / f"edit{i}_in.png", SHRINK[p["orientation"]])] + refs += [shrink(wd / "refs" / f"{n}.png", wd / "cand" / "refs_small" / f"{n}.png", SHRINK["ref"]) + for n in e["characters"]] + pick = review(d / f"edit{i}", "qwen_edit", e["prompt"], size, refs, skip=True) + keep(pick or scene, nxt) + scene = nxt + keep(scene, out) + print(f"page {p['n']} -> {out}") +``` + +- [ ] **Step 2: Wire `main`** + +Replace the end of `main` (after the `ap.error` line) so the whole function reads: + +```python +def main(): + ap = argparse.ArgumentParser(description="Turn story.md into a text-free illustrated book.") + ap.add_argument("story", nargs="?", type=Path, help="story idea in free prose") + ap.add_argument("--base", choices=("zimage", "krea2"), default="zimage", + help="base model for pages without a Model: line") + ap.add_argument("--selftest", action="store_true", help="run built-in checks and exit") + a = ap.parse_args() + if a.selftest: + return selftest() + if not a.story: + ap.error("story file required") + try: + lib = load_conf()["LIBRARY"] + wd = a.story.with_suffix("") + pj = plan_stage(a.story, wd) + chars, pages = todo(pj, wd, lib) + refs_stage(chars, wd) + for p in pages: + page_stage(p, wd, page_path(lib, pj, p["n"]), a.base) + print(f"\ndone: {page_path(lib, pj, 1).parent}") + except KeyboardInterrupt: + sys.exit("\ninterrupted, run again to resume") +``` + +- [ ] **Step 3: Selftest still passes** + +Run: `./fanfictioner --selftest` +Expected: `selftest ok` + +- [ ] **Step 4: End-to-end run (user at the keyboard, about 1 hour of GPU time)** + +Use the `/tmp/ff` story from Task 7. To keep it short, first edit `/tmp/ff/story/plan.md` down to one page with two characters (this makes plan.md newer than plan.json, so it is recompiled: confirm with `c`). + +Run in kitty: `./fanfictioner /tmp/ff/story.md` + +Check, in order: +1. One reference-sheet review per character; pick one. `/tmp/ff/story/refs/<Name>.png` appears. +2. Page 1 base review; press `r` once and see new seeds; pick. `cand/p001/base.png` appears. +3. Press `q` at the edit review. Rerun the same command: it skips refs and base and goes straight to the edit (resume works). +4. Edit review shows heads changed, poses kept; pick. `$LIBRARY/<Series>/<Book>/001.png` exists. +5. Rerun: prints `done:` with nothing generated. + +Then delete the test book from the library: `rm -r "$LIBRARY/<Series>/<Book>"` (look at the path first) and `rm -r /tmp/ff`. + +- [ ] **Step 5: Commit** + +```bash +git add fanfictioner +git commit -m "Generate reference sheets and pages into the library" +``` + +--- + +## Self-review notes + +- Spec coverage: pipeline (Tasks 7, 8), files and resume (2, 8), plan.md format and validation (3), Gemma draft/revise/compile with json_schema (7), edit groups at most 2 and chaining (4, 8), profiles, sizes, shrink, no offload on edit, all-three backends (5, 8), `-b 3` with seed labels (6), review keys 1-3/r/p/s/q (6), error handling: server unreachable (7), validation (3, 7), sd-cli failure and `[ERROR` (6), RAM watchdog (6), unload before every sd-cli call (6, 7), Ctrl-C (6, 8), CLI (1, 8), selftest coverage list (2-6), license header and README (1). +- Open check during Task 8: edit output size is the full page size (832x1216 / 1216x832). If the edit OOMs or the watchdog fires there, drop the edit output to the shrunk scene size and upscale is out of scope; report back instead of guessing. |
