# fanfictioner design Date: 2026-10-04 ## Goal Turn a free-prose story idea (`story.md`) into an illustrated, text-free book of page images, read with `../book-reader`. A local Gemma model plans the pages; the script turns the plan into image prompts from fixed templates; `sd-cli` (stable-diffusion.cpp) draws them. The user reviews and picks among candidates at every image stage. ## Constraints - Single Python 3 file, stdlib only. External tools: `sd-cli`, `magick`, `kitty +kitten icat`, `$EDITOR`. - GPU: Intel Arc B580, 12 GB VRAM (Vulkan0). Vulkan1 is the Radeon iGPU and must never be used. RAM: 30 GB, shared with the desktop. - llama-server runs in router mode on `localhost:8181`, model `Gemma4-12B-qat-mtp`, `--models-max 1`. Unload with `POST /models/unload {"model": ...}`; it reloads on the next request. - No text in images: no speech bubbles, captions or labels. Only exception: a page with a `Title:` line (the cover) gets that title as lettering. ## Pipeline ``` story.md ──Gemma draft──▶ plan.md ◀─┐ [c]onfirm / [e]dit in $EDITOR / [g]emma revise / [q]uit │ confirm ──script compile (templates, no LLM)──▶ plan.json │ unload Gemma │ refs: per character: sd-cli -b 3 turnaround sheet (Z-Image) → pick → refs/.png pages: per page: base: sd-cli -b 3 (Z-Image, or Krea2 on demand) → pick edit: sd-cli -b 3 (Qwen-Image 2.1 viggle-turbo v0.3, -r scene -r refs…) → pick → $LIBRARY///NNN.png ``` ## Files Work dir: a directory named after the story file, next to it (`stories/rooftop.md` → `stories/rooftop/`): - `plan.md`, user-facing plan (source of truth until confirmed) - `plan.json`, compiled from `plan.md` on confirm; recompiled when `plan.md` is newer than `plan.json` - `refs/.png`, chosen turnaround sheet per character - `cand//…`, candidates, montages and sd-cli logs Output: `$LIBRARY///NNN.png` (zero-padded page number), or `CC-NNN.png` when the story sets a chapter number, so all chapters of a book share its folder and read in order. `LIBRARY` is read from `~/.config/book-reader.conf` (shell `KEY="value"` lines). Series and Book come from plan.md. story.md may start with a `---` frontmatter (`series`, `book`, `chapter`, `chapter_title`, `title`, all optional). The script writes these into plan.md after every draft and revision and whenever story.md changes, overriding Gemma: Series/Book/Chapter header lines, the `# ` heading, and page 1's `Title:`. plan.json is rebuilt when plan.md or story.md is newer than it. Resume is derived from disk, no state file: a character with `refs/.png` is skipped; a page whose `NNN.png` exists in the library is skipped. ## plan.md format Written by Gemma, edited by the user. Fixed headings, free prose underneath: ```markdown # Series: <series folder name> Book: <book folder name> Style: <global art style line> ## Characters ### <Name> Pronoun: she|he|they <fixed visual description: who (young woman, old man...), age, hair, face, eyes, build, outfit> ## Pages ### 1. <short title> <what happens, setting, mood> Characters: <Name>, <Name> (or "none") Pose: <Name>: <position in frame, orientation, pose>; <Name>: <...> Framing: <camera angle, shot size>, portrait|landscape Model: krea2 (optional, overrides the base model for this page) Title: <book title> (cover only: lettering drawn on this page) ``` Validation before confirm: title, Series, Book, at least one character and one page; every character has `Pronoun:` she, he or they; every page has a `Characters:` line and every name in it exists under `## Characters`; `Pose:` names a pose for each listed character; `Framing:` says portrait or landscape; page numbers run 1..N; names are safe as folder and file names. Failures are listed; the user edits or asks Gemma to revise. ## Gemma calls All via `POST localhost:8181/v1/chat/completions`, `model: Gemma4-12B-qat-mtp`. 1. **Draft**: system prompt holds the plan.md format and rules (no text in images, fixed character descriptions starting with who they are, characters are people only while animals and objects belong to the scene, one page = one image, a scene may span several pages, page 1 is the cover with a `Title:` line). User message: `story.md`. Output: plan.md. 2. **Revise**: current plan.md + user's instruction -> complete new plan.md. A revision that fails validation is saved as `plan.rejected.md` and plan.md is kept; an accepted one keeps the previous version as `plan.md.bak`. ## Compile (script, no LLM) On confirm the script builds plan.json from the parsed plan.md: ```json { "title": "", "series": "", "book": "", "characters": [{"name": "", "turnaround_prompt": ""}], "pages": [{ "n": 1, "model": "", "orientation": "portrait", "base_prompt": "", "edits": [{"characters": ["Mara", "Jon"], "prompt": ""}] }] } ``` Templates: - `turnaround_prompt`: "Character turnaround reference sheet of one person: <description> The same person shown three times side by side, full body: front view, side view, back view. Neutral standing pose, arms relaxed. Plain light-grey background, even studio lighting. <Style>. No text, no labels." - `base_prompt`: Style, page text, then for each character present its description followed by its pose, then Framing, then "No text, no speech bubbles." A page with `Title:` instead ends with "Title lettering at the top of the image reading exactly "<title>". No other text, no speech bubbles.", and its edits also keep the title lettering unchanged. - `edits`: one entry per group of at most 2 characters (Qwen takes at most 3 reference images: the scene plus 2 refs), in `Characters:` order. Per character: "In image 1, change only <Name>'s head: give her/him/them the face, eyes and hair of the person in image N. Keep <Name>'s exact pose from image 1: <pose>." Then "Keep bodies, clothing, other people, background, lighting and art style of image 1 unchanged." Pages with no characters have no edits. - `orientation` comes from Framing; `model` from the page's `Model:` line (empty means `--base`). Gemma writes the plan only; the script owns every prompt and sd-cli setting. ## Model profiles A dict at the top of the script. Paths under `/data/LLM-models/SD`. Settings verified by hand on 2026-10-04: | profile | use | key args | |---|---|---| | `zimage` (default base, refs) | t2i | `--diffusion-model z_image_turbo-Q8_0.gguf --vae vae/flux1-ae.safetensors --llm Qwen3-4B-Instruct-2507-Q8_0.gguf --cfg-scale 1.0 --steps 8 --diffusion-fa --offload-to-cpu` | | `krea2` (on demand: `--base krea2` or `Model: krea2`) | t2i | `--diffusion-model Krea-2-Turbo-Q6_K.gguf --vae vae/wan_2.1_vae.safetensors --llm Qwen3VL-4B-Instruct-Q8_0.gguf --cfg-scale 1.0 --steps 8 --diffusion-fa --offload-to-cpu` | | `qwen_edit` | edit | `--diffusion-model Qwen-Image-2.1-viggle-turbo-v0.3-6step-Q8_0.gguf --vae vae/qwen_image_2.1_vae_bf16.safetensors --llm Qwen3VL-8B-Instruct-Q8_0.gguf --llm_vision mmproj-Qwen3VL-8B-Instruct-F16.gguf --cfg-scale 1.0 --steps 6 --sigmas 1.0,0.9375,0.875,0.75,0.5,0.25,0.0 --sampling-method euler --fa --backend te=cpu,diffusion=vulkan0,vae=vulkan0 --mmap --vae-tiling` | Sizes: portrait 832x1216, landscape 1216x832, turnaround 1216x832. Before edit, reference images are shrunk with `magick -resize`: scene to 576x832 (portrait) or 832x576 (landscape), turnaround refs to 560x384. `--backend te=cpu` without explicit `diffusion=`/`vae=` lands on the iGPU and crashes; always pass all three. Edit with `--offload-to-cpu` and full-size refs OOM-killed the desktop; do not use it there. Candidates: `-b 3` per call (one model load), seed random per call and recorded in the montage label so a pick can be reproduced. No negative prompts: all profiles run at cfg 1.0, where they have no effect. Measured timings: Z-Image ~150 s/image, Krea2 ~220 s/image, Qwen edit ~410 s (140 s of it CPU text encoding, done once per call). ## Review loop For each stage, the script builds a montage of the candidates with `magick montage` (labels 1..3 + seed) and shows it with `kitty +kitten icat`. Keys: - `1`-`3`: pick - `r`: regenerate with new seeds - `p`: edit this prompt in `$EDITOR`, then regenerate - `s` (edit stage only): skip the edit, keep the base image - `q`: quit (resume later) Chained edits (more than 2 characters): each edit group gets its own review; the pick from one group is the scene input of the next. ## Error handling - llama-server unreachable or model unknown: exit with a clear message, no files written. - plan.md validation failure: list problems, return to the review loop. - sd-cli non-zero exit, missing output, or `[ERROR` in its log: print the last log lines, offer `r` retry or `q` quit. Logs stay in `cand/`. - RAM watchdog: while sd-cli runs, poll `MemAvailable` every 2 s; below 2 GB, kill sd-cli and report it. - Unload Gemma before every sd-cli stage (idempotent). - Ctrl-C anywhere is safe; the next run resumes from disk. ## CLI ``` fanfictioner story.md [--base zimage|krea2] [--selftest] ``` ## Testing `fanfictioner --selftest`: plain asserts, no GPU, no server. Covers plan.md parsing and validation, edit grouping for more than 2 characters, plan.json templates, resume detection, and sd-cli argument building per profile. ## Out of scope Text, speech bubbles and captions; multi-panel page layouts; LoRA training; sd-server daemon; any GUI beyond the terminal. ## License GPLv2-only. LICENSE from gnu.org, header in the script, License section and Development Approach section in README.