From bb21fb90871dc631071a62ec9367c31fbda2623d Mon Sep 17 00:00:00 2001 From: "Danilo M." Date: Sun, 4 Oct 2026 20:12:17 +0200 Subject: Add GPLv2 license and design spec Design for fanfictioner: story.md -> Gemma plan.md -> sd-cli page images with per-stage review, character refs applied via Qwen-Image edit. Co-Authored-By: Claude Opus 5.5 --- .../specs/2026-10-04-fanfictioner-design.md | 198 +++++++++++++++++++++ 1 file changed, 198 insertions(+) create mode 100644 docs/superpowers/specs/2026-10-04-fanfictioner-design.md (limited to 'docs') diff --git a/docs/superpowers/specs/2026-10-04-fanfictioner-design.md b/docs/superpowers/specs/2026-10-04-fanfictioner-design.md new file mode 100644 index 0000000..cd620e8 --- /dev/null +++ b/docs/superpowers/specs/2026-10-04-fanfictioner-design.md @@ -0,0 +1,198 @@ +# fanfictioner design + +Date: 2026-10-04 + +## Goal + +Turn a free-prose story idea (`story.md`) into an illustrated, text-free book of +page images, read with `../book-reader`. A local Gemma model plans the pages and +writes the image prompts; `sd-cli` (stable-diffusion.cpp) draws them. The user +reviews and picks among candidates at every image stage. + +## Constraints + +- Single Python 3 file, stdlib only. External tools: `sd-cli`, `magick`, + `kitty +kitten icat`, `$EDITOR`. +- GPU: Intel Arc B580, 12 GB VRAM (Vulkan0). Vulkan1 is the Radeon iGPU and must + never be used. RAM: 30 GB, shared with the desktop. +- llama-server runs in router mode on `localhost:8181`, model + `Gemma4-12B-qat-mtp`, `--models-max 1`. Unload with + `POST /models/unload {"model": ...}`; it reloads on the next request. +- No text in images: no speech bubbles, captions or labels. + +## Pipeline + +``` +story.md ──Gemma draft──▶ plan.md ◀─┐ [c]onfirm / [e]dit in $EDITOR / [g]emma revise / [q]uit + │ + confirm ──Gemma compile (json_schema)──▶ plan.json + │ + unload Gemma + │ + refs: per character: sd-cli -b 3 turnaround sheet (Z-Image) → pick → refs/.png + pages: per page: + base: sd-cli -b 3 (Z-Image, or Krea2 on demand) → pick + edit: sd-cli -b 3 (Qwen-Image 2.1 viggle-turbo v0.3, -r scene -r refs…) → pick + → $LIBRARY///NNN.png +``` + +## Files + +Work dir: a directory named after the story file, next to it +(`stories/rooftop.md` → `stories/rooftop/`): + +- `plan.md`, user-facing plan (source of truth until confirmed) +- `plan.json`, compiled from `plan.md` on confirm; recompiled when `plan.md` + is newer than `plan.json` +- `refs/.png`, chosen turnaround sheet per character +- `cand//…`, candidates, montages and sd-cli logs + +Output: `$LIBRARY///NNN.png` (zero-padded page number), where +`LIBRARY` is read from `~/.config/book-reader.conf` (shell `KEY="value"` lines). +Series and Book come from plan.md. + +Resume is derived from disk, no state file: a character with `refs/.png` +is skipped; a page whose `NNN.png` exists in the library is skipped. + +## plan.md format + +Written by Gemma, edited by the user. Fixed headings, free prose underneath: + +```markdown +# +Series: <series folder name> +Book: <book folder name> +Style: <global art style line> + +## Characters + +### <Name> +<fixed visual description: age, hair, face, eyes, build, outfit> + +## Pages + +### 1. <short title> +<what happens, setting, mood> +Characters: <Name>, <Name> (or "none") +Pose: <body orientation and pose of each named character> +Framing: <camera angle, shot size, portrait|landscape> +Model: krea2 (optional, overrides the base model for this page) +``` + +Validation before confirm: title, Series, Book, at least one character and +one page; every name in `Characters:` exists under `## Characters`; every page +has `Pose:` when it lists characters. Failures are listed; the user edits or +asks Gemma to revise. + +## Gemma calls + +All via `POST localhost:8181/v1/chat/completions`, `model: Gemma4-12B-qat-mtp`. + +1. **Draft**: system prompt holds the plan.md format and rules (no text in + images, fixed character descriptions, one page = one image, a scene may span + several pages). User message: `story.md`. Output: plan.md. +2. **Revise**: current plan.md + user's instruction → complete new plan.md. +3. **Compile**: plan.md → plan.json with `response_format` `json_schema`: + +```json +{ + "title": "", "series": "", "book": "", + "characters": [{"name": "", "turnaround_prompt": ""}], + "pages": [{ + "n": 1, "model": "zimage", "orientation": "portrait", + "characters": ["Mara", "Jon"], + "base_prompt": "", + "edits": [{"characters": ["Mara", "Jon"], "prompt": ""}] + }] +} +``` + +Prompt-writing rules given to Gemma for compile: + +- `turnaround_prompt`: "Character turnaround reference sheet … the same <person> + shown three times side by side, full body: front view, side view, back view … + plain light-grey background … No text, no labels." +- `base_prompt`: Style line + scene + full description of every character + present + Pose + Framing + "No text, no speech bubbles." +- `edits`: one entry per group of at most 2 characters (Qwen takes at most 3 + reference images: the scene plus 2 refs). Prompt changes only heads/faces: + "In image 1, change only <who>'s head: give her/him the face, … and hair of + the <person> in image N. Keep <who>'s exact pose from image 1: <Pose>. Keep + bodies, clothing, other people, background, lighting and art style of image 1 + unchanged." Pages with no characters have no edits. + +Gemma writes prompts only; the script owns every sd-cli setting. + +## Model profiles + +A dict at the top of the script. Paths under `/data/LLM-models/SD`. Settings +verified by hand on 2026-10-04: + +| profile | use | key args | +|---|---|---| +| `zimage` (default base, refs) | t2i | `--diffusion-model z_image_turbo-Q8_0.gguf --vae vae/flux1-ae.safetensors --llm Qwen3-4B-Instruct-2507-Q8_0.gguf --cfg-scale 1.0 --steps 8 --diffusion-fa --offload-to-cpu` | +| `krea2` (on demand: `--base krea2` or `Model: krea2`) | t2i | `--diffusion-model Krea-2-Turbo-Q6_K.gguf --vae vae/wan_2.1_vae.safetensors --llm Qwen3VL-4B-Instruct-Q8_0.gguf --cfg-scale 1.0 --steps 8 --diffusion-fa --offload-to-cpu` | +| `qwen_edit` | edit | `--diffusion-model Qwen-Image-2.1-viggle-turbo-v0.3-6step-Q8_0.gguf --vae vae/qwen_image_2.1_vae_bf16.safetensors --llm Qwen3VL-8B-Instruct-Q8_0.gguf --llm_vision mmproj-Qwen3VL-8B-Instruct-F16.gguf --cfg-scale 1.0 --steps 6 --sigmas 1.0,0.9375,0.875,0.75,0.5,0.25,0.0 --sampling-method euler --fa --backend te=cpu,diffusion=vulkan0,vae=vulkan0 --mmap --vae-tiling` | + +Sizes: portrait 832x1216, landscape 1216x832, turnaround 1216x832. +Before edit, reference images are shrunk with `magick -resize`: scene to +576x832 (portrait) or 832x576 (landscape), turnaround refs to 560x384. +`--backend te=cpu` without explicit `diffusion=`/`vae=` lands on the iGPU and +crashes; always pass all three. Edit with `--offload-to-cpu` and full-size refs +OOM-killed the desktop; do not use it there. + +Candidates: `-b 3` per call (one model load), seed random per call and +recorded in the montage label so a pick can be reproduced. No negative prompts: +all profiles run at cfg 1.0, where they have no effect. + +Measured timings: Z-Image ~150 s/image, Krea2 ~220 s/image, Qwen edit ~410 s +(140 s of it CPU text encoding, done once per call). + +## Review loop + +For each stage, the script builds a montage of the candidates with +`magick montage` (labels 1..3 + seed) and shows it with `kitty +kitten icat`. +Keys: + +- `1`-`3`: pick +- `r`: regenerate with new seeds +- `p`: edit this prompt in `$EDITOR`, then regenerate +- `s` (edit stage only): skip the edit, keep the base image +- `q`: quit (resume later) + +Chained edits (more than 2 characters): each edit group gets its own review; +the pick from one group is the scene input of the next. + +## Error handling + +- llama-server unreachable or model unknown: exit with a clear message, no files + written. +- plan.md validation failure: list problems, return to the review loop. +- sd-cli non-zero exit, missing output, or `[ERROR` in its log: print the last + log lines, offer `r` retry or `q` quit. Logs stay in `cand/`. +- RAM watchdog: while sd-cli runs, poll `MemAvailable` every 2 s; below 2 GB, + kill sd-cli and report it. +- Unload Gemma before every sd-cli stage (idempotent). +- Ctrl-C anywhere is safe; the next run resumes from disk. + +## CLI + +``` +fanfictioner story.md [--base zimage|krea2] [--selftest] +``` + +## Testing + +`fanfictioner --selftest`: plain asserts, no GPU, no server. Covers plan.md +validation, edit grouping for more than 2 characters, resume detection, and +sd-cli argument building per profile. + +## Out of scope + +Text, speech bubbles and captions; multi-panel page layouts; LoRA training; +sd-server daemon; any GUI beyond the terminal. + +## License + +GPLv2-only. LICENSE from gnu.org, header in the script, License section and +Development Approach section in README. -- cgit v1.2.3