aboutsummaryrefslogtreecommitdiffstats
path: root/docs/superpowers/specs/2026-10-04-fanfictioner-design.md
blob: 6d054e6858ddd5c5282e76756d4005b4221d10c1 (plain)
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
# fanfictioner design

Date: 2026-10-04

## Goal

Turn a free-prose story idea (`story.md`) into an illustrated, text-free book of
page images, read with `../book-reader`. A local Gemma model plans the pages; the
script turns the plan into image prompts from fixed templates; `sd-cli`
(stable-diffusion.cpp) draws them. The user
reviews and picks among candidates at every image stage.

## Constraints

- Single Python 3 file, stdlib only. External tools: `sd-cli`, `magick`,
  `kitty +kitten icat`, `$EDITOR`.
- GPU: Intel Arc B580, 12 GB VRAM (Vulkan0). Vulkan1 is the Radeon iGPU and must
  never be used. RAM: 30 GB, shared with the desktop.
- llama-server runs in router mode on `localhost:8181`, model
  `Gemma4-12B-qat-mtp`, `--models-max 1`. Unload with
  `POST /models/unload {"model": ...}`; it reloads on the next request.
- No text in images: no speech bubbles, captions or labels.

## Pipeline

```
story.md ──Gemma draft──▶ plan.md ◀─┐ [c]onfirm / [e]dit in $EDITOR / [g]emma revise / [q]uit
                                    │
           confirm ──script compile (templates, no LLM)──▶ plan.json
                                    │
                         unload Gemma
                                    │
  refs:   per character: sd-cli -b 3 turnaround sheet (Z-Image) → pick → refs/<name>.png
  pages:  per page:
          base: sd-cli -b 3 (Z-Image, or Krea2 on demand) → pick
          edit: sd-cli -b 3 (Qwen-Image 2.1 viggle-turbo v0.3, -r scene -r refs…) → pick
          → $LIBRARY/<Series>/<Book>/NNN.png
```

## Files

Work dir: a directory named after the story file, next to it
(`stories/rooftop.md` → `stories/rooftop/`):

- `plan.md`, user-facing plan (source of truth until confirmed)
- `plan.json`, compiled from `plan.md` on confirm; recompiled when `plan.md`
  is newer than `plan.json`
- `refs/<Name>.png`, chosen turnaround sheet per character
- `cand/<stage>/…`, candidates, montages and sd-cli logs

Output: `$LIBRARY/<Series>/<Book>/NNN.png` (zero-padded page number), where
`LIBRARY` is read from `~/.config/book-reader.conf` (shell `KEY="value"` lines).
Series and Book come from plan.md.

Resume is derived from disk, no state file: a character with `refs/<Name>.png`
is skipped; a page whose `NNN.png` exists in the library is skipped.

## plan.md format

Written by Gemma, edited by the user. Fixed headings, free prose underneath:

```markdown
# <Title>
Series: <series folder name>
Book: <book folder name>
Style: <global art style line>

## Characters

### <Name>
Pronoun: she|he|they
<fixed visual description: who (young woman, old man...), age, hair, face, eyes, build, outfit>

## Pages

### 1. <short title>
<what happens, setting, mood>
Characters: <Name>, <Name>   (or "none")
Pose: <Name>: <position in frame, orientation, pose>; <Name>: <...>
Framing: <camera angle, shot size>, portrait|landscape
Model: krea2                 (optional, overrides the base model for this page)
```

Validation before confirm: title, Series, Book, at least one character and
one page; every character has `Pronoun:` she, he or they; every page has a
`Characters:` line and every name in it exists under `## Characters`; `Pose:`
names a pose for each listed character; `Framing:` says portrait or landscape;
page numbers run 1..N; names are safe as folder and file names. Failures are listed; the user edits or
asks Gemma to revise.

## Gemma calls

All via `POST localhost:8181/v1/chat/completions`, `model: Gemma4-12B-qat-mtp`.

1. **Draft**: system prompt holds the plan.md format and rules (no text in
   images, fixed character descriptions starting with who they are, characters
   are people only while animals and objects belong to the scene, one page = one
   image, a scene may span several pages). User message: `story.md`. Output: plan.md.
2. **Revise**: current plan.md + user's instruction -> complete new plan.md. A
   revision that fails validation is saved as `plan.rejected.md` and plan.md is
   kept; an accepted one keeps the previous version as `plan.md.bak`.

## Compile (script, no LLM)

On confirm the script builds plan.json from the parsed plan.md:

```json
{
  "title": "", "series": "", "book": "",
  "characters": [{"name": "", "turnaround_prompt": ""}],
  "pages": [{
    "n": 1, "model": "", "orientation": "portrait",
    "base_prompt": "",
    "edits": [{"characters": ["Mara", "Jon"], "prompt": ""}]
  }]
}
```

Templates:

- `turnaround_prompt`: "Character turnaround reference sheet of one person:
  <description> The same person shown three times side by side, full body: front
  view, side view, back view. Neutral standing pose, arms relaxed. Plain
  light-grey background, even studio lighting. <Style>. No text, no labels."
- `base_prompt`: Style, page text, then for each character present its
  description followed by its pose, then Framing, then "No text, no speech
  bubbles."
- `edits`: one entry per group of at most 2 characters (Qwen takes at most 3
  reference images: the scene plus 2 refs), in `Characters:` order. Per
  character: "In image 1, change only <Name>'s head: give her/him/them the face,
  eyes and hair of the person in image N. Keep <Name>'s exact pose from image 1:
  <pose>." Then "Keep bodies, clothing, other people, background, lighting and
  art style of image 1 unchanged." Pages with no characters have no edits.
- `orientation` comes from Framing; `model` from the page's `Model:` line
  (empty means `--base`).

Gemma writes the plan only; the script owns every prompt and sd-cli setting.

## Model profiles

A dict at the top of the script. Paths under `/data/LLM-models/SD`. Settings
verified by hand on 2026-10-04:

| profile | use | key args |
|---|---|---|
| `zimage` (default base, refs) | t2i | `--diffusion-model z_image_turbo-Q8_0.gguf --vae vae/flux1-ae.safetensors --llm Qwen3-4B-Instruct-2507-Q8_0.gguf --cfg-scale 1.0 --steps 8 --diffusion-fa --offload-to-cpu` |
| `krea2` (on demand: `--base krea2` or `Model: krea2`) | t2i | `--diffusion-model Krea-2-Turbo-Q6_K.gguf --vae vae/wan_2.1_vae.safetensors --llm Qwen3VL-4B-Instruct-Q8_0.gguf --cfg-scale 1.0 --steps 8 --diffusion-fa --offload-to-cpu` |
| `qwen_edit` | edit | `--diffusion-model Qwen-Image-2.1-viggle-turbo-v0.3-6step-Q8_0.gguf --vae vae/qwen_image_2.1_vae_bf16.safetensors --llm Qwen3VL-8B-Instruct-Q8_0.gguf --llm_vision mmproj-Qwen3VL-8B-Instruct-F16.gguf --cfg-scale 1.0 --steps 6 --sigmas 1.0,0.9375,0.875,0.75,0.5,0.25,0.0 --sampling-method euler --fa --backend te=cpu,diffusion=vulkan0,vae=vulkan0 --mmap --vae-tiling` |

Sizes: portrait 832x1216, landscape 1216x832, turnaround 1216x832.
Before edit, reference images are shrunk with `magick -resize`: scene to
576x832 (portrait) or 832x576 (landscape), turnaround refs to 560x384.
`--backend te=cpu` without explicit `diffusion=`/`vae=` lands on the iGPU and
crashes; always pass all three. Edit with `--offload-to-cpu` and full-size refs
OOM-killed the desktop; do not use it there.

Candidates: `-b 3` per call (one model load), seed random per call and
recorded in the montage label so a pick can be reproduced. No negative prompts:
all profiles run at cfg 1.0, where they have no effect.

Measured timings: Z-Image ~150 s/image, Krea2 ~220 s/image, Qwen edit ~410 s
(140 s of it CPU text encoding, done once per call).

## Review loop

For each stage, the script builds a montage of the candidates with
`magick montage` (labels 1..3 + seed) and shows it with `kitty +kitten icat`.
Keys:

- `1`-`3`: pick
- `r`: regenerate with new seeds
- `p`: edit this prompt in `$EDITOR`, then regenerate
- `s` (edit stage only): skip the edit, keep the base image
- `q`: quit (resume later)

Chained edits (more than 2 characters): each edit group gets its own review;
the pick from one group is the scene input of the next.

## Error handling

- llama-server unreachable or model unknown: exit with a clear message, no files
  written.
- plan.md validation failure: list problems, return to the review loop.
- sd-cli non-zero exit, missing output, or `[ERROR` in its log: print the last
  log lines, offer `r` retry or `q` quit. Logs stay in `cand/`.
- RAM watchdog: while sd-cli runs, poll `MemAvailable` every 2 s; below 2 GB,
  kill sd-cli and report it.
- Unload Gemma before every sd-cli stage (idempotent).
- Ctrl-C anywhere is safe; the next run resumes from disk.

## CLI

```
fanfictioner story.md [--base zimage|krea2] [--selftest]
```

## Testing

`fanfictioner --selftest`: plain asserts, no GPU, no server. Covers plan.md
parsing and validation, edit grouping for more than 2 characters, plan.json
templates, resume detection, and sd-cli argument building per profile.

## Out of scope

Text, speech bubbles and captions; multi-panel page layouts; LoRA training;
sd-server daemon; any GUI beyond the terminal.

## License

GPLv2-only. LICENSE from gnu.org, header in the script, License section and
Development Approach section in README.