aboutsummaryrefslogtreecommitdiffstats
path: root/docs/superpowers/specs/2026-10-04-fanfictioner-design.md
blob: cd620e866e126ec8136fc722b0d5797483b571ff (plain)
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
# fanfictioner design

Date: 2026-10-04

## Goal

Turn a free-prose story idea (`story.md`) into an illustrated, text-free book of
page images, read with `../book-reader`. A local Gemma model plans the pages and
writes the image prompts; `sd-cli` (stable-diffusion.cpp) draws them. The user
reviews and picks among candidates at every image stage.

## Constraints

- Single Python 3 file, stdlib only. External tools: `sd-cli`, `magick`,
  `kitty +kitten icat`, `$EDITOR`.
- GPU: Intel Arc B580, 12 GB VRAM (Vulkan0). Vulkan1 is the Radeon iGPU and must
  never be used. RAM: 30 GB, shared with the desktop.
- llama-server runs in router mode on `localhost:8181`, model
  `Gemma4-12B-qat-mtp`, `--models-max 1`. Unload with
  `POST /models/unload {"model": ...}`; it reloads on the next request.
- No text in images: no speech bubbles, captions or labels.

## Pipeline

```
story.md ──Gemma draft──▶ plan.md ◀─┐ [c]onfirm / [e]dit in $EDITOR / [g]emma revise / [q]uit
                                    │
           confirm ──Gemma compile (json_schema)──▶ plan.json
                                    │
                         unload Gemma
                                    │
  refs:   per character: sd-cli -b 3 turnaround sheet (Z-Image) → pick → refs/<name>.png
  pages:  per page:
          base: sd-cli -b 3 (Z-Image, or Krea2 on demand) → pick
          edit: sd-cli -b 3 (Qwen-Image 2.1 viggle-turbo v0.3, -r scene -r refs…) → pick
          → $LIBRARY/<Series>/<Book>/NNN.png
```

## Files

Work dir: a directory named after the story file, next to it
(`stories/rooftop.md` → `stories/rooftop/`):

- `plan.md`, user-facing plan (source of truth until confirmed)
- `plan.json`, compiled from `plan.md` on confirm; recompiled when `plan.md`
  is newer than `plan.json`
- `refs/<Name>.png`, chosen turnaround sheet per character
- `cand/<stage>/…`, candidates, montages and sd-cli logs

Output: `$LIBRARY/<Series>/<Book>/NNN.png` (zero-padded page number), where
`LIBRARY` is read from `~/.config/book-reader.conf` (shell `KEY="value"` lines).
Series and Book come from plan.md.

Resume is derived from disk, no state file: a character with `refs/<Name>.png`
is skipped; a page whose `NNN.png` exists in the library is skipped.

## plan.md format

Written by Gemma, edited by the user. Fixed headings, free prose underneath:

```markdown
# <Title>
Series: <series folder name>
Book: <book folder name>
Style: <global art style line>

## Characters

### <Name>
<fixed visual description: age, hair, face, eyes, build, outfit>

## Pages

### 1. <short title>
<what happens, setting, mood>
Characters: <Name>, <Name>   (or "none")
Pose: <body orientation and pose of each named character>
Framing: <camera angle, shot size, portrait|landscape>
Model: krea2                 (optional, overrides the base model for this page)
```

Validation before confirm: title, Series, Book, at least one character and
one page; every name in `Characters:` exists under `## Characters`; every page
has `Pose:` when it lists characters. Failures are listed; the user edits or
asks Gemma to revise.

## Gemma calls

All via `POST localhost:8181/v1/chat/completions`, `model: Gemma4-12B-qat-mtp`.

1. **Draft**: system prompt holds the plan.md format and rules (no text in
   images, fixed character descriptions, one page = one image, a scene may span
   several pages). User message: `story.md`. Output: plan.md.
2. **Revise**: current plan.md + user's instruction → complete new plan.md.
3. **Compile**: plan.md → plan.json with `response_format` `json_schema`:

```json
{
  "title": "", "series": "", "book": "",
  "characters": [{"name": "", "turnaround_prompt": ""}],
  "pages": [{
    "n": 1, "model": "zimage", "orientation": "portrait",
    "characters": ["Mara", "Jon"],
    "base_prompt": "",
    "edits": [{"characters": ["Mara", "Jon"], "prompt": ""}]
  }]
}
```

Prompt-writing rules given to Gemma for compile:

- `turnaround_prompt`: "Character turnaround reference sheet … the same <person>
  shown three times side by side, full body: front view, side view, back view …
  plain light-grey background … No text, no labels."
- `base_prompt`: Style line + scene + full description of every character
  present + Pose + Framing + "No text, no speech bubbles."
- `edits`: one entry per group of at most 2 characters (Qwen takes at most 3
  reference images: the scene plus 2 refs). Prompt changes only heads/faces:
  "In image 1, change only <who>'s head: give her/him the face, … and hair of
  the <person> in image N. Keep <who>'s exact pose from image 1: <Pose>. Keep
  bodies, clothing, other people, background, lighting and art style of image 1
  unchanged." Pages with no characters have no edits.

Gemma writes prompts only; the script owns every sd-cli setting.

## Model profiles

A dict at the top of the script. Paths under `/data/LLM-models/SD`. Settings
verified by hand on 2026-10-04:

| profile | use | key args |
|---|---|---|
| `zimage` (default base, refs) | t2i | `--diffusion-model z_image_turbo-Q8_0.gguf --vae vae/flux1-ae.safetensors --llm Qwen3-4B-Instruct-2507-Q8_0.gguf --cfg-scale 1.0 --steps 8 --diffusion-fa --offload-to-cpu` |
| `krea2` (on demand: `--base krea2` or `Model: krea2`) | t2i | `--diffusion-model Krea-2-Turbo-Q6_K.gguf --vae vae/wan_2.1_vae.safetensors --llm Qwen3VL-4B-Instruct-Q8_0.gguf --cfg-scale 1.0 --steps 8 --diffusion-fa --offload-to-cpu` |
| `qwen_edit` | edit | `--diffusion-model Qwen-Image-2.1-viggle-turbo-v0.3-6step-Q8_0.gguf --vae vae/qwen_image_2.1_vae_bf16.safetensors --llm Qwen3VL-8B-Instruct-Q8_0.gguf --llm_vision mmproj-Qwen3VL-8B-Instruct-F16.gguf --cfg-scale 1.0 --steps 6 --sigmas 1.0,0.9375,0.875,0.75,0.5,0.25,0.0 --sampling-method euler --fa --backend te=cpu,diffusion=vulkan0,vae=vulkan0 --mmap --vae-tiling` |

Sizes: portrait 832x1216, landscape 1216x832, turnaround 1216x832.
Before edit, reference images are shrunk with `magick -resize`: scene to
576x832 (portrait) or 832x576 (landscape), turnaround refs to 560x384.
`--backend te=cpu` without explicit `diffusion=`/`vae=` lands on the iGPU and
crashes; always pass all three. Edit with `--offload-to-cpu` and full-size refs
OOM-killed the desktop; do not use it there.

Candidates: `-b 3` per call (one model load), seed random per call and
recorded in the montage label so a pick can be reproduced. No negative prompts:
all profiles run at cfg 1.0, where they have no effect.

Measured timings: Z-Image ~150 s/image, Krea2 ~220 s/image, Qwen edit ~410 s
(140 s of it CPU text encoding, done once per call).

## Review loop

For each stage, the script builds a montage of the candidates with
`magick montage` (labels 1..3 + seed) and shows it with `kitty +kitten icat`.
Keys:

- `1`-`3`: pick
- `r`: regenerate with new seeds
- `p`: edit this prompt in `$EDITOR`, then regenerate
- `s` (edit stage only): skip the edit, keep the base image
- `q`: quit (resume later)

Chained edits (more than 2 characters): each edit group gets its own review;
the pick from one group is the scene input of the next.

## Error handling

- llama-server unreachable or model unknown: exit with a clear message, no files
  written.
- plan.md validation failure: list problems, return to the review loop.
- sd-cli non-zero exit, missing output, or `[ERROR` in its log: print the last
  log lines, offer `r` retry or `q` quit. Logs stay in `cand/`.
- RAM watchdog: while sd-cli runs, poll `MemAvailable` every 2 s; below 2 GB,
  kill sd-cli and report it.
- Unload Gemma before every sd-cli stage (idempotent).
- Ctrl-C anywhere is safe; the next run resumes from disk.

## CLI

```
fanfictioner story.md [--base zimage|krea2] [--selftest]
```

## Testing

`fanfictioner --selftest`: plain asserts, no GPU, no server. Covers plan.md
validation, edit grouping for more than 2 characters, resume detection, and
sd-cli argument building per profile.

## Out of scope

Text, speech bubbles and captions; multi-panel page layouts; LoRA training;
sd-server daemon; any GUI beyond the terminal.

## License

GPLv2-only. LICENSE from gnu.org, header in the script, License section and
Development Approach section in README.