aboutsummaryrefslogtreecommitdiffstats
path: root/docs/superpowers
diff options
context:
space:
mode:
Diffstat (limited to 'docs/superpowers')
-rw-r--r--docs/superpowers/specs/2026-09-18-voice-input-design.md18
1 files changed, 10 insertions, 8 deletions
diff --git a/docs/superpowers/specs/2026-09-18-voice-input-design.md b/docs/superpowers/specs/2026-09-18-voice-input-design.md
index aeb240f..58aaf58 100644
--- a/docs/superpowers/specs/2026-09-18-voice-input-design.md
+++ b/docs/superpowers/specs/2026-09-18-voice-input-design.md
@@ -115,14 +115,16 @@ the raw base64 payload. The image field stays a `data:` URL because images use
`image_url`; llama.cpp wants audio as `{"data": "<raw base64>", "format":
"wav"}`, with no `data:` prefix.
-- `classify()` recognises the common audio types (`audio/wav`, `audio/x-wav`,
- `audio/mpeg`, `audio/flac`, `audio/ogg`) and the extensions `.wav`, `.mp3`,
- `.flac`, `.ogg`, `.m4a`. Recording is the headline feature, but a dropped
- audio file is accepted and treated identically: same kind, same content
- part, same send gate. Sharing the kind means classification and the content
- builder have one audio path instead of two, so the file case costs nothing.
- A file the user dropped is theirs and is never deleted; only the recorder's
- temp clip is discarded after send.
+- `classify()` recognises WAV: the `audio/wav` and `audio/x-wav` types and
+ the `.wav` extension. Only WAV, because the content part's `format` field
+ is a fixed `"wav"` and sending MP3 or FLAC bytes under that label would
+ misrepresent them. A recorder produces WAV anyway. Recording is the
+ headline feature, but a dropped `.wav` file is accepted and treated
+ identically: same kind, same content part, same send gate. Sharing the kind
+ means classification and the content builder have one audio path instead of
+ two. A file the user dropped is theirs and is never deleted; only the
+ recorder's temp clip is discarded after send. (ponytail: WAV only; derive
+ the format per file if another type is ever needed.)
- `load_attachment()` fills `b64` for an audio file.
- `build_user_content()` gains an audio branch, parallel to the image branch.
It returns the multi-part array when either images or audio are present: