aboutsummaryrefslogtreecommitdiffstats
path: root/docs/superpowers/specs
diff options
context:
space:
mode:
authorDanilo M. <danix@danix.xyz>2026-09-18 20:07:37 +0200
committerDanilo M. <danix@danix.xyz>2026-09-18 20:07:37 +0200
commit8cbe904ae2c6c83210f580432f9c59079583aff3 (patch)
tree4cc9a74f4daee38dc35881ad28ea336b3f8b1453 /docs/superpowers/specs
parent650e633826acd6c9255ce64957b0288d49fecefe (diff)
downloadllamachat-8cbe904ae2c6c83210f580432f9c59079583aff3.tar.gz
llamachat-8cbe904ae2c6c83210f580432f9c59079583aff3.zip
docs: restrict voice input to WAV attachments
Diffstat (limited to 'docs/superpowers/specs')
-rw-r--r--docs/superpowers/specs/2026-09-18-voice-input-design.md18
1 files changed, 10 insertions, 8 deletions
diff --git a/docs/superpowers/specs/2026-09-18-voice-input-design.md b/docs/superpowers/specs/2026-09-18-voice-input-design.md
index aeb240f..58aaf58 100644
--- a/docs/superpowers/specs/2026-09-18-voice-input-design.md
+++ b/docs/superpowers/specs/2026-09-18-voice-input-design.md
@@ -115,14 +115,16 @@ the raw base64 payload. The image field stays a `data:` URL because images use
`image_url`; llama.cpp wants audio as `{"data": "<raw base64>", "format":
"wav"}`, with no `data:` prefix.
-- `classify()` recognises the common audio types (`audio/wav`, `audio/x-wav`,
- `audio/mpeg`, `audio/flac`, `audio/ogg`) and the extensions `.wav`, `.mp3`,
- `.flac`, `.ogg`, `.m4a`. Recording is the headline feature, but a dropped
- audio file is accepted and treated identically: same kind, same content
- part, same send gate. Sharing the kind means classification and the content
- builder have one audio path instead of two, so the file case costs nothing.
- A file the user dropped is theirs and is never deleted; only the recorder's
- temp clip is discarded after send.
+- `classify()` recognises WAV: the `audio/wav` and `audio/x-wav` types and
+ the `.wav` extension. Only WAV, because the content part's `format` field
+ is a fixed `"wav"` and sending MP3 or FLAC bytes under that label would
+ misrepresent them. A recorder produces WAV anyway. Recording is the
+ headline feature, but a dropped `.wav` file is accepted and treated
+ identically: same kind, same content part, same send gate. Sharing the kind
+ means classification and the content builder have one audio path instead of
+ two. A file the user dropped is theirs and is never deleted; only the
+ recorder's temp clip is discarded after send. (ponytail: WAV only; derive
+ the format per file if another type is ever needed.)
- `load_attachment()` fills `b64` for an audio file.
- `build_user_content()` gains an audio branch, parallel to the image branch.
It returns the multi-part array when either images or audio are present: