diff options
| author | Danilo M. <danix@danix.xyz> | 2026-09-18 20:07:37 +0200 |
|---|---|---|
| committer | Danilo M. <danix@danix.xyz> | 2026-09-18 20:07:37 +0200 |
| commit | 8cbe904ae2c6c83210f580432f9c59079583aff3 (patch) | |
| tree | 4cc9a74f4daee38dc35881ad28ea336b3f8b1453 /docs/superpowers/specs/2026-09-18-voice-input-design.md | |
| parent | 650e633826acd6c9255ce64957b0288d49fecefe (diff) | |
| download | llamachat-8cbe904ae2c6c83210f580432f9c59079583aff3.tar.gz llamachat-8cbe904ae2c6c83210f580432f9c59079583aff3.zip | |
docs: restrict voice input to WAV attachments
Diffstat (limited to 'docs/superpowers/specs/2026-09-18-voice-input-design.md')
| -rw-r--r-- | docs/superpowers/specs/2026-09-18-voice-input-design.md | 18 |
1 files changed, 10 insertions, 8 deletions
diff --git a/docs/superpowers/specs/2026-09-18-voice-input-design.md b/docs/superpowers/specs/2026-09-18-voice-input-design.md index aeb240f..58aaf58 100644 --- a/docs/superpowers/specs/2026-09-18-voice-input-design.md +++ b/docs/superpowers/specs/2026-09-18-voice-input-design.md @@ -115,14 +115,16 @@ the raw base64 payload. The image field stays a `data:` URL because images use `image_url`; llama.cpp wants audio as `{"data": "<raw base64>", "format": "wav"}`, with no `data:` prefix. -- `classify()` recognises the common audio types (`audio/wav`, `audio/x-wav`, - `audio/mpeg`, `audio/flac`, `audio/ogg`) and the extensions `.wav`, `.mp3`, - `.flac`, `.ogg`, `.m4a`. Recording is the headline feature, but a dropped - audio file is accepted and treated identically: same kind, same content - part, same send gate. Sharing the kind means classification and the content - builder have one audio path instead of two, so the file case costs nothing. - A file the user dropped is theirs and is never deleted; only the recorder's - temp clip is discarded after send. +- `classify()` recognises WAV: the `audio/wav` and `audio/x-wav` types and + the `.wav` extension. Only WAV, because the content part's `format` field + is a fixed `"wav"` and sending MP3 or FLAC bytes under that label would + misrepresent them. A recorder produces WAV anyway. Recording is the + headline feature, but a dropped `.wav` file is accepted and treated + identically: same kind, same content part, same send gate. Sharing the kind + means classification and the content builder have one audio path instead of + two. A file the user dropped is theirs and is never deleted; only the + recorder's temp clip is discarded after send. (ponytail: WAV only; derive + the format per file if another type is ever needed.) - `load_attachment()` fills `b64` for an audio file. - `build_user_content()` gains an audio branch, parallel to the image branch. It returns the multi-part array when either images or audio are present: |
