diff options
| author | Danilo M. <danix@danix.xyz> | 2026-09-18 20:06:13 +0200 |
|---|---|---|
| committer | Danilo M. <danix@danix.xyz> | 2026-09-18 20:06:13 +0200 |
| commit | 650e633826acd6c9255ce64957b0288d49fecefe (patch) | |
| tree | 4582e7428008850a870e45927e8854860596401f | |
| parent | 2f44135b5b84f805464491088d6c3d52c21e9210 (diff) | |
| download | llamachat-650e633826acd6c9255ce64957b0288d49fecefe.tar.gz llamachat-650e633826acd6c9255ce64957b0288d49fecefe.zip | |
docs: spec for voice input to audio-capable models
| -rw-r--r-- | docs/superpowers/specs/2026-09-18-voice-input-design.md | 254 |
1 files changed, 254 insertions, 0 deletions
diff --git a/docs/superpowers/specs/2026-09-18-voice-input-design.md b/docs/superpowers/specs/2026-09-18-voice-input-design.md new file mode 100644 index 0000000..aeb240f --- /dev/null +++ b/docs/superpowers/specs/2026-09-18-voice-input-design.md @@ -0,0 +1,254 @@ +# Voice input for audio-capable models + +Status: design approved + +## Problem + +llamachat can attach text and images to a message. It has no way to speak a +message. llama.cpp's OpenAI-compatible server accepts audio as a content +part, so a model served with an audio encoder (Qwen2-Audio, Qwen2.5-Omni, +Qwen3-Omni, Voxtral, Ultravox) can hear a recording directly. This adds a +record control that captures the microphone, sends the clip as +`input_audio`, and leaves a marker in history. + +This is native audio input, not speech-to-text. The model receives the +waveform and answers it; there is no transcription step and no transcript is +stored. Only models that declare audio support are offered the control. + +## Capability: `audio`, resolved in layers + +A model is audio-capable when its resolved `audio` flag is `True`. The flag +is tri-state (`True`, `False`, `None` for unknown) and is resolved most +specific first: + +1. **Router-reported modalities.** The llama.cpp model router's `/v1/models` + returns `architecture.input_modalities` per model, for example + `["text","image","audio"]`. When present it is authoritative: `audio` is + `"audio" in modalities`, and `vision` is `"image" in modalities`. This is + what lets a local model work with no manual flag, and it sharpens `vision` + for local models at the same time. +2. `models.ini`, written by the model settings dialog (`audio = true`). +3. The provider's `[providers.*] audio = true`, used as the prefill for cloud + models. +4. `presets.ini` (`vision` only, read from `mmproj`). There is no audio key: + a plain llama-server does not report an audio encoder in its preset + format, so without the router the manual flag is how a local audio model + declares itself. + +Layer 1 covers the router this app is built for. Layers 2 and 3 remain for +cloud providers, whose `/v1/models` does not carry modalities, and for a +plain llama-server. The `presets.ini` layer stays vision-only because that is +all its `mmproj` line can honestly say. + +Changes: + +- `backend.Client.models()` extracts each model's `architecture.input_modalities` + when the server reports it, and returns that mapping alongside the ids; + `MultiClient.models()` threads it through. A server that does not report it + yields an empty mapping, so a plain llama-server or a cloud provider is + unaffected. The parse tolerates both the `{"data": [...]}` envelope and a + bare array, as the existing listing does. +- `ui` holds the mapping and `model_info()` applies it above the stored and + provider flags, since a reported modality is a fact and a manual flag is a + guess. +- `models.ModelInfo` gains `audio: bool | None = None`; `get`, `save`, + `is_empty` and `resolve` handle it beside `vision`. +- `providers.Provider` gains `audio: bool | None = None`; `parse` reads it. +- `modeldialog` gains an "Accepts audio" checkbox. The vision checkbox's + tri-state rule applies unchanged: the box carries intent, and an unchecked + box over an unknown prefill stays `None` rather than writing `False`. + +Gate: the record control is shown only when the selected model resolves +`audio is True` and the capture module is available. Unknown and `False` both +hide it, so a recording can never be sent to a model whose audio support is +unconfirmed. An audio-capable model is rare and explicitly reported, so +requiring a positive signal costs nothing and avoids sending a clip to a +text-only model that happens to be unknown. + +## Capture: new guarded `llamachat/audio.py` + +QtMultimedia is not part of PySide6-Essentials, so the import is guarded and +the feature degrades to absent when the module is missing: + +``` +try: + from PySide6.QtMultimedia import QAudioSource, QAudioFormat, QMediaDevices + AVAILABLE = True +except ImportError: + AVAILABLE = False +``` + +`AVAILABLE` is the single gate the UI reads. The README's install section +gains a note that voice input needs `PySide6-Addons`, without making it a hard +requirement. + +`AudioRecorder`: + +- `QAudioSource` at a fixed `QAudioFormat`: 16000 Hz, 1 channel, `Int16`. + This is what speech encoders expect and keeps a clip small (about 2 MB per + minute). The format is not user-configurable in this version. +- Uses `QMediaDevices.defaultAudioInput()`. A device picker is deferred; the + default microphone is used. +- Writes PCM into a `QBuffer` wrapping a `QByteArray`. +- `start()` begins capture and returns `False` with a reason when there is no + input device or the backend refuses to start. +- `stop()` stops capture and returns the raw PCM `bytes`. + +`to_wav(pcm, rate, channels)` is a module-level function with no Qt +dependency. It prepends a RIFF/WAVE header using the standard library `wave` +module over a `BytesIO` and returns the complete `.wav` bytes. Keeping it +Qt-free is what makes it unit-testable. + +Safety rails: + +- Hard cap of 60 seconds. A `QTimer` stops capture and the UI reports that + the cap was hit. This bounds memory (about 2 MB) and the request size. +- A clip shorter than 0.3 seconds is discarded with a status note, so a + double-click does not send noise. + +No window is added. The recorder is a plain object the existing UI drives. + +## Send path + +`backend.Attachment` gains the kind `'audio'` and a `b64: str = ""` field for +the raw base64 payload. The image field stays a `data:` URL because images use +`image_url`; llama.cpp wants audio as `{"data": "<raw base64>", "format": +"wav"}`, with no `data:` prefix. + +- `classify()` recognises the common audio types (`audio/wav`, `audio/x-wav`, + `audio/mpeg`, `audio/flac`, `audio/ogg`) and the extensions `.wav`, `.mp3`, + `.flac`, `.ogg`, `.m4a`. Recording is the headline feature, but a dropped + audio file is accepted and treated identically: same kind, same content + part, same send gate. Sharing the kind means classification and the content + builder have one audio path instead of two, so the file case costs nothing. + A file the user dropped is theirs and is never deleted; only the recorder's + temp clip is discarded after send. +- `load_attachment()` fills `b64` for an audio file. +- `build_user_content()` gains an audio branch, parallel to the image branch. + It returns the multi-part array when either images or audio are present: + + ``` + [ + {"type": "text", "text": "<typed text or default instruction>"}, + {"type": "image_url", "image_url": {"url": "data:..."}}, + {"type": "input_audio", "input_audio": {"data": "<raw base64>", "format": "wav"}}, + ] + ``` + + The text part comes first, then audio, matching the order in llama.cpp's + own audio example. +- When a recording is present and no text was typed, the UI substitutes a + default instruction, `Listen to this audio and respond.`, so every audio + request carries a text part. A typed message is used as-is and the audio is + attached to it. +- `estimate_tokens()` currently counts every non-text part as one "image" and + adds a flat 600-token allowance. That branch is generalised to count any + non-text part, so an `input_audio` part gets the same flat allowance and + the context meter is not blind to a recording. The allowance is not tuned + per media type; it exists to show that media is present and roughly + expensive. + +The existing tool loop, search and skills are untouched. An audio turn is an +ordinary user message; tool calling still works on top of it if the model +supports both. + +## UI + +### Control + +A record button sits in the bottom bar next to "Attach…". It is hidden unless +the model is audio-capable and capture is available. + +- Click to start. The button switches to a recording state and the status + line reads that it is recording, so the state is never ambiguous. +- Click again to stop. The clip becomes a pending attachment and the button + returns to idle. + +This is toggle-then-send: stopping does not send. The user can type text +alongside the clip, then send or discard it with the existing Send and Clear +controls. This reuses the attachment lifecycle rather than adding a second +send path, and it lets a thought be recorded and annotated before it goes. + +### Pending clip + +`update_attach_label` shows `🎙 voice-note Ns` beside the existing `📄` and +`🖼` marks. The duration comes from the recorder. `send()` and +`clear_attachments()` are unchanged; they already carry whatever is in +`self.attachments`. + +### Model mismatch + +If a clip is pending and the selected model is not audio-capable, sending is +blocked with the same offer-to-switch flow as `_ensure_vision_model`: name an +audio-capable model, ask, and switch. When none is configured, say so rather +than letting the request fail at the server. + +## History + +Recordings are discarded after send. The turn stores: + +- `messages.content`: the text part only, exactly as image turns already + store the text part and not the image bytes. This is the typed message, or + the default instruction. +- One `attachments` row with `kind = 'audio'`, `mime = 'audio/wav'`, + `size`, `sha256`, and a synthetic `path` of `voice-note-<seconds>s.wav`. + The synthetic name carries the duration for display; it is not a path on + disk. `thumb` is NULL and `truncated` is `False`. + +No schema change. `attachments.path` is documented as the original path; for +a recording there is none, and the synthetic marker is the honest stand-in. +Reopening a chat shows the marker and cannot replay or resend the clip, which +is the accepted cost of not retaining audio. This matches images, which are +also not replayed when a conversation continues. + +## Error handling + +Every failure is a status-bar line, never a crash and never a partial send: + +- QtMultimedia missing: the control is absent. +- No default input device: the click reports it instead of starting. +- The backend refuses to start: same, with the reason. +- A clip shorter than 0.3 s: discarded with a note. +- The 60 s cap: capture stops automatically and the note says why. +- A pending clip on a non-audio model: blocked with the switch offer. + +## Testing + +Extend `test_llamachat.py` in its existing assert style: + +- `audio.to_wav`: bytes start with `RIFF`/`WAVE`, and reading the result back + with the `wave` module yields the same frame count, rate and channel count. +- `backend.build_user_content`: a recording and a typed text produce a + `text` part then an `input_audio` part with raw base64 and `format == "wav"`; + an image and audio together keep both parts; a typed message with no + attachment still returns a plain string. +- `backend.classify`: `.wav` and `audio/wav` classify as `audio`; existing + text and image cases are unchanged. +- `backend` listing: a `/v1/models` body carrying + `architecture.input_modalities` yields the expected id-to-modalities map, + a body without it yields an empty map, and both the enveloped and bare-array + shapes parse. This is the layer that detects the local audio model, so it is + tested directly. +- `models`: `ModelInfo.audio` round-trips through `save`/`get`, + `resolve` layers store over provider, and `is_empty` treats an audio-only + entry as non-empty. +- `providers.parse`: `audio = true` is read; a missing key is `None`. + +The `QAudioSource` plumbing is GUI and hardware, so it stays out of the unit +suite, consistent with the rest of the window. + +## Files + +New: `llamachat/audio.py`. + +Edited: `llamachat/backend.py`, `llamachat/models.py`, +`llamachat/providers.py`, `llamachat/modeldialog.py`, `llamachat/ui.py`, +`README.md`, `CHANGELOG.md`, `test_llamachat.py`. + +## Out of scope + +- Speech-to-text and the transcript of a recording. +- Retaining or replaying audio after a message is sent. +- A microphone device picker. +- Configurable sample rate, channels, or maximum duration. +- Server-side transcription tools; this is input to the chat model only. |
