aboutsummaryrefslogtreecommitdiffstats
diff options
context:
space:
mode:
-rw-r--r--docs/superpowers/specs/2026-09-18-voice-input-design.md254
1 files changed, 254 insertions, 0 deletions
diff --git a/docs/superpowers/specs/2026-09-18-voice-input-design.md b/docs/superpowers/specs/2026-09-18-voice-input-design.md
new file mode 100644
index 0000000..aeb240f
--- /dev/null
+++ b/docs/superpowers/specs/2026-09-18-voice-input-design.md
@@ -0,0 +1,254 @@
+# Voice input for audio-capable models
+
+Status: design approved
+
+## Problem
+
+llamachat can attach text and images to a message. It has no way to speak a
+message. llama.cpp's OpenAI-compatible server accepts audio as a content
+part, so a model served with an audio encoder (Qwen2-Audio, Qwen2.5-Omni,
+Qwen3-Omni, Voxtral, Ultravox) can hear a recording directly. This adds a
+record control that captures the microphone, sends the clip as
+`input_audio`, and leaves a marker in history.
+
+This is native audio input, not speech-to-text. The model receives the
+waveform and answers it; there is no transcription step and no transcript is
+stored. Only models that declare audio support are offered the control.
+
+## Capability: `audio`, resolved in layers
+
+A model is audio-capable when its resolved `audio` flag is `True`. The flag
+is tri-state (`True`, `False`, `None` for unknown) and is resolved most
+specific first:
+
+1. **Router-reported modalities.** The llama.cpp model router's `/v1/models`
+ returns `architecture.input_modalities` per model, for example
+ `["text","image","audio"]`. When present it is authoritative: `audio` is
+ `"audio" in modalities`, and `vision` is `"image" in modalities`. This is
+ what lets a local model work with no manual flag, and it sharpens `vision`
+ for local models at the same time.
+2. `models.ini`, written by the model settings dialog (`audio = true`).
+3. The provider's `[providers.*] audio = true`, used as the prefill for cloud
+ models.
+4. `presets.ini` (`vision` only, read from `mmproj`). There is no audio key:
+ a plain llama-server does not report an audio encoder in its preset
+ format, so without the router the manual flag is how a local audio model
+ declares itself.
+
+Layer 1 covers the router this app is built for. Layers 2 and 3 remain for
+cloud providers, whose `/v1/models` does not carry modalities, and for a
+plain llama-server. The `presets.ini` layer stays vision-only because that is
+all its `mmproj` line can honestly say.
+
+Changes:
+
+- `backend.Client.models()` extracts each model's `architecture.input_modalities`
+ when the server reports it, and returns that mapping alongside the ids;
+ `MultiClient.models()` threads it through. A server that does not report it
+ yields an empty mapping, so a plain llama-server or a cloud provider is
+ unaffected. The parse tolerates both the `{"data": [...]}` envelope and a
+ bare array, as the existing listing does.
+- `ui` holds the mapping and `model_info()` applies it above the stored and
+ provider flags, since a reported modality is a fact and a manual flag is a
+ guess.
+- `models.ModelInfo` gains `audio: bool | None = None`; `get`, `save`,
+ `is_empty` and `resolve` handle it beside `vision`.
+- `providers.Provider` gains `audio: bool | None = None`; `parse` reads it.
+- `modeldialog` gains an "Accepts audio" checkbox. The vision checkbox's
+ tri-state rule applies unchanged: the box carries intent, and an unchecked
+ box over an unknown prefill stays `None` rather than writing `False`.
+
+Gate: the record control is shown only when the selected model resolves
+`audio is True` and the capture module is available. Unknown and `False` both
+hide it, so a recording can never be sent to a model whose audio support is
+unconfirmed. An audio-capable model is rare and explicitly reported, so
+requiring a positive signal costs nothing and avoids sending a clip to a
+text-only model that happens to be unknown.
+
+## Capture: new guarded `llamachat/audio.py`
+
+QtMultimedia is not part of PySide6-Essentials, so the import is guarded and
+the feature degrades to absent when the module is missing:
+
+```
+try:
+ from PySide6.QtMultimedia import QAudioSource, QAudioFormat, QMediaDevices
+ AVAILABLE = True
+except ImportError:
+ AVAILABLE = False
+```
+
+`AVAILABLE` is the single gate the UI reads. The README's install section
+gains a note that voice input needs `PySide6-Addons`, without making it a hard
+requirement.
+
+`AudioRecorder`:
+
+- `QAudioSource` at a fixed `QAudioFormat`: 16000 Hz, 1 channel, `Int16`.
+ This is what speech encoders expect and keeps a clip small (about 2 MB per
+ minute). The format is not user-configurable in this version.
+- Uses `QMediaDevices.defaultAudioInput()`. A device picker is deferred; the
+ default microphone is used.
+- Writes PCM into a `QBuffer` wrapping a `QByteArray`.
+- `start()` begins capture and returns `False` with a reason when there is no
+ input device or the backend refuses to start.
+- `stop()` stops capture and returns the raw PCM `bytes`.
+
+`to_wav(pcm, rate, channels)` is a module-level function with no Qt
+dependency. It prepends a RIFF/WAVE header using the standard library `wave`
+module over a `BytesIO` and returns the complete `.wav` bytes. Keeping it
+Qt-free is what makes it unit-testable.
+
+Safety rails:
+
+- Hard cap of 60 seconds. A `QTimer` stops capture and the UI reports that
+ the cap was hit. This bounds memory (about 2 MB) and the request size.
+- A clip shorter than 0.3 seconds is discarded with a status note, so a
+ double-click does not send noise.
+
+No window is added. The recorder is a plain object the existing UI drives.
+
+## Send path
+
+`backend.Attachment` gains the kind `'audio'` and a `b64: str = ""` field for
+the raw base64 payload. The image field stays a `data:` URL because images use
+`image_url`; llama.cpp wants audio as `{"data": "<raw base64>", "format":
+"wav"}`, with no `data:` prefix.
+
+- `classify()` recognises the common audio types (`audio/wav`, `audio/x-wav`,
+ `audio/mpeg`, `audio/flac`, `audio/ogg`) and the extensions `.wav`, `.mp3`,
+ `.flac`, `.ogg`, `.m4a`. Recording is the headline feature, but a dropped
+ audio file is accepted and treated identically: same kind, same content
+ part, same send gate. Sharing the kind means classification and the content
+ builder have one audio path instead of two, so the file case costs nothing.
+ A file the user dropped is theirs and is never deleted; only the recorder's
+ temp clip is discarded after send.
+- `load_attachment()` fills `b64` for an audio file.
+- `build_user_content()` gains an audio branch, parallel to the image branch.
+ It returns the multi-part array when either images or audio are present:
+
+ ```
+ [
+ {"type": "text", "text": "<typed text or default instruction>"},
+ {"type": "image_url", "image_url": {"url": "data:..."}},
+ {"type": "input_audio", "input_audio": {"data": "<raw base64>", "format": "wav"}},
+ ]
+ ```
+
+ The text part comes first, then audio, matching the order in llama.cpp's
+ own audio example.
+- When a recording is present and no text was typed, the UI substitutes a
+ default instruction, `Listen to this audio and respond.`, so every audio
+ request carries a text part. A typed message is used as-is and the audio is
+ attached to it.
+- `estimate_tokens()` currently counts every non-text part as one "image" and
+ adds a flat 600-token allowance. That branch is generalised to count any
+ non-text part, so an `input_audio` part gets the same flat allowance and
+ the context meter is not blind to a recording. The allowance is not tuned
+ per media type; it exists to show that media is present and roughly
+ expensive.
+
+The existing tool loop, search and skills are untouched. An audio turn is an
+ordinary user message; tool calling still works on top of it if the model
+supports both.
+
+## UI
+
+### Control
+
+A record button sits in the bottom bar next to "Attach…". It is hidden unless
+the model is audio-capable and capture is available.
+
+- Click to start. The button switches to a recording state and the status
+ line reads that it is recording, so the state is never ambiguous.
+- Click again to stop. The clip becomes a pending attachment and the button
+ returns to idle.
+
+This is toggle-then-send: stopping does not send. The user can type text
+alongside the clip, then send or discard it with the existing Send and Clear
+controls. This reuses the attachment lifecycle rather than adding a second
+send path, and it lets a thought be recorded and annotated before it goes.
+
+### Pending clip
+
+`update_attach_label` shows `🎙 voice-note Ns` beside the existing `📄` and
+`🖼` marks. The duration comes from the recorder. `send()` and
+`clear_attachments()` are unchanged; they already carry whatever is in
+`self.attachments`.
+
+### Model mismatch
+
+If a clip is pending and the selected model is not audio-capable, sending is
+blocked with the same offer-to-switch flow as `_ensure_vision_model`: name an
+audio-capable model, ask, and switch. When none is configured, say so rather
+than letting the request fail at the server.
+
+## History
+
+Recordings are discarded after send. The turn stores:
+
+- `messages.content`: the text part only, exactly as image turns already
+ store the text part and not the image bytes. This is the typed message, or
+ the default instruction.
+- One `attachments` row with `kind = 'audio'`, `mime = 'audio/wav'`,
+ `size`, `sha256`, and a synthetic `path` of `voice-note-<seconds>s.wav`.
+ The synthetic name carries the duration for display; it is not a path on
+ disk. `thumb` is NULL and `truncated` is `False`.
+
+No schema change. `attachments.path` is documented as the original path; for
+a recording there is none, and the synthetic marker is the honest stand-in.
+Reopening a chat shows the marker and cannot replay or resend the clip, which
+is the accepted cost of not retaining audio. This matches images, which are
+also not replayed when a conversation continues.
+
+## Error handling
+
+Every failure is a status-bar line, never a crash and never a partial send:
+
+- QtMultimedia missing: the control is absent.
+- No default input device: the click reports it instead of starting.
+- The backend refuses to start: same, with the reason.
+- A clip shorter than 0.3 s: discarded with a note.
+- The 60 s cap: capture stops automatically and the note says why.
+- A pending clip on a non-audio model: blocked with the switch offer.
+
+## Testing
+
+Extend `test_llamachat.py` in its existing assert style:
+
+- `audio.to_wav`: bytes start with `RIFF`/`WAVE`, and reading the result back
+ with the `wave` module yields the same frame count, rate and channel count.
+- `backend.build_user_content`: a recording and a typed text produce a
+ `text` part then an `input_audio` part with raw base64 and `format == "wav"`;
+ an image and audio together keep both parts; a typed message with no
+ attachment still returns a plain string.
+- `backend.classify`: `.wav` and `audio/wav` classify as `audio`; existing
+ text and image cases are unchanged.
+- `backend` listing: a `/v1/models` body carrying
+ `architecture.input_modalities` yields the expected id-to-modalities map,
+ a body without it yields an empty map, and both the enveloped and bare-array
+ shapes parse. This is the layer that detects the local audio model, so it is
+ tested directly.
+- `models`: `ModelInfo.audio` round-trips through `save`/`get`,
+ `resolve` layers store over provider, and `is_empty` treats an audio-only
+ entry as non-empty.
+- `providers.parse`: `audio = true` is read; a missing key is `None`.
+
+The `QAudioSource` plumbing is GUI and hardware, so it stays out of the unit
+suite, consistent with the rest of the window.
+
+## Files
+
+New: `llamachat/audio.py`.
+
+Edited: `llamachat/backend.py`, `llamachat/models.py`,
+`llamachat/providers.py`, `llamachat/modeldialog.py`, `llamachat/ui.py`,
+`README.md`, `CHANGELOG.md`, `test_llamachat.py`.
+
+## Out of scope
+
+- Speech-to-text and the transcript of a recording.
+- Retaining or replaying audio after a message is sent.
+- A microphone device picker.
+- Configurable sample rate, channels, or maximum duration.
+- Server-side transcription tools; this is input to the chat model only.