# Voice input for audio-capable models Status: design approved ## Problem llamachat can attach text and images to a message. It has no way to speak a message. llama.cpp's OpenAI-compatible server accepts audio as a content part, so a model served with an audio encoder (Qwen2-Audio, Qwen2.5-Omni, Qwen3-Omni, Voxtral, Ultravox) can hear a recording directly. This adds a record control that captures the microphone, sends the clip as `input_audio`, and leaves a marker in history. This is native audio input, not speech-to-text. The model receives the waveform and answers it; there is no transcription step and no transcript is stored. Only models that declare audio support are offered the control. ## Capability: `audio`, resolved in layers A model is audio-capable when its resolved `audio` flag is `True`. The flag is tri-state (`True`, `False`, `None` for unknown) and is resolved most specific first: 1. **Router-reported modalities.** The llama.cpp model router's `/v1/models` returns `architecture.input_modalities` per model, for example `["text","image","audio"]`. When present it is authoritative: `audio` is `"audio" in modalities`, and `vision` is `"image" in modalities`. This is what lets a local model work with no manual flag, and it sharpens `vision` for local models at the same time. 2. `models.ini`, written by the model settings dialog (`audio = true`). 3. The provider's `[providers.*] audio = true`, used as the prefill for cloud models. 4. `presets.ini` (`vision` only, read from `mmproj`). There is no audio key: a plain llama-server does not report an audio encoder in its preset format, so without the router the manual flag is how a local audio model declares itself. Layer 1 covers the router this app is built for. Layers 2 and 3 remain for cloud providers, whose `/v1/models` does not carry modalities, and for a plain llama-server. The `presets.ini` layer stays vision-only because that is all its `mmproj` line can honestly say. Changes: - `backend.Client.models()` extracts each model's `architecture.input_modalities` when the server reports it, and returns that mapping alongside the ids; `MultiClient.models()` threads it through. A server that does not report it yields an empty mapping, so a plain llama-server or a cloud provider is unaffected. The parse tolerates both the `{"data": [...]}` envelope and a bare array, as the existing listing does. - `ui` holds the mapping and `model_info()` applies it above the stored and provider flags, since a reported modality is a fact and a manual flag is a guess. - `models.ModelInfo` gains `audio: bool | None = None`; `get`, `save`, `is_empty` and `resolve` handle it beside `vision`. - `providers.Provider` gains `audio: bool | None = None`; `parse` reads it. - `modeldialog` gains an "Accepts audio" checkbox. The vision checkbox's tri-state rule applies unchanged: the box carries intent, and an unchecked box over an unknown prefill stays `None` rather than writing `False`. Gate: the record control is shown only when the selected model resolves `audio is True` and the capture module is available. Unknown and `False` both hide it, so a recording can never be sent to a model whose audio support is unconfirmed. An audio-capable model is rare and explicitly reported, so requiring a positive signal costs nothing and avoids sending a clip to a text-only model that happens to be unknown. ## Capture: new guarded `llamachat/audio.py` QtMultimedia is not part of PySide6-Essentials, so the import is guarded and the feature degrades to absent when the module is missing: ``` try: from PySide6.QtMultimedia import QAudioSource, QAudioFormat, QMediaDevices AVAILABLE = True except ImportError: AVAILABLE = False ``` `AVAILABLE` is the single gate the UI reads. The README's install section gains a note that voice input needs `PySide6-Addons`, without making it a hard requirement. `AudioRecorder`: - `QAudioSource` at a fixed `QAudioFormat`: 16000 Hz, 1 channel, `Int16`. This is what speech encoders expect and keeps a clip small (about 2 MB per minute). The format is not user-configurable in this version. - Uses `QMediaDevices.defaultAudioInput()`. A device picker is deferred; the default microphone is used. - Writes PCM into a `QBuffer` wrapping a `QByteArray`. - `start()` begins capture and returns `False` with a reason when there is no input device or the backend refuses to start. - `stop()` stops capture and returns the raw PCM `bytes`. `to_wav(pcm, rate, channels)` is a module-level function with no Qt dependency. It prepends a RIFF/WAVE header using the standard library `wave` module over a `BytesIO` and returns the complete `.wav` bytes. Keeping it Qt-free is what makes it unit-testable. Safety rails: - Hard cap of 60 seconds. A `QTimer` stops capture and the UI reports that the cap was hit. This bounds memory (about 2 MB) and the request size. - A clip shorter than 0.3 seconds is discarded with a status note, so a double-click does not send noise. No window is added. The recorder is a plain object the existing UI drives. ## Send path `backend.Attachment` gains the kind `'audio'` and a `b64: str = ""` field for the raw base64 payload. The image field stays a `data:` URL because images use `image_url`; llama.cpp wants audio as `{"data": "", "format": "wav"}`, with no `data:` prefix. - `classify()` recognises the common audio types (`audio/wav`, `audio/x-wav`, `audio/mpeg`, `audio/flac`, `audio/ogg`) and the extensions `.wav`, `.mp3`, `.flac`, `.ogg`, `.m4a`. Recording is the headline feature, but a dropped audio file is accepted and treated identically: same kind, same content part, same send gate. Sharing the kind means classification and the content builder have one audio path instead of two, so the file case costs nothing. A file the user dropped is theirs and is never deleted; only the recorder's temp clip is discarded after send. - `load_attachment()` fills `b64` for an audio file. - `build_user_content()` gains an audio branch, parallel to the image branch. It returns the multi-part array when either images or audio are present: ``` [ {"type": "text", "text": ""}, {"type": "image_url", "image_url": {"url": "data:..."}}, {"type": "input_audio", "input_audio": {"data": "", "format": "wav"}}, ] ``` The text part comes first, then audio, matching the order in llama.cpp's own audio example. - When a recording is present and no text was typed, the UI substitutes a default instruction, `Listen to this audio and respond.`, so every audio request carries a text part. A typed message is used as-is and the audio is attached to it. - `estimate_tokens()` currently counts every non-text part as one "image" and adds a flat 600-token allowance. That branch is generalised to count any non-text part, so an `input_audio` part gets the same flat allowance and the context meter is not blind to a recording. The allowance is not tuned per media type; it exists to show that media is present and roughly expensive. The existing tool loop, search and skills are untouched. An audio turn is an ordinary user message; tool calling still works on top of it if the model supports both. ## UI ### Control A record button sits in the bottom bar next to "Attach…". It is hidden unless the model is audio-capable and capture is available. - Click to start. The button switches to a recording state and the status line reads that it is recording, so the state is never ambiguous. - Click again to stop. The clip becomes a pending attachment and the button returns to idle. This is toggle-then-send: stopping does not send. The user can type text alongside the clip, then send or discard it with the existing Send and Clear controls. This reuses the attachment lifecycle rather than adding a second send path, and it lets a thought be recorded and annotated before it goes. ### Pending clip `update_attach_label` shows `🎙 voice-note Ns` beside the existing `📄` and `🖼` marks. The duration comes from the recorder. `send()` and `clear_attachments()` are unchanged; they already carry whatever is in `self.attachments`. ### Model mismatch If a clip is pending and the selected model is not audio-capable, sending is blocked with the same offer-to-switch flow as `_ensure_vision_model`: name an audio-capable model, ask, and switch. When none is configured, say so rather than letting the request fail at the server. ## History Recordings are discarded after send. The turn stores: - `messages.content`: the text part only, exactly as image turns already store the text part and not the image bytes. This is the typed message, or the default instruction. - One `attachments` row with `kind = 'audio'`, `mime = 'audio/wav'`, `size`, `sha256`, and a synthetic `path` of `voice-note-s.wav`. The synthetic name carries the duration for display; it is not a path on disk. `thumb` is NULL and `truncated` is `False`. No schema change. `attachments.path` is documented as the original path; for a recording there is none, and the synthetic marker is the honest stand-in. Reopening a chat shows the marker and cannot replay or resend the clip, which is the accepted cost of not retaining audio. This matches images, which are also not replayed when a conversation continues. ## Error handling Every failure is a status-bar line, never a crash and never a partial send: - QtMultimedia missing: the control is absent. - No default input device: the click reports it instead of starting. - The backend refuses to start: same, with the reason. - A clip shorter than 0.3 s: discarded with a note. - The 60 s cap: capture stops automatically and the note says why. - A pending clip on a non-audio model: blocked with the switch offer. ## Testing Extend `test_llamachat.py` in its existing assert style: - `audio.to_wav`: bytes start with `RIFF`/`WAVE`, and reading the result back with the `wave` module yields the same frame count, rate and channel count. - `backend.build_user_content`: a recording and a typed text produce a `text` part then an `input_audio` part with raw base64 and `format == "wav"`; an image and audio together keep both parts; a typed message with no attachment still returns a plain string. - `backend.classify`: `.wav` and `audio/wav` classify as `audio`; existing text and image cases are unchanged. - `backend` listing: a `/v1/models` body carrying `architecture.input_modalities` yields the expected id-to-modalities map, a body without it yields an empty map, and both the enveloped and bare-array shapes parse. This is the layer that detects the local audio model, so it is tested directly. - `models`: `ModelInfo.audio` round-trips through `save`/`get`, `resolve` layers store over provider, and `is_empty` treats an audio-only entry as non-empty. - `providers.parse`: `audio = true` is read; a missing key is `None`. The `QAudioSource` plumbing is GUI and hardware, so it stays out of the unit suite, consistent with the rest of the window. ## Files New: `llamachat/audio.py`. Edited: `llamachat/backend.py`, `llamachat/models.py`, `llamachat/providers.py`, `llamachat/modeldialog.py`, `llamachat/ui.py`, `README.md`, `CHANGELOG.md`, `test_llamachat.py`. ## Out of scope - Speech-to-text and the transcript of a recording. - Retaining or replaying audio after a message is sent. - A microphone device picker. - Configurable sample rate, channels, or maximum duration. - Server-side transcription tools; this is input to the chat model only.