aboutsummaryrefslogtreecommitdiffstats
path: root/docs/superpowers/specs/2026-09-18-voice-input-design.md
blob: 58aaf588defc874ce38335cdf26684c6272252ac (plain)
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
# Voice input for audio-capable models

Status: design approved

## Problem

llamachat can attach text and images to a message. It has no way to speak a
message. llama.cpp's OpenAI-compatible server accepts audio as a content
part, so a model served with an audio encoder (Qwen2-Audio, Qwen2.5-Omni,
Qwen3-Omni, Voxtral, Ultravox) can hear a recording directly. This adds a
record control that captures the microphone, sends the clip as
`input_audio`, and leaves a marker in history.

This is native audio input, not speech-to-text. The model receives the
waveform and answers it; there is no transcription step and no transcript is
stored. Only models that declare audio support are offered the control.

## Capability: `audio`, resolved in layers

A model is audio-capable when its resolved `audio` flag is `True`. The flag
is tri-state (`True`, `False`, `None` for unknown) and is resolved most
specific first:

1. **Router-reported modalities.** The llama.cpp model router's `/v1/models`
   returns `architecture.input_modalities` per model, for example
   `["text","image","audio"]`. When present it is authoritative: `audio` is
   `"audio" in modalities`, and `vision` is `"image" in modalities`. This is
   what lets a local model work with no manual flag, and it sharpens `vision`
   for local models at the same time.
2. `models.ini`, written by the model settings dialog (`audio = true`).
3. The provider's `[providers.*] audio = true`, used as the prefill for cloud
   models.
4. `presets.ini` (`vision` only, read from `mmproj`). There is no audio key:
   a plain llama-server does not report an audio encoder in its preset
   format, so without the router the manual flag is how a local audio model
   declares itself.

Layer 1 covers the router this app is built for. Layers 2 and 3 remain for
cloud providers, whose `/v1/models` does not carry modalities, and for a
plain llama-server. The `presets.ini` layer stays vision-only because that is
all its `mmproj` line can honestly say.

Changes:

- `backend.Client.models()` extracts each model's `architecture.input_modalities`
  when the server reports it, and returns that mapping alongside the ids;
  `MultiClient.models()` threads it through. A server that does not report it
  yields an empty mapping, so a plain llama-server or a cloud provider is
  unaffected. The parse tolerates both the `{"data": [...]}` envelope and a
  bare array, as the existing listing does.
- `ui` holds the mapping and `model_info()` applies it above the stored and
  provider flags, since a reported modality is a fact and a manual flag is a
  guess.
- `models.ModelInfo` gains `audio: bool | None = None`; `get`, `save`,
  `is_empty` and `resolve` handle it beside `vision`.
- `providers.Provider` gains `audio: bool | None = None`; `parse` reads it.
- `modeldialog` gains an "Accepts audio" checkbox. The vision checkbox's
  tri-state rule applies unchanged: the box carries intent, and an unchecked
  box over an unknown prefill stays `None` rather than writing `False`.

Gate: the record control is shown only when the selected model resolves
`audio is True` and the capture module is available. Unknown and `False` both
hide it, so a recording can never be sent to a model whose audio support is
unconfirmed. An audio-capable model is rare and explicitly reported, so
requiring a positive signal costs nothing and avoids sending a clip to a
text-only model that happens to be unknown.

## Capture: new guarded `llamachat/audio.py`

QtMultimedia is not part of PySide6-Essentials, so the import is guarded and
the feature degrades to absent when the module is missing:

```
try:
    from PySide6.QtMultimedia import QAudioSource, QAudioFormat, QMediaDevices
    AVAILABLE = True
except ImportError:
    AVAILABLE = False
```

`AVAILABLE` is the single gate the UI reads. The README's install section
gains a note that voice input needs `PySide6-Addons`, without making it a hard
requirement.

`AudioRecorder`:

- `QAudioSource` at a fixed `QAudioFormat`: 16000 Hz, 1 channel, `Int16`.
  This is what speech encoders expect and keeps a clip small (about 2 MB per
  minute). The format is not user-configurable in this version.
- Uses `QMediaDevices.defaultAudioInput()`. A device picker is deferred; the
  default microphone is used.
- Writes PCM into a `QBuffer` wrapping a `QByteArray`.
- `start()` begins capture and returns `False` with a reason when there is no
  input device or the backend refuses to start.
- `stop()` stops capture and returns the raw PCM `bytes`.

`to_wav(pcm, rate, channels)` is a module-level function with no Qt
dependency. It prepends a RIFF/WAVE header using the standard library `wave`
module over a `BytesIO` and returns the complete `.wav` bytes. Keeping it
Qt-free is what makes it unit-testable.

Safety rails:

- Hard cap of 60 seconds. A `QTimer` stops capture and the UI reports that
  the cap was hit. This bounds memory (about 2 MB) and the request size.
- A clip shorter than 0.3 seconds is discarded with a status note, so a
  double-click does not send noise.

No window is added. The recorder is a plain object the existing UI drives.

## Send path

`backend.Attachment` gains the kind `'audio'` and a `b64: str = ""` field for
the raw base64 payload. The image field stays a `data:` URL because images use
`image_url`; llama.cpp wants audio as `{"data": "<raw base64>", "format":
"wav"}`, with no `data:` prefix.

- `classify()` recognises WAV: the `audio/wav` and `audio/x-wav` types and
  the `.wav` extension. Only WAV, because the content part's `format` field
  is a fixed `"wav"` and sending MP3 or FLAC bytes under that label would
  misrepresent them. A recorder produces WAV anyway. Recording is the
  headline feature, but a dropped `.wav` file is accepted and treated
  identically: same kind, same content part, same send gate. Sharing the kind
  means classification and the content builder have one audio path instead of
  two. A file the user dropped is theirs and is never deleted; only the
  recorder's temp clip is discarded after send. (ponytail: WAV only; derive
  the format per file if another type is ever needed.)
- `load_attachment()` fills `b64` for an audio file.
- `build_user_content()` gains an audio branch, parallel to the image branch.
  It returns the multi-part array when either images or audio are present:

  ```
  [
    {"type": "text", "text": "<typed text or default instruction>"},
    {"type": "image_url", "image_url": {"url": "data:..."}},
    {"type": "input_audio", "input_audio": {"data": "<raw base64>", "format": "wav"}},
  ]
  ```

  The text part comes first, then audio, matching the order in llama.cpp's
  own audio example.
- When a recording is present and no text was typed, the UI substitutes a
  default instruction, `Listen to this audio and respond.`, so every audio
  request carries a text part. A typed message is used as-is and the audio is
  attached to it.
- `estimate_tokens()` currently counts every non-text part as one "image" and
  adds a flat 600-token allowance. That branch is generalised to count any
  non-text part, so an `input_audio` part gets the same flat allowance and
  the context meter is not blind to a recording. The allowance is not tuned
  per media type; it exists to show that media is present and roughly
  expensive.

The existing tool loop, search and skills are untouched. An audio turn is an
ordinary user message; tool calling still works on top of it if the model
supports both.

## UI

### Control

A record button sits in the bottom bar next to "Attach…". It is hidden unless
the model is audio-capable and capture is available.

- Click to start. The button switches to a recording state and the status
  line reads that it is recording, so the state is never ambiguous.
- Click again to stop. The clip becomes a pending attachment and the button
  returns to idle.

This is toggle-then-send: stopping does not send. The user can type text
alongside the clip, then send or discard it with the existing Send and Clear
controls. This reuses the attachment lifecycle rather than adding a second
send path, and it lets a thought be recorded and annotated before it goes.

### Pending clip

`update_attach_label` shows `🎙 voice-note Ns` beside the existing `📄` and
`🖼` marks. The duration comes from the recorder. `send()` and
`clear_attachments()` are unchanged; they already carry whatever is in
`self.attachments`.

### Model mismatch

If a clip is pending and the selected model is not audio-capable, sending is
blocked with the same offer-to-switch flow as `_ensure_vision_model`: name an
audio-capable model, ask, and switch. When none is configured, say so rather
than letting the request fail at the server.

## History

Recordings are discarded after send. The turn stores:

- `messages.content`: the text part only, exactly as image turns already
  store the text part and not the image bytes. This is the typed message, or
  the default instruction.
- One `attachments` row with `kind = 'audio'`, `mime = 'audio/wav'`,
  `size`, `sha256`, and a synthetic `path` of `voice-note-<seconds>s.wav`.
  The synthetic name carries the duration for display; it is not a path on
  disk. `thumb` is NULL and `truncated` is `False`.

No schema change. `attachments.path` is documented as the original path; for
a recording there is none, and the synthetic marker is the honest stand-in.
Reopening a chat shows the marker and cannot replay or resend the clip, which
is the accepted cost of not retaining audio. This matches images, which are
also not replayed when a conversation continues.

## Error handling

Every failure is a status-bar line, never a crash and never a partial send:

- QtMultimedia missing: the control is absent.
- No default input device: the click reports it instead of starting.
- The backend refuses to start: same, with the reason.
- A clip shorter than 0.3 s: discarded with a note.
- The 60 s cap: capture stops automatically and the note says why.
- A pending clip on a non-audio model: blocked with the switch offer.

## Testing

Extend `test_llamachat.py` in its existing assert style:

- `audio.to_wav`: bytes start with `RIFF`/`WAVE`, and reading the result back
  with the `wave` module yields the same frame count, rate and channel count.
- `backend.build_user_content`: a recording and a typed text produce a
  `text` part then an `input_audio` part with raw base64 and `format == "wav"`;
  an image and audio together keep both parts; a typed message with no
  attachment still returns a plain string.
- `backend.classify`: `.wav` and `audio/wav` classify as `audio`; existing
  text and image cases are unchanged.
- `backend` listing: a `/v1/models` body carrying
  `architecture.input_modalities` yields the expected id-to-modalities map,
  a body without it yields an empty map, and both the enveloped and bare-array
  shapes parse. This is the layer that detects the local audio model, so it is
  tested directly.
- `models`: `ModelInfo.audio` round-trips through `save`/`get`,
  `resolve` layers store over provider, and `is_empty` treats an audio-only
  entry as non-empty.
- `providers.parse`: `audio = true` is read; a missing key is `None`.

The `QAudioSource` plumbing is GUI and hardware, so it stays out of the unit
suite, consistent with the rest of the window.

## Files

New: `llamachat/audio.py`.

Edited: `llamachat/backend.py`, `llamachat/models.py`,
`llamachat/providers.py`, `llamachat/modeldialog.py`, `llamachat/ui.py`,
`README.md`, `CHANGELOG.md`, `test_llamachat.py`.

## Out of scope

- Speech-to-text and the transcript of a recording.
- Retaining or replaying audio after a message is sent.
- A microphone device picker.
- Configurable sample rate, channels, or maximum duration.
- Server-side transcription tools; this is input to the chat model only.