aboutsummaryrefslogtreecommitdiffstats
path: root/docs/superpowers/plans
diff options
context:
space:
mode:
authorDanilo M. <danix@danix.xyz>2026-09-18 20:09:05 +0200
committerDanilo M. <danix@danix.xyz>2026-09-18 20:09:05 +0200
commit446c0d537c60b39200e36f9ebc79242668c08d3e (patch)
treeb02af2cc61cf982617620ac69b0cef4a76d67698 /docs/superpowers/plans
parent8cbe904ae2c6c83210f580432f9c59079583aff3 (diff)
downloadllamachat-446c0d537c60b39200e36f9ebc79242668c08d3e.tar.gz
llamachat-446c0d537c60b39200e36f9ebc79242668c08d3e.zip
docs: implementation plan for voice input
Diffstat (limited to 'docs/superpowers/plans')
-rw-r--r--docs/superpowers/plans/2026-09-18-voice-input.md1309
1 files changed, 1309 insertions, 0 deletions
diff --git a/docs/superpowers/plans/2026-09-18-voice-input.md b/docs/superpowers/plans/2026-09-18-voice-input.md
new file mode 100644
index 0000000..6715cf1
--- /dev/null
+++ b/docs/superpowers/plans/2026-09-18-voice-input.md
@@ -0,0 +1,1309 @@
+# Voice Input Implementation Plan
+
+> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking.
+
+**Goal:** Let a recording from the microphone be sent as an `input_audio` content part to a model that reports audio input.
+
+**Architecture:** A guarded `audio.py` records via QtMultimedia at 16 kHz mono and returns WAV bytes. The bytes ride the existing attachment path as a third kind (`audio`), which `build_user_content` turns into the OpenAI `input_audio` part. Capability is detected from the router's per-model `architecture.input_modalities`, with the model settings dialog and provider table as fallbacks. Recordings are discarded after send; history keeps a marker row.
+
+**Tech Stack:** Python 3.11+, PySide6 (QtMultimedia from Addons, optional), httpx, sqlite3, stdlib `wave`.
+
+**Spec:** `docs/superpowers/specs/2026-09-18-voice-input-design.md`
+
+## Global Constraints
+
+- No new Python dependency. Capture uses QtMultimedia when importable; when it is not (PySide6-Essentials only), the record control is absent and nothing else changes.
+- Recording format is fixed: 16000 Hz, 1 channel, Int16, wrapped as WAV. Not configurable.
+- The record control shows only when the resolved `audio` flag is `True`. Unknown and `False` hide it.
+- `input_audio` payload is raw base64 with `format: "wav"`, never a `data:` URL.
+- Recordings are discarded after send. No database schema change.
+- Only WAV files are classified as audio. Do not add MP3/FLAC: the wire `format` is fixed to `"wav"`.
+- Tests are plain assert functions in `test_llamachat.py`, each invoked from its `if __name__ == "__main__":` block. Run the suite with `./test_llamachat.py`.
+- New files carry the GPL-2.0-only SPDX header and copyright line used by every existing module.
+- No em dashes in prose, comments, or commit messages.
+- Commits are GPG-signed by global config. Never pass `-c commit.gpgsign=false`.
+
+## File Structure
+
+- `llamachat/audio.py` (new): WAV encoding and the QtMultimedia recorder.
+- `llamachat/backend.py`: audio attachment kind, content part, attachment factory, token estimate.
+- `llamachat/models.py`: `ModelInfo.audio`, `apply_inputs`.
+- `llamachat/providers.py`: `Provider.audio`.
+- `llamachat/modeldialog.py`: the "Accepts audio" checkbox.
+- `llamachat/ui.py`: modalities storage, record control, send gate.
+- `test_llamachat.py`: tests for all of the above.
+- `README.md`, `CHANGELOG.md`: docs.
+
+---
+
+### Task 1: Audio capture module
+
+**Files:**
+- Create: `llamachat/audio.py`
+- Modify: `test_llamachat.py` (add `import io`, a test, and its call)
+
+**Interfaces:**
+- Produces: `audio.AVAILABLE` (bool), `audio.RATE`, `audio.CHANNELS`, `audio.SAMPLE_BYTES`, `audio.MAX_SECONDS`, `audio.MIN_SECONDS`, `audio.to_wav(pcm: bytes, rate: int = RATE, channels: int = CHANNELS) -> bytes`, and (when `AVAILABLE`) `audio.AudioRecorder(parent=None)` with `.start() -> bool`, `.stop() -> bytes`, `.timed_out: bool`, and signal `.capped`.
+
+- [ ] **Step 1: Write the failing test**
+
+Add `import io` to the imports at the top of `test_llamachat.py` (after `import contextlib`).
+
+Add this test near `test_user_content`:
+
+```python
+def test_audio_wav():
+ """Recorded PCM becomes a readable WAVE file."""
+ import wave
+
+ from llamachat import audio
+
+ # Eight bytes: four Int16 samples.
+ pcm = b"\x01\x00\x02\x00\x03\x00\x04\x00"
+ wav = audio.to_wav(pcm)
+ assert wav[:4] == b"RIFF"
+ assert wav[8:12] == b"WAVE"
+
+ with wave.open(io.BytesIO(wav), "rb") as fh:
+ assert fh.getnchannels() == 1
+ assert fh.getsampwidth() == 2
+ assert fh.getframerate() == 16000
+ assert fh.getnframes() == 4
+ assert fh.readframes(4) == pcm
+
+ # The recorder is optional: AVAILABLE must exist as a bool either way.
+ assert isinstance(audio.AVAILABLE, bool)
+ print("ok audio wav encoding")
+```
+
+Add `test_audio_wav()` to the `if __name__ == "__main__":` block, after `test_user_content()`.
+
+- [ ] **Step 2: Run the test to verify it fails**
+
+Run: `./test_llamachat.py`
+Expected: FAIL with `ModuleNotFoundError: No module named 'llamachat.audio'`.
+
+- [ ] **Step 3: Write `llamachat/audio.py`**
+
+```python
+# SPDX-License-Identifier: GPL-2.0-only
+#
+# llamachat - a small native chat client for a local llama.cpp router
+# Copyright (C) 2026 Danilo M. <danix@danix.xyz>
+#
+# This program is free software; you can redistribute it and/or modify
+# it under the terms of the GNU General Public License version 2 as
+# published by the Free Software Foundation.
+#
+# This program is distributed in the hope that it will be useful,
+# but WITHOUT ANY WARRANTY; without even the implied warranty of
+# MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE. See the
+# GNU General Public License for more details.
+"""Microphone capture for audio-capable models."""
+
+import io
+import wave
+
+try:
+ from PySide6.QtCore import (
+ QBuffer,
+ QByteArray,
+ QIODevice,
+ QObject,
+ QTimer,
+ Signal,
+ )
+ from PySide6.QtMultimedia import QAudioFormat, QAudioSource, QMediaDevices
+
+ AVAILABLE = True
+except ImportError:
+ # PySide6-Essentials ships without QtMultimedia. The feature is optional:
+ # the record control is hidden and nothing else changes.
+ AVAILABLE = False
+
+# Speech encoders expect 16 kHz mono. Int16 is the sample format the WAVE
+# header below declares, so SAMPLE_BYTES and the QString format must agree.
+RATE = 16000
+CHANNELS = 1
+SAMPLE_BYTES = 2
+MAX_SECONDS = 60
+MIN_SECONDS = 0.3
+
+
+def to_wav(pcm: bytes, rate: int = RATE, channels: int = CHANNELS) -> bytes:
+ """Wrap raw Int16 PCM in a RIFF/WAVE header.
+
+ No Qt in here on purpose: this is the one piece of the recorder that can
+ be tested without a sound device.
+ """
+ buf = io.BytesIO()
+ with wave.open(buf, "wb") as wav:
+ wav.setnchannels(channels)
+ wav.setsampwidth(SAMPLE_BYTES)
+ wav.setframerate(rate)
+ wav.writeframes(pcm)
+ return buf.getvalue()
+
+
+if AVAILABLE:
+
+ class AudioRecorder(QObject):
+ """Records the default input device to raw PCM at 16 kHz mono.
+
+ The UI drives it: start(), then stop() for the PCM. A clip that hits
+ MAX_SECONDS is halted by the recorder, which emits `capped` so the UI
+ can finalize it without waiting for a click.
+ """
+
+ capped = Signal()
+
+ def __init__(self, parent=None):
+ super().__init__(parent)
+ self._buffer = QByteArray()
+ self._io = None
+ self._source = None
+ self.timed_out = False
+ self._timer = QTimer(self)
+ self._timer.setSingleShot(True)
+ self._timer.timeout.connect(self._on_cap)
+
+ def start(self) -> bool:
+ """Begin capture. False when there is no input device."""
+ device = QMediaDevices.defaultAudioInput()
+ if device.isNull():
+ return False
+ fmt = QAudioFormat()
+ fmt.setSampleRate(RATE)
+ fmt.setChannelCount(CHANNELS)
+ fmt.setSampleFormat(QAudioFormat.Int16)
+ self._buffer = QByteArray()
+ self.timed_out = False
+ self._io = QBuffer(self._buffer)
+ self._io.open(QIODevice.WriteOnly)
+ self._source = QAudioSource(device, fmt, self)
+ self._source.start(self._io)
+ self._timer.start(MAX_SECONDS * 1000)
+ return True
+
+ def stop(self) -> bytes:
+ """Stop capture and return the raw PCM. Safe to call twice."""
+ self._halt()
+ return bytes(self._buffer)
+
+ def _halt(self) -> None:
+ if self._source is None:
+ return
+ self._timer.stop()
+ self._source.stop()
+ self._source = None
+ self._io.close()
+
+ def _on_cap(self) -> None:
+ self.timed_out = True
+ self._halt()
+ self.capped.emit()
+```
+
+- [ ] **Step 4: Run the test to verify it passes**
+
+Run: `./test_llamachat.py`
+Expected: The full suite passes and prints `ok audio wav encoding`.
+
+- [ ] **Step 5: Commit**
+
+```bash
+git add llamachat/audio.py test_llamachat.py
+git commit -m "feat: audio capture and WAV encoding"
+```
+
+---
+
+### Task 2: Audio attachment and content part
+
+**Files:**
+- Modify: `llamachat/backend.py` (`Attachment`, constants, `classify`, `load_attachment`, `build_user_content`, `estimate_tokens`, new `audio_attachment`)
+- Modify: `test_llamachat.py` (extend `test_classify` and `test_user_content`, extend `test_token_estimate`, add `test_audio_load`)
+
+**Interfaces:**
+- Consumes: `audio.to_wav` from Task 1 (not directly; callers pass bytes).
+- Produces: `backend.AUDIO_MIMES`, `backend.AUDIO_SUFFIXES`, `backend.AUDIO_PROMPT` (str), `backend.Attachment.b64` (str), `backend.audio_attachment(wav: bytes, seconds: int) -> Attachment`, `backend.classify` returning `'audio'`, and `build_user_content` emitting `{"type": "input_audio", "input_audio": {"data": <raw b64>, "format": "wav"}}`.
+
+- [ ] **Step 1: Write the failing tests**
+
+In `test_classify`, after the image asserts, add:
+
+```python
+ assert backend.classify(Path("a.wav")) == "audio"
+ assert backend.classify(Path("a.WAV")) == "audio"
+```
+
+In `test_user_content`, inside the `with tempfile.TemporaryDirectory()` block, after the image asserts, add:
+
+```python
+ # Audio -> the input_audio part llama.cpp expects, raw base64. The
+ # data must not carry a data: prefix, unlike the image URL above.
+ rec = backend.audio_attachment(b"\x00\x01" * 8, seconds=2)
+ assert rec.kind == "audio"
+ assert rec.path.name == "voice-note-2s.wav"
+ assert rec.sha256 and rec.size == 16
+
+ content = backend.build_user_content("listen", [rec])
+ assert isinstance(content, list)
+ assert content[0] == {"type": "text", "text": "listen"}
+ assert content[1]["type"] == "input_audio"
+ assert content[1]["input_audio"]["format"] == "wav"
+ assert content[1]["input_audio"]["data"] == rec.b64
+ assert not content[1]["input_audio"]["data"].startswith("data:")
+
+ # Image and audio together keep both parts, text first.
+ both = backend.build_user_content("look and listen", [img, rec])
+ assert [p["type"] for p in both] == ["text", "image_url", "input_audio"]
+```
+
+Add this new test after `test_attachment_truncation`:
+
+```python
+def test_audio_load():
+ """A .wav file loads as an audio attachment, not as text or unknown."""
+ with tempfile.TemporaryDirectory() as tmp:
+ src = Path(tmp) / "note.wav"
+ src.write_bytes(b"RIFF\x00\x00\x00\x00WAVEfmt ")
+ att = backend.load_attachment(src, 1000)
+ assert att.kind == "audio"
+ assert att.b64
+ assert not att.b64.startswith("data:")
+ assert att.thumb is None
+ assert att.text == ""
+ print("ok audio attachment load")
+```
+
+Add `test_audio_load()` to `__main__` after `test_attachment_truncation()`.
+
+In `test_token_estimate`, after the `with_image` assertion, add:
+
+```python
+ # An audio part costs like an image, not free.
+ with_audio = [
+ {
+ "role": "user",
+ "content": [
+ {"type": "text", "text": "listen"},
+ {"type": "input_audio", "input_audio": {"data": "AAA", "format": "wav"}},
+ ],
+ }
+ ]
+ assert backend.estimate_tokens(with_audio, 3.5) > 500
+```
+
+- [ ] **Step 2: Run the tests to verify they fail**
+
+Run: `./test_llamachat.py`
+Expected: FAIL. `test_classify` reports `AssertionError` on the `.wav` case (classify returns `'unknown'`); `audio_attachment` is undefined.
+
+- [ ] **Step 3: Add the constants and the `Attachment.b64` field**
+
+In `backend.py`, after `IMAGE_MIMES`, add:
+
+```python
+AUDIO_MIMES = {"audio/wav", "audio/x-wav"}
+AUDIO_SUFFIXES = {".wav"}
+
+# Used when a recording is sent with no typed text. llama.cpp pairs audio with
+# a text part, and this is the least presumptuous instruction that still tells
+# the model what the clip is.
+AUDIO_PROMPT = "Listen to this audio and respond."
+```
+
+In `Attachment`, after `data_url`, add:
+
+```python
+ b64: str = "" # kind == 'audio': raw base64 WAV (no data: prefix)
+```
+
+- [ ] **Step 4: Update `classify` and `load_attachment`**
+
+Replace `classify` with:
+
+```python
+def classify(path: Path) -> str:
+ """Decide whether a path is a text, image or audio attachment."""
+ mime, _ = mimetypes.guess_type(path.name)
+ if mime in IMAGE_MIMES:
+ return "image"
+ if mime in AUDIO_MIMES or path.suffix.lower() in AUDIO_SUFFIXES:
+ return "audio"
+ if path.suffix in CODE_SUFFIXES:
+ return "text"
+ if mime and mime.startswith("text/"):
+ return "text"
+ return "unknown"
+```
+
+In `load_attachment`, insert an audio branch after the text branch and before the image lines (`raw = path.read_bytes()`):
+
+```python
+ if kind == "audio":
+ raw = path.read_bytes()
+ att.b64 = base64.b64encode(raw).decode("ascii")
+ return att
+```
+
+- [ ] **Step 5: Update `build_user_content` and add `audio_attachment`**
+
+Replace the tail of `build_user_content` (from `images = ...` to `return content`) with:
+
+```python
+ images = [a for a in attachments if a.kind == "image"]
+ audios = [a for a in attachments if a.kind == "audio"]
+ if not images and not audios:
+ return joined
+
+ content: list[dict] = [{"type": "text", "text": joined}]
+ for att in images:
+ content.append(
+ {"type": "image_url", "image_url": {"url": att.data_url}}
+ )
+ for att in audios:
+ content.append(
+ {
+ "type": "input_audio",
+ "input_audio": {"data": att.b64, "format": "wav"},
+ }
+ )
+ return content
+```
+
+After `build_user_content`, add:
+
+```python
+def audio_attachment(wav: bytes, seconds: int) -> Attachment:
+ """An Attachment for a recorded clip.
+
+ The bytes are never written to disk: the model gets them in the request,
+ and the history row records only their size and hash. The name is
+ synthetic and carries the duration for display.
+ """
+ return Attachment(
+ path=Path(f"voice-note-{seconds}s.wav"),
+ kind="audio",
+ mime="audio/wav",
+ size=len(wav),
+ sha256=hashlib.sha256(wav).hexdigest(),
+ b64=base64.b64encode(wav).decode("ascii"),
+ )
+```
+
+- [ ] **Step 6: Generalize `estimate_tokens`**
+
+Replace the body's counting variable and loop:
+
+```python
+ chars = 0
+ media = 0
+ for message in messages:
+ chars += len(str(message.get("role", "")))
+ content = message.get("content")
+ if isinstance(content, str):
+ chars += len(content)
+ continue
+ for part in content or []:
+ if part.get("type") == "text":
+ chars += len(part.get("text", ""))
+ else:
+ media += 1
+ # A few tokens per message go to the chat template's own markup.
+ overhead = 4 * len(messages)
+ return int(chars / max(chars_per_token, 1.0)) + overhead + media * 600
+```
+
+Update the docstring line `Images are counted as a flat allowance` to `A non-text part (image or audio) is counted as a flat allowance`.
+
+- [ ] **Step 7: Run the tests to verify they pass**
+
+Run: `./test_llamachat.py`
+Expected: PASS, including `ok audio attachment load`, `ok user content assembly`, `ok token estimate`.
+
+- [ ] **Step 8: Commit**
+
+```bash
+git add llamachat/backend.py test_llamachat.py
+git commit -m "feat: audio attachment kind and input_audio content part"
+```
+
+---
+
+### Task 3: Audio capability in metadata and the model dialog
+
+**Files:**
+- Modify: `llamachat/models.py` (`ModelInfo`, `get`, `save`, `resolve`, new `apply_inputs`)
+- Modify: `llamachat/providers.py` (`Provider`, `parse`)
+- Modify: `llamachat/modeldialog.py` (`to_info`, `to_fields`, `ModelDialog`)
+- Modify: `test_llamachat.py` (extend `test_models_store`, `test_provider_parsing`, `test_model_dialog_values`; add `test_audio_metadata_layers`)
+
+**Interfaces:**
+- Produces: `ModelInfo.audio: bool | None`, `models.apply_inputs(info: ModelInfo, inputs) -> ModelInfo`, `Provider.audio: bool | None`, `modeldialog.to_info(..., audio=False, audio_prefill=None)` and `modeldialog.to_fields(info) -> tuple[str, bool, str, str, bool]`.
+
+- [ ] **Step 1: Write the failing tests**
+
+In `test_models_store`, after the `novision` assertion, add:
+
+```python
+ # Audio is the same tri-state and must survive the same round trip.
+ again.save("together:hear", models.ModelInfo(audio=True))
+ assert models.ModelStore(path).get("together:hear").audio is True
+ assert "audio = true" in path.read_text(encoding="utf-8")
+ assert models.ModelInfo(audio=True).is_empty() is False
+ assert models.ModelInfo(audio=False).is_empty() is False
+```
+
+In `test_provider_parsing`, add `"audio": True,` to the `together` entry, then after the `replay_reasoning` asserts add:
+
+```python
+ assert parsed["together"].audio is True
+ assert parsed["local"].audio is None
+```
+
+Add this new test after `test_models_store`:
+
+```python
+def test_audio_metadata_layers():
+ """Audio resolves store over provider; router overlay is Task 4's job."""
+ from llamachat import models, providers
+
+ with tempfile.TemporaryDirectory() as tmp:
+ table = providers.parse(
+ {"providers": {"local": {"base_url": "http://x", "audio": True}}}
+ )
+ store = models.ModelStore(Path(tmp) / "models.ini")
+
+ assert models.resolve("m", table, store).audio is True
+ store.save("m", models.ModelInfo(audio=False))
+ # The store is more specific than the provider.
+ assert models.resolve("m", table, store).audio is False
+ print("ok audio metadata layers")
+```
+
+Add `test_audio_metadata_layers()` to `__main__` after `test_models_store()`.
+
+In `test_model_dialog_values`, replace the two `to_fields` assertions and add audio asserts:
+
+```python
+ # Audio is tri-state in the same way.
+ assert modeldialog.to_info(
+ ctx_text="", vision=False, in_text="", out_text="", audio=True
+ ).audio is True
+ assert modeldialog.to_info(
+ ctx_text="", vision=False, in_text="", out_text="", audio=False
+ ).audio is None
+ assert modeldialog.to_info(
+ ctx_text="", vision=False, in_text="", out_text="",
+ audio=False, audio_prefill=True,
+ ).audio is False
+
+ # Prefill is the inverse: unknown becomes an empty field.
+ assert modeldialog.to_fields(models.ModelInfo()) == ("", False, "", "", False)
+ assert modeldialog.to_fields(
+ models.ModelInfo(ctx_size=4096, vision=True, price_in=0.5, audio=True)
+ ) == ("4096", True, "0.5", "", True)
+```
+
+- [ ] **Step 2: Run the tests to verify they fail**
+
+Run: `./test_llamachat.py`
+Expected: FAIL. `ModelInfo(audio=True)` raises `TypeError: unexpected keyword argument 'audio'`.
+
+- [ ] **Step 3: Add `audio` to `models.ModelInfo` and its layers**
+
+In `models.py`, change the import to:
+
+```python
+from dataclasses import dataclass, replace
+```
+
+Add the field:
+
+```python
+@dataclass
+class ModelInfo:
+ """What a model can tell us. Every field may be unknown."""
+
+ ctx_size: int | None = None
+ vision: bool | None = None
+ audio: bool | None = None
+ price_in: float | None = None
+ price_out: float | None = None
+
+ def is_empty(self) -> bool:
+ return all(
+ v is None
+ for v in (
+ self.ctx_size,
+ self.vision,
+ self.audio,
+ self.price_in,
+ self.price_out,
+ )
+ )
+```
+
+In `get`, add `audio=_get_bool(section, "audio"),` after the `vision` line. In `save`, add `("audio", info.audio),` after the `("vision", info.vision),` line. In `resolve`, add `audio=pick(stored.audio, provider.audio),` after `vision=pick(...)`.
+
+After `resolve`, add:
+
+```python
+def apply_inputs(info: ModelInfo, inputs) -> ModelInfo:
+ """Overlay router-reported input modalities on a ModelInfo.
+
+ `inputs` is the set from the server's `architecture.input_modalities`, or
+ None when the server reports none. A reported modality is a fact, so it
+ overrides a stored or provider guess for vision and audio at once.
+ """
+ if inputs is None:
+ return info
+ return replace(
+ info,
+ vision="image" in inputs,
+ audio="audio" in inputs,
+ )
+```
+
+- [ ] **Step 4: Add `audio` to `providers.Provider` and `parse`**
+
+In `providers.py`, add the field after `vision`:
+
+```python
+ vision: bool | None = None
+ audio: bool | None = None
+```
+
+In `parse`, after `vision = entry.get("vision")`, add `audio = entry.get("audio")`, and in the `Provider(...)` call after `vision=None if vision is None else bool(vision),` add:
+
+```python
+ audio=None if audio is None else bool(audio),
+```
+
+- [ ] **Step 5: Add the dialog checkbox**
+
+In `modeldialog.py`, replace `to_info` and `to_fields` with:
+
+```python
+def to_info(
+ ctx_text, vision, in_text, out_text,
+ vision_prefill=None, audio=False, audio_prefill=None,
+):
+ """Build a ModelInfo from the dialog's raw field values.
+
+ Vision and audio are the tri-state fields, and a binary checkbox cannot
+ hold three states on its own. The checkbox carries the user's intent
+ (checked means "yes"), and the prefill carries what was known before the
+ dialog opened: an unchecked box over an unknown prefill stays unknown
+ rather than writing False, which would shadow a provider-level True
+ through models.resolve's pick(). An unchecked box over a known prefill,
+ True or False, is a real "no".
+ """
+ if vision:
+ resolved_vision = True
+ elif vision_prefill is None:
+ resolved_vision = None
+ else:
+ resolved_vision = False
+
+ if audio:
+ resolved_audio = True
+ elif audio_prefill is None:
+ resolved_audio = None
+ else:
+ resolved_audio = False
+
+ return ModelInfo(
+ ctx_size=_number(ctx_text, int),
+ vision=resolved_vision,
+ audio=resolved_audio,
+ price_in=_number(in_text, float),
+ price_out=_number(out_text, float),
+ )
+
+
+def to_fields(info: ModelInfo) -> tuple[str, bool, str, str, bool]:
+ """The inverse, for prefilling. Unknown becomes an empty field."""
+ return (
+ "" if info.ctx_size is None else str(info.ctx_size),
+ bool(info.vision),
+ "" if info.price_in is None else str(info.price_in),
+ "" if info.price_out is None else str(info.price_out),
+ bool(info.audio),
+ )
+```
+
+In `ModelDialog.__init__`, add the prefill tracker next to `self._vision_prefill`:
+
+```python
+ self._audio_prefill = info.audio
+```
+
+Replace the unpack and checkbox wiring:
+
+```python
+ ctx, vision, price_in, price_out, audio = to_fields(info)
+ self.ctx = QLineEdit(ctx)
+ self.ctx.setPlaceholderText("unknown")
+ self.vision = QCheckBox("Accepts images")
+ self.vision.setChecked(vision)
+ self.audio = QCheckBox("Accepts audio")
+ self.audio.setChecked(audio)
+ self.price_in = QLineEdit(price_in)
+ self.price_in.setPlaceholderText("unpriced")
+ self.price_out = QLineEdit(price_out)
+ self.price_out.setPlaceholderText("unpriced")
+```
+
+In the form, add a row after the vision row:
+
+```python
+ form.addRow("", self.vision)
+ form.addRow("", self.audio)
+```
+
+Replace `info()` with:
+
+```python
+ def info(self) -> ModelInfo:
+ """What the user entered."""
+ return to_info(
+ self.ctx.text(),
+ self.vision.isChecked(),
+ self.price_in.text(),
+ self.price_out.text(),
+ vision_prefill=self._vision_prefill,
+ audio=self.audio.isChecked(),
+ audio_prefill=self._audio_prefill,
+ )
+```
+
+- [ ] **Step 6: Run the tests to verify they pass**
+
+Run: `./test_llamachat.py`
+Expected: PASS, including `ok audio metadata layers` and `ok model dialog value conversion`.
+
+- [ ] **Step 7: Commit**
+
+```bash
+git add llamachat/models.py llamachat/providers.py llamachat/modeldialog.py test_llamachat.py
+git commit -m "feat: audio model capability in metadata and dialog"
+```
+
+---
+
+### Task 4: Router-reported input modalities
+
+**Files:**
+- Modify: `llamachat/backend.py` (`Client.models`, `MultiClient.models`)
+- Modify: `llamachat/ui.py` (`__init__`, `refresh_models`, `model_info`)
+- Modify: `test_llamachat.py` (`test_client_auth_header`, `test_multi_client`, new `test_model_inputs`)
+
+**Interfaces:**
+- Consumes: `models.apply_inputs` from Task 3.
+- Produces: `Client.models() -> tuple[list[str], dict[str, frozenset[str]]]`; `MultiClient.models() -> tuple[list[str], list[str], dict[str, frozenset[str]]]`; `ChatWindow.modalities: dict[str, frozenset[str]]`.
+
+- [ ] **Step 1: Write the failing tests**
+
+Add this new test after `test_multi_client`:
+
+```python
+def test_model_inputs():
+ """The router's per-model input modalities are read from the listing."""
+ class _Resp:
+ status_code = 200
+
+ def __init__(self, payload):
+ self._payload = payload
+
+ def raise_for_status(self):
+ return None
+
+ def json(self):
+ return self._payload
+
+ payload = {
+ "data": [
+ {
+ "id": "gemma4",
+ "architecture": {"input_modalities": ["text", "image", "audio"]},
+ },
+ {"id": "qwen", "architecture": {"input_modalities": ["text"]}},
+ {"id": "plain"},
+ ]
+ }
+ import httpx
+ with _patched(httpx, "get", lambda *a, **k: _Resp(payload)):
+ ids, modalities = backend.Client("http://x.example.org").models()
+
+ assert ids == ["gemma4", "qwen", "plain"]
+ assert modalities["gemma4"] == frozenset({"text", "image", "audio"})
+ assert modalities["qwen"] == frozenset({"text"})
+ # A server that reports nothing yields an empty map, not a guess.
+ assert "plain" not in modalities
+ print("ok model input modalities")
+```
+
+Add `test_model_inputs()` to `__main__` after `test_multi_client()`.
+
+In `test_client_auth_header`, change the two `.models()` assertions:
+
+```python
+ assert backend.Client("http://x.example.org").models() == (["m1"], {})
+```
+
+and
+
+```python
+ assert backend.Client("http://x.example.org").models() == (["bare"], {})
+```
+
+In `test_multi_client`, change the stub to return a pair and unpack three values:
+
+```python
+ def models(self, list_timeout=30):
+ if self.base_url not in listings:
+ raise backend.BackendError(f"cannot reach {self.base_url}")
+ return listings[self.base_url], {}
+```
+
+```python
+ listed, problems, modalities = multi.models()
+```
+
+After the `listed ==` assertion, add:
+
+```python
+ # The stub reports no modalities, so the map is empty.
+ assert modalities == {}
+```
+
+and change the empty-filter unpack:
+
+```python
+ _, notes, _ = empty.models()
+```
+
+- [ ] **Step 2: Run the tests to verify they fail**
+
+Run: `./test_llamachat.py`
+Expected: FAIL. `test_client_auth_header` asserts a list but gets a tuple; `test_multi_client` raises `ValueError: not enough values to unpack`; `test_model_inputs` fails on `Client.models()` returning a list.
+
+- [ ] **Step 3: Update `Client.models`**
+
+Replace the tail of `Client.models` (from `data = payload ...` to the `return`) with:
+
+```python
+ # Some OpenAI-compatible providers return a bare array instead of
+ # the standard {"data": [...]} envelope. Accept both.
+ data = payload if isinstance(payload, list) else payload.get("data", [])
+ ids: list[str] = []
+ modalities: dict[str, frozenset[str]] = {}
+ for entry in data:
+ if "id" not in entry:
+ continue
+ ids.append(entry["id"])
+ # The llama.cpp model router reports which inputs a model accepts.
+ # A plain llama-server and every cloud provider omit it, which is
+ # why the map is allowed to be empty rather than assumed.
+ arch = entry.get("architecture")
+ inputs = arch.get("input_modalities") if isinstance(arch, dict) else None
+ if isinstance(inputs, list):
+ modalities[entry["id"]] = frozenset(str(x) for x in inputs)
+ return ids, modalities
+```
+
+Update the method signature and docstring:
+
+```python
+ def models(
+ self, list_timeout: int = 30
+ ) -> tuple[list[str], dict[str, frozenset[str]]]:
+ """Model ids the router currently offers, plus reported input modalities.
+
+ `list_timeout` is the overall deadline for the listing call, separate
+ from the per-turn `self.timeout` used by chat: the picker has to
+ populate fast, and a dead provider must not hang it for minutes.
+ The connect deadline is shorter still, so an unreachable host is
+ reported quickly while a reachable one still gets the full window.
+
+ The second return value maps a model id to the set of inputs the
+ server reports it accepts. It is empty for a server that does not
+ report modalities.
+ """
+```
+
+- [ ] **Step 4: Update `MultiClient.models`**
+
+Replace `MultiClient.models` with:
+
+```python
+ def models(self) -> tuple[list[str], list[str], dict[str, frozenset[str]]]:
+ """Every offered model id, plus notes and reported input modalities.
+
+ A provider that is unreachable or whose filter matched nothing must
+ not stop the others being listed: local models have to stay usable
+ when the network is down.
+
+ The fan-out is serial on purpose: `KeyResolver._cache` is unlocked
+ and only safe while every caller resolves on the GUI thread. Threading
+ this would spawn two pinentry prompts for one hardware token.
+ """
+ listed: list[str] = []
+ problems: list[str] = []
+ modalities: dict[str, frozenset[str]] = {}
+ for name, provider in self.table.items():
+ try:
+ available, reported = self._listing_client(provider).models(
+ list_timeout=self.list_timeout
+ )
+ except (BackendError, providers_mod.KeyResolutionError) as exc:
+ problems.append(f"{name}: {exc}")
+ continue
+ kept = providers_mod.apply_filter(provider, available)
+ if available and not kept:
+ problems.append(
+ f"{name}: 0 of {len(available)} models matched filter"
+ )
+ for model in kept:
+ qualified = providers_mod.qualify(name, model)
+ listed.append(qualified)
+ if model in reported:
+ modalities[qualified] = reported[model]
+ return listed, problems, modalities
+```
+
+- [ ] **Step 5: Store and apply modalities in the window**
+
+In `ui.py`, in `ChatWindow.__init__`, after `self.loaded_skills: list[str] = []`, add:
+
+```python
+ # Router-reported input modalities per model id, refreshed on listing.
+ self.modalities: dict[str, frozenset[str]] = {}
+```
+
+In `refresh_models`, change the call and assign the map before the empty check, so a failed listing still clears stale modalities:
+
+```python
+ available, listing_problems, modalities = self.client.models()
+ self.modalities = modalities
+ if not available:
+ self.show_status(
+ "; ".join(listing_problems) or "No models available", error=True
+ )
+ return
+```
+
+In `model_info`, change the final `return info` to:
+
+```python
+ # A reported modality is a fact and overrides the stored, provider and
+ # preset guesses above.
+ return models_mod.apply_inputs(info, self.modalities.get(model_id))
+```
+
+- [ ] **Step 6: Run the tests to verify they pass**
+
+Run: `./test_llamachat.py`
+Expected: PASS, including `ok model input modalities`, `ok multi-provider client`, and `ok client auth header`.
+
+- [ ] **Step 7: Commit**
+
+```bash
+git add llamachat/backend.py llamachat/ui.py test_llamachat.py
+git commit -m "feat: detect model input modalities from the router"
+```
+
+---
+
+### Task 5: Record control and voice-note marker
+
+**Files:**
+- Modify: `llamachat/ui.py` (import, `_button_icon`, `__init__`, `_build_ui`, new recording methods, `update_attach_label`, `_bubble_html`)
+- Modify: `test_llamachat.py` (new `test_audio_bubble_marker`)
+
+**Interfaces:**
+- Consumes: `audio.AudioRecorder`, `backend.audio_attachment`, `backend.classify`.
+- Produces: `ChatWindow.recording: bool`, `ChatWindow.recorder`, `_update_record_button()`, `_toggle_record()`, `_finish_recording()`.
+
+- [ ] **Step 1: Write the failing test**
+
+Add after `test_markdown_rendering`:
+
+```python
+def test_audio_bubble_marker():
+ """A discarded recording shows as a voice note, not a missing file."""
+ import os
+
+ os.environ.setdefault("QT_QPA_PLATFORM", "offscreen")
+ from PySide6.QtWidgets import QApplication
+
+ from llamachat.ui import _bubble_html
+
+ app = QApplication.instance() or QApplication([])
+ assert app is not None
+
+ rendered = _bubble_html(
+ "user",
+ "Listen to this audio and respond.",
+ saved_rows=[{"kind": "audio", "path": "voice-note-4s.wav"}],
+ )
+ assert "πŸŽ™ voice-note-4s.wav" in rendered
+ # The synthetic path is not on disk, so it must not read as missing.
+ assert "missing" not in rendered
+ print("ok audio bubble marker")
+```
+
+Add `test_audio_bubble_marker()` to `__main__` after `test_markdown_rendering()`.
+
+- [ ] **Step 2: Run the test to verify it fails**
+
+Run: `./test_llamachat.py`
+Expected: FAIL. The rendered HTML contains `voice-note-4s.wav (missing)` and no microphone marker.
+
+- [ ] **Step 3: Add the module import and the mic glyph**
+
+At the top of `ui.py`, where sibling modules are imported, add:
+
+```python
+from . import audio
+```
+
+In `_button_icon`, before `elif kind == "send":`, add:
+
+```python
+ elif kind == "record":
+ # A microphone: capsule, cradle arc and stand.
+ capsule = QPainterPath()
+ capsule.addRoundedRect(QRectF(9.5, 4.0, 5.0, 9.5), 2.5, 2.5)
+ painter.drawPath(capsule)
+ painter.drawArc(QRectF(7.0, 7.0, 10.0, 9.5), 200 * 16, 140 * 16)
+ painter.drawLine(12, 17, 12, 20)
+ painter.drawLine(9, 20, 15, 20)
+```
+
+- [ ] **Step 4: Create the recorder and connect its cap signal**
+
+In `ChatWindow.__init__`, after the `self.modalities` line added in Task 4, add:
+
+```python
+ self.recording = False
+ self.recorder = audio.AudioRecorder(self) if audio.AVAILABLE else None
+```
+
+Immediately after `self.recorder = ...`, add:
+
+```python
+ if self.recorder is not None:
+ self.recorder.capped.connect(self._finish_recording)
+```
+
+- [ ] **Step 5: Add the button to the bottom bar**
+
+In `_build_ui`, immediately before `self.attach_button = QPushButton("Attach…")`, add:
+
+```python
+ self.record_button = QPushButton("Record")
+ self.record_button.setIcon(_button_icon("record"))
+ self.record_button.setToolTip("Record audio for an audio-capable model")
+ self.record_button.clicked.connect(self._toggle_record)
+ self.record_button.hide()
+ buttons.addWidget(self.record_button)
+```
+
+- [ ] **Step 6: Add the recording methods**
+
+Add a new section after `update_attach_label`:
+
+```python
+ # -- recording --------------------------------------------------------
+
+ def _update_record_button(self) -> None:
+ """Show the record control only for an audio-capable model."""
+ if not hasattr(self, "record_button"):
+ return # still building the window
+ can = self.recorder is not None and self.current_info().audio is True
+ self.record_button.setVisible(can)
+ self.record_button.setText("Stop" if self.recording else "Record")
+ self.record_button.setIcon(
+ _button_icon("stop" if self.recording else "record")
+ )
+
+ def _toggle_record(self) -> None:
+ """Start a recording, or finish the one in progress."""
+ if self.recording:
+ self._finish_recording()
+ return
+ if self.recorder is None or self.current_info().audio is not True:
+ return
+ if not self.recorder.start():
+ self.show_status("No microphone available.", error=True)
+ return
+ self.recording = True
+ self.show_status("Recording. Click Stop when done.")
+ self._update_record_button()
+
+ @Slot()
+ def _finish_recording(self) -> None:
+ """Stop capture and hold the clip as a pending attachment."""
+ if not self.recording:
+ return
+ pcm = self.recorder.stop()
+ self.recording = False
+ self._update_record_button()
+
+ seconds = len(pcm) / (audio.RATE * audio.CHANNELS * audio.SAMPLE_BYTES)
+ if seconds < audio.MIN_SECONDS:
+ self.show_status("Recording too short, discarded.", error=True)
+ return
+ if self.recorder.timed_out:
+ self.show_status(
+ f"Recording stopped at the {audio.MAX_SECONDS}s limit."
+ )
+ else:
+ self.hide_status()
+ self.attachments.append(
+ backend.audio_attachment(audio.to_wav(pcm), round(seconds))
+ )
+ self.update_attach_label()
+```
+
+- [ ] **Step 7: Show the audio marker in the pending label and in history**
+
+In `update_attach_label`, replace the `mark = ...` line with:
+
+```python
+ mark = {"image": "πŸ–Ό", "audio": "πŸŽ™"}.get(att.kind, "πŸ“„")
+```
+
+In `_bubble_html`, replace the `elif saved_rows:` block body with:
+
+```python
+ elif saved_rows:
+ parts = []
+ for row in saved_rows:
+ if row["kind"] == "audio":
+ # A recording is discarded, so there is no file to check for.
+ parts.append("πŸŽ™ " + html.escape(row["path"]))
+ continue
+ missing = "" if Path(row["path"]).exists() else " (missing)"
+ parts.append(html.escape(row["path"]) + missing)
+ files = f"<br><i>files: {', '.join(parts)}</i>"
+```
+
+- [ ] **Step 8: Wire the button refresh into model changes and listing**
+
+In `_on_model_changed`, add `self._update_record_button()` after `self.update_meter()`.
+
+In `refresh_models`, after `self.update_cost()`, add `self._update_record_button()`.
+
+- [ ] **Step 9: Run the tests to verify they pass**
+
+Run: `./test_llamachat.py`
+Expected: PASS, including `ok audio bubble marker`.
+
+- [ ] **Step 10: Commit**
+
+```bash
+git add llamachat/ui.py test_llamachat.py
+git commit -m "feat: record control and voice-note history marker"
+```
+
+---
+
+### Task 6: Gate audio sends and supply the default prompt
+
+**Files:**
+- Modify: `llamachat/ui.py` (new `audio_models`, new `_ensure_audio_model`, `send`)
+
+**Interfaces:**
+- Consumes: `backend.AUDIO_PROMPT`, `self.attachments`, `self.model_info`.
+- Produces: `ChatWindow.audio_models() -> list[str]`, `ChatWindow._ensure_audio_model() -> bool`.
+
+- [ ] **Step 1: Add the audio-model helpers**
+
+In `ui.py`, after `vision_models`, add:
+
+```python
+ def audio_models(self) -> list[str]:
+ names = []
+ for i in range(self.model_box.count()):
+ name = self.model_box.itemData(i)
+ if self.model_info(name).audio is True:
+ names.append(name)
+ return names
+
+ def _ensure_audio_model(self) -> bool:
+ """Offer to switch to an audio-capable model. True when one is active.
+
+ Unlike vision, an unknown model is not given the benefit of the
+ doubt: audio input is rare and explicitly reported, so only a model
+ known to accept it is allowed to receive a recording.
+ """
+ if self.current_info().audio is True:
+ return True
+
+ candidates = self.audio_models()
+ if not candidates:
+ QMessageBox.warning(
+ self,
+ "No audio model",
+ "No model on the router reports audio input, so a recording "
+ "cannot be sent.",
+ )
+ return False
+
+ target = candidates[0]
+ answer = QMessageBox.question(
+ self,
+ "Switch model?",
+ f"Audio needs a model that accepts voice input.\n\n"
+ f"Switch to {target}?\n\n"
+ "The router unloads the current model to do this, so the next "
+ "reply will take several seconds to start.",
+ QMessageBox.Yes | QMessageBox.No,
+ QMessageBox.Yes,
+ )
+ if answer != QMessageBox.Yes:
+ return False
+
+ index = self.model_box.findData(target)
+ if index >= 0:
+ self.model_box.setCurrentIndex(index)
+ return True
+```
+
+- [ ] **Step 2: Gate `send` on the audio model**
+
+In `send()`, insert between the first empty guard and `model = self.current_model()`:
+
+```python
+ if any(a.kind == "audio" for a in self.attachments):
+ # A switch here changes the model, so this must run before the
+ # model is read below.
+ if not self._ensure_audio_model():
+ return
+ if not text:
+ text = backend.AUDIO_PROMPT
+```
+
+The result:
+
+```python
+ text = self.input.toPlainText().strip()
+ if not text and not self.attachments:
+ return
+
+ if any(a.kind == "audio" for a in self.attachments):
+ if not self._ensure_audio_model():
+ return
+ if not text:
+ text = backend.AUDIO_PROMPT
+
+ model = self.current_model()
+ if not model:
+ self.show_status("No model selected.", error=True)
+ return
+```
+
+- [ ] **Step 3: Run the suite to confirm nothing regressed**
+
+Run: `./test_llamachat.py`
+Expected: PASS (no unit test covers the window's send path; this confirms the imports and construction still work).
+
+- [ ] **Step 4: Manual smoke against the real router**
+
+Start the app (`llamachat --quit; llamachat --daemon`), select `Gemma4-12B-qat-mtp`, confirm the Record button appears. Record a short clip, confirm it becomes a pending `πŸŽ™ voice-note Ns` item, type a question, send, and confirm a reply. Then select `Qwen3.8-9B`, confirm the button disappears, and confirm a pending clip blocks with the switch offer.
+
+- [ ] **Step 5: Commit**
+
+```bash
+git add llamachat/ui.py
+git commit -m "feat: gate audio sends and add the default audio prompt"
+```
+
+---
+
+### Task 7: Documentation
+
+**Files:**
+- Modify: `README.md` (Features, Requirements, a new Voice input section)
+- Modify: `CHANGELOG.md` (Unreleased, Added)
+
+**Interfaces:** None.
+
+- [ ] **Step 1: Add the README feature and requirement notes**
+
+In the Features list, after the Skills bullet, add:
+
+```markdown
+- **Voice input** for models that accept audio. A Record button appears when
+ the selected model reports audio input; a clip is sent as an `input_audio`
+ part alongside any typed text and is discarded after sending. Requires
+ `PySide6-Addons` (QtMultimedia), which the base install does not pull in.
+```
+
+In Requirements, after the PySide6 line, add:
+
+```markdown
+Voice input needs QtMultimedia, which ships in `PySide6-Addons`. Without it
+the Record button simply does not appear and everything else works unchanged.
+```
+
+- [ ] **Step 2: Add a README section**
+
+After the `### Skills` section and before `### External providers`, add:
+
+```markdown
+### Voice input
+
+A model that reports audio input gets a Record button next to Attach. Click
+to start, click Stop to finish; the clip becomes a pending attachment you can
+send on its own or with a typed message. Filling a clip with no text sends a
+short default instruction so the request always carries a text part.
+
+Capability comes from the router: llama-server's model router reports each
+model's accepted inputs in `/v1/models` (`architecture.input_modalities`), and
+a model listing `"audio"` there gets the control. Cloud models, whose
+endpoints do not report modalities, can be marked by hand with the
+**Accepts audio** box in the model settings dialog.
+
+Recording is 16 kHz mono WAV, the format speech encoders expect. The clip is
+sent as an `input_audio` content part and is **not** written to disk: history
+records only a `voice-note Ns` marker, so a recording cannot be replayed or
+resent later. Attaching an existing `.wav` file is supported the same way;
+only WAV, because the request labels the audio as WAV.
+```
+
+- [ ] **Step 3: Add the CHANGELOG entry**
+
+Under `## [Unreleased]` -> `### Added`, after the Skills bullet, add:
+
+```markdown
+- Voice input for models that accept audio. A Record button appears when the
+ selected model reports audio input (`architecture.input_modalities` from the
+ router, or the per-model "Accepts audio" setting); the clip is captured at
+ 16 kHz mono, sent as an `input_audio` content part, and discarded after
+ send, leaving only a `voice-note Ns` marker in history. Capture needs
+ QtMultimedia from PySide6-Addons and is absent without it.
+```
+
+- [ ] **Step 4: Verify and commit**
+
+Run: `./test_llamachat.py`
+Expected: PASS, including `ok version matches changelog` (the Unreleased section is not a version bump, so this check is unaffected).
+
+```bash
+git add README.md CHANGELOG.md
+git commit -m "docs: document voice input"
+```
+
+---
+
+## Self-Review Notes
+
+- Spec coverage: capture (Task 1), content part and discard-after-send (Tasks 2, 5), capability via router/stored/provider/presets (Tasks 3, 4), UI control and default prompt (Tasks 5, 6), history marker (Tasks 2, 5), error handling (Tasks 1, 2, 5, 6), testing (each task), docs (Task 7). All spec sections map to a task.
+- Type consistency: `apply_inputs` is defined in Task 3 and consumed in Task 4; `audio_attachment` in Task 2 and consumed in Task 5; `AUDIO_PROMPT` in Task 2 and consumed in Task 6; `_update_record_button` defined and called within Task 5; `modalities` initialized in Task 4 and read in Task 4 before Task 5 adds the button refresh.
+- Known intentional gap: the QtMultimedia recorder itself and the window's send path are not unit tested, matching the repo's "self-checks for everything except the GUI" convention. Task 6 covers them with a manual smoke against the real router.