diff options
| author | Danilo M. <danix@danix.xyz> | 2026-09-18 20:30:55 +0200 |
|---|---|---|
| committer | Danilo M. <danix@danix.xyz> | 2026-09-18 20:30:55 +0200 |
| commit | f40274657bef103f6ac132b3e0f029221044f155 (patch) | |
| tree | f43a74c04d97413ba21612a06c06fd5401a911f6 | |
| parent | 4edb09f7b9737e0db89e227803c09017fe669599 (diff) | |
| download | llamachat-f40274657bef103f6ac132b3e0f029221044f155.tar.gz llamachat-f40274657bef103f6ac132b3e0f029221044f155.zip | |
docs: document voice input
| -rw-r--r-- | CHANGELOG.md | 6 | ||||
| -rw-r--r-- | README.md | 26 |
2 files changed, 32 insertions, 0 deletions
diff --git a/CHANGELOG.md b/CHANGELOG.md index c4f6473..9eb1ee5 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -80,6 +80,12 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0 mode; one-shot mode keeps them in memory until New or the chip is clicked, never in history. The backend tool loop now dispatches between `web_search` and `load_skill`; `max_searches` caps total tool rounds per turn. +- Voice input for models that accept audio. A Record button appears when the + selected model reports audio input (`architecture.input_modalities` from the + router, or the per-model "Accepts audio" setting); the clip is captured at + 16 kHz mono, sent as an `input_audio` content part, and discarded after + send, leaving only a `voice-note Ns` marker in history. Capture needs + QtMultimedia from PySide6-Addons and is absent without it. ### Fixed @@ -48,6 +48,10 @@ persistent process and toggles like a scratchpad from a Hyprland keybind. mention a skill's name. Loaded skills show as removable chips, stay for the whole session in chat mode, and in one-shot mode stay in memory until New or the chip is clicked (never written to history). +- **Voice input** for models that accept audio. A Record button appears when + the selected model reports audio input; a clip is sent as an `input_audio` + part alongside any typed text and is discarded after sending. Requires + `PySide6-Addons` (QtMultimedia), which the base install does not pull in. - **Tray icon** for show/hide/quit, hosted by waybar's tray module. Not in this version: RAG or embedding search over history, multi-user @@ -61,6 +65,9 @@ support, remote access. Pillow is only used for attachment thumbnails. +Voice input needs QtMultimedia, which ships in `PySide6-Addons`. Without it +the Record button simply does not appear and everything else works unchanged. + ## Installation PySide6 is not in SlackBuilds, so the simplest route on Slackware is a venv. @@ -503,6 +510,25 @@ user-authored files in a directory you control, trusted the same way the prompts in `prompts/` are, but treat a skill downloaded from elsewhere like any other untrusted instructions. +### Voice input + +A model that reports audio input gets a Record button next to Attach. Click +to start, click Stop to finish; the clip becomes a pending attachment you can +send on its own or with a typed message. Filling a clip with no text sends a +short default instruction so the request always carries a text part. + +Capability comes from the router: llama-server's model router reports each +model's accepted inputs in `/v1/models` (`architecture.input_modalities`), and +a model listing `"audio"` there gets the control. Cloud models, whose +endpoints do not report modalities, can be marked by hand with the +**Accepts audio** box in the model settings dialog. + +Recording is 16 kHz mono WAV, the format speech encoders expect. The clip is +sent as an `input_audio` content part and is **not** written to disk: history +records only a `voice-note Ns` marker, so a recording cannot be replayed or +resent later. Attaching an existing `.wav` file is supported the same way; +only WAV, because the request labels the audio as WAV. + ### External providers llamachat can talk to OpenAI-compatible cloud providers alongside the local |
