aboutsummaryrefslogtreecommitdiffstats
diff options
context:
space:
mode:
authorDanilo M. <danix@danix.xyz>2026-09-18 20:30:55 +0200
committerDanilo M. <danix@danix.xyz>2026-09-18 20:30:55 +0200
commitf40274657bef103f6ac132b3e0f029221044f155 (patch)
treef43a74c04d97413ba21612a06c06fd5401a911f6
parent4edb09f7b9737e0db89e227803c09017fe669599 (diff)
downloadllamachat-f40274657bef103f6ac132b3e0f029221044f155.tar.gz
llamachat-f40274657bef103f6ac132b3e0f029221044f155.zip
docs: document voice input
-rw-r--r--CHANGELOG.md6
-rw-r--r--README.md26
2 files changed, 32 insertions, 0 deletions
diff --git a/CHANGELOG.md b/CHANGELOG.md
index c4f6473..9eb1ee5 100644
--- a/CHANGELOG.md
+++ b/CHANGELOG.md
@@ -80,6 +80,12 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
mode; one-shot mode keeps them in memory until New or the chip is clicked,
never in history. The backend tool loop now dispatches between `web_search`
and `load_skill`; `max_searches` caps total tool rounds per turn.
+- Voice input for models that accept audio. A Record button appears when the
+ selected model reports audio input (`architecture.input_modalities` from the
+ router, or the per-model "Accepts audio" setting); the clip is captured at
+ 16 kHz mono, sent as an `input_audio` content part, and discarded after
+ send, leaving only a `voice-note Ns` marker in history. Capture needs
+ QtMultimedia from PySide6-Addons and is absent without it.
### Fixed
diff --git a/README.md b/README.md
index 3e05f00..b1ad066 100644
--- a/README.md
+++ b/README.md
@@ -48,6 +48,10 @@ persistent process and toggles like a scratchpad from a Hyprland keybind.
mention a skill's name. Loaded skills show as removable chips, stay for
the whole session in chat mode, and in one-shot mode stay in memory until
New or the chip is clicked (never written to history).
+- **Voice input** for models that accept audio. A Record button appears when
+ the selected model reports audio input; a clip is sent as an `input_audio`
+ part alongside any typed text and is discarded after sending. Requires
+ `PySide6-Addons` (QtMultimedia), which the base install does not pull in.
- **Tray icon** for show/hide/quit, hosted by waybar's tray module.
Not in this version: RAG or embedding search over history, multi-user
@@ -61,6 +65,9 @@ support, remote access.
Pillow is only used for attachment thumbnails.
+Voice input needs QtMultimedia, which ships in `PySide6-Addons`. Without it
+the Record button simply does not appear and everything else works unchanged.
+
## Installation
PySide6 is not in SlackBuilds, so the simplest route on Slackware is a venv.
@@ -503,6 +510,25 @@ user-authored files in a directory you control, trusted the same way the
prompts in `prompts/` are, but treat a skill downloaded from elsewhere like
any other untrusted instructions.
+### Voice input
+
+A model that reports audio input gets a Record button next to Attach. Click
+to start, click Stop to finish; the clip becomes a pending attachment you can
+send on its own or with a typed message. Filling a clip with no text sends a
+short default instruction so the request always carries a text part.
+
+Capability comes from the router: llama-server's model router reports each
+model's accepted inputs in `/v1/models` (`architecture.input_modalities`), and
+a model listing `"audio"` there gets the control. Cloud models, whose
+endpoints do not report modalities, can be marked by hand with the
+**Accepts audio** box in the model settings dialog.
+
+Recording is 16 kHz mono WAV, the format speech encoders expect. The clip is
+sent as an `input_audio` content part and is **not** written to disk: history
+records only a `voice-note Ns` marker, so a recording cannot be replayed or
+resent later. Attaching an existing `.wav` file is supported the same way;
+only WAV, because the request labels the audio as WAV.
+
### External providers
llamachat can talk to OpenAI-compatible cloud providers alongside the local