aboutsummaryrefslogtreecommitdiffstats
path: root/README.md
diff options
context:
space:
mode:
Diffstat (limited to 'README.md')
-rw-r--r--README.md26
1 files changed, 26 insertions, 0 deletions
diff --git a/README.md b/README.md
index 3e05f00..b1ad066 100644
--- a/README.md
+++ b/README.md
@@ -48,6 +48,10 @@ persistent process and toggles like a scratchpad from a Hyprland keybind.
mention a skill's name. Loaded skills show as removable chips, stay for
the whole session in chat mode, and in one-shot mode stay in memory until
New or the chip is clicked (never written to history).
+- **Voice input** for models that accept audio. A Record button appears when
+ the selected model reports audio input; a clip is sent as an `input_audio`
+ part alongside any typed text and is discarded after sending. Requires
+ `PySide6-Addons` (QtMultimedia), which the base install does not pull in.
- **Tray icon** for show/hide/quit, hosted by waybar's tray module.
Not in this version: RAG or embedding search over history, multi-user
@@ -61,6 +65,9 @@ support, remote access.
Pillow is only used for attachment thumbnails.
+Voice input needs QtMultimedia, which ships in `PySide6-Addons`. Without it
+the Record button simply does not appear and everything else works unchanged.
+
## Installation
PySide6 is not in SlackBuilds, so the simplest route on Slackware is a venv.
@@ -503,6 +510,25 @@ user-authored files in a directory you control, trusted the same way the
prompts in `prompts/` are, but treat a skill downloaded from elsewhere like
any other untrusted instructions.
+### Voice input
+
+A model that reports audio input gets a Record button next to Attach. Click
+to start, click Stop to finish; the clip becomes a pending attachment you can
+send on its own or with a typed message. Filling a clip with no text sends a
+short default instruction so the request always carries a text part.
+
+Capability comes from the router: llama-server's model router reports each
+model's accepted inputs in `/v1/models` (`architecture.input_modalities`), and
+a model listing `"audio"` there gets the control. Cloud models, whose
+endpoints do not report modalities, can be marked by hand with the
+**Accepts audio** box in the model settings dialog.
+
+Recording is 16 kHz mono WAV, the format speech encoders expect. The clip is
+sent as an `input_audio` content part and is **not** written to disk: history
+records only a `voice-note Ns` marker, so a recording cannot be replayed or
+resent later. Attaching an existing `.wav` file is supported the same way;
+only WAV, because the request labels the audio as WAV.
+
### External providers
llamachat can talk to OpenAI-compatible cloud providers alongside the local