diff options
Diffstat (limited to 'README.md')
| -rw-r--r-- | README.md | 26 |
1 files changed, 26 insertions, 0 deletions
@@ -48,6 +48,10 @@ persistent process and toggles like a scratchpad from a Hyprland keybind. mention a skill's name. Loaded skills show as removable chips, stay for the whole session in chat mode, and in one-shot mode stay in memory until New or the chip is clicked (never written to history). +- **Voice input** for models that accept audio. A Record button appears when + the selected model reports audio input; a clip is sent as an `input_audio` + part alongside any typed text and is discarded after sending. Requires + `PySide6-Addons` (QtMultimedia), which the base install does not pull in. - **Tray icon** for show/hide/quit, hosted by waybar's tray module. Not in this version: RAG or embedding search over history, multi-user @@ -61,6 +65,9 @@ support, remote access. Pillow is only used for attachment thumbnails. +Voice input needs QtMultimedia, which ships in `PySide6-Addons`. Without it +the Record button simply does not appear and everything else works unchanged. + ## Installation PySide6 is not in SlackBuilds, so the simplest route on Slackware is a venv. @@ -503,6 +510,25 @@ user-authored files in a directory you control, trusted the same way the prompts in `prompts/` are, but treat a skill downloaded from elsewhere like any other untrusted instructions. +### Voice input + +A model that reports audio input gets a Record button next to Attach. Click +to start, click Stop to finish; the clip becomes a pending attachment you can +send on its own or with a typed message. Filling a clip with no text sends a +short default instruction so the request always carries a text part. + +Capability comes from the router: llama-server's model router reports each +model's accepted inputs in `/v1/models` (`architecture.input_modalities`), and +a model listing `"audio"` there gets the control. Cloud models, whose +endpoints do not report modalities, can be marked by hand with the +**Accepts audio** box in the model settings dialog. + +Recording is 16 kHz mono WAV, the format speech encoders expect. The clip is +sent as an `input_audio` content part and is **not** written to disk: history +records only a `voice-note Ns` marker, so a recording cannot be replayed or +resent later. Attaching an existing `.wav` file is supported the same way; +only WAV, because the request labels the audio as WAV. + ### External providers llamachat can talk to OpenAI-compatible cloud providers alongside the local |
