1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
|
# Original project brief
This is the brief llamachat was built from, kept for historical reference.
It describes the v1 target; see the README for what the code actually does
and the CHANGELOG for what changed since.
## Context
The backend is a llama.cpp server running in router mode, configured through
a `presets.ini` file that defines multiple models, including at least one
vision-capable model paired with an mmproj file. The server exposes an
OpenAI-compatible API on localhost.
The target desktop is Slackware with Hyprland on Wayland, using qt6ct and
Kvantum for Qt theming.
## Goals
- A lightweight PySide6 (Qt6) desktop app, not a browser-based or Electron
solution. It should pick up the system Qt theme rather than fighting it.
- Toggled via a global keyboard shortcut bound in Hyprland, behaving like a
scratchpad: press once to show, press again to hide. This requires the
app to run as a single persistent background process that listens on a
Unix domain socket for a toggle signal, rather than spawning a new
process/window per invocation.
- Two chat modes:
1. **One-shot**: quick single question/answer, no multi-turn context.
2. **Chat**: persistent multi-turn conversation, resumable.
Both are saved to history and searchable. The UI should make the mode
switch obvious and easy (e.g. a toggle or tab, not buried in a menu).
- **Model selection**: a dropdown/picker populated from the models
currently available on the llama.cpp router (query `/v1/models` or
equivalent at runtime, do not hardcode model names).
- **File attachment**: both drag-and-drop onto the chat window and a
standard file picker dialog.
- Text/code files: read content and inject into context (with reasonable
size limits and a truncation warning if the file is too large for the
active model's context window).
- Image files: only valid when a vision-capable (mmproj-paired) model is
selected. If an image is dropped while a non-vision model is active,
prompt to switch to the vision preset (this triggers a router-side
unload/reload, so show a "loading model..." state, expect a
multi-second delay, do not treat it as instant).
- **History**: stored in SQLite, covering both one-shot and chat-mode
sessions. Needs:
- A browsable list (most recent first) that can be scrolled and reopened.
- Full-text search across past conversations (use SQLite FTS5).
- Reopening a past chat-mode conversation should allow continuing it
(append new messages with prior context intact). Reopening a one-shot
entry is read-only reference, no continuation needed.
- No automatic "model searches its own history" / RAG feature in this
version. Explicitly out of scope for v1.
- **Streaming responses**: use the OpenAI-compatible streaming endpoint so
responses render token-by-token, not as one blocking call.
## Non-goals for v1
- No RAG / embedding-based history search.
- No multi-user support.
- No cloud fallback or non-local model support.
- No mobile/remote access, this is a local desktop tool only.
## Technical constraints
- Python 3, PySide6 for GUI.
- `httpx` (with streaming support) for talking to the llama.cpp OpenAI-
compatible API.
- SQLite (stdlib `sqlite3`) with an FTS5 virtual table for search.
- Unix domain socket for IPC between the Hyprland keybind and the running
app instance (toggle show/hide). Provide the small client-side command
the Hyprland keybind should call.
- Config (backend URL, socket path, default model, etc.) in a simple file,
not hardcoded, e.g. `~/.config/llamachat/config.toml`.
- Target platform is Slackware/Hyprland/Wayland specifically, so avoid
anything X11-only or assuming GNOME/KDE session services are present.
- The Hyprland config is written in **Lua** (Hyprland 0.55+, which replaced
the older hyprlang `.conf` format). Use Lua syntax throughout (keybinds,
window rules, and the autostart entry as Lua function calls), not
old-style `key = value` hyprlang lines. Do not assume backward
compatibility with hyprlang is present, target Lua only.
## Process
1. Confirm the llama.cpp router port/endpoint and presets.ini model naming
convention before assuming them.
2. Propose the SQLite schema (tables for sessions, messages, attachments)
and the FTS5 setup before implementing.
3. Propose the Unix socket protocol (message format for toggle/show/hide,
and any future commands) before implementing.
4. Provide the Hyprland config snippets needed (window rule for
floating/scratchpad behavior, autostart line, keybind) as part of the
deliverable, not just the Python app.
5. Keep the UI minimal: model picker, mode toggle, chat/input area,
drag-drop zone, history sidebar or panel. No feature creep beyond what
is listed above.
Work through this iteratively: propose the architecture and schema first,
review, then implement.
|