aboutsummaryrefslogtreecommitdiffstats
path: root/test_llamachat.py
AgeCommit message (Collapse)AuthorFilesLines
4 daysfix: web search on DeepSeek and other reasoning modelsfeature/external-providersDanilo M.1-10/+318
A searched turn on a cloud reasoning model ended at the thinking: the model emitted the tool call, but finish_reason: "tool_calls" landed on the same SSE line as the include_usage block, so the single-event parser returned that line as a usage chunk and the loop never saw the tool finish. The parser now emits every event a line carries, so the search fires. Also in this change: - Replay each round's reasoning_content on the assistant tool-call message, which interleaved-thinking models require to keep going. - Add a per-provider replay_reasoning option (DeepSeek, SiliconFlow GLM-4.7+) to carry prior turns' reasoning_content when search is on. - Add an on-demand diagnostic log gated by $LLAMACHAT_DEBUG_LOG. - Record provider, usage_json and reported_cost_usd per reply, so a searched turn keeps every round's billed usage for external consumers.
2026-08-11fix: prevent cloud reasoning models from exhausting default max_tokensDanilo M.1-6/+68
- Send max_tokens = 32768 for non-local providers so reasoning models have room for both chain-of-thought and answer. - Add per-provider thinking_budget option for endpoints (currently SiliconFlow) that cap reasoning tokens separately. - Accept bare JSON arrays from /v1/models; some OpenAI-compatible endpoints omit the {"data": [...]} envelope.
2026-08-10feat: list and route models through every configured providerDanilo M.1-4/+8
Metadata now resolves models.ini over provider defaults over presets.ini, so a local model keeps getting its context and vision from the preset while remaining overridable. A model whose vision support is unknown no longer blocks an attachment: unknown is not the same as no, and the API will reject an image if it really cannot take one. Key resolution moves to the worker thread so a pass: key's pinentry never freezes the GUI. KeyResolver gains a lock since StreamWorker and TitleWorker can now overlap. Skipped providers surface in the status bar rather than only on stderr. model_info uses dataclasses.replace to avoid mutating the ModelInfo that resolve() may return from the store. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-10feat: cost label showing spend and the next send's projectionDanilo M.1-0/+31
Three states stay distinct: blank for a free local model, ? for a cloud model whose prices were never entered, and a figure when they were. Collapsing the first two would make an unpriced cloud model read as free. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-10feat: dialog for per-model context size, vision and pricesDanilo M.1-0/+54
Field conversion is separated from the widget so the part with the edge cases is testable without a running Qt application. Vision is tri-state through a binary checkbox: an unchecked box with an unknown prefill stays unknown rather than writing False, which would shadow a provider-level vision=True through models.resolve's pick(). Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-10feat: fan model listing out across providers and route by idDanilo M.1-25/+97
A provider that is unreachable or whose filter matched nothing is reported rather than fatal: local models must stay usable when the network is down. Clients are built on demand so a key is resolved only when that provider is really used. The model listing carries its own timeout, separate from the per-turn one: the picker has to populate fast, and a dead cloud provider must not hang it for minutes. Client.models() takes list_timeout (default 30) and uses httpx.Timeout(list_timeout, connect=5), and MultiClient threads its own list_timeout through to each listing call. The fan-out stays serial on purpose: KeyResolver._cache is unlocked and only safe while every caller resolves on the GUI thread. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-09feat: send bearer auth when a provider needs a keyDanilo M.1-2/+98
The local router needs none, so the header is omitted entirely rather than sent empty. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-09feat: record token counts and the producing model per messageDanilo M.1-0/+97
Cost has to survive reopening a conversation, which means storing what each reply used. The model goes on the message rather than the session because sessions records only the current one, and a conversation that switched models would otherwise be priced entirely at whichever is selected now. All three columns are nullable so old rows read as unknown. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-09fix: reject non-finite provider prices, and read them as unpricedDanilo M.1-0/+47
A nan or inf in [providers.*] survived float() and reached the cost arithmetic, rendering $nan or, for -inf, a negative cost. Worse than a display wart: is_priced answered True, so the readout took the priced branch and showed a poisoned figure in the one state the ? is there to admit it cannot say. Guarded at both ends. providers._number mirrors models._get, which already covered the models.ini writer. is_priced states its own predicate completely rather than trusting its callers, since the dialog is a third writer and the file is documented for hand-editing. Records why float money stays, and that exact-equality assertions in the cost test hold for the chosen prices rather than in general. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-09feat: resolve model metadata in layers and compute costDanilo M.1-0/+183
Resolution merges per field, not per source, so an entry setting only ctx_size still inherits the provider's prices. Cost prices each reply at the model that produced it, and rows predating the token columns contribute zero rather than a fabricated number. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-09fix: reject non-finite prices and escape a model named DEFAULTDanilo M.1-24/+42
float() accepts "nan" and "inf", so a hand-edited models.ini could feed either into the cost arithmetic and render "$nan". Rejected once in _get rather than at each downstream call site. A model whose id is literally DEFAULT wrote a [DEFAULT] section, which a stock configparser reader treats as defaults inherited by every other model. Since the file is documented for hand-editing, that reintroduces the contamination the renamed default_section exists to prevent, so the id is escaped on write and mapped back on read. Also drops the marker branch in save(), leaving mark_skipped() as the sole writer of the "marker present iff no real keys" invariant, and prunes two test blocks that could not fail. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-09feat: store per-model metadata in models.iniDanilo M.1-0/+98
Cloud models have no presets.ini entry, so context size, vision and prices are recorded per model in a file the app writes. A cancelled dialog leaves a marker so a model tried once never asks again, which is distinct from an absent section meaning never asked. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-09fix: warn when [[providers.local]] discards its fieldsDanilo M.1-0/+42
A malformed local entry was the one skip that said nothing, on the reasoning that local still ends up working. It does, but the merge rebuilds it from the bare base_url, so an api_key, filter or ctx_size set on that entry is dropped without a word. The app then comes up looking healthy, which is the strongest possible signal that nothing is wrong, making this the case that most needs saying, not least. Local surviving at all depends on DEFAULTS supplying base_url, since that is what the rebuild reads. Noted in config.py so the key is not removed as redundant, and pinned by a test. The not-a-table wording now follows what was written: a list gets the single-bracket fix by name, while a scalar entry does not, since telling someone to fix a [[...]] they never typed points at the wrong line. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0176tzAW6H1i2Kz8vm2XXVGV
2026-08-09fix: survive malformed provider config at startupDanilo M.1-0/+92
Eight config shapes, all valid TOML, crashed config.load() before any window existed, so the user got a traceback on a terminal a desktop launch does not have. Skip malformed providers instead, keeping the local path working, and say on stderr which one was dropped and why. The guards are isinstance tests rather than a try/except because the shapes raise three different exception types, and the local entry is checked the same way: [[providers.local]] is a truthy list that raised during the bare base_url merge, before the loop could skip it. A config that cannot load at all now reports the file and exits rather than half-starting on defaults the user never configured. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0176tzAW6H1i2Kz8vm2XXVGV
2026-08-09feat: expose the provider table from configDanilo M.1-0/+42
base_url stays populated so existing callers are untouched; the provider table is built alongside it. models.ini sits beside config.toml and state.ini, following the same pattern as the window layout. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-09docs: stop crediting from None with secret hygieneDanilo M.1-2/+3
A TimeoutExpired chained without from None already prints no secret, since its __str__ shows only the command and the duration. The suppressed display buys a tidier traceback, not a narrower leak, and the comments claimed otherwise. The from None calls stay. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-09fix: sharpen key resolution errors, rename KeyResolutionErrorDanilo M.1-4/+44
Raise KEY_TIMEOUT to 120s: a graphical pinentry plus a hardware token the user has to find and touch makes 30s a plausible successful unlock, and failing one converts a slow success into an error the user can only retry against the same clock. Point the timeout message at a pinentry that never appeared, which is the silent failure, rather than at one the user is already looking at. Name the binary when pass is missing instead of reporting a bare Errno 2. Raise from None so a partial secret on the failed process stdout stays out of printed tracebacks, and record on the cache that it is unlocked only because every caller resolves on the GUI thread today. Rename KeyError_ to KeyResolutionError before later tasks import it. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-09feat: resolve provider API keys from pass, env or literalDanilo M.1-0/+103
One prefix-dispatched field. Resolution is lazy so a local-only session never triggers a pinentry, cached for the process lifetime, and bounded by a timeout so a stuck pinentry surfaces as an error instead of a frozen send. Failures name the provider. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-09fix: skip unusable provider names, name the split ceilingDanilo M.1-1/+18
A provider name containing a colon builds ids that split back to a different provider, and an empty name builds ':model'; both routed to the local router under a nonsense name with no error, so parse() now skips them. split() stays non-injective by design, since bare local ids are the premise here, so the residual risk is recorded as a ponytail comment with its upgrade path rather than hidden. Also pin the trailing-colon form Task 9 relies on, and the empty-needle filter, which were correct but untested. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-09feat: namespace model ids by provider and filter model listsDanilo M.1-0/+67
Cloud models are addressed as provider:model; local ones stay bare so existing sessions keep resolving. An unknown prefix is treated as part of the model name rather than a provider. Filters are case-insensitive substrings because provider ids capitalise inconsistently. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-09fix: stop a partial [providers.local] from deleting the local providerDanilo M.1-0/+20
The synthesis guard tested key membership, so a [providers.local] that set only an api_key claimed the slot, blocked the bare base_url from filling it, then failed the URL check and vanished. Losing the local provider is the one outcome this feature cannot have. The guard now tests the URL and merges, so the table adds detail to the local provider rather than replacing it. A filter given as a bare string was iterated character-wise, turning filter = "qwen" into four needles that match almost every model id. Both failures were silent, which is what made them worth fixing now. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-09feat: parse provider definitions from configDanilo M.1-0/+57
Providers come from a [providers.*] table. A bare top-level base_url synthesizes the local provider so existing configs keep working, and an explicit [providers.local] wins over it. Unset numbers stay None so 'unknown' never collapses into zero. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-01feat: model-written titles, fix empty replies after searchDanilo M.1-5/+156
Titles: once a reply lands in a session still wearing its placeholder title, the exchange goes back to the model in a short side request asking for six words or fewer. Keyed on the placeholder rather than the turn number, so an empty first reply does not forfeit titling for the session and a one-shot window is titled from the question that was answered. Thinking is disabled for that request, or a reasoning model spends the whole budget thinking and returns nothing. Empty replies: the final search round now says so in the tool result. Withdrawing the tool schema is invisible to the model, which asks for another search regardless; the request then surfaces as literal <tool_call> text or vanishes into the thinking block, leaving the reply empty either way. The note rides on the tool result because Qwen3.5's template rejects a trailing system message outright. max_searches now defaults to 2. The other half of that bug was llama-server's quantized KV cache: with cache-type-k/v = q8_0 three of five turns broke, and f16 answered six of six. Measured against a live router, not mocked.
2026-07-31fix: cap searches at one per turnDanilo M.1-0/+6
A second tool call in the same turn, issued after the model has seen the first set of results, does not arrive as a tool_calls delta. It comes back as literal <tool_call><function=web_search> text inside reasoning_content, with finish_reason stop, no tool_calls and empty content. There is no structured call for the loop to act on, so the turn ends and the user is shown a blank reply. Reproduced 3/3 with Qwen3.5-9B through the router; the first call of a turn parses correctly every time, so this is specific to a call that follows a tool result. A retry with the tool withdrawn was tried and rejected: it produced an answer only some of the time and could surface the raw XML in the reply, which is worse than the blank it replaces. Capping at one search avoids reaching the broken state at all. Empty replies drop from 3/3 to about 1/3 and the XML no longer leaks, but the failure is not eliminated. Raise max_searches once the chat template parses follow-up calls. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-31feat: web search via SearXNGDanilo M.1-6/+583
The model is offered a web_search tool and decides when a question needs current information. llamachat runs the query against SearXNG's JSON API, feeds the results back, and the model answers from them. Off unless both search_enabled and search_url are set, so an upgrade never starts talking to the network on its own. The loop lives in backend.stream_chat: a round ending in tool_calls is searched for, the result appended, and the request re-sent. After max_searches rounds the tool is withdrawn, which forces an answer rather than letting an uncertain model search forever. Searches show as a collapsible block above the thinking block, listing each query and its sources. The queries are visible on purpose: when an answer is wrong it is usually the query that was wrong, and without seeing it a bad search and a bad answer look identical. Two failures found while testing against the live stack shaped the design: The model has no clock, so it falls back on its training cutoff and writes that year into the query itself ("latest kernel ... 2025"), poisoning the results before they are fetched. The current date now goes into the system prompt whenever search is on, with an instruction not to date its own queries. A SearXNG whose engines are all rate-limited or CAPTCHA'd returns a valid response with zero results. Reporting that as "no results" tells the model the web is empty and invites a confident answer from stale training data, so a search where every engine failed is now an error naming the engines. Results are attacker-influenced text entering the model's context. Only title, url and the snippet survive, snippets are truncated, result text is escaped on display, result links are never fetched automatically, and a reply forging the search block's URL scheme has it defused as the reasoning scheme already was. None of that stops a poisoned snippet from influencing the answer, which is why the sources stay visible. Adds eight test groups, 23 to 31, all hermetic behind a fake HTTP layer. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-31fix: keep reasoning out of the reply bodyDanilo M.1-0/+51
A delta carrying both reasoning_content and an empty content was classified by truthiness, so the reasoning fell through to the content branch. The tail of the model's thinking was stored and displayed as the answer, splitting mid-sentence across the two fields. Classify on the presence of the field instead, and return nothing for an empty reasoning delta rather than falling through. The visible symptom was a code block swallowing the message: leaked thinking is dense with backticks, and an odd count leaves a run open to the end of the reply. Close an unterminated fence before parsing, with a closer matching the opener's length, and drop a lone dangling inline backtick. Streaming needs this regardless, since a fence is unclosed on nearly every frame. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-31feat: add Ctrl+N and Ctrl+F shortcutsDanilo M.1-0/+108
Ctrl+N starts a fresh conversation and puts the cursor in the input box. Ctrl+F jumps to the search field. The search field sits in the top bar, but its results render in the history panel, so focusing it reveals a hidden panel rather than leaving the search with nowhere to show its hits. Escape now backs out of the search field before it hides the window: first press clears the text, second returns to the input, third hides. Previously a stray Escape while filtering dismissed the whole window. That check reads focusWidget() rather than hasFocus(), so it still holds when the window is not the active one. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-31feat: add a history panel toggleDanilo M.1-0/+76
The history panel now hides, on the ☰ button in the top bar and on Ctrl+\. Its width is captured before hiding and reapplied on show, so toggling does not snap the panel back to a default. Both the width and whether it was hidden persist between runs. Layout state lives in state.ini beside the config rather than at QSettings' default path, which keeps it obvious where it is and lets the checks point it somewhere temporary instead of writing to the real one. The button's checked state, not isVisible(), is the authority for whether the panel is shown: isVisible() is False for every child of a window that has not been mapped yet, so consulting it during startup would disagree with what the user sees. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-31feat: add context meter and file-based system promptsDanilo M.1-0/+175
Two additions to the chat window. A context meter in the top bar shows how much of the active model's window the next request will occupy: system prompt, prior turns, attachments and the current draft. The router does not report a usable context size in router mode (/props returns n_ctx 0), so the limit comes from ctx-size in presets.ini. Streaming requests now ask for usage statistics, which makes the figure exact after each reply at no extra round trip; before that it is an estimate marked with a leading ~. The estimate reads low on models with a reasoning budget, since thinking tokens cannot be known before the reply arrives. The tooltip says so. System prompts are markdown files in the prompts/ directory beside the config, one per file, so they can be edited in an editor and kept in version control. default.md is global, any other file is a named preset that replaces it rather than adding to it, and a conversation may instead carry one-off text or opt out entirely. The choice is stored per session so reopening a chat restores the prompt it was built with. Prompt names are reduced to safe filenames, so a name like ../../etc/passwd cannot write outside the prompts directory, and a name that reduces to nothing is rejected rather than creating a dotfile. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-31feat: initial release of llamachat 0.1.0v0.1.0Danilo M.1-0/+514
A native PySide6 chat client for a local llama.cpp server in router mode. Runs as a single persistent process with a Unix-socket control channel, so a Hyprland keybind toggles the window with a socket round trip rather than a process start. The control commands do not import Qt and answer in under a tenth of a second. Features: one-shot and multi-turn chat modes, runtime model discovery from /v1/models, streaming replies rendered as markdown, collapsible reasoning for models with a thinking budget, drag-and-drop file and image attachment with context-aware truncation, and SQLite history with FTS5 search. Markdown is parsed with MarkdownNoHTML so markup in a reply is displayed rather than interpreted, and links a reply produces cannot drive the interface. Includes the Hyprland Lua snippets for autostart, keybind and window rules, and a self-check suite covering everything except the GUI. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>