| Age | Commit message (Collapse) | Author | Files | Lines |
|
The synthesis guard tested key membership, so a [providers.local] that set
only an api_key claimed the slot, blocked the bare base_url from filling
it, then failed the URL check and vanished. Losing the local provider is
the one outcome this feature cannot have. The guard now tests the URL and
merges, so the table adds detail to the local provider rather than
replacing it.
A filter given as a bare string was iterated character-wise, turning
filter = "qwen" into four needles that match almost every model id. Both
failures were silent, which is what made them worth fixing now.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|
|
Sixteen TDD tasks. Providers, model-id namespacing and key resolution go
in a new providers.py; metadata and cost math in models.py; the settings
dialog in modeldialog.py, keeping ui.py from growing further.
The last task is manual: it needs a real key and spends real money, and
it is where cloud tool-calling behaviour gets found out rather than
guessed at.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|
|
The name is behaviour, not config: it decides which provider a bare
base_url synthesizes, which one renders without a prefix, which one
skips filtering, and which one is exempt from the model dialog. All
four live in code, so the name had to be picked before implementation.
"default" was doing double duty as "the fallback" and "the local one",
and only the second is true.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|
|
Design for talking to OpenAI-compatible cloud providers alongside the
local router. Providers are config table entries, local becomes the
"default" one, and a bare base_url still synthesizes it so existing
configs keep working.
Cloud models carry no presets.ini entry and cost money, so missing
metadata (context size, vision, prices) is entered through a per-model
dialog stored in models.ini, and an approximate per-conversation cost
sits next to the context meter. Token counts and the producing model
move onto the messages table so cost survives reopening a chat and
prices each reply at whatever produced it.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|
|
The comments and README stated flatly that a second tool call always comes
back as literal <tool_call> XML. On llama.cpp b10208 that is no longer true:
forcing a follow-up by starving the first search gave 1/5 parsed as a
structured call, 1/5 leaking XML, 3/5 answered without searching again.
Earlier builds never parsed it.
Still not reliable enough to raise max_searches, since the failure mode is
a blank reply, but the wording should not read as permanent when it was
measured once against one build.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|
|
A second tool call in the same turn, issued after the model has seen the
first set of results, does not arrive as a tool_calls delta. It comes back
as literal <tool_call><function=web_search> text inside reasoning_content,
with finish_reason stop, no tool_calls and empty content. There is no
structured call for the loop to act on, so the turn ends and the user is
shown a blank reply.
Reproduced 3/3 with Qwen3.5-9B through the router; the first call of a turn
parses correctly every time, so this is specific to a call that follows a
tool result.
A retry with the tool withdrawn was tried and rejected: it produced an
answer only some of the time and could surface the raw XML in the reply,
which is worse than the blank it replaces. Capping at one search avoids
reaching the broken state at all. Empty replies drop from 3/3 to about 1/3
and the XML no longer leaks, but the failure is not eliminated.
Raise max_searches once the chat template parses follow-up calls.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|
|
The model is offered a web_search tool and decides when a question needs
current information. llamachat runs the query against SearXNG's JSON API,
feeds the results back, and the model answers from them. Off unless both
search_enabled and search_url are set, so an upgrade never starts talking
to the network on its own.
The loop lives in backend.stream_chat: a round ending in tool_calls is
searched for, the result appended, and the request re-sent. After
max_searches rounds the tool is withdrawn, which forces an answer rather
than letting an uncertain model search forever.
Searches show as a collapsible block above the thinking block, listing
each query and its sources. The queries are visible on purpose: when an
answer is wrong it is usually the query that was wrong, and without
seeing it a bad search and a bad answer look identical.
Two failures found while testing against the live stack shaped the
design:
The model has no clock, so it falls back on its training cutoff and
writes that year into the query itself ("latest kernel ... 2025"),
poisoning the results before they are fetched. The current date now goes
into the system prompt whenever search is on, with an instruction not to
date its own queries.
A SearXNG whose engines are all rate-limited or CAPTCHA'd returns a
valid response with zero results. Reporting that as "no results" tells
the model the web is empty and invites a confident answer from stale
training data, so a search where every engine failed is now an error
naming the engines.
Results are attacker-influenced text entering the model's context. Only
title, url and the snippet survive, snippets are truncated, result text
is escaped on display, result links are never fetched automatically, and
a reply forging the search block's URL scheme has it defused as the
reasoning scheme already was. None of that stops a poisoned snippet from
influencing the answer, which is why the sources stay visible.
Adds eight test groups, 23 to 31, all hermetic behind a fake HTTP layer.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|
|
Model-decided tool calling against a SearXNG instance, with the loop
inside stream_chat, results shown as a collapsible block, and a cap of
two searches per turn.
Verified against the live stack before writing: Qwen3.5-9B through the
router emits well-formed tool_calls and answers correctly from a
role: "tool" result. No llama-server flags are involved; the shipped
--tools option runs server-side tools and none of them search the web.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|