diff options
| author | Danilo M. <danix@danix.xyz> | 2026-07-31 19:09:44 +0200 |
|---|---|---|
| committer | Danilo M. <danix@danix.xyz> | 2026-07-31 19:09:44 +0200 |
| commit | 0e35416128e38f158af49935262ecbf8dbf585e6 (patch) | |
| tree | 2af059613853634dc8968ad700e98651bb383514 /docs/superpowers | |
| parent | 283fc9ae0371f4d318a539f021a4b88e05fac190 (diff) | |
| download | llamachat-0e35416128e38f158af49935262ecbf8dbf585e6.tar.gz llamachat-0e35416128e38f158af49935262ecbf8dbf585e6.zip | |
fix: cap searches at one per turn
A second tool call in the same turn, issued after the model has seen the
first set of results, does not arrive as a tool_calls delta. It comes back
as literal <tool_call><function=web_search> text inside reasoning_content,
with finish_reason stop, no tool_calls and empty content. There is no
structured call for the loop to act on, so the turn ends and the user is
shown a blank reply.
Reproduced 3/3 with Qwen3.5-9B through the router; the first call of a turn
parses correctly every time, so this is specific to a call that follows a
tool result.
A retry with the tool withdrawn was tried and rejected: it produced an
answer only some of the time and could surface the raw XML in the reply,
which is worse than the blank it replaces. Capping at one search avoids
reaching the broken state at all. Empty replies drop from 3/3 to about 1/3
and the XML no longer leaks, but the failure is not eliminated.
Raise max_searches once the chat template parses follow-up calls.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Diffstat (limited to 'docs/superpowers')
| -rw-r--r-- | docs/superpowers/specs/2026-07-31-web-search-design.md | 16 |
1 files changed, 16 insertions, 0 deletions
diff --git a/docs/superpowers/specs/2026-07-31-web-search-design.md b/docs/superpowers/specs/2026-07-31-web-search-design.md index 7e14b2d..46ebc91 100644 --- a/docs/superpowers/specs/2026-07-31-web-search-design.md +++ b/docs/superpowers/specs/2026-07-31-web-search-design.md @@ -235,6 +235,22 @@ at the end, outside the suite. - **Citation formatting in replies** — depends on 9B instruction-following; the search block already shows sources +## Implementation note: the cap shipped as 1, not 2 + +The table above chose a cap of two searches per turn. It shipped as one. + +The first tool call of a turn arrives as a well-formed `tool_calls` delta, +exactly as designed. A *second* call, issued after the model has seen the +first set of results, comes back as literal +`<tool_call><function=web_search>` text inside `reasoning_content`, with +`finish_reason: stop`, no `tool_calls`, and empty `content`. There is no +structured call for the loop to act on, so the turn ends with a blank reply. + +Reproduced 3/3 with Qwen3.5-9B through the router. A retry with the tool +withdrawn was tried and rejected: it recovered an answer only some of the +time and could surface the raw XML to the user, which is worse than the +blank it replaces. The cap is 1 until the template parses follow-up calls. + ## Notes `reasoning-budget = 1024` in the active `Qwen3.5-9B` preset is tight for tool |
