aboutsummaryrefslogtreecommitdiffstats
path: root/docs
diff options
context:
space:
mode:
authorDanilo M. <danix@danix.xyz>2026-08-09 12:15:48 +0200
committerDanilo M. <danix@danix.xyz>2026-08-09 12:15:48 +0200
commit573c04262e814ac1d72c56487ac95300758af8e8 (patch)
tree0ba7c151d5fdd27a9e4bb2abef6416ae860459d1 /docs
parent4df1523aae7f99b36d9d3696e4406e58ce095fe4 (diff)
downloadllamachat-573c04262e814ac1d72c56487ac95300758af8e8.tar.gz
llamachat-573c04262e814ac1d72c56487ac95300758af8e8.zip
docs: implementation plan for external cloud providers
Sixteen TDD tasks. Providers, model-id namespacing and key resolution go in a new providers.py; metadata and cost math in models.py; the settings dialog in modeldialog.py, keeping ui.py from growing further. The last task is manual: it needs a real key and spends real money, and it is where cloud tool-calling behaviour gets found out rather than guessed at. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Diffstat (limited to 'docs')
-rw-r--r--docs/superpowers/plans/2026-08-09-external-providers.md2521
1 files changed, 2521 insertions, 0 deletions
diff --git a/docs/superpowers/plans/2026-08-09-external-providers.md b/docs/superpowers/plans/2026-08-09-external-providers.md
new file mode 100644
index 0000000..a17c1d9
--- /dev/null
+++ b/docs/superpowers/plans/2026-08-09-external-providers.md
@@ -0,0 +1,2521 @@
+# External Cloud Providers Implementation Plan
+
+> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking.
+
+**Goal:** Let llamachat use OpenAI-compatible cloud providers (together.ai, siliconflow) alongside the local llama.cpp router, with per-model metadata entered through a dialog and an approximate per-conversation cost shown beside the context meter.
+
+**Architecture:** Providers become entries in a `[providers.*]` config table; the local router is one of them, named `local`, and a bare top-level `base_url` synthesizes it so existing configs keep working. Model ids are `provider:model`, except `local` which stays bare. Cloud models have no `presets.ini` entry, so context size, vision and prices come from a `models.ini` the app writes, prefilled from per-provider defaults. Token counts and the producing model move onto the `messages` table so cost survives reopening a conversation.
+
+**Tech Stack:** Python 3.11+ standard library (`tomllib`, `configparser`, `subprocess`, `sqlite3`), `httpx` for HTTP, PySide6 for the dialog. No new dependencies.
+
+**Spec:** `docs/superpowers/specs/2026-08-09-external-providers-design.md`
+
+---
+
+## Conventions for this codebase
+
+Read this before Task 1. It is not optional context.
+
+**Tests are not pytest.** `test_llamachat.py` is a single executable file of plain
+functions. Each test ends with `print("ok <short description>")`. Every test must
+be registered by name in the `if __name__ == "__main__":` block at the bottom of the
+file, in the order it should run. A test that is written but not registered never
+runs, and the suite will still say "all checks passed".
+
+Run the whole suite with:
+
+```bash
+./test_llamachat.py
+```
+
+There is no way to run a single test from the command line. To verify one test
+fails or passes in isolation, run the suite and read that test's line. Adding a
+temporary `if __name__` entry for just the new test is acceptable during a
+red/green cycle but must be restored before committing.
+
+**Every new source file needs the GPL header.** Copy it verbatim from the top of
+`llamachat/config.py`, changing nothing but the docstring on the last line.
+
+**Commits are GPG-signed automatically.** Do not pass `-c commit.gpgsign=false`.
+Two git hooks scan for personal data and secrets; treat a rejection as correct.
+Use `example.org`, fake keys like `sk-test-not-a-real-key`, and generic paths in
+tests and fixtures. Never put a real API key anywhere, including a test.
+
+**Style:** comments explain why, not what. Deliberate simplifications get a
+`# ponytail:` comment naming the ceiling and the upgrade path. Match the
+surrounding code's density.
+
+---
+
+## File Structure
+
+**New files:**
+
+| File | Responsibility |
+| --- | --- |
+| `llamachat/providers.py` | Parsing `[providers.*]`, splitting/joining model ids, filtering model lists, resolving API keys (`pass:`/`env:`/literal) with caching. No Qt, no HTTP. |
+| `llamachat/models.py` | `models.ini` read/write, three-layer metadata resolution, cost arithmetic and formatting. No Qt, no HTTP. |
+| `llamachat/modeldialog.py` | The Qt dialog for entering ctx size, vision and prices. Only file in this feature that imports PySide6. |
+
+**Modified files:**
+
+| File | Change |
+| --- | --- |
+| `llamachat/config.py` | Parse `[providers.*]`, synthesize `local` from bare `base_url`, add `models_path`. |
+| `llamachat/backend.py` | `Client` gains `api_key` and sends `Authorization`. New `MultiClient` fans `models()` out across providers and routes requests by model id. |
+| `llamachat/db.py` | Three nullable columns on `messages`, migration, and storing them. |
+| `llamachat/ui.py` | Model list from `MultiClient`, metadata via `models.py` instead of `presets` alone, dialog triggers, cost label. |
+| `test_llamachat.py` | New tests, each registered in the `__main__` block. |
+
+**Dependency direction:** `providers.py` and `models.py` depend on nothing in the
+project. `backend.py` imports `providers`. `ui.py` imports all three plus
+`modeldialog`. Nothing imports `ui`.
+
+---
+
+## Task 1: Provider config parsing
+
+**Files:**
+- Create: `llamachat/providers.py`
+- Modify: `test_llamachat.py`
+
+- [ ] **Step 1: Write the failing test**
+
+Add to `test_llamachat.py`, after `test_config_defaults()`:
+
+```python
+def test_provider_parsing():
+ """Providers come from [providers.*]; a bare base_url synthesizes local."""
+ from llamachat import providers
+
+ # A modern config with two providers.
+ parsed = providers.parse(
+ {
+ "providers": {
+ "local": {"base_url": "http://localhost:8181/"},
+ "together": {
+ "base_url": "https://api.example.org",
+ "api_key": "env:TEST_KEY_NAME",
+ "filter": ["qwen", "deepseek"],
+ "ctx_size": 32768,
+ "price_in": 0.6,
+ "price_out": 0.9,
+ },
+ }
+ }
+ )
+ assert set(parsed) == {"local", "together"}
+ # Trailing slashes are stripped so URL joining stays predictable.
+ assert parsed["local"].base_url == "http://localhost:8181"
+ assert parsed["local"].api_key == ""
+ assert parsed["together"].filter == ["qwen", "deepseek"]
+ assert parsed["together"].ctx_size == 32768
+ assert parsed["together"].price_in == 0.6
+ assert parsed["together"].price_out == 0.9
+
+ # An old config: bare base_url, no providers table at all.
+ legacy = providers.parse({"base_url": "http://localhost:8181"})
+ assert set(legacy) == {"local"}
+ assert legacy["local"].base_url == "http://localhost:8181"
+
+ # Both present: the explicit entry wins over the bare key.
+ both = providers.parse(
+ {
+ "base_url": "http://ignored.example.org",
+ "providers": {"local": {"base_url": "http://explicit.example.org"}},
+ }
+ )
+ assert both["local"].base_url == "http://explicit.example.org"
+
+ # A provider with no base_url is skipped rather than half-configured.
+ broken = providers.parse(
+ {"providers": {"local": {"base_url": "http://x.example.org"},
+ "bad": {"api_key": "literal"}}}
+ )
+ assert set(broken) == {"local"}
+
+ # Unset numbers stay None so "unknown" is distinguishable from zero.
+ assert parsed["local"].ctx_size is None
+ assert parsed["local"].price_in is None
+ print("ok provider config parsing")
+```
+
+Register it in the `__main__` block immediately after `test_config_defaults()`:
+
+```python
+ test_config_defaults()
+ test_provider_parsing()
+```
+
+- [ ] **Step 2: Run the suite to verify it fails**
+
+Run: `./test_llamachat.py`
+Expected: FAIL with `ModuleNotFoundError: No module named 'llamachat.providers'`
+
+- [ ] **Step 3: Write the implementation**
+
+Create `llamachat/providers.py`. Copy the 14-line GPL header verbatim from the top
+of `llamachat/config.py`, then:
+
+```python
+"""Provider definitions, model-id namespacing and API key resolution."""
+
+import os
+import subprocess
+from dataclasses import dataclass, field
+
+# The local llama.cpp router. Its models are shown and stored without a
+# prefix, so an existing session pointing at a local model still resolves.
+LOCAL = "local"
+
+# A stuck pinentry must not freeze the worker thread forever.
+KEY_TIMEOUT = 30
+
+
+class KeyError_(Exception):
+ """An api_key that could not be resolved, phrased for the user."""
+
+
+@dataclass
+class Provider:
+ name: str
+ base_url: str
+ api_key: str = ""
+ filter: list[str] = field(default_factory=list)
+ # None rather than 0: unset must stay distinguishable from "zero".
+ ctx_size: int | None = None
+ vision: bool | None = None
+ price_in: float | None = None
+ price_out: float | None = None
+
+ @property
+ def is_local(self) -> bool:
+ return self.name == LOCAL
+
+
+def _number(raw, cast):
+ if raw is None or raw == "":
+ return None
+ try:
+ return cast(raw)
+ except (TypeError, ValueError):
+ return None
+
+
+def parse(values: dict) -> dict[str, Provider]:
+ """Build the provider table from already-loaded config values.
+
+ Takes the raw dict rather than a path so config.py owns file reading and
+ this stays testable without touching disk.
+ """
+ table = dict(values.get("providers") or {})
+
+ # An old config has only a bare base_url. Synthesize the local provider
+ # from it so nothing needs migrating, but never override an explicit one.
+ bare = values.get("base_url")
+ if bare and LOCAL not in table:
+ table[LOCAL] = {"base_url": bare}
+
+ out: dict[str, Provider] = {}
+ for name, entry in table.items():
+ entry = entry or {}
+ base_url = str(entry.get("base_url") or "").rstrip("/")
+ if not base_url:
+ # ponytail: a provider with no URL is misconfigured, not a
+ # partial one. Skipping beats inventing a default endpoint.
+ continue
+ vision = entry.get("vision")
+ out[str(name)] = Provider(
+ name=str(name),
+ base_url=base_url,
+ api_key=str(entry.get("api_key") or ""),
+ filter=[str(f) for f in (entry.get("filter") or [])],
+ ctx_size=_number(entry.get("ctx_size"), int),
+ vision=None if vision is None else bool(vision),
+ price_in=_number(entry.get("price_in"), float),
+ price_out=_number(entry.get("price_out"), float),
+ )
+ return out
+```
+
+- [ ] **Step 4: Run the suite to verify it passes**
+
+Run: `./test_llamachat.py`
+Expected: PASS, including the line `ok provider config parsing`
+
+- [ ] **Step 5: Commit**
+
+```bash
+git add llamachat/providers.py test_llamachat.py
+git commit -m "feat: parse provider definitions from config
+
+Providers come from a [providers.*] table. A bare top-level base_url
+synthesizes the local provider so existing configs keep working, and an
+explicit [providers.local] wins over it. Unset numbers stay None so
+'unknown' never collapses into zero.
+
+Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>"
+```
+
+---
+
+## Task 2: Model id namespacing and filtering
+
+**Files:**
+- Modify: `llamachat/providers.py`
+- Modify: `test_llamachat.py`
+
+- [ ] **Step 1: Write the failing test**
+
+Add after `test_provider_parsing()`:
+
+```python
+def test_model_ids_and_filtering():
+ """Ids are provider:model, local stays bare, filters are substrings."""
+ from llamachat import providers
+
+ table = providers.parse(
+ {
+ "providers": {
+ "local": {"base_url": "http://localhost:8181"},
+ "together": {
+ "base_url": "https://api.example.org",
+ "filter": ["qwen", "deepseek"],
+ },
+ "unfiltered": {"base_url": "https://api2.example.org"},
+ }
+ }
+ )
+
+ # Local models carry no prefix, in the dropdown and in the database.
+ assert providers.qualify("local", "gemma4") == "gemma4"
+ assert providers.qualify("together", "Qwen/Qwen2.5") == "together:Qwen/Qwen2.5"
+
+ # Splitting is the inverse, and only for providers that exist.
+ assert providers.split("gemma4", table) == ("local", "gemma4")
+ assert providers.split("together:Qwen/Qwen2.5", table) == (
+ "together",
+ "Qwen/Qwen2.5",
+ )
+ # An unknown prefix is part of the model name, not a provider. This is
+ # what keeps a local model whose name contains a colon working.
+ assert providers.split("weird:name", table) == ("local", "weird:name")
+ # Only the first colon splits.
+ assert providers.split("together:a:b", table) == ("together", "a:b")
+
+ # Filtering is case-insensitive substring, any match wins.
+ listed = [
+ "Qwen/Qwen2.5-72B-Instruct-Turbo",
+ "deepseek-ai/DeepSeek-V3",
+ "meta-llama/Llama-3.3-70B",
+ ]
+ kept = providers.apply_filter(table["together"], listed)
+ assert kept == [
+ "Qwen/Qwen2.5-72B-Instruct-Turbo",
+ "deepseek-ai/DeepSeek-V3",
+ ]
+
+ # No filter means everything.
+ assert providers.apply_filter(table["unfiltered"], listed) == listed
+ # The local provider is never filtered even if one is configured.
+ table["local"].filter = ["nothing-matches-this"]
+ assert providers.apply_filter(table["local"], listed) == listed
+ # A filter matching nothing yields nothing, it does not fall back to all.
+ table["together"].filter = ["zzz"]
+ assert providers.apply_filter(table["together"], listed) == []
+ print("ok model ids and filtering")
+```
+
+Register it after `test_provider_parsing()` in the `__main__` block.
+
+- [ ] **Step 2: Run the suite to verify it fails**
+
+Run: `./test_llamachat.py`
+Expected: FAIL with `AttributeError: module 'llamachat.providers' has no attribute 'qualify'`
+
+- [ ] **Step 3: Write the implementation**
+
+Append to `llamachat/providers.py`:
+
+```python
+def qualify(provider: str, model: str) -> str:
+ """The stored, displayed id for one model. Local models stay bare."""
+ if provider == LOCAL:
+ return model
+ return f"{provider}:{model}"
+
+
+def split(model_id: str, table: dict[str, Provider]) -> tuple[str, str]:
+ """Inverse of qualify, resolved against the configured providers.
+
+ A prefix that is not a configured provider is treated as part of the
+ model name, which keeps a bare local model containing a colon working.
+ """
+ prefix, sep, rest = model_id.partition(":")
+ if sep and prefix in table and prefix != LOCAL:
+ return prefix, rest
+ return LOCAL, model_id
+
+
+def apply_filter(provider: Provider, listed: list[str]) -> list[str]:
+ """Keep models matching any of the provider's substrings.
+
+ Case-insensitive, because provider ids capitalise inconsistently:
+ "qwen" has to match "Qwen/Qwen2.5-72B-Instruct-Turbo".
+ """
+ if provider.is_local or not provider.filter:
+ return list(listed)
+ needles = [f.lower() for f in provider.filter if f]
+ if not needles:
+ return list(listed)
+ return [m for m in listed if any(n in m.lower() for n in needles)]
+```
+
+- [ ] **Step 4: Run the suite to verify it passes**
+
+Run: `./test_llamachat.py`
+Expected: PASS, including `ok model ids and filtering`
+
+- [ ] **Step 5: Commit**
+
+```bash
+git add llamachat/providers.py test_llamachat.py
+git commit -m "feat: namespace model ids by provider and filter model lists
+
+Cloud models are addressed as provider:model; local ones stay bare so
+existing sessions keep resolving. An unknown prefix is treated as part of
+the model name rather than a provider. Filters are case-insensitive
+substrings because provider ids capitalise inconsistently.
+
+Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>"
+```
+
+---
+
+## Task 3: API key resolution
+
+**Files:**
+- Modify: `llamachat/providers.py`
+- Modify: `test_llamachat.py`
+
+- [ ] **Step 1: Write the failing test**
+
+Add after `test_model_ids_and_filtering()`:
+
+```python
+def test_key_resolution():
+ """api_key is prefix-dispatched, resolved lazily and cached."""
+ from llamachat import providers
+
+ resolver = providers.KeyResolver()
+
+ # No key configured: no Authorization header, and nothing is run.
+ empty = providers.Provider(name="local", base_url="http://x.example.org")
+ assert resolver.resolve(empty) == ""
+
+ # A literal key is used as-is.
+ literal = providers.Provider(
+ name="p", base_url="http://x.example.org", api_key="sk-test-not-a-real-key"
+ )
+ assert resolver.resolve(literal) == "sk-test-not-a-real-key"
+
+ # env: reads the environment.
+ os.environ["LLAMACHAT_TEST_KEY"] = "from-env"
+ env = providers.Provider(
+ name="e", base_url="http://x.example.org",
+ api_key="env:LLAMACHAT_TEST_KEY",
+ )
+ assert resolver.resolve(env) == "from-env"
+ del os.environ["LLAMACHAT_TEST_KEY"]
+
+ # A missing env var is an error naming the provider, not a silent "".
+ missing = providers.Provider(
+ name="gone", base_url="http://x.example.org",
+ api_key="env:LLAMACHAT_ABSENT_VAR",
+ )
+ try:
+ resolver.resolve(missing)
+ assert False, "a missing env var must raise"
+ except providers.KeyError_ as exc:
+ assert "gone" in str(exc)
+
+ # pass: shells out. Substitute the runner rather than requiring gpg.
+ calls = []
+
+ def fake_run(cmd, timeout):
+ calls.append((cmd, timeout))
+ return "line-one\nline-two\n"
+
+ passed = providers.Provider(
+ name="together", base_url="http://x.example.org",
+ api_key="pass:api/together",
+ )
+ cached = providers.KeyResolver(runner=fake_run)
+ assert cached.resolve(passed) == "line-one" # first line only
+ assert calls[0][0] == ["pass", "show", "api/together"]
+ assert calls[0][1] == providers.KEY_TIMEOUT
+
+ # Cached: a second resolve must not shell out again.
+ assert cached.resolve(passed) == "line-one"
+ assert len(calls) == 1
+
+ # A failing pass is reported, naming the provider.
+ def boom(cmd, timeout):
+ raise OSError("pass: entry not found")
+
+ try:
+ providers.KeyResolver(runner=boom).resolve(passed)
+ assert False, "a failing pass must raise"
+ except providers.KeyError_ as exc:
+ assert "together" in str(exc)
+
+ # Empty output is a failure too: an empty key would 401 confusingly.
+ try:
+ providers.KeyResolver(runner=lambda cmd, timeout: " \n").resolve(passed)
+ assert False, "empty pass output must raise"
+ except providers.KeyError_:
+ pass
+ print("ok api key resolution")
+```
+
+Register it after `test_model_ids_and_filtering()`.
+
+- [ ] **Step 2: Run the suite to verify it fails**
+
+Run: `./test_llamachat.py`
+Expected: FAIL with `AttributeError: module 'llamachat.providers' has no attribute 'KeyResolver'`
+
+- [ ] **Step 3: Write the implementation**
+
+Append to `llamachat/providers.py`:
+
+```python
+def _run_pass(cmd: list[str], timeout: int) -> str:
+ """Run `pass show NAME` and return its stdout."""
+ result = subprocess.run(
+ cmd, capture_output=True, text=True, timeout=timeout, check=True
+ )
+ return result.stdout
+
+
+class KeyResolver:
+ """Resolves api_key fields, lazily and once per process.
+
+ Lazy matters: `pass` needs the GPG key, so a session that only touches
+ local models must never trigger a pinentry. Resolved values stay in
+ memory and are never written anywhere.
+ """
+
+ def __init__(self, runner=_run_pass):
+ self._runner = runner
+ self._cache: dict[str, str] = {}
+
+ def resolve(self, provider: Provider) -> str:
+ """The bearer token for this provider, or '' when it needs none."""
+ spec = provider.api_key
+ if not spec:
+ return ""
+ if provider.name in self._cache:
+ return self._cache[provider.name]
+
+ if spec.startswith("env:"):
+ value = os.environ.get(spec[4:], "")
+ if not value:
+ raise KeyError_(
+ f"{provider.name}: environment variable {spec[4:]} is not set"
+ )
+ elif spec.startswith("pass:"):
+ value = self._from_pass(provider, spec[5:])
+ else:
+ value = spec
+
+ self._cache[provider.name] = value
+ return value
+
+ def _from_pass(self, provider: Provider, entry: str) -> str:
+ try:
+ out = self._runner(["pass", "show", entry], timeout=KEY_TIMEOUT)
+ except subprocess.TimeoutExpired:
+ raise KeyError_(
+ f"{provider.name}: `pass show {entry}` timed out after "
+ f"{KEY_TIMEOUT}s. Is a pinentry waiting for input?"
+ )
+ except Exception as exc:
+ raise KeyError_(f"{provider.name}: `pass show {entry}` failed: {exc}")
+ # A password store entry keeps the secret on the first line and
+ # metadata below it.
+ first = (out or "").strip().splitlines()
+ if not first or not first[0].strip():
+ raise KeyError_(f"{provider.name}: `pass show {entry}` returned nothing")
+ return first[0].strip()
+```
+
+- [ ] **Step 4: Run the suite to verify it passes**
+
+Run: `./test_llamachat.py`
+Expected: PASS, including `ok api key resolution`
+
+- [ ] **Step 5: Commit**
+
+```bash
+git add llamachat/providers.py test_llamachat.py
+git commit -m "feat: resolve provider API keys from pass, env or literal
+
+One prefix-dispatched field. Resolution is lazy so a local-only session
+never triggers a pinentry, cached for the process lifetime, and bounded
+by a timeout so a stuck pinentry surfaces as an error instead of a frozen
+send. Failures name the provider.
+
+Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>"
+```
+
+---
+
+## Task 4: Wire providers into config
+
+**Files:**
+- Modify: `llamachat/config.py:24-56` (DEFAULTS), `llamachat/config.py:76-95` (Config), `llamachat/config.py:110-152` (load)
+- Modify: `test_llamachat.py`
+
+- [ ] **Step 1: Write the failing test**
+
+Add after `test_key_resolution()`:
+
+```python
+def test_config_providers():
+ """config.load exposes the provider table and the models.ini path."""
+ with tempfile.TemporaryDirectory() as tmp:
+ path = Path(tmp) / "config.toml"
+
+ # A legacy config: bare base_url only.
+ path.write_text('base_url = "http://localhost:9999"\n')
+ cfg = config.load(path)
+ assert set(cfg.providers) == {"local"}
+ assert cfg.providers["local"].base_url == "http://localhost:9999"
+ # base_url stays populated: existing code still reads it.
+ assert cfg.base_url == "http://localhost:9999"
+ assert cfg.models_path == path.parent / "models.ini"
+
+ # A config with an explicit cloud provider.
+ path.write_text(
+ 'base_url = "http://localhost:9999"\n'
+ "\n"
+ "[providers.together]\n"
+ 'base_url = "https://api.example.org"\n'
+ 'api_key = "pass:api/together"\n'
+ 'filter = ["qwen"]\n'
+ "ctx_size = 32768\n"
+ "price_in = 0.6\n"
+ "price_out = 0.9\n"
+ )
+ cfg = config.load(path)
+ assert set(cfg.providers) == {"local", "together"}
+ assert cfg.providers["together"].api_key == "pass:api/together"
+ assert cfg.providers["together"].filter == ["qwen"]
+ assert cfg.providers["together"].price_out == 0.9
+
+ # A config with no base_url and no providers still loads, with the
+ # built-in default synthesizing local.
+ path.write_text("request_timeout = 60\n")
+ cfg = config.load(path)
+ assert set(cfg.providers) == {"local"}
+ assert cfg.providers["local"].base_url == config.DEFAULTS["base_url"]
+ print("ok config provider table")
+```
+
+Register it after `test_key_resolution()`.
+
+- [ ] **Step 2: Run the suite to verify it fails**
+
+Run: `./test_llamachat.py`
+Expected: FAIL with `AttributeError: 'Config' object has no attribute 'providers'`
+
+- [ ] **Step 3: Write the implementation**
+
+In `llamachat/config.py`, add the import beside the existing ones:
+
+```python
+from . import providers as providers_mod
+```
+
+Add two fields to the `Config` dataclass, after `max_searches`:
+
+```python
+ max_searches: int
+ providers: dict
+ models_path: Path
+```
+
+In `load()`, after the `search_url`/`search_enabled` lines and before the
+`return Config(`, add:
+
+```python
+ # Providers are built from the raw values so a bare base_url still
+ # synthesizes the local entry. DEFAULTS supplies base_url when the file
+ # names neither, which keeps a config with no network settings working.
+ provider_table = providers_mod.parse(values)
+```
+
+Then add the two arguments to the `return Config(...)` call, after
+`max_searches=int(values["max_searches"]),`:
+
+```python
+ providers=provider_table,
+ models_path=path.parent / "models.ini",
+```
+
+- [ ] **Step 4: Run the suite to verify it passes**
+
+Run: `./test_llamachat.py`
+Expected: PASS, including `ok config provider table`
+
+- [ ] **Step 5: Commit**
+
+```bash
+git add llamachat/config.py test_llamachat.py
+git commit -m "feat: expose the provider table from config
+
+base_url stays populated so existing callers are untouched; the provider
+table is built alongside it. models.ini sits beside config.toml and
+state.ini, following the same pattern as the window layout.
+
+Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>"
+```
+
+---
+
+## Task 5: models.ini storage
+
+**Files:**
+- Create: `llamachat/models.py`
+- Modify: `test_llamachat.py`
+
+- [ ] **Step 1: Write the failing test**
+
+Add after `test_config_providers()`:
+
+```python
+def test_models_store():
+ """models.ini round-trips per-model metadata and the cancel record."""
+ from llamachat import models
+
+ with tempfile.TemporaryDirectory() as tmp:
+ path = Path(tmp) / "models.ini"
+ store = models.ModelStore(path)
+
+ # Nothing recorded yet.
+ assert store.get("together:Qwen/Qwen2.5") is None
+ assert store.was_offered("together:Qwen/Qwen2.5") is False
+
+ store.save(
+ "together:Qwen/Qwen2.5",
+ models.ModelInfo(
+ ctx_size=32768, vision=False, price_in=1.2, price_out=1.2
+ ),
+ )
+ # A cancelled dialog records that it was offered, nothing more.
+ store.mark_skipped("together:Llama-Vision-Free")
+
+ # Reread from disk, not from memory: this is the round trip.
+ fresh = models.ModelStore(path)
+ info = fresh.get("together:Qwen/Qwen2.5")
+ assert info.ctx_size == 32768
+ assert info.vision is False
+ assert info.price_in == 1.2
+ assert info.price_out == 1.2
+ assert fresh.was_offered("together:Qwen/Qwen2.5") is True
+
+ assert fresh.get("together:Llama-Vision-Free") is None
+ assert fresh.was_offered("together:Llama-Vision-Free") is True
+
+ # Partial entries are legal: prices may be left blank.
+ fresh.save("together:cheap", models.ModelInfo(ctx_size=8192))
+ again = models.ModelStore(path)
+ partial = again.get("together:cheap")
+ assert partial.ctx_size == 8192
+ assert partial.price_in is None
+ assert partial.vision is None
+
+ # A model id with a colon must survive being an ini section name.
+ again.save("together:org/name:v2", models.ModelInfo(ctx_size=4096))
+ assert models.ModelStore(path).get("together:org/name:v2").ctx_size == 4096
+ print("ok models.ini storage")
+```
+
+Register it after `test_config_providers()`.
+
+- [ ] **Step 2: Run the suite to verify it fails**
+
+Run: `./test_llamachat.py`
+Expected: FAIL with `ModuleNotFoundError: No module named 'llamachat.models'`
+
+- [ ] **Step 3: Write the implementation**
+
+Create `llamachat/models.py` with the GPL header copied from `config.py`, then:
+
+```python
+"""Per-model metadata: what presets.ini cannot answer for a cloud model."""
+
+import configparser
+from dataclasses import dataclass
+from pathlib import Path
+
+# Recorded when the dialog is cancelled, so a model tried once never nags.
+SKIPPED = "configured"
+
+
+@dataclass
+class ModelInfo:
+ """What a model can tell us. Every field may be unknown."""
+ ctx_size: int | None = None
+ vision: bool | None = None
+ price_in: float | None = None
+ price_out: float | None = None
+
+ def is_empty(self) -> bool:
+ return all(
+ v is None
+ for v in (self.ctx_size, self.vision, self.price_in, self.price_out)
+ )
+
+
+def _get(section, key, cast):
+ raw = section.get(key, "").strip()
+ if not raw:
+ return None
+ try:
+ return cast(raw)
+ except ValueError:
+ return None
+
+
+def _get_bool(section, key):
+ raw = section.get(key, "").strip().lower()
+ if raw in ("true", "yes", "1", "on"):
+ return True
+ if raw in ("false", "no", "0", "off"):
+ return False
+ return None
+
+
+class ModelStore:
+ """models.ini, keyed by full model id.
+
+ configparser rather than TOML because this file is written by the app,
+ and tomllib is read-only in the standard library.
+ """
+
+ def __init__(self, path: Path):
+ self.path = path
+ # Model ids contain colons and slashes, so no key/value delimiter
+ # may be inferred from a section name. Sections are safe as-is.
+ self.parser = configparser.ConfigParser(interpolation=None)
+ if path.exists():
+ try:
+ self.parser.read(path, encoding="utf-8")
+ except configparser.Error:
+ # ponytail: a corrupt file reads as empty; the dialog can
+ # rewrite it. Failing to start over metadata is worse.
+ self.parser = configparser.ConfigParser(interpolation=None)
+
+ def get(self, model_id: str) -> ModelInfo | None:
+ """Stored metadata, or None when there is none worth having."""
+ if not self.parser.has_section(model_id):
+ return None
+ section = self.parser[model_id]
+ info = ModelInfo(
+ ctx_size=_get(section, "ctx_size", int),
+ vision=_get_bool(section, "vision"),
+ price_in=_get(section, "price_in", float),
+ price_out=_get(section, "price_out", float),
+ )
+ return None if info.is_empty() else info
+
+ def was_offered(self, model_id: str) -> bool:
+ """Whether the dialog has already been shown for this model."""
+ return self.parser.has_section(model_id)
+
+ def save(self, model_id: str, info: ModelInfo) -> None:
+ section = self._section(model_id)
+ for key, value in (
+ ("ctx_size", info.ctx_size),
+ ("vision", info.vision),
+ ("price_in", info.price_in),
+ ("price_out", info.price_out),
+ ):
+ if value is None:
+ section.pop(key, None)
+ elif isinstance(value, bool):
+ section[key] = "true" if value else "false"
+ else:
+ section[key] = str(value)
+ section.pop(SKIPPED, None)
+ self._write()
+
+ def mark_skipped(self, model_id: str) -> None:
+ """Record a cancelled dialog: offered, declined, do not ask again."""
+ section = self._section(model_id)
+ if not any(k != SKIPPED for k in section):
+ section[SKIPPED] = "false"
+ self._write()
+
+ def _section(self, model_id: str):
+ if not self.parser.has_section(model_id):
+ self.parser.add_section(model_id)
+ return self.parser[model_id]
+
+ def _write(self) -> None:
+ self.path.parent.mkdir(parents=True, exist_ok=True)
+ with open(self.path, "w", encoding="utf-8") as fh:
+ self.parser.write(fh)
+```
+
+- [ ] **Step 4: Run the suite to verify it passes**
+
+Run: `./test_llamachat.py`
+Expected: PASS, including `ok models.ini storage`
+
+- [ ] **Step 5: Commit**
+
+```bash
+git add llamachat/models.py test_llamachat.py
+git commit -m "feat: store per-model metadata in models.ini
+
+Cloud models have no presets.ini entry, so context size, vision and
+prices are recorded per model in a file the app writes. A cancelled
+dialog leaves a marker so a model tried once never asks again, which is
+distinct from an absent section meaning never asked.
+
+Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>"
+```
+
+---
+
+## Task 6: Metadata resolution and cost arithmetic
+
+**Files:**
+- Modify: `llamachat/models.py`
+- Modify: `test_llamachat.py`
+
+- [ ] **Step 1: Write the failing test**
+
+Add after `test_models_store()`:
+
+```python
+def test_metadata_and_cost():
+ """models.ini beats provider defaults beats unknown; cost sums per model."""
+ from llamachat import models, providers
+
+ table = providers.parse(
+ {
+ "providers": {
+ "local": {"base_url": "http://localhost:8181"},
+ "together": {
+ "base_url": "https://api.example.org",
+ "api_key": "env:X",
+ "ctx_size": 32768,
+ "price_in": 0.6,
+ "price_out": 0.9,
+ },
+ "free": {"base_url": "https://api3.example.org"},
+ }
+ }
+ )
+
+ with tempfile.TemporaryDirectory() as tmp:
+ store = models.ModelStore(Path(tmp) / "models.ini")
+ store.save(
+ "together:specific",
+ models.ModelInfo(ctx_size=8192, vision=True, price_in=5.0),
+ )
+
+ # Layer 1: models.ini wins where it has a value.
+ info = models.resolve("together:specific", table, store)
+ assert info.ctx_size == 8192
+ assert info.vision is True
+ assert info.price_in == 5.0
+ # Layer 2 fills the gap models.ini left: price_out was never set.
+ assert info.price_out == 0.9
+
+ # Layer 2 alone for a model with no models.ini entry.
+ other = models.resolve("together:other", table, store)
+ assert other.ctx_size == 32768
+ assert other.price_in == 0.6
+ assert other.vision is None # layer 3: still unknown
+
+ # Layer 3 throughout for a provider that configured nothing.
+ bare = models.resolve("free:anything", table, store)
+ assert bare.ctx_size is None
+ assert bare.price_in is None
+
+ # Cost: prompt at the input rate, completion at the output rate.
+ rows = [
+ {"model": "together:other", "prompt_tokens": 1_000_000,
+ "completion_tokens": 1_000_000},
+ # A pre-migration row: no counts, no model. Contributes zero.
+ {"model": None, "prompt_tokens": None, "completion_tokens": None},
+ ]
+ assert models.conversation_cost(rows, table, store) == 1.5
+
+ # A mixed conversation prices each reply at what produced it.
+ mixed = [
+ {"model": "together:other", "prompt_tokens": 1_000_000,
+ "completion_tokens": 0},
+ {"model": "together:specific", "prompt_tokens": 1_000_000,
+ "completion_tokens": 0},
+ ]
+ assert models.conversation_cost(mixed, table, store) == 5.6
+
+ # An unpriced model contributes nothing rather than guessing.
+ assert models.conversation_cost(
+ [{"model": "free:anything", "prompt_tokens": 1_000_000,
+ "completion_tokens": 0}], table, store
+ ) == 0.0
+
+ # Whether a model can be priced at all decides ? versus blank.
+ assert models.is_priced("together:other", table, store) is True
+ assert models.is_priced("free:anything", table, store) is False
+ # Local is free, never unpriced.
+ assert models.is_billable("gemma4", table) is False
+ assert models.is_billable("free:anything", table) is False # no api_key
+ assert models.is_billable("together:other", table) is True
+
+ # Formatting: cents matter, so three decimals below a dollar.
+ assert models.format_cost(0.0) == "$0.000"
+ assert models.format_cost(1.5) == "$1.50"
+ assert models.format_cost(12.345) == "$12.35"
+ print("ok metadata resolution and cost")
+```
+
+Register it after `test_models_store()`.
+
+- [ ] **Step 2: Run the suite to verify it fails**
+
+Run: `./test_llamachat.py`
+Expected: FAIL with `AttributeError: module 'llamachat.models' has no attribute 'resolve'`
+
+- [ ] **Step 3: Write the implementation**
+
+Append to `llamachat/models.py`. Add the import at the top of the file beside the
+existing ones:
+
+```python
+from . import providers as providers_mod
+```
+
+Then append:
+
+```python
+def resolve(model_id: str, table: dict, store: "ModelStore") -> ModelInfo:
+ """Merge the three metadata layers, most specific first.
+
+ Per field, not per source: a models.ini entry that sets only ctx_size
+ still inherits the provider's prices.
+ """
+ stored = store.get(model_id) or ModelInfo()
+ name, _ = providers_mod.split(model_id, table)
+ provider = table.get(name)
+
+ def pick(from_store, from_provider):
+ return from_store if from_store is not None else from_provider
+
+ if provider is None:
+ return stored
+ return ModelInfo(
+ ctx_size=pick(stored.ctx_size, provider.ctx_size),
+ vision=pick(stored.vision, provider.vision),
+ price_in=pick(stored.price_in, provider.price_in),
+ price_out=pick(stored.price_out, provider.price_out),
+ )
+
+
+def is_billable(model_id: str, table: dict) -> bool:
+ """Whether this model costs money, regardless of prices being known.
+
+ Keyed on the provider having an api_key: that is what distinguishes a
+ free local model from a cloud one whose price was never entered.
+ """
+ name, _ = providers_mod.split(model_id, table)
+ provider = table.get(name)
+ return bool(provider and not provider.is_local and provider.api_key)
+
+
+def is_priced(model_id: str, table: dict, store: "ModelStore") -> bool:
+ """Whether a cost can be computed for this model."""
+ info = resolve(model_id, table, store)
+ return info.price_in is not None or info.price_out is not None
+
+
+def message_cost(row, table: dict, store: "ModelStore") -> float:
+ """Cost of one stored assistant row, priced at the model that made it."""
+ model_id = row["model"] if row["model"] else ""
+ if not model_id:
+ return 0.0
+ info = resolve(model_id, table, store)
+ prompt = row["prompt_tokens"] or 0
+ completion = row["completion_tokens"] or 0
+ total = 0.0
+ if info.price_in is not None:
+ total += prompt * info.price_in
+ if info.price_out is not None:
+ total += completion * info.price_out
+ return total / 1_000_000
+
+
+def conversation_cost(rows, table: dict, store: "ModelStore") -> float:
+ """Everything spent in one conversation so far.
+
+ Rows predating the token columns carry NULLs and contribute zero, so an
+ old conversation reads as free rather than as a fabricated number.
+ """
+ return sum(message_cost(row, table, store) for row in rows)
+
+
+def projected_cost(tokens: int, model_id: str, table: dict, store) -> float:
+ """What sending `tokens` of input to this model would cost.
+
+ Only the input rate: the length of the reply is unknowable in advance.
+ """
+ info = resolve(model_id, table, store)
+ if info.price_in is None:
+ return 0.0
+ return tokens * info.price_in / 1_000_000
+
+
+def format_cost(amount: float) -> str:
+ """Money, with enough precision to see a cheap turn move the number."""
+ if amount < 1:
+ return f"${amount:.3f}"
+ return f"${amount:.2f}"
+```
+
+- [ ] **Step 4: Run the suite to verify it passes**
+
+Run: `./test_llamachat.py`
+Expected: PASS, including `ok metadata resolution and cost`
+
+- [ ] **Step 5: Commit**
+
+```bash
+git add llamachat/models.py test_llamachat.py
+git commit -m "feat: resolve model metadata in layers and compute cost
+
+Resolution merges per field, not per source, so an entry setting only
+ctx_size still inherits the provider's prices. Cost prices each reply at
+the model that produced it, and rows predating the token columns
+contribute zero rather than a fabricated number.
+
+Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>"
+```
+
+---
+
+## Task 7: Store token counts and the producing model
+
+**Files:**
+- Modify: `llamachat/db.py:35-43` (SCHEMA), `llamachat/db.py:101-124` (_migrate), `llamachat/db.py:210-231` (update_message)
+- Modify: `test_llamachat.py`
+
+- [ ] **Step 1: Write the failing test**
+
+Add after `test_metadata_and_cost()`:
+
+```python
+def test_token_column_migration():
+ """A pre-token database opens, and new rows record counts and model."""
+ import sqlite3
+
+ with tempfile.TemporaryDirectory() as tmp:
+ path = Path(tmp) / "old.db"
+ conn = sqlite3.connect(path)
+ conn.executescript(
+ "CREATE TABLE sessions (id INTEGER PRIMARY KEY, mode TEXT,"
+ " title TEXT, model TEXT, created_at INTEGER, updated_at INTEGER);"
+ "CREATE TABLE messages (id INTEGER PRIMARY KEY, session_id INTEGER,"
+ " role TEXT NOT NULL, content TEXT NOT NULL,"
+ " created_at INTEGER NOT NULL);"
+ "INSERT INTO sessions VALUES (1,'chat','old','m',0,0);"
+ "INSERT INTO messages VALUES (1,1,'assistant','older reply',0);"
+ )
+ conn.commit()
+ conn.close()
+
+ history = db.History(path)
+ rows = history.messages(1)
+ # The pre-migration row survives and reads as unknown, not as zero.
+ assert rows[0]["content"] == "older reply"
+ assert rows[0]["prompt_tokens"] is None
+ assert rows[0]["completion_tokens"] is None
+ assert rows[0]["model"] is None
+
+ mid = history.add_message(1, "assistant", "")
+ history.update_message(
+ mid, "new reply",
+ prompt_tokens=1200, completion_tokens=340,
+ model="together:Qwen/Qwen2.5",
+ )
+ fresh = history.messages(1)[1]
+ assert fresh["prompt_tokens"] == 1200
+ assert fresh["completion_tokens"] == 340
+ assert fresh["model"] == "together:Qwen/Qwen2.5"
+
+ # Omitting them leaves stored values alone, as with reasoning.
+ history.update_message(mid, "edited")
+ kept = history.messages(1)[1]
+ assert kept["content"] == "edited"
+ assert kept["prompt_tokens"] == 1200
+ assert kept["model"] == "together:Qwen/Qwen2.5"
+ history.close()
+ print("ok token column migration")
+```
+
+Register it after `test_metadata_and_cost()`.
+
+- [ ] **Step 2: Run the suite to verify it fails**
+
+Run: `./test_llamachat.py`
+Expected: FAIL with `IndexError: No item with that key` (sqlite3.Row has no
+`prompt_tokens` column)
+
+- [ ] **Step 3: Write the implementation**
+
+In `llamachat/db.py`, extend the `messages` table in `SCHEMA` so a fresh database
+gets the columns directly. The block becomes:
+
+```sql
+CREATE TABLE IF NOT EXISTS messages (
+ id INTEGER PRIMARY KEY,
+ session_id INTEGER NOT NULL REFERENCES sessions(id) ON DELETE CASCADE,
+ role TEXT NOT NULL,
+ content TEXT NOT NULL,
+ reasoning TEXT NOT NULL DEFAULT '',
+ searches TEXT,
+ prompt_tokens INTEGER,
+ completion_tokens INTEGER,
+ model TEXT,
+ created_at INTEGER NOT NULL
+);
+```
+
+In `_migrate()`, extend the `messages` entry of the `added` dict:
+
+```python
+ "messages": {
+ "reasoning": "TEXT NOT NULL DEFAULT ''",
+ # Nullable rather than defaulted: NULL means "no searches",
+ # which is exactly what every pre-migration row wants.
+ "searches": "TEXT",
+ # Nullable for the same reason: an old row has no counts,
+ # and zero would read as a reply that cost nothing.
+ "prompt_tokens": "INTEGER",
+ "completion_tokens": "INTEGER",
+ # The model on `sessions` is the current one, which prices a
+ # switched conversation wrongly. Record what actually replied.
+ "model": "TEXT",
+ },
+```
+
+Replace `update_message` with:
+
+```python
+ def update_message(
+ self,
+ message_id: int,
+ content: str,
+ reasoning: str | None = None,
+ searches: str | None = None,
+ prompt_tokens: int | None = None,
+ completion_tokens: int | None = None,
+ model: str | None = None,
+ ) -> None:
+ """Fill in a streamed reply. Omitted fields keep their stored value."""
+ columns = ["content = ?"]
+ values: list = [content]
+ for column, value in (
+ ("reasoning", reasoning),
+ ("searches", searches),
+ ("prompt_tokens", prompt_tokens),
+ ("completion_tokens", completion_tokens),
+ ("model", model),
+ ):
+ if value is not None:
+ columns.append(f"{column} = ?")
+ values.append(value)
+ values.append(message_id)
+ self.conn.execute(
+ f"UPDATE messages SET {', '.join(columns)} WHERE id = ?", values
+ )
+ self.conn.commit()
+```
+
+- [ ] **Step 4: Run the suite to verify it passes**
+
+Run: `./test_llamachat.py`
+Expected: PASS, including `ok token column migration`. `ok reasoning column
+migration` and `ok search storage` must still pass, proving the rewritten
+`update_message` kept its old behaviour.
+
+- [ ] **Step 5: Commit**
+
+```bash
+git add llamachat/db.py test_llamachat.py
+git commit -m "feat: record token counts and the producing model per message
+
+Cost has to survive reopening a conversation, which means storing what
+each reply used. The model goes on the message rather than the session
+because sessions records only the current one, and a conversation that
+switched models would otherwise be priced entirely at whichever is
+selected now. All three columns are nullable so old rows read as unknown.
+
+Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>"
+```
+
+---
+
+## Task 8: Authorization header on the client
+
+**Files:**
+- Modify: `llamachat/backend.py:169-176` (Client.__init__), `:178-190` (models), `:192-224` (complete), `:299-337` (_stream_once)
+- Modify: `test_llamachat.py`
+
+- [ ] **Step 1: Write the failing test**
+
+Add after `test_token_column_migration()`:
+
+```python
+def test_client_auth_header():
+ """A client with a key sends Bearer auth; one without sends no header."""
+ sent = {}
+
+ class _HttpxResponse:
+ """Enough of an httpx response for Client.models().
+
+ The existing _FakeResponse in this file wraps bytes for urlopen and
+ has neither .json() nor .raise_for_status(), so it cannot stand in
+ for an httpx call.
+ """
+
+ status_code = 200
+
+ def __init__(self, payload):
+ self._payload = payload
+
+ def raise_for_status(self):
+ return None
+
+ def json(self):
+ return self._payload
+
+ def _recorder(url, timeout=None, headers=None):
+ sent["url"] = url
+ sent["headers"] = headers or {}
+ return _HttpxResponse({"data": [{"id": "m1"}]})
+
+ import httpx
+ original = httpx.get
+ try:
+ httpx.get = _recorder
+
+ assert backend.Client("http://x.example.org").models() == ["m1"]
+ assert "Authorization" not in sent["headers"]
+
+ backend.Client(
+ "http://x.example.org", api_key="sk-test-not-a-real-key"
+ ).models()
+ assert sent["headers"]["Authorization"] == "Bearer sk-test-not-a-real-key"
+ finally:
+ httpx.get = original
+ print("ok client authorization header")
+```
+
+Register it after `test_token_column_migration()` in the `__main__` block.
+
+- [ ] **Step 2: Run the suite to verify it fails**
+
+Run: `./test_llamachat.py`
+Expected: FAIL with `TypeError: Client.__init__() got an unexpected keyword argument 'api_key'`
+
+- [ ] **Step 3: Write the implementation**
+
+In `llamachat/backend.py`, replace `Client.__init__` with:
+
+```python
+ def __init__(self, base_url: str, timeout: int = 300, api_key: str = ""):
+ self.base_url = base_url.rstrip("/")
+ self.timeout = timeout
+ self.api_key = api_key
+
+ def _headers(self) -> dict:
+ """Bearer auth when the provider needs it, nothing when it does not."""
+ return {"Authorization": f"Bearer {self.api_key}"} if self.api_key else {}
+```
+
+Then add `headers=self._headers()` to each of the three outbound calls:
+
+In `models()`:
+
+```python
+ resp = httpx.get(
+ f"{self.base_url}/v1/models", timeout=10, headers=self._headers()
+ )
+```
+
+In `complete()`, add the argument after `json={...}`:
+
+```python
+ timeout=httpx.Timeout(self.timeout, connect=10),
+ headers=self._headers(),
+ )
+```
+
+In `_stream_once()`, add it to the `httpx.stream` call:
+
+```python
+ timeout=httpx.Timeout(self.timeout, connect=10),
+ headers=self._headers(),
+ ) as resp:
+```
+
+- [ ] **Step 4: Run the suite to verify it passes**
+
+Run: `./test_llamachat.py`
+Expected: PASS, including `ok client authorization header`
+
+- [ ] **Step 5: Commit**
+
+```bash
+git add llamachat/backend.py test_llamachat.py
+git commit -m "feat: send bearer auth when a provider needs a key
+
+The local router needs none, so the header is omitted entirely rather
+than sent empty.
+
+Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>"
+```
+
+---
+
+## Task 9: MultiClient routing
+
+**Files:**
+- Modify: `llamachat/backend.py` (append after `Client`)
+- Modify: `test_llamachat.py`
+
+- [ ] **Step 1: Write the failing test**
+
+Add after `test_client_auth_header()`:
+
+```python
+def test_multi_client():
+ """Models fan out across providers; requests route by model id."""
+ from llamachat import providers
+
+ table = providers.parse(
+ {
+ "providers": {
+ "local": {"base_url": "http://localhost:8181"},
+ "together": {
+ "base_url": "https://api.example.org",
+ "api_key": "env:MULTI_TEST_KEY",
+ "filter": ["qwen"],
+ },
+ "down": {"base_url": "https://dead.example.org"},
+ }
+ }
+ )
+ os.environ["MULTI_TEST_KEY"] = "sk-test-not-a-real-key"
+
+ listings = {
+ "http://localhost:8181": ["gemma4", "qwen3.5-9b"],
+ "https://api.example.org": [
+ "Qwen/Qwen2.5-72B",
+ "meta-llama/Llama-3.3-70B",
+ ],
+ }
+ built = []
+
+ class _StubClient:
+ def __init__(self, base_url, timeout=300, api_key=""):
+ self.base_url = base_url
+ self.api_key = api_key
+ built.append(self)
+
+ def models(self):
+ if self.base_url not in listings:
+ raise backend.BackendError(f"cannot reach {self.base_url}")
+ return listings[self.base_url]
+
+ multi = backend.MultiClient(
+ table, timeout=300, resolver=providers.KeyResolver(),
+ client_factory=_StubClient,
+ )
+ listed, problems = multi.models()
+
+ # Local models stay bare, cloud ones are prefixed, and the filter cut
+ # the Llama model out of together's listing.
+ assert listed == ["gemma4", "qwen3.5-9b", "together:Qwen/Qwen2.5-72B"]
+
+ # The unreachable provider is reported, and did not break the rest.
+ assert any("down" in p for p in problems)
+
+ # Routing: the client for a cloud model carries that provider's key.
+ client = multi.client_for("together:Qwen/Qwen2.5-72B")
+ assert client.base_url == "https://api.example.org"
+ assert client.api_key == "sk-test-not-a-real-key"
+
+ # And a local model gets the local client with no key at all.
+ local = multi.client_for("gemma4")
+ assert local.base_url == "http://localhost:8181"
+ assert local.api_key == ""
+
+ # The bare model name is what goes on the wire, not the prefixed id.
+ assert multi.wire_name("together:Qwen/Qwen2.5-72B") == "Qwen/Qwen2.5-72B"
+ assert multi.wire_name("gemma4") == "gemma4"
+
+ # A filter that matches nothing is reported by name with counts.
+ table["together"].filter = ["zzz"]
+ empty = backend.MultiClient(
+ table, timeout=300, resolver=providers.KeyResolver(),
+ client_factory=_StubClient,
+ )
+ _, notes = empty.models()
+ assert any("together: 0 of 2" in n for n in notes)
+
+ del os.environ["MULTI_TEST_KEY"]
+ print("ok multi-provider client")
+```
+
+Register it after `test_client_auth_header()`.
+
+- [ ] **Step 2: Run the suite to verify it fails**
+
+Run: `./test_llamachat.py`
+Expected: FAIL with `AttributeError: module 'llamachat.backend' has no attribute 'MultiClient'`
+
+- [ ] **Step 3: Write the implementation**
+
+In `llamachat/backend.py`, add the import beside the existing `from . import search`:
+
+```python
+from . import providers as providers_mod
+```
+
+Append after the `Client` class:
+
+```python
+class MultiClient:
+ """One façade over every configured provider.
+
+ Holds a `Client` per provider, built on demand so a key is resolved only
+ when that provider is actually used. The UI talks in prefixed model ids
+ and never needs to know which endpoint one lives on.
+ """
+
+ def __init__(self, table, timeout=300, resolver=None, client_factory=Client):
+ self.table = table
+ self.timeout = timeout
+ self.resolver = resolver or providers_mod.KeyResolver()
+ self._factory = client_factory
+ self._clients: dict[str, Client] = {}
+
+ def client_for(self, model_id: str) -> Client:
+ """The client that serves this model, resolving its key on first use."""
+ name, _ = providers_mod.split(model_id, self.table)
+ if name not in self._clients:
+ provider = self.table[name]
+ self._clients[name] = self._factory(
+ provider.base_url,
+ timeout=self.timeout,
+ api_key=self.resolver.resolve(provider),
+ )
+ return self._clients[name]
+
+ def wire_name(self, model_id: str) -> str:
+ """The model name the provider itself expects, without our prefix."""
+ _, model = providers_mod.split(model_id, self.table)
+ return model
+
+ def models(self) -> tuple[list[str], list[str]]:
+ """Every offered model id, plus notes about what went wrong.
+
+ A provider that is unreachable or whose filter matched nothing must
+ not stop the others being listed: local models have to stay usable
+ when the network is down.
+ """
+ listed: list[str] = []
+ problems: list[str] = []
+ for name, provider in self.table.items():
+ try:
+ available = self._listing_client(provider).models()
+ except (BackendError, providers_mod.KeyError_) as exc:
+ problems.append(f"{name}: {exc}")
+ continue
+ kept = providers_mod.apply_filter(provider, available)
+ if available and not kept:
+ problems.append(
+ f"{name}: 0 of {len(available)} models matched filter"
+ )
+ listed.extend(providers_mod.qualify(name, m) for m in kept)
+ return listed, problems
+
+ def _listing_client(self, provider) -> Client:
+ """Listing needs a client too, and needs the key for a private API."""
+ return self.client_for(providers_mod.qualify(provider.name, ""))
+```
+
+- [ ] **Step 4: Run the suite to verify it passes**
+
+Run: `./test_llamachat.py`
+Expected: PASS, including `ok multi-provider client`
+
+- [ ] **Step 5: Commit**
+
+```bash
+git add llamachat/backend.py test_llamachat.py
+git commit -m "feat: fan model listing out across providers and route by id
+
+A provider that is unreachable or whose filter matched nothing is
+reported rather than fatal: local models must stay usable when the
+network is down. Clients are built on demand so a key is resolved only
+when that provider is really used.
+
+Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>"
+```
+
+---
+
+## Task 10: The model dialog
+
+**Files:**
+- Create: `llamachat/modeldialog.py`
+- Modify: `test_llamachat.py`
+
+- [ ] **Step 1: Write the failing test**
+
+The dialog's layout needs a running Qt application, but its field conversion does
+not. Test the conversion, which is where the bugs live. Add after
+`test_multi_client()`:
+
+```python
+def test_model_dialog_values():
+ """The dialog's field text converts to ModelInfo, blanks meaning unknown."""
+ from llamachat import models, modeldialog
+
+ # Everything filled in.
+ info = modeldialog.to_info(
+ ctx_text="32768", vision=True, in_text="1.2", out_text="0.9"
+ )
+ assert info.ctx_size == 32768
+ assert info.vision is True
+ assert info.price_in == 1.2
+ assert info.price_out == 0.9
+
+ # Blank prices are legal and mean unpriced, not free.
+ blank = modeldialog.to_info(
+ ctx_text="8192", vision=False, in_text="", out_text=" "
+ )
+ assert blank.ctx_size == 8192
+ assert blank.price_in is None
+ assert blank.price_out is None
+
+ # Garbage reads as unknown rather than crashing the dialog.
+ junk = modeldialog.to_info(
+ ctx_text="not a number", vision=False, in_text="free", out_text=""
+ )
+ assert junk.ctx_size is None
+ assert junk.price_in is None
+
+ # Prefill is the inverse: unknown becomes an empty field.
+ assert modeldialog.to_fields(models.ModelInfo()) == ("", False, "", "")
+ assert modeldialog.to_fields(
+ models.ModelInfo(ctx_size=4096, vision=True, price_in=0.5)
+ ) == ("4096", True, "0.5", "")
+ print("ok model dialog value conversion")
+```
+
+Register it after `test_multi_client()`.
+
+- [ ] **Step 2: Run the suite to verify it fails**
+
+Run: `./test_llamachat.py`
+Expected: FAIL with `ModuleNotFoundError: No module named 'llamachat.modeldialog'`
+
+- [ ] **Step 3: Write the implementation**
+
+Create `llamachat/modeldialog.py` with the GPL header copied from `config.py`, then:
+
+```python
+"""Dialog for entering what presets.ini cannot answer about a model."""
+
+from PySide6.QtWidgets import (
+ QCheckBox,
+ QDialog,
+ QDialogButtonBox,
+ QFormLayout,
+ QLabel,
+ QLineEdit,
+ QVBoxLayout,
+)
+
+from .models import ModelInfo
+
+
+def _number(text: str, cast):
+ """Field text to a number, treating blank and garbage alike as unknown."""
+ text = (text or "").strip()
+ if not text:
+ return None
+ try:
+ return cast(text)
+ except ValueError:
+ return None
+
+
+def to_info(ctx_text: str, vision: bool, in_text: str, out_text: str) -> ModelInfo:
+ """Build a ModelInfo from the dialog's raw field values."""
+ return ModelInfo(
+ ctx_size=_number(ctx_text, int),
+ vision=bool(vision),
+ price_in=_number(in_text, float),
+ price_out=_number(out_text, float),
+ )
+
+
+def to_fields(info: ModelInfo) -> tuple[str, bool, str, str]:
+ """The inverse, for prefilling. Unknown becomes an empty field."""
+ return (
+ "" if info.ctx_size is None else str(info.ctx_size),
+ bool(info.vision),
+ "" if info.price_in is None else str(info.price_in),
+ "" if info.price_out is None else str(info.price_out),
+ )
+
+
+class ModelDialog(QDialog):
+ """Context size, vision and prices for one model.
+
+ Prefilled from the provider's defaults, so the common case is checking
+ the numbers rather than typing them.
+ """
+
+ def __init__(self, model_id: str, info: ModelInfo, parent=None):
+ super().__init__(parent)
+ self.setWindowTitle("Model settings")
+ self.model_id = model_id
+
+ ctx, vision, price_in, price_out = to_fields(info)
+ self.ctx = QLineEdit(ctx)
+ self.ctx.setPlaceholderText("unknown")
+ self.vision = QCheckBox("Accepts images")
+ self.vision.setChecked(vision)
+ self.price_in = QLineEdit(price_in)
+ self.price_in.setPlaceholderText("unpriced")
+ self.price_out = QLineEdit(price_out)
+ self.price_out.setPlaceholderText("unpriced")
+
+ layout = QVBoxLayout(self)
+ heading = QLabel(f"<b>{model_id}</b>")
+ heading.setTextInteractionFlags(heading.textInteractionFlags())
+ layout.addWidget(heading)
+
+ form = QFormLayout()
+ form.addRow("Context size (tokens)", self.ctx)
+ form.addRow("", self.vision)
+ form.addRow("Input price (per 1M tokens)", self.price_in)
+ form.addRow("Output price (per 1M tokens)", self.price_out)
+ layout.addLayout(form)
+
+ note = QLabel(
+ "Leave prices empty if you do not want a cost estimate.\n"
+ "Context size drives the meter and the attachment budget."
+ )
+ note.setWordWrap(True)
+ layout.addWidget(note)
+
+ buttons = QDialogButtonBox(
+ QDialogButtonBox.Save | QDialogButtonBox.Cancel
+ )
+ buttons.accepted.connect(self.accept)
+ buttons.rejected.connect(self.reject)
+ layout.addWidget(buttons)
+
+ def info(self) -> ModelInfo:
+ """What the user entered."""
+ return to_info(
+ self.ctx.text(),
+ self.vision.isChecked(),
+ self.price_in.text(),
+ self.price_out.text(),
+ )
+```
+
+- [ ] **Step 4: Run the suite to verify it passes**
+
+Run: `./test_llamachat.py`
+Expected: PASS, including `ok model dialog value conversion`
+
+- [ ] **Step 5: Commit**
+
+```bash
+git add llamachat/modeldialog.py test_llamachat.py
+git commit -m "feat: dialog for per-model context size, vision and prices
+
+Field conversion is separated from the widget so the part with the edge
+cases is testable without a running Qt application. Blank and unparseable
+both read as unknown, which is what an empty price field has to mean.
+
+Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>"
+```
+
+---
+
+## Task 11: Cost label widget
+
+**Files:**
+- Modify: `llamachat/ui.py` (add after `_short`, around line 135)
+- Modify: `test_llamachat.py`
+
+- [ ] **Step 1: Write the failing test**
+
+Add after `test_model_dialog_values()`:
+
+```python
+def test_cost_label_text():
+ """The label distinguishes free, unpriced, and a real figure."""
+ from llamachat import ui
+
+ # A local model costs nothing, so the label says nothing.
+ assert ui.cost_text(spent=0.0, projected=0.0, billable=False, priced=False) == ""
+
+ # A cloud model whose price was never entered: ? rather than blank, so
+ # it cannot be mistaken for free.
+ assert ui.cost_text(
+ spent=0.0, projected=0.0, billable=True, priced=False
+ ) == "?"
+
+ # Spent so far, with nothing composed yet.
+ assert ui.cost_text(
+ spent=0.043, projected=0.0, billable=True, priced=True
+ ) == "$0.043"
+
+ # Spent plus what sending the draft would add, kept visually separate.
+ assert ui.cost_text(
+ spent=0.043, projected=0.011, billable=True, priced=True
+ ) == "$0.043 +$0.011"
+
+ # A fresh conversation on a priced model still shows the projection.
+ assert ui.cost_text(
+ spent=0.0, projected=0.002, billable=True, priced=True
+ ) == "$0.000 +$0.002"
+ print("ok cost label text")
+```
+
+Register it after `test_model_dialog_values()`.
+
+- [ ] **Step 2: Run the suite to verify it fails**
+
+Run: `./test_llamachat.py`
+Expected: FAIL with `AttributeError: module 'llamachat.ui' has no attribute 'cost_text'`
+
+- [ ] **Step 3: Write the implementation**
+
+In `llamachat/ui.py`, add the import beside the existing project imports:
+
+```python
+from . import models as models_mod
+```
+
+Add after the `_short` function (around line 135):
+
+```python
+def cost_text(spent: float, projected: float, billable: bool, priced: bool) -> str:
+ """The cost label beside the context meter.
+
+ Three states, deliberately distinct: a local model shows nothing, a
+ cloud model with no price entered shows '?', and a priced one shows
+ what it has cost plus what the composed draft would add. Blank and '?'
+ must not collapse into each other, or an unpriced cloud model reads as
+ free.
+ """
+ if not billable:
+ return ""
+ if not priced:
+ return "?"
+ text = models_mod.format_cost(spent)
+ if projected > 0:
+ text += f" +{models_mod.format_cost(projected)}"
+ return text
+
+
+class CostLabel(QLabel):
+ """A one-line money readout that hides itself when there is nothing to say."""
+
+ def __init__(self):
+ super().__init__("")
+ self.setToolTip("")
+
+ def set_cost(
+ self, spent: float, projected: float, billable: bool, priced: bool
+ ) -> None:
+ text = cost_text(spent, projected, billable, priced)
+ self.setText(text)
+ self.setVisible(bool(text))
+ if not billable:
+ self.setToolTip("")
+ elif not priced:
+ self.setToolTip(
+ "No prices set for this model.\n"
+ "Use Model settings to enter them."
+ )
+ else:
+ tip = f"{models_mod.format_cost(spent)} spent in this conversation"
+ if projected > 0:
+ tip += (
+ f"\n+{models_mod.format_cost(projected)} to send what is "
+ "composed now"
+ )
+ tip += "\nApproximate: based on the prices you entered."
+ self.setToolTip(tip)
+```
+
+Confirm `QLabel` is in the `PySide6.QtWidgets` import list at the top of `ui.py`;
+it is already imported, so no import change is needed beyond `models_mod`.
+
+- [ ] **Step 4: Run the suite to verify it passes**
+
+Run: `./test_llamachat.py`
+Expected: PASS, including `ok cost label text`
+
+- [ ] **Step 5: Commit**
+
+```bash
+git add llamachat/ui.py test_llamachat.py
+git commit -m "feat: cost label showing spend and the next send's projection
+
+Three states stay distinct: blank for a free local model, ? for a cloud
+model whose prices were never entered, and a figure when they were.
+Collapsing the first two would make an unpriced cloud model read as free.
+
+Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>"
+```
+
+---
+
+## Task 12: Wire the window to providers
+
+**Files:**
+- Modify: `llamachat/__main__.py:155-165`
+- Modify: `llamachat/ui.py:356-406` (`ChatWindow.__init__`), `:669-717` (model handling)
+
+This task is wiring, not new logic, and its behaviour is covered by the tests
+already written plus a manual check. No new automated test.
+
+- [ ] **Step 1: Build the MultiClient at startup**
+
+In `llamachat/__main__.py`, replace the client construction (line 159) with:
+
+```python
+ client = MultiClient(cfg.providers, cfg.request_timeout)
+ presets = config.parse_presets(cfg.presets_path)
+ store = models.ModelStore(cfg.models_path)
+ window = ChatWindow(cfg, history, client, presets, store)
+```
+
+Update the imports at the top of `__main__.py`: replace the `Client` import with
+`MultiClient` and add `from llamachat import models`. Check the existing import
+lines and keep their style.
+
+- [ ] **Step 2: Accept the store in the window**
+
+In `llamachat/ui.py`, change `ChatWindow.__init__` (line 359) to take the new
+argument and keep it:
+
+```python
+ def __init__(self, cfg, history, client, presets, store):
+ ...
+ self.presets = presets
+ self.store = store
+```
+
+Add it right after the existing `self.presets = presets` line at 364.
+
+- [ ] **Step 3: List models from every provider**
+
+Replace `refresh_models` (line 669) with:
+
+```python
+ def refresh_models(self) -> None:
+ """Repopulate the picker from every provider, keeping the selection."""
+ previous = self.model_box.currentText()
+ available, problems = self.client.models()
+ if not available:
+ self.show_status(
+ "; ".join(problems) or "No models available", error=True
+ )
+ return
+
+ self.model_box.blockSignals(True)
+ self.model_box.clear()
+ for name in available:
+ info = self.model_info(name)
+ label = f"{name} 👁" if info.vision else name
+ self.model_box.addItem(label, name)
+ self.model_box.blockSignals(False)
+
+ target = previous or self.cfg.default_model
+ if target:
+ index = self.model_box.findData(_strip_marker(target))
+ if index < 0:
+ index = self.model_box.findText(target)
+ if index >= 0:
+ self.model_box.setCurrentIndex(index)
+ # A provider that failed is worth saying so even when others worked.
+ if problems:
+ self.show_status("; ".join(problems), error=True)
+ else:
+ self.hide_status()
+```
+
+- [ ] **Step 4: Resolve metadata through the three layers**
+
+Replace `current_preset`, `char_budget` and `vision_models` (lines 699-717) with:
+
+```python
+ def model_info(self, model_id: str):
+ """Metadata for one model: models.ini, then provider, then presets.
+
+ Local models get their context and vision from presets.ini, which
+ the first two layers can still override if the user entered values.
+ """
+ info = models_mod.resolve(model_id, self.cfg.providers, self.store)
+ preset = self.presets.get(model_id)
+ if preset is not None:
+ if info.ctx_size is None:
+ info.ctx_size = preset.ctx_size
+ if info.vision is None:
+ info.vision = preset.vision
+ return info
+
+ def current_info(self):
+ return self.model_info(self.current_model())
+
+ def current_preset(self):
+ return self.presets.get(self.current_model())
+
+ def char_budget(self) -> int:
+ ctx = self.current_info().ctx_size or 4096
+ return int(ctx * self.cfg.chars_per_token * self.cfg.attach_ctx_fraction)
+
+ def vision_models(self) -> list[str]:
+ names = []
+ for i in range(self.model_box.count()):
+ name = self.model_box.itemData(i)
+ if self.model_info(name).vision:
+ names.append(name)
+ return names
+```
+
+- [ ] **Step 5: Permit attachments on unknown-capability models**
+
+In `_ensure_vision_model` (line 842), replace the opening check:
+
+```python
+ info = self.current_info()
+ if info.vision:
+ return True
+ # Unknown is not the same as "no": a cloud model we know nothing
+ # about may well accept images, and the API will say so if it does
+ # not. Only a model known to lack vision gets stopped here.
+ if info.vision is None:
+ return True
+```
+
+- [ ] **Step 6: Route requests through the right client**
+
+Every call that currently does `self.client.<method>(model, ...)` must become a
+call on the routed client with the wire name. In `send()` (line 919) and
+`_start_stream()` (line 1040), the model passed to `StreamWorker` must be split.
+Change `_start_stream` to resolve both up front:
+
+```python
+ def _start_stream(self, model: str, messages: list[dict]) -> None:
+ try:
+ client = self.client.client_for(model)
+ except providers_mod.KeyError_ as exc:
+ self.show_status(str(exc), error=True)
+ return
+ wire = self.client.wire_name(model)
+```
+
+and pass `client` and `wire` to `StreamWorker` in place of `self.client` and
+`model`. Add the import at the top of `ui.py`:
+
+```python
+from . import providers as providers_mod
+```
+
+Apply the same treatment in `_start_titling` (line 1133), which builds a
+`TitleWorker`: resolve `client` and `wire` the same way, and on `KeyError_` skip
+titling silently rather than showing an error, since titling is never the user's
+own turn.
+
+- [ ] **Step 7: Verify the whole suite still passes**
+
+Run: `./test_llamachat.py`
+Expected: PASS, all checks, ending in `all checks passed`
+
+- [ ] **Step 8: Verify the app still starts against the local router**
+
+Run: `./llamachat.py`
+Expected: the window opens, the model dropdown lists the local models exactly as
+before with no prefix, and sending a message works. Close it.
+
+This is the check that the wiring is right; the unit tests cannot see it.
+
+- [ ] **Step 9: Commit**
+
+```bash
+git add llamachat/ui.py llamachat/__main__.py
+git commit -m "feat: list and route models through every configured provider
+
+Metadata now resolves models.ini over provider defaults over presets.ini,
+so a local model keeps getting its context and vision from the preset
+while remaining overridable. A model whose vision support is unknown no
+longer blocks an attachment: unknown is not the same as no, and the API
+will reject an image if it really cannot take one.
+
+Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>"
+```
+
+---
+
+## Task 13: Dialog triggers
+
+**Files:**
+- Modify: `llamachat/ui.py:407-572` (`_build_ui`), `:719-725` (`_on_model_changed`)
+
+- [ ] **Step 1: Add the on-demand button**
+
+In `_build_ui`, immediately after the `model_box` is added to its layout, add:
+
+```python
+ self.model_settings_button = QToolButton()
+ self.model_settings_button.setText("⚙")
+ self.model_settings_button.setToolTip(
+ "Context size, vision and prices for the selected model"
+ )
+ self.model_settings_button.clicked.connect(self.edit_model_settings)
+```
+
+Add it to the same layout the model box lives in, directly after it. Confirm
+`QToolButton` is in the `PySide6.QtWidgets` import list at the top of `ui.py` and
+add it if it is not.
+
+- [ ] **Step 2: Add the trigger and editor methods**
+
+Add after `vision_models` (around line 717):
+
+```python
+ def edit_model_settings(self, model_id: str | None = None) -> bool:
+ """Open the dialog for one model. True when values were saved."""
+ model_id = model_id or self.current_model()
+ if not model_id:
+ return False
+ current = models_mod.resolve(model_id, self.cfg.providers, self.store)
+ dialog = ModelDialog(model_id, current, self)
+ if dialog.exec() != QDialog.Accepted:
+ self.store.mark_skipped(model_id)
+ return False
+ self.store.save(model_id, dialog.info())
+ self.update_meter()
+ return True
+
+ def _maybe_offer_model_settings(self, model_id: str) -> None:
+ """Ask once, on first selection of an unconfigured cloud model.
+
+ Local models are exempt: presets.ini already answers context and
+ vision for them and they cost nothing. Cancelling records that the
+ offer was made, so a model tried once never asks again.
+ """
+ if not model_id or not models_mod.is_billable(model_id, self.cfg.providers):
+ return
+ if self.store.was_offered(model_id):
+ return
+ self.edit_model_settings(model_id)
+```
+
+- [ ] **Step 3: Fire it on selection**
+
+In `_on_model_changed` (line 719), add the offer before the existing body:
+
+```python
+ def _on_model_changed(self, _text: str) -> None:
+ self._maybe_offer_model_settings(self.current_model())
+```
+
+Keep everything the method already does after that line.
+
+- [ ] **Step 4: Add the imports**
+
+At the top of `ui.py`, beside the other project imports:
+
+```python
+from .modeldialog import ModelDialog
+```
+
+Confirm `QDialog` is already imported from `PySide6.QtWidgets`; it is, since
+`PromptDialog` subclasses it.
+
+- [ ] **Step 5: Verify the suite still passes**
+
+Run: `./test_llamachat.py`
+Expected: PASS, all checks
+
+- [ ] **Step 6: Verify the local path is unchanged**
+
+Run: `./llamachat.py`
+Expected: switching between local models opens no dialog. The ⚙ button opens the
+dialog for the current local model, and Cancel leaves it unchanged. Close it.
+
+- [ ] **Step 7: Commit**
+
+```bash
+git add llamachat/ui.py
+git commit -m "feat: offer model settings on first use of a cloud model
+
+Fires on selection rather than on send, so the interruption lands while
+the user is already changing settings instead of mid-thought. Local
+models never trigger it, and a cancelled dialog is recorded so a model
+tried once never asks again.
+
+Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>"
+```
+
+---
+
+## Task 14: Show the cost
+
+**Files:**
+- Modify: `llamachat/ui.py:407-572` (`_build_ui`), `:993-1029` (`update_meter`), `:1075-1082` (`_on_usage`), `:1110-1123` (`_on_stream_finished`)
+
+- [ ] **Step 1: Put the label beside the meter**
+
+In `_build_ui`, immediately after the line that adds `self.meter` to its layout,
+add:
+
+```python
+ self.cost = CostLabel()
+```
+
+and add it to the same layout directly after the meter.
+
+- [ ] **Step 2: Track the reply's token counts**
+
+In `_on_usage` (line 1075), record the counts so the finished reply can be stored
+with them. Replace the method with:
+
+```python
+ @Slot(int, int)
+ def _on_usage(self, prompt_tokens: int, total_tokens: int) -> None:
+ """Replace the estimate with the counts the server reported."""
+ info = self.current_info()
+ limit = info.ctx_size or 0
+ # What the next turn starts from is everything sent plus the reply.
+ self.exact_tokens = total_tokens or prompt_tokens
+ self.turn_prompt_tokens = prompt_tokens
+ self.turn_completion_tokens = max(total_tokens - prompt_tokens, 0)
+ self.meter.set_usage(self.exact_tokens, limit, exact=True)
+ self.update_cost()
+```
+
+Initialise both counters to 0 in `_teardown_stream` (line 1185) and in
+`ChatWindow.__init__` beside the other per-turn state:
+
+```python
+ self.turn_prompt_tokens = 0
+ self.turn_completion_tokens = 0
+```
+
+- [ ] **Step 3: Store them with the finished reply**
+
+In `_on_stream_finished` (line 1110), extend the `update_message` call:
+
+```python
+ self.history.update_message(
+ self.assistant_message_id,
+ self.assistant_buffer,
+ self.reasoning_buffer,
+ self._searches_json(),
+ prompt_tokens=self.turn_prompt_tokens or None,
+ completion_tokens=self.turn_completion_tokens or None,
+ model=self.current_model(),
+ )
+```
+
+- [ ] **Step 4: Compute and show the cost**
+
+Add after `update_meter` (around line 1029):
+
+```python
+ def update_cost(self) -> None:
+ """Refresh the money readout from stored counts plus the draft."""
+ if not hasattr(self, "cost"):
+ return # still building the window
+ model_id = self.current_model()
+ billable = models_mod.is_billable(model_id, self.cfg.providers)
+ priced = models_mod.is_priced(model_id, self.cfg.providers, self.store)
+ if not billable or not priced:
+ self.cost.set_cost(0.0, 0.0, billable, priced)
+ return
+
+ rows = (
+ self.history.messages(self.session_id)
+ if self.session_id is not None
+ else []
+ )
+ spent = models_mod.conversation_cost(rows, self.cfg.providers, self.store)
+
+ # Reopening a conversation resends its whole history, so the
+ # projection has to price everything that would go out, not just
+ # what was typed. That is what makes an expensive turn visible
+ # before it is paid rather than after.
+ pending = backend.estimate_tokens(
+ self._chat_context(
+ backend.build_user_content(
+ self.input.toPlainText(), self.attachments
+ )
+ ),
+ self.cfg.chars_per_token,
+ ) if self.session_id is not None or self.input.toPlainText() else 0
+ projected = models_mod.projected_cost(
+ pending, model_id, self.cfg.providers, self.store
+ )
+ self.cost.set_cost(spent, projected, billable, priced)
+```
+
+- [ ] **Step 5: Refresh it wherever the meter refreshes**
+
+At the end of `update_meter` (line 1029), add:
+
+```python
+ self.update_cost()
+```
+
+Three more call sites, so switching model or reopening a conversation updates the
+figure:
+
+- At the end of `refresh_models`. **Task 12 rewrote this method**, so add the call
+ to that rewritten version, at the very end of both the `if problems:` and `else:`
+ branches, or on a single line after the whole `if/else`.
+- At the end of the method that loads a session from the sidebar (around line
+ 1343, the one starting `session = self.history.get_session(session_id)`).
+- At the end of `new_session` (line 789), so starting a fresh conversation clears
+ the previous one's figure.
+
+- [ ] **Step 6: Verify the suite still passes**
+
+Run: `./test_llamachat.py`
+Expected: PASS, all checks
+
+- [ ] **Step 7: Verify against the local router**
+
+Run: `./llamachat.py`
+Expected: no cost label appears for a local model, the context meter behaves
+exactly as before, and sending a message still works. Close it.
+
+- [ ] **Step 8: Commit**
+
+```bash
+git add llamachat/ui.py
+git commit -m "feat: show conversation cost and the next send's projection
+
+The projection prices the whole request, not just the draft, because
+reopening a conversation resends its entire history and that is billed
+per turn on a cloud provider. Seeing it before sending is the whole point
+of the readout.
+
+Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>"
+```
+
+---
+
+## Task 15: Config template and documentation
+
+**Files:**
+- Modify: `llamachat/config.py:155-215` (`write_default`)
+- Modify: `README.md`
+- Modify: `CHANGELOG.md`
+
+- [ ] **Step 1: Document providers in the generated config**
+
+In `write_default`, append to the written text, before the closing parenthesis:
+
+```python
+ '\n'
+ '# External providers. The local router is a provider named "local",\n'
+ '# synthesized from base_url above when no [providers.local] exists.\n'
+ '# Any OpenAI-compatible endpoint works.\n'
+ '#\n'
+ '# Cloud models appear in the picker as provider:model. Local ones\n'
+ '# stay bare, so nothing about the local setup changes.\n'
+ '#\n'
+ '# api_key accepts three forms:\n'
+ '# "pass:api/together" read from the password store (preferred)\n'
+ '# "env:TOGETHER_KEY" read from the environment\n'
+ '# "sk-..." the key itself, in this file\n'
+ '# It is read lazily, on the first request to that provider, so a\n'
+ '# local-only session never unlocks the password store.\n'
+ '#\n'
+ '# filter keeps only models whose id contains one of these strings,\n'
+ '# case-insensitively. Providers list hundreds of models; without a\n'
+ '# filter the picker is unusable. Omit it to list them all.\n'
+ '#\n'
+ '# ctx_size, vision, price_in and price_out prefill the per-model\n'
+ '# dialog. Prices are US dollars per million tokens. Everything the\n'
+ '# dialog saves goes to models.ini beside this file, so none of\n'
+ '# these has to be set here.\n'
+ '#\n'
+ '# [providers.together]\n'
+ '# base_url = "https://api.together.xyz"\n'
+ '# api_key = "pass:api/together"\n'
+ '# filter = ["qwen", "deepseek"]\n'
+ '# ctx_size = 32768\n'
+ '# price_in = 0.60\n'
+ '# price_out = 0.60\n'
+```
+
+- [ ] **Step 2: Verify the generated config still parses**
+
+Run:
+
+```bash
+python3 -c "
+import tempfile, tomllib
+from pathlib import Path
+import sys; sys.path.insert(0, '.')
+from llamachat import config
+with tempfile.TemporaryDirectory() as t:
+ p = config.write_default(Path(t) / 'config.toml')
+ tomllib.loads(p.read_text())
+ cfg = config.load(p)
+ assert set(cfg.providers) == {'local'}, cfg.providers
+ print('generated config parses, providers:', list(cfg.providers))
+"
+```
+
+Expected: `generated config parses, providers: ['local']`
+
+The commented-out provider block must stay commented, or a fresh install would
+try to reach an endpoint the user never configured.
+
+- [ ] **Step 3: Document it in the README**
+
+Add a section after the existing web search documentation, matching its tone and
+depth. It must cover: the `[providers.*]` table with a worked together.ai example,
+the three `api_key` forms and why `pass:` is preferred, lazy resolution meaning no
+pinentry for local-only sessions, filtering and why it is needed, the per-model
+dialog and `models.ini`, and the cost readout being an approximation based on
+hand-entered prices.
+
+State plainly, as the README already does for search: **a conversation resends its
+whole history every turn, so a long conversation on a cloud provider is billed for
+all of it on each message.** That is the behaviour most likely to surprise, and the
+projection in the cost label exists to make it visible.
+
+Also note that tool calling and reasoning output are known to vary between
+providers, so web search on a cloud model may not work as it does locally.
+
+- [ ] **Step 4: Add the changelog entry**
+
+Add an `## [Unreleased]` section at the top of `CHANGELOG.md`, following the
+existing format, describing: external OpenAI-compatible providers, `pass`/env/
+literal key resolution, model filtering, the per-model settings dialog and
+`models.ini`, per-conversation cost with projection, and the three new `messages`
+columns. Note the cloud tool-calling caveat under a "Known limitations" line if
+the file's format has one; otherwise state it in the entry itself.
+
+- [ ] **Step 5: Verify the version test still passes**
+
+`test_version_matches_changelog` matches `^## \[(\d+\.\d+\.\d+)\]` and asserts the
+first hit equals `__version__`. `## [Unreleased]` does not match that pattern, so
+it is skipped and `## [0.3.0]` stays the first hit. The entry is safe as long as
+the heading is exactly `## [Unreleased]` and no version number is invented for it.
+
+Run: `./test_llamachat.py`
+Expected: PASS, including `ok version matches changelog`
+
+- [ ] **Step 6: Commit**
+
+```bash
+git add llamachat/config.py README.md CHANGELOG.md
+git commit -m "docs: document external providers, keys and cost
+
+The generated config ships the provider block commented out so a fresh
+install never reaches an endpoint nobody configured. The README states
+plainly that every turn resends the whole conversation, which is free
+locally and billed per message on a cloud provider.
+
+Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>"
+```
+
+---
+
+## Task 16: End-to-end verification against a real provider
+
+This task needs the user: it spends real money and needs a real API key. Do not
+attempt it autonomously.
+
+- [ ] **Step 1: Ask the user to configure one provider**
+
+The user adds a `[providers.together]` block (or siliconflow) to their
+`config.toml` with a `pass:` key and a filter, then starts llamachat.
+
+- [ ] **Step 2: Verify listing and the dialog**
+
+Expected: local models appear bare, cloud models appear as `provider:model` and
+only those matching the filter. Selecting a cloud model for the first time opens
+the settings dialog once. Cancelling it does not reopen it on reselection.
+
+- [ ] **Step 3: Verify a cloud reply and the cost**
+
+Send a short message to a cloud model with prices entered. Expected: the reply
+streams, the cost label shows a figure, and it grows on the next message.
+
+- [ ] **Step 4: Verify cost survives reopening**
+
+Close and reopen the conversation from the history sidebar. Expected: the cost
+label shows the accumulated figure, not zero.
+
+- [ ] **Step 5: Verify the local path is untouched**
+
+Switch back to a local model. Expected: no cost label, the meter behaves as
+before, and web search still works exactly as it did in 0.3.0.
+
+- [ ] **Step 6: Try web search on a cloud model and record what happens**
+
+This is the known risk. Expected: unknown. Record the outcome, whether the tool
+call is emitted, whether the reply arrives, and whether `reasoning_content` shows
+up. If it fails, capture the shape of the response and open it as its own piece of
+work rather than fixing it inside this one.
+
+- [ ] **Step 7: Commit any fixes found**
+
+Only if the earlier steps surfaced defects. Each fix gets a test first, following
+the pattern of every task above.
+
+---
+
+## Notes for the implementer
+
+**The local path is the one that must not regress.** Every task keeps local
+models bare, unfiltered, and free of the dialog. If a change makes a local model
+behave differently than it did in 0.3.0, that is a bug in the change, not an
+acceptable cost.
+
+**Never put a real API key in a test, a fixture or a commit message.** Two git
+hooks scan for exactly this and will reject the commit. Use
+`sk-test-not-a-real-key` and `example.org`.
+
+**Cloud tool calling is unverified.** Task 16 step 6 is where that gets found out.
+Nothing before it should assume search works on a cloud provider, and nothing
+should assume it is broken either.