diff options
Diffstat (limited to 'AGENTS.md')
| -rw-r--r-- | AGENTS.md | 62 |
1 files changed, 62 insertions, 0 deletions
diff --git a/AGENTS.md b/AGENTS.md new file mode 100644 index 0000000..3919287 --- /dev/null +++ b/AGENTS.md @@ -0,0 +1,62 @@ +# AGENTS.md + +Guidance for AI coding agents working in this repository. + +## What this is + +`bunkr_dl`: single-file Python script that downloads Bunkr albums and search +results. Runtime dependencies: `requests`, `beautifulsoup4`, `lxml`. Keep it a +single file and do not add dependencies. + +On the maintainer's machine the script is installed as a symlink into the repo, +so edits are live immediately: + + ~/bin/bunkr_dl -> bunkr_dl + +## Download chain + +Bunkr breaks this regularly (domain rotation, layout changes). When the script +finds nothing, re-inspect the live pages with `curl` before changing code. + +1. Album page `https://<domain>/a/<id>`: relative `/f/<slug>` anchors. Large + albums are paginated as `?page=N` (100 files each). An out-of-range page + returns the last page again, so pagination stops when a page adds no new + links, not when a page is empty. The page also contains a JS template with + `href="/f/' + file.slug + '"`, which the parser correctly ignores. +2. File page `/f/<slug>`: anchor to `https://dl.<domain>/file/<numeric id>`. +3. `POST https://dl.<domain>/api/_001_v2` with JSON `{"id": "<id>"}` returns + `{"mediafiles", "path", "original"}`. +4. `GET https://glb-apisign.cdn.cr/sign?path=<path>` returns `{"token", "ex"}`. + Download `<mediafiles><path>?token=..&ex=..`. + +## Conventions + +- New options go in the `argparse` setup in `main()` and in the README usage + table. +- `download_file` resumes with a per-request `Range` header. Never mutate the + global `headers` dict. 206 appends, 200 overwrites, 416 means already complete. +- Use placeholder URLs (`https://bunkr.cr/a/XXXXXXXX`) in docs, comments and + commit messages, never real album URLs or titles. + +## Testing + +No test suite. Verify against a real album without writing to the default +output folder, by loading the script as a module and stubbing the download +step: + + python - <<'EOF' + import importlib.util, importlib.machinery + l = importlib.machinery.SourceFileLoader('b', 'bunkr_dl') + b = importlib.util.module_from_spec(importlib.util.spec_from_loader('b', l)); l.exec_module(b) + u = 'https://bunkr.cr/a/XXXXXXXX' + soup = b.get_soup(u); links = b.find_media_links(u, soup); print(len(links)) + fm = b.get_final_media_link(b.find_download_link(links[0])); print(fm[1]) + EOF + +Compare the link count with the "N files" shown on the album page. For a full +run use `-o` pointing at a temp dir. + +## Repo facts + +- License: GPLv2 only (`LICENSE`, header in `bunkr_dl`). +- `origin` is the personal git server (cgit section "Linux"). |
