aboutsummaryrefslogtreecommitdiffstats
path: root/AGENTS.md
diff options
context:
space:
mode:
authorDanilo M. <danix@danix.xyz>2026-09-10 11:23:27 +0200
committerDanilo M. <danix@danix.xyz>2026-09-10 11:23:27 +0200
commit2611c6485fa733e627f2c62c0369260f3a96d0bc (patch)
treeb4d55c13042d40d50146854ea82689245b1882e0 /AGENTS.md
parentf2fea43813ccebb1cf60e8afa5d4c7cdc52d0a20 (diff)
downloadabusectl-2611c6485fa733e627f2c62c0369260f3a96d0bc.tar.gz
abusectl-2611c6485fa733e627f2c62c0369260f3a96d0bc.zip
docs: record the report spec and both sweeps
Sweep A over 92 real messages after task 1 changed parse.py: 1685 indicators, 92 bodies, 0 crashes, 0 empty parses, no address from a raw source in the IOC output and no recipient address in any generated body. What it measured is the limit the spec already accepts. Subject is published verbatim, and 14 of the 92 messages carried the recipient's local part inside it because the kit personalises the lure. None carried it in the From display name, and the envelope recipient was cut from the boundary Received line in every message. That is the whitelist governing which headers travel rather than what is inside one, which the spec's "attacker-controlled free text is published unfiltered" section states outright and names the sweep as the cover for. personalised-subject.eml pins all three behaviours, the accepted one included, so the number cannot drift unnoticed. The for-clause test is mutation-checked: stop cutting the clause and all three fail. Sweep B, 12 hand-picked public targets and none from the corpus: 12 of 12, 0 failures, all five RIRs parseable, IPv6 live, the label walk and the multi-part suffix both correct. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01LByBnw83xr9YP85nskzkyE
Diffstat (limited to 'AGENTS.md')
-rw-r--r--AGENTS.md27
1 files changed, 27 insertions, 0 deletions
diff --git a/AGENTS.md b/AGENTS.md
index 2dc08cc..e627e50 100644
--- a/AGENTS.md
+++ b/AGENTS.md
@@ -332,6 +332,27 @@ failures: AFRINIC and `nic.cz` publish a handle but no `abuse` role, and
`example.museum` has no RDAP server for the TLD. The IANA bootstrap held 5 IPv4
services, 5 IPv6 and 590 DNS.
+`report` was swept the same way on 2026-09-10, after `parse.py` changed.
+Sweep A: 92 messages, 1685 indicators, 92 bodies generated, 0 crashes, 0 empty
+parses, 475 query targets. No address from a raw source reached the IOC output,
+and no recipient address reached a generated body.
+
+That sweep measured the limit the spec accepts rather than finding a defect.
+The `Subject` header is published verbatim, and 14 of the 92 messages carried
+the recipient's LOCAL PART inside it, because the kit personalises the lure.
+None carried it in the `From` display name. The envelope recipient was cut
+from the boundary `Received` line in every message. `tests/fixtures/`
+`personalised-subject.eml` pins all three behaviours, including the accepted
+one, so that number cannot change silently. Re-measure it when the whitelist
+changes.
+
+Sweep B on the same date: 12 of 12 targets completed, 0 failures, all five
+RIRs returning a parseable jCard, IPv6 live, the label walk reducing
+`www.ripe.net` and `a.b.c.example.org`, and `nic.uk` resolving the multi-part
+suffix. Three non-resolutions were correct: AFRINIC and `nic.cz` publish no
+abuse role, and `.museum` has no RDAP server. Bootstrap held 5 IPv4 services,
+5 IPv6 and 590 DNS.
+
**Sweep B must never draw its targets from the user's own spam corpus.** A query
tells a registrar which of their customers someone is investigating, and for a
phishing domain that registrar may be the attacker's own. Pick targets that are
@@ -359,6 +380,12 @@ settles only what they share.
touching RDAP, the bootstrap cache, or anything that issues a query. The
fourth property it introduced is stated above in its own right; the spec
carries the reasoning behind the rest of the module.
+- `docs/specs/2026-09-09-report.md`, the `report` spec. Read it before
+ changing the publishable-header whitelist, the freeze rule, or what a
+ report body contains. It records what is deliberately NOT filtered and
+ why, which is the first thing to read if a sweep result looks like a leak.
+- `docs/plans/2026-09-09-report.md`, the plan `report` was built from.
+ Historical in the same way, and it records the defect found in each task.
- `docs/plans/2026-09-08-parse.md`, the plan `init` and `parse` were built
from. Historical once built, but it records why each test exists.
- `docs/plans/2026-09-09-contacts.md`, the plan `contacts` and `rdap` were