aboutsummaryrefslogtreecommitdiffstats
path: root/docs/specs/2026-09-09-contacts.md
blob: c751bcb3e3aeba90750cd5666d8965617ec6145c (plain)
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
322
323
324
325
326
327
328
329
330
331
332
333
334
335
336
337
338
339
340
341
342
343
344
345
346
347
348
349
350
# abusectl `contacts`: IOCs to abuse contacts, over RDAP

Status: **agreed 2026-09-09**, in one brainstorming session with the user.
This spec settles the part the umbrella design left open, named there as
"RDAP bootstrap and referral chasing, caching policy, rate limits, and what
happens for a netblock that publishes no abuse contact".

Read `docs/specs/2026-09-08-abusectl-design.md` first. This document assumes
its manifest format, its case directory and its ordering between the parts,
and changes none of them.

## What it does

```
abusectl contacts <case>
```

Reads a case manifest, resolves an abuse contact for every IP and domain
indicator via RDAP, and rewrites the manifest with a `contacts[]` array. It
is read-only against the network: it queries registries and sends no mail,
files no report and changes nothing a third party can observe beyond the
queries themselves.

It is the FIRST network module in this tool. Everything below that looks
defensive is defensive because of that.

## Modules

```
abusectl/
  contacts.py     IOCs -> contacts[]            network, read-only
  rdap.py         bootstrap, query, jCard       network, read-only
```

`rdap.py` is the PROTOCOL: fetching and caching the IANA bootstrap files,
selecting a server for an address or a TLD, issuing a query, and pulling an
abuse address out of a jCard. `contacts.py` is the POLICY: which indicators
are resolvable, how hosts fold together, what is written into the manifest.

**The split is the same one that puts `redact.py` beside `parse.py`.** The
protocol dirt is where the defects will be, and a module of its own gets
tests that name it rather than tests that reach it incidentally through case
handling.

### The offline guarantee survives, for the whole suite

Both modules take a `fetch` callable as an ARGUMENT, defaulting to the real
urllib transport. Tests pass a fake and never construct the real one.

The suite must keep passing with `socket.socket`, `socket.create_connection`
and `socket.getaddrinfo` all raising, exactly as it does today, and that
property stays a property of the WHOLE suite. Not "the whole suite except
contacts". A network module that can only be tested with a network is a
module that stops being tested.

This is the pattern the umbrella design promised for exactly this moment: if
a later part needs the network, it goes in its own module with an injected
transport.

## THE FOURTH PROPERTY

The umbrella design states three properties that are not negotiable. This
part adds a fourth, because it is the first part that can violate it.

### Never disclose more than the host under investigation

An RDAP query tells a third party what the user is looking at. Queries carry
BARE HOSTS AND IP ADDRESSES ONLY, never a full URL.

A URL path can carry recipient identity. `parse.suspect_path_segments()`
FLAGS those rather than redacting them, deliberately, because a path segment
may be the thing being reported. That decision is safe only while the URL
stays local. Sending a URL to a registry would leak precisely what property 1
exists to prevent, through a channel property 1 does not cover: property 1
governs what is PUBLISHED, and a query is a disclosure that never appears in
any report.

Concretely: a `url` indicator contributes its HOST to the domain worklist and
nothing else. Path, query and fragment never leave the machine.

**This is a trap, not a theoretical concern.** The obvious implementation
resolves "a contact for each indicator" by reading each indicator's `value`,
and for a `url` indicator that value is an entire URL. The sweep described
under Testing exists to catch exactly that mistake against real mail.

## What gets resolved

| Indicator | Query | Why |
|---|---|---|
| `ipv4`, `ipv6` | RDAP IP lookup | the sending IP is the primary takedown target |
| `domain` | RDAP domain lookup | the registrar holds the abuse desk |
| `url` | its HOST, as a domain | folded into the domain worklist, see the fourth property |
| `sha256` | none | no registry owns a hash |
| `observation` | none | not a network object |

Indicators carrying `confidence: untrusted-hop` ARE resolved. A forged chain's
IP may still be the real sender's, and refusing to resolve it would discard
the case the tool exists for. The confidence marker stays in the manifest, so
review shows the user that a contact came from an attacker-supplied hop
before any report goes out.

**Hosts fold.** Twenty URLs on one host produce one worklist entry and one
query. The resulting contact entry lists every indicator id that contributed,
so nothing is lost by folding.

## RDAP mechanics

### Bootstrap

Fetch `ipv4.json`, `ipv6.json` and `dns.json` from `data.iana.org/rdap/`.
Cache them under `$XDG_CACHE_HOME/abusectl/rdap/`, defaulting to
`~/.cache/abusectl/rdap/`. TTL 7 days.

The cache is NOT in the case directory. It is not evidence, it is a copy of a
public map.

**A failed refetch falls back to the stale copy**, with a warning, rather
than failing the run. Last week's map is almost certainly still correct, and
IANA being unreachable should not stop the user filing a report. A stale
bootstrap is safe to serve: the worst case is querying a server that has
moved, which fails and reads as "no contact".

### Server selection

- IP: longest-prefix match over the bootstrap ranges, using `ipaddress`.
- Domain: exact TLD match in `dns.json`.
- No match: `abuse: []` with the reason recorded. Many TLDs publish no RDAP
  service at all, and that is a normal outcome rather than a defect.

### The registrable domain, by walking up

RDAP domain lookup wants the registrable domain. `mail.deep.example.invalid`
is not registrable and most registries answer it with a 404.

Query the full host, then drop one label and retry, until a server answers or
the labels run out. Cap at 5 attempts and NEVER query a bare TLD.

**This replaces a Public Suffix List deliberately.** Taking the last two
labels is wrong for every multi-part suffix, `co.uk` and `com.au` among them,
and a bundled PSL snapshot is a transcribed table that goes stale weekly,
which is the failure `init.PROVIDERS` already documents at length. The
registry is the authority on what is registrable, so the walk asks it instead
of approximating it.

The in-memory dedup below makes the walk cheap in the common case: a
campaign's twenty subdomains under one registrable domain collapse to one
successful query plus a few cheap misses.

### Transport

`urllib.request` from the standard library. The tool still declares no
dependencies, and there is no `requirements.txt` in the repository to add one
to.

- `Accept: application/rdap+json, application/json;q=0.9`. The weaker
  alternative is offered because some registries serve plain
  `application/json`, and a strict header would lose a contact that was
  published.
- **timeout mandatory**, 10 seconds. `urllib` with no timeout blocks forever,
  and a hung registry would hang a review.
- **redirects capped at 5 hops.** RIRs redirect to each other and the default
  handler follows without a limit of ours.
- **an https to http downgrade is refused.**

The cap and the downgrade refusal follow the discipline `_MAX_REDIRECT_DEPTH`
already sets in `parse.py`, and for the same reason: this is remote data
directing our next move.

`requests` was considered and rejected. It is present in slackware64-current,
so packaging was not the deciding factor. It simply buys little here: the two
things this module needs beyond a plain GET are the redirect cap and the
downgrade refusal, and neither is a knob on `requests` either, both needing a
`Session` plus custom wiring of comparable size. Declaring no dependencies at
all means the tool runs and tests with no venv step.

### Caching policy

Bootstrap: on disk, 7 day TTL, as above.

Responses: **in memory only, keyed by query URL, for the duration of one
run.** Nothing on disk.

The asymmetry is deliberate and is about which staleness hurts. A stale
bootstrap fails loudly enough, producing a miss. **A stale abuse address
sends a report into a mailbox nobody reads, silently**, and the user learns
nothing came of it only by never hearing back. RDAP rate limits are generous
for a human-paced tool filing a handful of reports, so the saving is not
worth that.

In-run dedup is free and still worth having: it is what stops a case with
forty URLs on one host from issuing forty identical queries.

## Extracting the address

Strict role match, and nothing else. An entity whose `roles` contains `abuse`
yields the `email` entries of its jCard. Recurse into nested `entities`,
depth capped at 4, because the abuse entity is usually a child of the
organisation entity.

Four rules, each with a victim:

**1. No fallback to a non-abuse entity.** A `technical` or `registrant`
contact is a named human who never volunteered to receive abuse mail. Mailing
them is useless and is a small privacy harm to an uninvolved third party.

**2. No `abuse@<domain>` by convention.** RFC 2142 says the mailbox should
exist; for a phishing domain it belongs to the ATTACKER. Constructing it
would mail the attacker to tell them their campaign was caught and that the
user's address is live, which is the exact outcome property 2 exists to
prevent. This is the single worst thing this module could do, and it looks
like a helpful fallback, which is why it is written down.

**3. Never follow a `links` referral to fetch a contact.** Some responses
point elsewhere for an entity. Following that is an outbound fetch to a URL
the response chose for us, which is SSRF-shaped. If the address is not in the
response, there is no address.

**4. Validate the address before storing it.** `email.utils.parseaddr`,
exactly one address, and reject any control character or newline. It becomes
a mail recipient in `report` and `submit`; a CRLF there is header injection
into mail this tool sends.

**Multiple abuse addresses are all kept.** Some netblocks publish two, and
picking one arbitrarily can drop the desk that would have answered.

**A missing contact is a normal outcome, not a failure.** The umbrella design
already settles this: the indicator keeps an empty `abuse` with the reason,
no email destination is created for it, and review shows what could not be
resolved. The indicator still reaches MISP and the vendor feeds.

## The manifest

```json
"contacts": [
  { "iocs": ["ioc-1"], "query": "198.51.100.7",
    "abuse": ["abuse@example.invalid"],
    "source": "rdap", "handle": "AS64496",
    "server": "rdap.example.invalid" },
  { "iocs": ["ioc-3", "ioc-7"], "query": "example.invalid",
    "abuse": [], "source": "rdap",
    "error": "no abuse role published" }
]
```

Two departures from the sketch in the umbrella design, both deliberate:

- **`iocs` is a list**, because hosts fold and one contact can serve several
  indicators.
- **`abuse` is a list**, because multiple desks are real.

`query` records what was actually asked. Review can then see that a URL
indicator was resolved by its host, which is the fourth property made visible
rather than merely promised.

`case.py` remains the only writer, and the write stays atomic.

### Re-running

**A re-run overwrites `contacts[]` wholesale**, in a single atomic write
after every lookup has finished.

Never a merge. A merge lets a contact resolved a week ago survive into a
report filed today, which is the stale-address hazard the response caching
policy already refuses. Overwriting makes a re-run always safe and always
current, which matters because a partial network failure makes re-running the
natural next step.

### Failure is per query

A timeout or an error on one indicator records that reason against that entry
and the run continues. The manifest is written once at the end regardless. A
run that resolved nothing at all still writes, so that review shows why
rather than showing an absent section.

## Testing

TDD, as everywhere in this repository. Write the failing test, watch it fail,
implement, watch it pass.

`rdap.py` is tested with a fake `fetch` returning canned responses in real RIR
SHAPES, with every address, handle and range replaced by `example.invalid`
and RFC 5737 documentation values, per the fixture rule in AGENTS.md.

Tests that have a right answer:

- longest-prefix server selection, including an address matching two ranges
- the label walk stops at the first server that answers, and never queries a
  bare TLD
- nested entity recursion finds an abuse entity under an organisation
- a response carrying only a `technical` role yields no address, not the
  technical one
- a `links` referral is NOT fetched: the fake transport asserts it was never
  called for that URL
- a CRLF in an email field is rejected
- **a `url` indicator contributes only its host**: the fake transport asserts
  no query ever contained the path
- a stale bootstrap is used when the refetch fails
- the whole suite still passes with sockets raising

### Two sweeps of the user's real spam

The `parse` sweep found a header the parser never read and a spoofed
`Reply-To` no fixture had. The same corpus is worth sweeping here, but this
module sends traffic, so it splits in two.

**Sweep A, offline, the whole corpus.** Build the worklist from all messages
with a fake transport that RECORDS queries and returns canned responses. No
packets. This carries the assertion:

> no query string ever contained anything but a bare host or IP address

That is the fourth property checked against real mail rather than against
fixtures, over the shapes a synthetic fixture misses: a URL with a path, with
userinfo, an IDN, a trailing dot, a bracketed IPv6 literal, a host that is
itself an address. It also yields free counts: unique hosts, how many
indicators fold per host, how deep the label walks go.

**Sweep B, online, a deliberate handful.** Ten to twenty indicators picked by
hand across distinct netblocks and TLDs, against the real registries. This
confirms that real jCard shapes parse, that the walk terminates, and that the
selection works. Twenty queries is polite and unremarkable.

Sweep A carries the assertion. Sweep B is a shape check.

**Ask before either, and sweep B needs an explicit go-ahead**, because unlike
every test before it in this repository it is OUTWARD-FACING: it discloses to
registries, and to a registrar who may be the attacker's, which netblocks and
domains the user is investigating, from their address, at a known time. That
cannot be undone, which is why it stays small and hand-picked rather than
being run over the corpus.

What may leave the sweep scripts: counts, tallies, TLDs, error reasons, walk
depths, and whether any query was ever malformed. What may not: an address, a
real abuse contact, a real domain written into a test, or a URL from a real
message. A defect found by a sweep is reproduced as a synthetic fixture and
THAT is committed. The real messages stay in the scratchpad, which is
per-session and outside the repository.

## Deliberately absent

- **No response cache on disk.** See the caching policy above.
- **No PSL.** See the label walk above.
- **No retry or backoff here.** Rate-limit handling belongs to `submit` and
  `retry`, which own the deadline logic the umbrella design specifies. A
  read-only lookup that fails is simply re-run.
- **No `whois` fallback.** It is a different protocol with an unparseable
  free-text response per registry, and its absence costs a contact rather
  than a report.
- **No contact for `sha256` or `observation` indicators.** No registry owns
  them.