aboutsummaryrefslogtreecommitdiffstats
path: root/docs/specs
diff options
context:
space:
mode:
Diffstat (limited to 'docs/specs')
-rw-r--r--docs/specs/2026-09-09-contacts.md25
-rw-r--r--docs/specs/2026-09-09-report.md443
2 files changed, 464 insertions, 4 deletions
diff --git a/docs/specs/2026-09-09-contacts.md b/docs/specs/2026-09-09-contacts.md
index c751bcb..d45cf83 100644
--- a/docs/specs/2026-09-09-contacts.md
+++ b/docs/specs/2026-09-09-contacts.md
@@ -234,19 +234,36 @@ resolved. The indicator still reaches MISP and the vendor feeds.
"contacts": [
{ "iocs": ["ioc-1"], "query": "198.51.100.7",
"abuse": ["abuse@example.invalid"],
- "source": "rdap", "handle": "AS64496",
- "server": "rdap.example.invalid" },
- { "iocs": ["ioc-3", "ioc-7"], "query": "example.invalid",
+ "source": "rdap", "handle": "AS64496" },
+ { "iocs": ["ioc-3", "ioc-7"], "query": "a.b.c.example.invalid",
+ "queried": "example.invalid",
"abuse": [], "source": "rdap",
"error": "no abuse role published" }
]
```
-Two departures from the sketch in the umbrella design, both deliberate:
+Three departures from the sketch in the umbrella design, all deliberate:
- **`iocs` is a list**, because hosts fold and one contact can serve several
indicators.
- **`abuse` is a list**, because multiple desks are real.
+- **`queried` is present only when the label walk shortened the name.** It
+ records WHAT was asked about, not what the message contained: a contact
+ found for `a.b.c.example.invalid` at `example.invalid` belongs to the
+ registered domain rather than the exact host. Absent when the query and
+ the answer are the same name, and never present on an IP, which is always
+ asked as itself.
+
+**`server` is specified but NOT built.** An earlier draft of this example
+carried it and nothing ever wrote it. It records WHO was asked, the RDAP
+endpoint the bootstrap selected, which is a different fact from `queried`
+and independent of it: one server answers thousands of names, and the same
+name would move to another server if the bootstrap changed. Unlike `queried`
+it is meaningful on the IP branch too, where longest-prefix selection picks
+an endpoint. It is worth building when a desk disputes a report and the
+answer is "this is the registry that published the address"; until `report`
+needs that, the bootstrap cache on disk makes the mapping reproducible and
+the field is dead weight.
`query` records what was actually asked. Review can then see that a URL
indicator was resolved by its host, which is the fourth property made visible
diff --git a/docs/specs/2026-09-09-report.md b/docs/specs/2026-09-09-report.md
new file mode 100644
index 0000000..12d2157
--- /dev/null
+++ b/docs/specs/2026-09-09-report.md
@@ -0,0 +1,443 @@
+# abusectl `report`: IOCs and contacts to report bodies
+
+Status: **agreed 2026-09-09**, in one brainstorming session with the user.
+This spec settles the part the umbrella design left open, named there as
+"the X-ARF (RFC 5965) schema version, which fields the user's reporting
+identity fills, and the plain-text alternative for desks that do not parse
+X-ARF".
+
+Read `docs/specs/2026-09-08-abusectl-design.md` first. This document assumes
+its manifest format, its case directory and its ordering between the parts.
+It changes one thing there and says so where it does: the re-run guard is a
+content hash rather than a timestamp.
+
+## What it does
+
+```
+abusectl report <case>
+```
+
+Reads a case manifest, groups the resolved contacts into destinations, writes
+a report body per destination under `bodies/`, and rewrites the manifest with
+a `destinations[]` array. It is OFFLINE and PURE: it opens no socket, sends
+no mail, and reads no file outside the case directory except the config.
+
+It is the last step before anything irreversible happens. What it produces is
+a document the user reads, edits and approves, so its output is written for a
+human first and a parser second.
+
+## Modules
+
+```
+abusectl/
+ report.py IOCs + contacts -> bodies + destinations[] pure
+```
+
+One module. There is no protocol/policy split here of the kind that puts
+`rdap.py` beside `contacts.py`, because there is no protocol: X-ARF is a MIME
+document and `email.message` in the standard library already is that layer.
+
+`report.py` takes the reporting identity as an ARGUMENT, the way `parse.py`
+takes the trust boundary. `cli.py` reads the config and passes it in. That is
+what keeps the module testable with no files on disk, and it matters more here
+than it did for `parse`: the identity is the one thing in a report that is
+disclosed deliberately, and a module that reaches for it itself is a module
+that can disclose it in a code path nobody reviewed.
+
+## The report
+
+Each email destination gets one MIME document, `multipart/report` with
+`report-type=feedback-report`, per RFC 5965. Three parts, in order.
+
+### Part 1, `text/plain`: what a human reads
+
+This is the part that decides whether the report is acted on. Desks triage a
+queue; a report whose ask is buried is a report that waits.
+
+```
+Phishing message received 2026-09-08, reporting infrastructure on your
+network. Requesting takedown and customer notification.
+
+Observed on your infrastructure:
+
+ 198.51.100.7 sending IP, first hop outside our trust boundary
+ example.invalid domain in message links, via a.b.c.example.invalid
+
+Message as declared:
+
+ Date: Mon, 08 Sep 2026 09:12:44 +0000
+ From: "Example Bank" <phish@example.invalid>
+ Subject: Your account requires verification
+
+Authentication results:
+
+ SPF: fail DKIM: none DMARC: fail
+
+URLs, redacted:
+
+ http://login-example.invalid/verify?id=REDACTED&src=REDACTED
+
+Recipient identifiers have been removed from this report by policy.
+Parameter names are preserved, parameter values are not. Full evidence is
+retained locally and is available on request.
+
+Reported by: Danilo M., Example Consulting <reporter@example.org>
+Generated by abusectl/<version>.
+```
+
+Four decisions in that shape.
+
+**The ask is the first sentence.** Not the evidence, not the identity. A desk
+reading one line must know what happened and what is wanted.
+
+**Only the recipient's own indicators appear.** Destinations are grouped per
+abuse address (below), and a desk shown three IPs that are not theirs stops
+reading. The `queried` field from `contacts` earns its keep here: "via
+`a.b.c.example.invalid`" tells the desk why they are being mailed about a name
+that is not literally in the message.
+
+**The redaction note is ALWAYS present, never conditional on whether anything
+was redacted.** A desk that sees `?id=REDACTED` with no explanation may read
+the report as malformed or doctored. One sentence turns that into a report
+that looks careful, and it opens the door for a desk that genuinely needs more
+to ask for it, which is the hand-paste route during review.
+
+**Plain text, hard-wrapped at 72 columns, with no HTML alternative.** Abuse
+desks run ticketing systems and many strip HTML. An HTML part would be a
+second body to keep in sync with the first for no reader.
+
+### Part 2, `message/feedback-report`: what a parser reads
+
+An RFC 5965 envelope carrying x-arf fields inside it.
+
+```
+Feedback-Type: abuse
+User-Agent: abusectl/0.1.0
+Version: 1
+Report-Type: phishing
+Source: 198.51.100.7
+Source-IP: 198.51.100.7
+Reported-Domain: example.invalid
+Arrival-Date: Mon, 08 Sep 2026 09:12:44 +0000
+Reported-Uri: http://login-example.invalid/verify?id=REDACTED
+```
+
+`Feedback-Type`, `User-Agent` and `Version` are the three fields RFC 5965
+requires. The rest are optional there or come from x-arf.
+
+**Why an RFC 5965 envelope with x-arf fields inside, rather than either
+alone.** RFC 5965 is an IETF standard and universally understood, but it was
+designed for feedback loops, where a report is ABOUT A MESSAGE. These reports
+are about INDICATORS, and 5965 has no natural field for "this specific host is
+the thing being reported". x-arf's `Source` does. The envelope is the
+standard's own extension point: the part is key/value, so a 5965 parser reads
+the fields it knows and ignores the rest, and x-arf tooling finds what it
+wants.
+
+**This choice deliberately does not depend on which of the two is more widely
+deployed**, which is a number nobody publishes and which this document does
+not claim to know. It was chosen so that the answer does not matter: a
+standards parser works, x-arf tooling works, and the human part works
+regardless of both.
+
+**`Source` is singular and a destination may carry several indicators.** The
+primary indicator fills it; the full list appears in the text part and in
+repeated `Reported-Uri` and `Source-IP` fields. RFC 5965 permits one report
+part per indicator instead, and that was rejected as heavier for no reader:
+the desk acts on the incident, not on each row.
+
+### Part 3, `text/rfc822-headers`: the message itself, almost
+
+**The original message is NOT attached.** `source.eml` carries every
+identifier the first property exists to keep out: `To`, `Cc`, `Delivered-To`,
+unredacted URLs whose query and path segments encode the recipient, the user's
+own Message-IDs, maildir paths and account keys. An abuse desk forwards a
+report to the abused customer, and for a phishing domain that customer may be
+the attacker; URLhaus is a public feed. Attaching it would deanonymise the
+reporter to the attacker, and for a consultant the tracking parameter may
+carry a CLIENT's identifier rather than the user's own.
+
+RFC 5965 provides `text/rfc822-headers` for exactly the case where the full
+message cannot be included, so this is the standard's own answer and not a
+deviation from it.
+
+The headers included are a WHITELIST:
+
+```
+Received (down to the untrusted hop only, never below)
+From, Subject, Date, Message-ID, Reply-To, Return-Path
+Authentication-Results, Received-SPF
+MIME-Version, Content-Type
+```
+
+**A whitelist, never a blacklist.** A blacklist means every header the parser
+learns to read later is a leak waiting for someone to remember. This is the
+same reasoning that has `parse.py` not reading `To` at all rather than
+stripping it afterwards.
+
+`Subject` and the `From` display name are attacker-controlled free text, and a
+sweep has already found a spoofed `Reply-To` display name. They are kept: they
+are the message's own content rather than the recipient's identity, and they
+are what lets a desk recognise a campaign they have seen before.
+
+**Considered and not built: a redacted body text part.** A desk analysing a
+campaign wants the lure, the impersonated brand and the pretext, and none of
+that survives headers-only. The body's INDICATORS already survive as IOCs
+regardless, so what is lost is the prose. It is not built because the prose is
+an unbounded attacker-supplied string, and deciding what is safe inside free
+text is a judgement rather than a whitelist, which is the shape of every leak
+this project has had. Build it when a desk actually asks for the lure, and
+build it as a redaction rule with its own tests, not as a passthrough.
+
+## Where the headers come from
+
+**`parse` stores the whitelisted headers in the manifest, and `report` never
+opens `source.eml`.**
+
+This is the same structural argument as the first property, applied one level
+down: `report` cannot disclose a header it was never given. The alternative,
+re-reading `source.eml` at report time and filtering there, would put a second
+"what may be disclosed" decision in a second module, away from `parse.py`
+where that decision currently lives, and two places to remember is how the
+fourth property leaked three times.
+
+The cost is real and is accepted: this is a change to `parse.py`, a new
+`headers` block in the manifest, and **both sweeps must be re-run**, per
+`AGENTS.md`. The sweep assertion is what proves the whitelist does not carry
+an address, and a whitelist written by hand is exactly the kind of thing a
+sweep catches being wrong.
+
+### Two accepted disclosures, named so they are not mistaken for leaks
+
+The first property reads as an unqualified "recipient identifiers must never
+reach a report". These are the deliberate exceptions the whitelist creates,
+recorded here rather than left to be rediscovered in a test comment.
+
+**Our own receiving relay's hostname is published.** The boundary `Received`
+line names it in its `by` clause and `Authentication-Results` names it as the
+authserv-id, so `mx.example.org` travels with every report. That is the
+user's mail host, not the user's identity, and an abuse desk learns it from
+the report's own `From` regardless. It is accepted because removing it would
+mean rewriting the inside of two headers whose value to a desk is precisely
+that they are the receiving server's own verbatim words. The consequence is
+that the manifest-wide "no bare `example.org`" assertion cannot hold over the
+`headers` block; `tests/test_cli.py` narrows it there and asserts the
+ADDRESS is still absent, which is the part that matters.
+
+**Attacker-controlled free text is published unfiltered.** `Subject` and the
+`From` display name are kept deliberately, because they are what lets a desk
+recognise a campaign. An attacker who writes the recipient's address into
+one, plainly or obfuscated as `you%40example.org`, gets it published: the
+whitelist governs WHICH headers travel, never what is inside one. This is
+not fixed by filtering free text, which is the judgement-shaped problem that
+`AGENTS.md` names as the origin of every leak this project has had. The sweep
+over real mail is what covers this class, which is one more reason it is not
+optional here.
+
+The envelope recipient is NOT in this list. Our own relay writes it into the
+boundary `Received` line's optional `for` clause, and that clause is cut
+before the line is stored, in every shape the grammar allows.
+
+## The reporting identity
+
+Three config keys, all under a `[reporter]` section:
+
+```toml
+[reporter]
+name = "Danilo M."
+org = "Example Consulting"
+email = "reporter@example.org"
+```
+
+They fill the report's `From`, the `Reported by:` line in the text part, and
+nothing else. `User-Agent` is `abusectl/<version>` and is not configurable.
+
+**The reporter's identity is disclosed DELIBERATELY, and that is what makes it
+different from every other identifier this tool refuses to publish.** The same
+address recovered from a `To` header is a leak; supplied in a config file it
+is the user choosing to be identified, and a report with no reply address is
+one a desk deprioritises. The distinction is provenance, so it is enforced by
+provenance: the identity comes from config ONLY, and `parse` must never supply
+it. If the two ever became one path, the distinction would be a comment rather
+than a guarantee.
+
+A skipped answer is ABSENT from the config, never an empty string, the same
+rule the rest of the config follows. `init` grows prompts for these three, and
+they are hand-tested like the rest of the prompts.
+
+**Noted, not built: a per-case reporting identity.** Reporting a campaign that
+targeted a client, under the user's own name, tells the abuse desk which
+consultant is working that incident, which is a disclosure about the
+engagement rather than about the mail. The escape hatch already exists without
+new machinery: review edits the bodies, and a `From` in a body is a line the
+user can change. A config-level override belongs to the first real engagement
+where it matters, not to this spec.
+
+## Destinations
+
+**One destination per abuse ADDRESS.** Every IOC whose contact resolved to
+`abuse@example.invalid` is grouped into one report to that desk.
+
+Contacts already fold by host, but two different contacts can still resolve to
+the same address, an IP and a domain both at one hoster being the common case.
+Grouping per contact would send that desk two mails about one incident, which
+is the duplicate-mail behaviour desks complain about. Grouping per IOC would
+be a storm.
+
+`report` writes three kinds of destination:
+
+| kind | body | who builds the payload |
+|---|---|---|
+| `email` | `bodies/<id>.xarf` | `report`, the MIME document above |
+| `api` | `null` for now, see below | `submit`, once its spec settles the shape |
+| `misp` | `null` | `submit`, via PyMISP |
+
+MISP is a destination like the others with a null body, per the umbrella
+design, so `submit` stays one loop with one ordering rule.
+
+**Only `email` bodies are built by THIS spec.** The `api` kinds are created as
+destination entries with their IOC lists and a `pending` status, and their
+bodies stay null until the `submit` spec settles each vendor's payload shape.
+Writing a vendor's JSON now would mean guessing an endpoint's contract from
+memory, which is the mistake the provider table already records: the first
+draft was written from memory and every range was wrong. The destination
+entries exist from the start so that `submit` fills bodies rather than
+inventing rows, and so review can already see which vendors a case will reach.
+
+### An indicator with no abuse contact
+
+A missing contact is a normal outcome, not an error: the umbrella design
+already settles that, and RDAP publishes no abuse role for many netblocks.
+
+`report` writes an `unreportable[]` array into the manifest, one entry per IOC
+that reached no email destination, carrying the reason from its contact entry:
+
+```json
+"unreportable": [
+ { "ioc": "ioc-4", "reason": "no abuse role published" },
+ { "ioc": "ioc-9", "reason": "not ASCII, and we do not guess at an IDN encoding" }
+]
+```
+
+Two reasons for making it explicit. Review becomes honest: the user sees that
+four indicators are going to MISP and the vendors but no desk was found for
+them, which is a fact they may want to act on by hand, and finding it any other
+way means diffing the IOC list against every destination's IOC list. And it is
+the same instinct as `suspect_path_segments` flagging rather than redacting: a
+failure that is visible beats a failure that is merely absent.
+
+**It is not an error and does not affect the exit code.** A case where nothing
+resolved still produces MISP and vendor destinations and is a perfectly good
+report. `report` exits non-zero only when it could not write.
+
+## Re-running, and the frozen case
+
+The umbrella design says `report` refuses to run again on a case whose bodies
+were modified after generation, by a timestamp check. **The intent stands and
+the mechanism changes**: it is a content hash.
+
+An mtime is a poor witness in both directions. A `git checkout`, an `rsync`, a
+backup restore or an editor that writes-and-renames all move mtime with no
+human having edited anything, and an editor that preserves mtime hides a real
+edit. So `report` records the SHA-256 of each body it writes, in that body's
+destination entry, and compares content rather than a rumour about content.
+`hashlib` is standard library, so this costs a field and no dependency.
+
+The hash has a SECOND job, and the spec states it so a later change does not
+drop it as redundant: for a destination that has been sent, the hash is the
+record of what was actually disclosed to a third party.
+
+### The rule
+
+**A case where ANY destination has been sent is FROZEN.** `report` refuses,
+and there is no `--force` override.
+
+The destinations are not independent artifacts, they are one incident reported
+in parallel. If one desk holds the report and a body for another desk is then
+regenerated with different content, two desks hold contradictory accounts of
+the same case, and a desk that forwards to the other finds the reporter
+unreliable. Regeneration is also not as isolated as it looks: the shared parts,
+the identity, the header block, the IOC list, come from the manifest, so
+regenerating one body after the manifest has changed produces a case whose
+bodies were built from two different states.
+
+Freezing is recorded EXPLICITLY, written by `submit` at its first success:
+
+```json
+"frozen": { "at": "2026-09-08T12:40:11Z", "by": "abusedb" }
+```
+
+Absent means not frozen, the same convention the config follows for a skipped
+answer. It is an optional field, so existing manifests stay loadable and the
+`format` version does not change.
+
+**Explicit rather than derived from the statuses**, because the marker is
+write-once and monotonic. A status corrected by hand, or a status added by a
+later schema that nobody remembered to add to a frozen set, would quietly
+unfreeze a derived check. For a rule protecting an evidence record, a field
+that can only be turned on is the right shape.
+
+`submit` must write the marker in the SAME atomic manifest write as the first
+`sent` status. A separate write leaves a window where a crash produces a case
+that has been disclosed and does not know it.
+
+### The whole rule, in order
+
+| case state | `report` does |
+|---|---|
+| `frozen` present | refuses, no override, names the destination that landed |
+| bodies modified, not frozen | refuses without `--force` |
+| bodies modified, `--force` | backs each up to `<name>.<timestamp>.orig`, regenerates |
+| bodies missing or unmodified | regenerates |
+
+A deleted body regenerates silently. The only reasons to delete one are a
+mistake or a deliberate start-over, and regeneration is what both want; the
+case that looked like it needed protecting, a body deleted after it was sent,
+is caught by the freeze rather than by the file check.
+
+Because `--force` can now only ever touch a case that nothing has left, it is
+a far safer flag than it first appears. The backup is kept anyway: the
+umbrella design's point about a review that took twenty minutes applies, and
+`init` already backs up a config it is about to replace.
+
+A user who needs to change a body on a frozen or deferred case still can, by
+editing it during review. That route is unaffected and is the right one: it is
+a deliberate act with the user looking at the text.
+
+## Testing
+
+TDD, and the same rule as the rest of the repository: test what has a right
+answer.
+
+Tested, because there is one:
+
+- the MIME structure: three parts, the right types, `report-type=feedback-report`
+- the header whitelist keeps what it should and, more importantly, DROPS
+ `To`, `Cc`, `Delivered-To` and `X-Original-To` when a fixture carries them
+- `Received` is truncated at the untrusted hop and never includes the ones
+ below it, against `forged-chain.eml`
+- destinations group per address, including the two-contacts-one-address case
+- an IOC with no contact lands in `unreportable[]` and creates no destination
+- the body hash detects a modified body, and does not fire on an untouched one
+- a frozen case refuses even with `--force`
+- `--force` writes the timestamped backup before overwriting
+
+Not unit-tested: whether the text part READS well to an abuse desk. That has
+no assertion, and it is the same category as the interactive prompts. The user
+hand-tests it by reading a generated report.
+
+**The leak sweep must be re-run**, both A and B, because this spec changes
+`parse.py`. Sweep A gains a third assertion: no address from the raw source
+appears in any generated report body. That is the assertion that actually
+proves the whitelist, and it should be as broad as the existing one, every
+address in the source against every byte of every body.
+
+## What this leaves for `submit`
+
+Named here because this spec creates them, not to settle them:
+
+- writing the `frozen` marker atomically with the first `sent` status
+- the vendor JSON shapes, and which side writes them. `report` creates the
+ `api` destination rows; whether it also learns to write their bodies is a
+ decision for that spec. If a vendor's payload turns out to need a value the
+ manifest does not hold, that is a change here, not a workaround there.