aboutsummaryrefslogtreecommitdiffstats
path: root/docs/specs
diff options
context:
space:
mode:
Diffstat (limited to 'docs/specs')
-rw-r--r--docs/specs/2026-09-08-abusectl-design.md423
1 files changed, 423 insertions, 0 deletions
diff --git a/docs/specs/2026-09-08-abusectl-design.md b/docs/specs/2026-09-08-abusectl-design.md
new file mode 100644
index 0000000..7f72b5c
--- /dev/null
+++ b/docs/specs/2026-09-08-abusectl-design.md
@@ -0,0 +1,423 @@
+# abusectl: abuse reporting sidecar, umbrella design
+
+Status: **agreed 2026-09-08**, in one brainstorming session with the user.
+This is the UMBRELLA spec. Each part gets its own spec before it is built;
+this one settles the decisions the parts share and would be expensive to
+change later.
+
+Backlog: item 194 in qtmaildir's
+`docs/superpowers/plans/2026-08-03-post-0.1.0-usability.md`, which carries the
+qtmaildir half and its ordering against items 187 and 190. This document is
+the sidecar, and the two repositories are coupled only by the manifest format
+below and the command name in qtmaildir's config.
+
+## Why this exists
+
+The user is a cybersecurity consultant. They have been phished, and they want
+to act quickly and effectively on a campaign that targets them: parse a
+flagged message, extract its indicators, find who to report each one to, and
+file the result both to their own MISP instance and to the public abuse
+channels. Filling a database with phishing attempts is a feature no other mail
+client offers, which is a reason this one exists.
+
+Two of the user's own constraints are SAFETY PROPERTIES rather than
+preferences, stated when the item was first written and not to be traded away
+for convenience. They are specified in full below: recipient identifiers are
+never captured, and no remote content is ever fetched during parsing.
+
+## Why it is not in qtmaildir
+
+qtmaildir does no network protocol work at all, by design: fetching and
+sending are external scripts, which is why there is no IMAP client, no SMTP
+client, and why send is a per-account `send_command` on stdin. This tool needs
+RDAP, three vendor REST APIs and mail to abuse desks, which is four outbound
+protocols. Building it into `src/` would repeal that rule rather than extend
+it. Its ecosystem is Python besides, where `assets/hooks/` already puts the
+non-GUI parts of this mail system.
+
+**The split is architectural and is not a judgement on the feature.** The
+review UI lives in qtmaildir, because a reviewable document is exactly what
+makes a GUI review possible.
+
+**A separate repository, not a submodule.** qtmaildir invokes `abusectl` as a
+subprocess by name, the way it already invokes `mailsync.sh` and the
+`send_command`; nothing in `src/` includes or imports it. A submodule pins a
+commit that is not the one that runs, since what executes is whatever is
+installed on `$PATH`, and it would cost a two-step clone, an ordering
+constraint on every push across two remotes, and an empty directory in the
+release tarball, because `git archive` does not include submodule contents.
+Compatibility, if it ever matters, is a `format` version in the manifest, the
+same discipline `rules.json` already uses across two readers in two languages.
+
+## Shape
+
+One tool, subcommands, one repository. Each subcommand reads and writes a
+**case directory**, which is where state lives between steps. A review can
+therefore take a week and survive a reboot.
+
+```
+abusectl parse msg.eml -> case dir, IOCs offline, pure
+abusectl contacts <case> -> + abuse contacts network, read-only
+abusectl report <case> -> + report bodies offline, pure
+ [ review, in qtmaildir or $EDITOR ]
+abusectl submit <case> -> MISP, then vendors network, writes
+```
+
+Every subcommand is independently runnable and independently useful. `parse`
+alone triages a message. `parse` + `contacts` + `report` produces a document
+that can be sent by hand, with no API key configured anywhere.
+
+**The tool stays usable from a terminal.** The qtmaildir dialog is one front
+end, never the only one: a message that never came through qtmaildir, or a
+session over SSH, must still be workable. Same reasoning as `mailsync.sh`
+printing to stdout as well as its log.
+
+## The one hard ordering rule
+
+`submit` writes **MISP first** and aborts the vendor fan-out if that write
+fails.
+
+MISP is the user's own instance and is correctable through its REST API
+(update, soft-delete, hard-delete); a submission to AbuseIPDB, URLhaus or
+VirusTotal cannot be recalled. So the reversible step gates the irreversible
+ones. Reporting externally with no local record of having done so, and no
+dedup entry to stop a duplicate later, is worse than not reporting at all.
+
+MISP is also the DEDUP ORACLE. It correlates attributes by value across
+events and answers `/attributes/restSearch`, so `submit` can ask whether the
+infrastructure has been reported before and skip or attach rather than
+duplicate. Writing it before the fan-out is what makes that answer available.
+
+**Correction is MISP's own web UI**, not a subcommand here. A destructive verb
+in a tool run often, guarding data the user cares about, is not worth building
+for something MISP already does well. A correction subcommand is possible
+later.
+
+## The case directory
+
+```
+<cases>/2026-09-08-a3f1/
+ source.eml the original, unredacted
+ manifest.json IOCs, contacts, destinations, per-destination status
+ bodies/
+ abusedb.json
+ rdap-1.xarf
+```
+
+The path is configurable. **Nothing deletes cases**: they are the user's
+evidence, and a tool that silently bins a report sent last month is worse than
+a directory that grows. Cleanup is not built.
+
+`manifest.json` carries a `format` version. Bodies are separate files because
+they are text a human edits; putting them in the manifest would mean editing
+escaped strings inside a JSON document, or growing an extract/apply pair,
+which is a directory reinvented with extra steps.
+
+**`source.eml` is unredacted local evidence, and the submit path must never
+attach it wholesale.** This is the one route by which the redaction guarantee
+below could leak, so it is stated rather than left obvious. The case directory
+is sensitive at rest.
+
+## Redaction, structurally
+
+**The parser never captures recipient identifiers.** Not stored, not hashed,
+not written to the manifest:
+
+- `To`, `Cc`, `Delivered-To`, `X-Original-To`
+- the user's own Message-IDs
+- maildir paths and account keys
+
+The guarantee is that **the tool cannot disclose an identifier it was never
+given**. It holds no matter what a later subcommand does, which is why it is
+enforced at extraction rather than at submission: a strip-on-submit rule
+depends on every future sending path remembering, and this one does not.
+
+**The report is reviewable and hand-editable before anything is sent.** What
+the tool derives and what the user chooses to disclose are separate things.
+The user can add facts during review that the tool never held, including how
+many of their addresses a campaign hit, which is a fact they supply from their
+own knowledge.
+
+### Tracking parameters are recipient identifiers
+
+A phishing URL commonly carries the recipient's identity in its query string:
+`?e=you@example.org` plaintext, `?u=<base64 of the address>`, or
+`?id=<md5 of the address>`, which a wordlist or a targeted guess reverses.
+Publishing that to four third parties is the same leak as publishing the `To`
+header, hidden one level down.
+
+It also **deanonymises the reporter to the attacker**. Abuse desks forward
+reports to their customers and URLhaus is a public feed, so a kit operator
+watching for their own URLs learns which target reported them. For a
+consultant the parameter may carry a CLIENT's identifier rather than the
+user's own.
+
+**So: keep scheme, host and path; keep parameter NAMES; redact parameter
+VALUES.**
+
+```
+http://login-example.invalid/verify?id=REDACTED&src=REDACTED
+```
+
+This costs the report almost nothing. Parameter names are part of the kit's
+fingerprint and campaigns correlate on infrastructure, not on per-victim
+tokens. The token is unique per recipient BY DESIGN, so it is the one part of
+the URL that cannot correlate anything, and keeping it makes dedup actively
+worse: two messages from one campaign would look like different URLs.
+
+**Path segments carry the token too.** `/verify/ZGFuaXhAZXhhbXBsZS5vcmc/` is
+common and a query-string-only rule misses it entirely. A path segment that
+looks like base64 or a long hex string is FLAGGED FOR REVIEW rather than
+silently redacted, since a path may also be meaningful.
+
+The full URL survives in `source.eml` either way, so the evidence exists
+locally; it is simply not what gets published by default, and the review gate
+allows pasting one in by hand when a particular desk genuinely needs it.
+
+## Components
+
+```
+abusectl/
+ cli.py argparse dispatch, exit codes. No logic.
+ case.py case dir: create, load, save manifest, atomic writes
+ parse.py .eml -> IOCs stdlib only, pure
+ contacts.py IOCs -> abuse contacts (RDAP) network, read-only
+ report.py IOCs + contacts -> bodies pure
+ submit.py bodies -> MISP, then vendors network, writes
+ destinations/ one module per target
+ misp.py abusedb.py urlhaus.py virustotal.py email.py
+ config.py ~/.config/abusectl/config.toml
+```
+
+**`case.py` is the spine.** Every subcommand goes through it and nothing else
+writes the case directory. It owns the manifest schema and its version, and it
+writes ATOMICALLY, temp file plus rename, the way `mailrules.py` saves
+`rules.json`. A half-written manifest during a review is a corrupted evidence
+record.
+
+**`parse.py` is stdlib-only and pure.** Bytes in, IOC list out; no socket, no
+config read. That is what makes it testable against fixtures with no setup.
+
+**`contacts.py` and `submit.py` are the only network modules**, and both take
+an INJECTED TRANSPORT. Not an abstraction for its own sake: it is the seam
+that lets the irreversible path be tested without sending anything.
+
+**`destinations/` is one module per target**, each exposing the same two
+functions, build a body and send one. A fifth vendor is a new file rather than
+an edit to `submit.py`. This is the one place a plugin shape earns itself,
+because there are four known members with genuinely different APIs. There is
+no discovery mechanism; it is a package with four members.
+
+**`config.py`** reads TOML through stdlib `tomllib`, no dependency. It holds
+the cases path, the MISP URL and key, vendor keys, the user's reporting
+identity for X-ARF, and the trusted-relay boundary described below. Secrets
+live in a file with mode `0600`, checked on load, and **never reach the
+manifest or a log**.
+
+### Deliberately absent
+
+No database of its own: MISP is the database and case directories are the
+local record. No daemon, no queue, no retry scheduler, since retry is the user
+running `submit` again.
+
+### Dependencies by part
+
+| Part | Needs |
+|---|---|
+| `parse` | stdlib only |
+| `contacts` | an HTTP client |
+| `report` | stdlib only |
+| `submit` | HTTP client, PyMISP |
+
+`requirements.txt` is empty until `contacts` is built. **Development runs from
+a venv in the repository**, gitignored; packaging and distribution are
+deliberately out of scope until the tool does something worth installing. A
+SlackBuild in `my-slackbuilds` is the eventual shape, with an nvchecker
+stanza, matching every other tool the user runs.
+
+## Data flow
+
+```
+msg.eml
+ | parse creates case, writes source.eml + iocs[]
+ v
+case dir --------------------------------------------------+
+ | contacts RDAP per IOC, writes contacts[] |
+ v | every step
+case dir | reads and
+ | report writes bodies/, destinations[] | rewrites
+ v | manifest.json
+case dir -- [ REVIEW: qtmaildir dialog or $EDITOR ] -------+
+ | submit |
+ +--> MISP ------ fails? STOP, nothing sent --------------+
+ +--> vendors + abuse desks, per-destination status
+```
+
+**Review edits the BODIES, not the IOCs.** Once `report` has run the bodies
+are the artifact, so `report` refuses to run again on a case whose bodies were
+modified after generation unless forced, by a timestamp check in the manifest.
+Silently discarding a review that took twenty minutes is what makes a tool
+untrustworthy.
+
+**`submit` is resumable.** It reads per-destination status and skips anything
+already `sent`, so running it twice is safe. This matters because the failure
+that will actually be hit is a rate limit, not a bug.
+
+## Manifest schema
+
+```json
+{
+ "format": 1,
+ "case_id": "2026-09-08-a3f1",
+ "created": "2026-09-08T12:31:04Z",
+ "source": "source.eml",
+
+ "iocs": [
+ { "id": "ioc-1", "type": "ipv4", "value": "203.0.113.42",
+ "origin": "received-chain", "confidence": "untrusted-hop" },
+ { "id": "ioc-2", "type": "domain", "value": "login-example.invalid",
+ "origin": "href" },
+ { "id": "ioc-3", "type": "url",
+ "value": "http://login-example.invalid/verify?id=REDACTED",
+ "origin": "href-html" },
+ { "id": "ioc-4", "type": "sha256", "value": "e3b0c442...",
+ "origin": "attachment", "filename": "invoice.pdf" }
+ ],
+
+ "auth": { "spf": "fail", "dkim": "none", "dmarc": "fail" },
+
+ "contacts": [
+ { "ioc": "ioc-1", "abuse": "abuse@example.invalid",
+ "source": "rdap", "handle": "AS64496" },
+ { "ioc": "ioc-2", "abuse": null, "source": "rdap",
+ "error": "no abuse contact published" }
+ ],
+
+ "destinations": [
+ { "id": "misp", "kind": "misp", "iocs": ["ioc-1","ioc-2","ioc-3"],
+ "body": null, "status": "pending" },
+ { "id": "abusedb", "kind": "api", "iocs": ["ioc-1"],
+ "body": "bodies/abusedb.json", "status": "pending" },
+ { "id": "rdap-1", "kind": "email", "iocs": ["ioc-1"],
+ "target": "abuse@example.invalid",
+ "body": "bodies/rdap-1.xarf", "status": "pending" }
+ ]
+}
+```
+
+After a submit in which one destination failed:
+
+```json
+{ "id": "abusedb", "status": "sent",
+ "sent_at": "2026-09-08T12:40:11Z", "receipt": "8891234" },
+{ "id": "virustotal", "status": "failed",
+ "attempted_at": "2026-09-08T12:40:12Z", "error": "429 rate limited" }
+```
+
+### Five schema decisions, and why
+
+**IOCs carry an `id` and everything references it.** Contacts and destinations
+point at `ioc-2` rather than repeating the value, so there is one place to
+correct it and a destination cannot drift from the IOC it reports.
+
+**`origin` on every IOC** names where it came from: `received-chain`, `href`,
+`attachment`. During review the user needs to know whether an IP came from a
+header to trust or one the attacker wrote. Without it, review is guesswork.
+
+**`confidence: untrusted-hop`** exists because of the `Received` trap below.
+An IP from beneath the trusted-relay boundary is attacker-supplied, and the
+manifest says so rather than presenting it as fact.
+
+**`status` is per destination**, never one status for the case. There is no
+single answer when four destinations disagree.
+
+**MISP is a destination like the others**, with `body: null` because PyMISP
+builds its own payload. That keeps `submit` one loop with one ordering rule
+rather than a special case beside a loop.
+
+## Error handling
+
+**Partial failure is the normal case, not an exception.** Four destinations
+can disagree, so each carries its own status and a retry sends only what did
+not land. All-or-nothing would be wrong here: the successful submissions
+really happened and cannot be recalled, so recording them as failed would make
+the next attempt duplicate real reports.
+
+**A MISP failure aborts before anything irreversible runs**, per the ordering
+rule above.
+
+**A missing abuse contact is not an error.** RDAP publishes none for many
+netblocks. The IOC keeps `"abuse": null` with the reason, the destination is
+simply not created, and review shows what could not be resolved.
+
+## Testing
+
+`parse` is fixtures in, JSON out, so the whole reversible half is testable
+offline with no network and no keys. The network parts get their transport
+stubbed through the injected seam.
+
+**Fixtures use `example.org` and `.invalid` addresses and generic
+placeholders**, per the user's standing personal-data rule. Real phishing
+samples do not go in the repository: they carry the recipient identifiers this
+tool exists to keep out of reports, and a repository is potentially public.
+
+Python conventions follow `assets/hooks/` in qtmaildir: GPLv2 header on every
+file, `unittest`, a `test_<module>.py` beside each module, no framework.
+
+## `parse` in detail: the first part to build
+
+Extracted:
+
+- **Sending IP** from the `Received` chain: the last UNTRUSTED hop
+- **Sender domains**: `Return-Path`, `From`, `Reply-To` where it differs
+- **Href domains and full URLs** from HTML and text parts, decoded, redacted
+ per the rule above, never fetched
+- **Redirect chains** as DECLARED in headers and href text, never followed
+- **Auth results**: `Authentication-Results`, and the SPF/DKIM/DMARC verdicts
+ as the receiving server recorded them
+- **Attachment filenames and SHA-256 hashes**, since a hash is reportable and
+ a filename is an indicator
+
+Never captured: recipients, the user's Message-IDs, maildir paths, account
+keys.
+
+### Two traps, which is where this kind of parser goes wrong
+
+**The `Received` chain is attacker-controlled below the user's own
+infrastructure.** Headers can be forged wholesale, and only the hops the
+user's own MTA added are trustworthy. Without a configured trusted-relay
+boundary, the "sending IP" is whatever the attacker chose to write. The
+boundary is config; IOCs below it are marked `untrusted-hop` rather than
+presented as fact.
+
+**Never resolve and never fetch.** Not the URLs, not the redirects, not remote
+images. Following a link confirms the address is live to the sender and fires
+exactly the tracker the message wanted. This is a parser over bytes, and that
+is a safety property rather than a performance choice.
+
+## The qtmaildir half
+
+Specified in item 194 and built after 187 and 190, which settle what Mark spam
+does and put it on the message bar. It is a review dialog over the case
+directory: show the bodies, allow editing, show what could not be resolved,
+and call `submit` on approval, then show per-destination results with a retry
+for the failures. Sized M in the backlog, not the S a plain button would be:
+a review dialog is real UI, and the user asked for the review to happen in the
+application rather than in an editor.
+
+The wire contract is this document's manifest plus the `abusectl` command name
+in qtmaildir's config, in the shape `[sync] command` and the per-account
+`send_command` already use.
+
+## What each later spec has to settle
+
+- **`contacts`**: RDAP bootstrap and referral chasing, caching policy, rate
+ limits, and what happens for a netblock that publishes no abuse contact.
+- **`report`**: the X-ARF (RFC 5965) schema version, which fields the user's
+ reporting identity fills, and the plain-text alternative for desks that do
+ not parse X-ARF.
+- **`submit`**: per-vendor auth and payload shapes, MISP event structure
+ (one event per case, or per campaign, and how the dedup query decides),
+ and how mail to an abuse desk is sent, which most likely reuses qtmaildir's
+ own `send_command` rather than adding an SMTP client.
+- **The qtmaildir dialog**: its own item, after 187 and 190.