diff options
Diffstat (limited to 'docs/BACKLOG.md')
| -rw-r--r-- | docs/BACKLOG.md | 67 |
1 files changed, 67 insertions, 0 deletions
diff --git a/docs/BACKLOG.md b/docs/BACKLOG.md index 79c0ab1..780cbf5 100644 --- a/docs/BACKLOG.md +++ b/docs/BACKLOG.md @@ -7,6 +7,8 @@ number and gains a status rather than being renumbered. |---|------|------|--------| | 1 | Skip boilerplate namespace URLs | XS | open | | 2 | An IDN indicator resolves to no contact | S | open | +| 3 | Expose kept cases so qtmaildir can tag spam | ? | open, unsized | +| 4 | `Report-Type: phishing` is unverified against x-arf | XS | open | ## 1. Skip boilerplate namespace URLs @@ -67,3 +69,68 @@ where the attacker wants the tool to normalise on their behalf, and a consultant chasing one indicator by hand is a smaller cost than a query made about a name the user never saw. Wait for a real IDN indicator in a sweep before building it. + +## 3. Expose kept cases so qtmaildir can tag spam + +**Source.** The author's idea note, not a defect found in the code. Unlike +items 1 and 2 the cause here has NOT been verified against the code, because +there is nothing built yet to verify: this is a feature request, and it is +recorded unsized on purpose. + +**Observed.** Case directories are permanent by design, so over time they +become a local corpus of messages the user has already judged to be phishing. +Nothing reads them back. The idea is that qtmaildir could ask this tool +whether an incoming message resembles one, and tag it as spam when it does. + +**Approach.** Undecided, and the shape matters more than the code. The +umbrella design already fixes the coupling between the two repositories: the +manifest format and a command name in qtmaildir's config, with no submodule. +A read-only subcommand answering a question about one message fits that +contract; a daemon, a socket or a shared database does not, and the umbrella +design rules out a database of this tool's own. + +**Constraints, and the real tension.** Deciding a message is spam by +resemblance is a classifier, and this tool has so far been deliberately +mechanical: it reports what a message declared, and refuses rather than +guesses when the trust boundary is unset. A resemblance score is the first +thing here that would be an opinion rather than an observation, and a wrong +one either hides real mail or teaches the user to distrust the tag. + +There is also a quieter question about what a match is allowed to be based on. +The obvious signals are the ones already in a manifest, a sending IP, a +domain, a URL shape, an attachment hash. Those are safe. Matching on the +message body would mean holding attacker-supplied text against new mail, and +`source.eml` is unredacted, so anything built here must not become a route by +which a stored recipient identifier reaches a comparison that is later +reported or logged. Property 1 governs what may be published, and a tag is not +a report, but the path from one to the other is short. + +**Before building.** Ask the author what "fits certain requisites" means to +him concretely, since that phrase is doing all the work in the note, and +whether he wants a judgement or only the facts, for instance a subcommand that +answers "this IP appears in three kept cases" and leaves the tagging decision +to qtmaildir. The second is much more in keeping with the rest of the tool. + +## 4. `Report-Type: phishing` is unverified against x-arf + +**Observed.** `report.feedback_fields()` emits `Report-Type: phishing` in the +machine-readable part. Every other field there was verified against RFC 5965 +itself; this one was not, because no primary source for x-arf's own field +semantics could be reached while building it. The abusix README documents only +the v3 to v4 deprecation and does not define the field. + +**Cause.** Not a defect found in the code. The value follows the worked +example in `docs/specs/2026-09-09-report.md`, so it is internally consistent, +and the hybrid envelope means a strict RFC 5965 parser ignores the field +either way (the RFC requires implementors ignore fields they do not support). +The exposure is limited to x-arf tooling reading a field name or value that +does not exist in the version it implements. + +**Approach.** Find a primary source for x-arf v4 field names, confirm or +correct the value, and record what it was checked against. If x-arf turns out +to name the field differently, the fix is one string and one test. + +**Constraints.** Low urgency: nothing here is a leak, and the failure mode is +a field an x-arf parser skips rather than acts on wrongly. Worth doing before +the first real report is filed, so a desk running x-arf tooling gets what it +expects. |
