aboutsummaryrefslogtreecommitdiffstats
path: root/docs/plans/2026-09-08-parse.md
AgeCommit message (Collapse)AuthorFilesLines
33 hoursplan: add Namecheap Private Email to the provider tableDanilo M.1-0/+26
Its SPF is a tree of includes rather than a flat list: spf.privateemail.com carries no addresses at all, only includes, and one branch nests a further level. Two of the branches live on jellyfish.systems. The entry here is the flattened union of the five leaf records, deduplicated, 18 networks. The re-verification command lists the leaves rather than the top-level name, and a note says why: querying spf.privateemail.com and finding no ip4 entries looks like a stale record and is not one. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01KphFXTc2QajxXsHWyvGJ4R
33 hoursplan: take the provider ranges from SPF, not from memoryDanilo M.1-16/+88
Every range in the first draft was wrong. They had been written from memory, and the correct source is each provider's own SPF record, which is the list of addresses it declares it sends from. Gmail is the clearest case: the table claimed eleven IPv4 ranges and _spf.google.com publishes two. Fastmail, Proton and the rest were wrong in the same way. Outlook and Zoho are added since they were queried anyway, and the transcription commands are recorded in a comment so the next check is a copy-paste rather than a search. Google's goog.json is deliberately NOT the source, though it looks like one: it lists all Google infrastructure, over a hundred ranges, and using it would trust every Google-hosted service as part of the user's own mail path. _spf.google.com is the mail-sending answer. The table is IPv4 only. An IPv6 hop from one of these providers does not match and the user is asked instead, which is the safe direction: a hop wrongly trusted means the real sender is never reported. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01KphFXTc2QajxXsHWyvGJ4R
33 hoursplan: implementation plan for init and parseDanilo M.1-0/+2322
Thirteen tasks, TDD throughout, stdlib only. parse is pure and offline: the trusted-relay boundary arrives as an argument rather than a config read, so the whole extractor is testable against fixtures with no setup. The plan carries three checks that are not ordinary unit tests. The Received-chain task has a mutation step, because walking one hop too far reports an innocent third party named in a header the attacker wrote, and a test that cannot fail would not protect against it. The URL task runs the suite with sockets refused, so the never-fetch rule is verified rather than read. And every fixture is asserted to leave no recipient address anywhere in the manifest. init exists because parse refuses to guess the trust boundary. It asks for CIDRs, offers a static table of known provider ranges, or reads the chain of a known-good sample and lets the user pick their own hops. A pure builder with the prompts and the flags as two front ends over it, so --non-interactive covers agent-driven setup and the config writing is tested without a terminal. Re-running shows what is already configured and asks; either route backs the old file up first and preserves sections this run does not set, so a later init cannot silently drop a MISP key. The prompts themselves are hand-tested rather than driven from stdin: a test there would assert the wording it was written against and break on a rewording that improved it. Task 12 is the checklist, weighted towards wrong answers. Redirect chains were missing from the first draft of the plan and are now specified: a parameter whose value is itself a URL is recovered as an indicator while every other value stays redacted, which resolves the conflict between reporting the destination and never publishing a tracking token. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01KphFXTc2QajxXsHWyvGJ4R