| Age | Commit message (Collapse) | Author | Files | Lines |
|
report must never open source.eml, so parse decides once what may be
published and report formats only what it is given. To, Cc, Delivered-To and
X-Original-To are absent by construction rather than stripped.
Received is cut to the boundary hop alone, in both directions. Above it are
our own relays; below it is the attacker's own writing, and a forged chain
names an innocent third party there, so publishing a hop below the boundary
puts someone else's address into a report a desk will act on. That is the
third property applied to disclosure rather than to sending_ip().
Truncating the chain was not sufficient on its own: the surviving line is
written by our own relay and records the envelope recipient in its optional
"for <addr>" clause, so the whitelist alone would have published the
victim's address verbatim in the one header a report reproduces in full.
The clause is removed and the rest of the hop kept.
The manifest-wide "no example.org" assertion is narrowed to the headers
block only, where the whitelist deliberately publishes our receiving relay's
name in a by/authserv-id clause. The address itself is still barred there,
asserted separately.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Xj1ayFRSUQ2u7cwb3S4axE
|
|
A sweep of the user's real spam found List-Unsubscribe naming a domain
that appeared nowhere else in the message. It is attacker infrastructure
and was going unreported.
Every url from that header goes through redact.url() like a body url: an
unsubscribe link has to say who is unsubscribing, which makes it one of
the likeliest carriers of a recipient token. mailto: entries are skipped
rather than redacted, since the address is the whole value and nothing
useful survives removing it.
Sender is collected on the same terms as Reply-To, included only when it
differs from From. One repeating From is noise; one naming a separate
relay is the infrastructure behind the run.
Also drops the unused urlencode import left in redact.py when
_redact_kv_string stopped using urllib to rebuild the query string.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019NHaqA1Rz5ybed7wFUeQbK
|
|
_ADDR_DOMAIN.search() returned the first @domain anywhere in the raw
header text. A display name sits before the angle brackets and is
attacker-controlled, so it won.
Two ways that reached a published report. The sender was misattributed:
"Billing at billing@innocent.example" <phish@sender.example.invalid>
filed the report against a third party who sent nothing. And it defeated
the structural guarantee in sender_domains(): the module reads no
recipient header, but an attacker who writes the victim's own address
into the display name hands it one anyway, and it came back out as a
sender domain.
parseaddr() parses the header grammar rather than scanning it, so a
quoted display name cannot supply the address.
leaky.eml's display name now carries the recipient address. The existing
test_no_ioc_holds_a_recipient_address assertion catches this class; it
was green before only because the fixture used a harmless domain.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019NHaqA1Rz5ybed7wFUeQbK
|
|
test_no_ioc_holds_a_recipient_address grepped for the literal
"example.org", and every existing fixture hides the recipient address
as base64, so URL redaction could be disabled entirely and both this
test and test_cli's counterpart stayed green.
leaky.eml carries you@example.org in five URL shapes plus a From
display-name trap. Adding it to the fixture list, plus a new test
asserting on the raw address and its percent-encoded form, turns the
suite red: the valueless-query-parameter defect (redact.py) currently
lets ?victim@example.org through as a kept parameter NAME. Left
failing on purpose; the next commit fixes redact.py and turns it
green.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KphFXTc2QajxXsHWyvGJ4R
|
|
Four hand-written messages using example.org, example.invalid and the
RFC 5737 documentation IP ranges. No real phishing sample goes in this
repository: it would carry the recipient identifiers this tool exists to
keep out of reports, and a repository is potentially public.
forged-chain.eml is the one that matters. The attacker prepends two
Received headers naming an innocent third party, so a parser that walks
past the trust boundary reports 198.51.100.7 rather than 203.0.113.99.
Weekdays verified with date(1) rather than written from memory, since an
RFC2822 parser validates the day against the date and a wrong one is
indistinguishable from a malformed header.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KphFXTc2QajxXsHWyvGJ4R
|