Two rounds in, the analyzer had a sound scoring model, real delivery-path analysis, and output built for someone to actually act on. This pass came from asking a blunter question about the part I'd leaned on the most: how much do I actually trust the Authentication-Results header, and have I earned that trust or just assumed it.
The header I'd been taking on faith
Every SPF, DKIM, and DMARC finding this tool produces comes from one place: reading whatever the Authentication-Results header says and reporting it back. That's the correct design, because this tool makes no DNS queries of its own and never will, that's the whole "nothing leaves this server" premise. But it means the tool has been trusting that header completely, and I'd never actually sat with what that trust was built on.
Nothing stops someone from writing their own Authentication-Results header into a message before it ever reaches a real mail server. Craft an email, add a line claiming spf=pass; dkim=pass; dmarc=pass from some plausible-looking hostname, and send it. If the tool reading that message never checks who actually wrote the header, a forged pass looks identical to a real one. This is a known technique against automated triage tools, and my own tool had no defense against it whatsoever.
I can't verify the claim cryptographically without making a network call, which I'm not going to start doing. But I don't need to verify it to notice when it doesn't add up. The header names a server that supposedly performed the check, the authserv-id. The message's own Received chain, which I already parse in detail, shows which server actually handled final delivery. Those two should point at the same organization. When they don't, something is off: either a legitimate security gateway sits in the path and stamps its own identity, or the header was written by someone who never touched a real mail server at all.
I compare at the level of the registrable domain rather than the exact hostname, specifically so this doesn't fire constantly on organizations that route mail through several internally-named servers. Gmail's Authentication-Results says mx.google.com; an actual delivery hop might show something like a specific mail-sor-f41.google.com relay. Different hostname, same company, no finding. What does fire is a header naming a server that has nothing to do with the delivery chain at all.
I kept this at medium severity rather than higher, on purpose. A real filtering service in the path is a completely ordinary reason for this mismatch to appear, and I don't want a legitimate corporate mail setup reading as a forgery. It's evidence worth weighing, not a verdict on its own.
Links with nothing for a URL check to look at
Every link check I'd built so far, typosquats, homographs, deceptive anchor text, all of it, works by examining a domain. That's a real blind spot if a link doesn't have one.
An HTML email can link to a data: URI, which encodes content directly into the href itself rather than pointing anywhere. A phishing kit can build an entire fake login page as one of these and put it behind a link with no hostname at all for any reputation-style check to ever examine, mine included. javascript: links are a smaller, blunter version of the same gap: no domain either, and clicking one just runs code directly.
Checking for these turned out to be almost embarrassingly simple once I framed it correctly. I wasn't looking for a bad domain, I was looking for the absence of one in a place a domain should be. A link scanning for href="data: or href="javascript: closes a gap that had nothing to do with how good my domain analysis already was; those links were never in scope for it to begin with.
The gap in my own keyword matching
While I was looking hard at how attackers get past automated tools, I went back and checked whether my own phrase-matching held up the same way, and it didn't.
Every urgency, credential-request, and financial-request phrase this tool looks for gets matched as a plain substring against the message text. That works fine until someone inserts a zero-width character in the middle of a word. "Verify your account" with an invisible character sitting between "ver" and "ify" reads identically to a person and passes right through a mail client's rendering, but it no longer contains the substring "verify" at all, so my own scanning missed it completely.
The frustrating part is that I'd already built the fix, just not aimed at the right target. stripDangerousUnicode existed from the very first version of this tool, stripping exactly these characters from header text before it gets displayed, so a filename can't visually lie about its own extension. I'd built the defense and then never pointed it at the one place it was actually needed for detection to work rather than for display to be honest. The text that gets pattern-matched against my urgency and credential-request phrase lists was reaching that check completely unfiltered.
Once I saw it that way, the fix was a one-line change. The lesson is less about the fix and more about the blind spot: I'd built real infrastructure for exactly this problem and then only wired it into half the places it mattered.
A CSS fix that looked right and did nothing
The last thing this round wasn't a detection at all, it was a bug in the tool's own output, and it's the one that took the longest to actually get right, because my first attempt looked completely reasonable and simply didn't work.
The analyst section of the report, delivery path, indicators, the Sigma rule, is collapsed by default so the page doesn't bury the verdict under detail nobody asked for yet. The obvious problem is that if someone wants to print the report or save it as a PDF for a ticket, a browser only renders what's currently visible on screen. Collapsed content doesn't print. Someone printing an incident report would get the summary and lose everything underneath it.
I wrote what looked like the standard fix: a print-specific CSS rule forcing the hidden content to display: block regardless of whether the section was expanded. It read correctly, I could point to the exact selector doing the exact thing it needed to do, and I almost moved on without checking it against anything real.
I didn't, mostly out of habit at this point, and rendered an actual result page to PDF with the section still collapsed. The delivery path, the indicators, the Sigma rule, all of it, simply weren't there. The CSS I'd written had done nothing.
It turns out collapsed <details> content isn't hidden the way an ordinary display: none element is. Browsers treat it as something closer to not being in the render tree at all when collapsed, and that isn't something an author stylesheet reliably overrides just by matching the selector and adding !important, however correct the selector looks on paper. The actual fix has to reach for the real mechanism: toggle the element's open attribute directly. A small script now does exactly that the moment printing starts, and puts it back the way it was afterward.
What stuck with me here wasn't the CSS trivia, it's that I trusted a fix I never actually watched work. Everywhere else in this project the habit has been to generate a real crafted email and read the real output before believing anything was fixed. This was the one place I skipped that step first, because the change felt too small and too standard to need it, and it was exactly the one that turned out to be wrong.
Where this leaves it
None of this changes the shape of the tool. It still runs entirely offline, still stores nothing, still makes no calls anywhere on its own. What changed is how much weight I'm willing to put on a single header without another line of evidence behind it, and how many of the places I already claim to defend against evasion actually do.