Blog

The Week Phish-Signals Stopped Being a Mirror

← Back to Blog

Last week I wrote about pulling the phish-report engine into its own package, @farksecurity/phish-signals, with a git-subtree mirror quietly keeping a public repo in sync with the private one behind the scenes. That setup worked. It also turned out to be the wrong shape for where this was actually going, and in the three days since, I tore it out and replaced it with something better.

Killing the mirror instead of building around it

The mirror existed to solve a problem that, on inspection, wasn't really there. Nothing in this site's live app actually imports @farksecurity/phish-signals. routes/ and lib/ run their own copy of the detection logic, same as before, same as after. The mirror wasn't a real dependency relationship, it was a publish vehicle: a way to get code that lived inside this site's private monorepo out onto npm without maintaining two copies by hand. That's a fine thing to want, but it also meant the public repo's entire history was an artifact of a subtree split, one step removed from anything a contributor could actually work against directly.

So PR #13 deleted it. packages/phish-signals/ is gone from this repo, and so is the mirror workflow. github.com/MasonC-402/phish-signals is now the actual source of truth, with its own git history that isn't a derived artifact of anything else. It's a real project now, not a publish target.

Building a second implementation that has to agree with the first

Once it was standalone, the obvious next question was whether it should stay TypeScript-only. A lot of the people who'd want a phishing-detection engine as a library are working in Python, so I built a full port: python/ alongside typescript/, each a complete, independently installable package, not a thin binding wrapped around one shared core.

"Port" undersells what that actually took. This isn't translating syntax from one language to another and calling it done. It's writing a second implementation, by hand, using a different language's libraries for punycode decoding and MIME parsing and everything else, that has to land on the exact same verdict as the first: the same signal ids, the same categories, the same severities, on the same input. Two people can each reason correctly about how to count Received headers or which Unicode scripts count as confusable, and still land in slightly different places on some edge case neither of them thought to test. That's not a bug in either implementation individually. It's just what happens when the same logic gets written twice, and it's exactly the kind of drift that's invisible until someone happens to run the same input through both and notice they disagree.

The fix was conformance/: a set of language-neutral JSON vectors, input in, expected findings out, checked into the repo and run by a small test harness in each language. A vector that only one language's harness satisfies is a failing test in CI, the same day it's written, not something discovered months later when someone compares outputs by hand and finds a mismatch. The Python port went through the same PR-reviewed sequence the original did: scaffolding first, then the rule engine, then the checks, aggregation, and parsing layers with a full test suite behind each one (PRs #1, #2, and #4). Getting the TypeScript side to full conformance parity was its own separate piece of work (PR #5), which is worth saying plainly: the original wasn't fully conformant with the standard I was now holding both implementations to either. Building the second one forced real fixes in the first.

A feature that quietly came along for the ride

Both languages also picked up KQL query generation: typescript/src/kqlQuery.ts and python/src/phish_signals/kql_query.py, each with full test coverage. Give it whatever a message analysis actually found, and it generates a KQL query for Microsoft 365 Defender or Sentinel Advanced Hunting, targeting whichever of EmailEvents, EmailUrlInfo, EmailAttachmentInfo, UrlClickEvents, DeviceNetworkEvents, and DeviceFileEvents are relevant to what was found.

This isn't new. It started life on this site directly, in /phish-report and the /tools/kql-builder and /tools/kql-library pages, before getting ported into the package the same way the rest of the engine originally was. I'm flagging it here specifically because it's easy to miss: the one-line description on npm and PyPI only calls out Sigma rule generation. If you went looking at the package listing alone, you'd have no reason to know KQL generation exists at all.

The security patch, as an example of what maintenance actually looks like

On the 24th, npm audit flagged something real: deepmerge-ts, a transitive dependency pulled in through mailparser by way of html-to-text, had a disclosed stack-exhaustion denial-of-service vulnerability on recursive object graphs (GHSA-ggr8-5vv4-36mx). That's not a hypothetical concern for this specific project. mailparser is what parses attacker-controlled raw .eml content, which is the entire input surface this engine exists to analyze. A crafted malicious message is exactly the shape of input that could trigger something like this.

The fix ended up being small: bump mailparser from 3.9.15 to 3.9.16, already satisfied by the existing ^3.9.15 range in package.json, so no manifest change, no --force, no major version bump to justify. That pulled in the patched html-to-text 10.0.1 and deepmerge-ts 8.0.2 as transitive updates. npm audit went to zero vulnerabilities, and the full test suite, 186 tests, still passed clean afterward. I'm writing this one down not because it was hard, it wasn't, but because "small, fast, boring" is exactly what a dependency patch should look like when the surrounding test suite and audit tooling are doing their job.

Docs that don't drift either

The last piece was documentation, and it had the same drift problem as the code, just made of prose instead of logic. python/docs/rules.md was originally the only place the rule engine got explained at all, and a lot of what it said, what a Rule, RuleContext, and Ruleset mean, what the engine guarantees, the declarative JSON rule format, how it interacts with scoring, wasn't actually Python-specific. It was a description of the engine itself, sitting in one language's docs folder, with no mechanism keeping it in sync with what the TypeScript engine actually did.

So I split it apart. rules/CONCEPTS.md now holds the engine model written once, shared by both languages. python/docs/rules.md got trimmed down to just Python's API surface. A new typescript/docs/rules.md mirrors that same structure for the TypeScript side, and every code example in it is verified to actually run rather than hand-derived from memory. It also says plainly that the rule-engine types themselves, Rule, RuleContext, Ruleset, evaluateRuleset, loadRuleFile, loadRuleset, aren't part of the published npm package's public API yet. package.json's exports map only exposes the top-level entry point. That's a real, true-today limitation, and it's stated as one rather than glossed over.

Both sets of docs then got shipped inside the actual installed packages, the npm tarball and the Python wheel, not just left sitting in the repo where only someone who went looking on GitHub would find them. That happened in the two most recent commits, which is also where the npm package landed at 0.2.1 and the PyPI package at 0.2.2.

Where it stands now

@farksecurity/phish-signals on npm is at 0.2.1, with a full version history from 0.1.0 running from the 16th through the 24th. phish-signals on PyPI is at 0.2.2, published the 24th, one step ahead of npm right now: 0.1.0 on the 22nd, 0.2.0 on the 23rd, then 0.2.1 and 0.2.2 both on the 24th. Both registries use OIDC trusted publishing with no stored tokens, and each language releases off its own tag prefix, npm-v* for TypeScript and pypi-v* for Python, specifically so a release in one language never fires the other's pipeline by accident.

None of this changes what the engine finds when you point it at a phishing email. What changed is everything underneath that: a second implementation that has to keep agreeing with the first instead of just resembling it, a dependency patched the same day it mattered, and documentation that explains the engine once instead of twice. Less new detection logic this stretch, more of the infrastructure that makes two languages and one project actually trustworthy together. Same theme as last week, just one layer further down.