email-phishing-detector: The Phishing Filter That Trusts Both a Model and a Red Flag Checklist
A full-stack email scanner that pairs TF-IDF logistic regression with frontend heuristics, then wraps the whole thing in a cyber-terminal interface that explains risk instead of just naming it.
- The repo treats phishing detection as a two-layer judgment: a backend classifier and a frontend threat score that explains why an email looks dangerous.
- Its frontend is not decorative, because regex-based cues like urgency and credential requests actively shape the user’s risk perception.
- The machine learning pipeline stays intentionally simple with TF-IDF and logistic regression, which makes the system lightweight and defensible.
- The cyber-terminal UI turns a prediction into triage, which is the real product distinction here.
Most phishing tools stop at a score. This one adds a second opinion, and that changes the product. In email-phishing-detector, the backend predicts, but the frontend also interprets, so the user gets something closer to an analyst’s readout than a bare model result.
The real product is a two-layer judgment call
The architecture is simple in the best way. One path runs the classifier. The other path runs a visible heuristic pass that looks for red flags and turns them into a separate threat score. That split matters because phishing review is rarely just about whether an email is malicious. It is about why it looks suspicious right now.
Why the frontend does more than decorate
The frontend’s analysis layer is the surprise. According to the repo’s `generateMockAnalysis` logic, it scans for signals such as urgency, credential requests, and social engineering patterns, then assigns points to a threat score. That is not fake theater. It is a lightweight interpretation layer that makes the verdict legible.
const hasUrgency = /urgent|immediately|verify now/i.test(text);
const hasCredentialRequest = /password|login|account verification/i.test(text);
const hasSocialEngineering = /click here|limited time|suspended/i.test(text);
let threatScore = 0;
if (hasUrgency) threatScore += 20;
if (hasCredentialRequest) threatScore += 25;
if (hasSocialEngineering) threatScore += 30;
The ML pipeline is classic on purpose
Under the hood, the model stack is intentionally familiar: regex cleanup, TF-IDF vectorization, and logistic regression. The vectorizer caps vocabulary at 5,000 features, which keeps the model lightweight and avoids letting rare noise dominate the decision. In a repo like this, that restraint is a feature. The point is to provide a dependable baseline that the UI can explain, not to win a benchmark.
| Approach | Strength | Weakness | What email-phishing-detector does differently |
|---|---|---|---|
| Plain TF-IDF + logistic regression spam detector | Fast and easy to train | Hard to explain to a user | Adds a visible heuristic layer that interprets risk |
| Black-box phishing demo | Can look impressive | Users cannot see why an email was flagged | Shows cues like urgency and credential requests in the UI |
| Rule-only phishing checklist | Easy to inspect | Misses statistical patterns | Uses ML for classification and rules for triage |
A terminal UI turns classification into triage
The cyber-terminal styling is not cosmetic fluff. It changes the user’s job from “trust this prediction” to “review this assessment.” Sender, subject, content, and links all sit inside one analyst-style workspace, which makes the product feel built for fast review under uncertainty. That framing is the reason the hybrid scoring system lands. It looks like a tool for decision-making, not a demo for model inference.
What the repo reveals about practical security tools
The best security tools often mix statistical prediction, hand-authored heuristics, and a presentation layer that makes uncertainty obvious. This repository does exactly that. It is small, but it captures a real pattern: the model is only half the product. The other half is how you tell a person what to do with the result.