email-phishing-detector: The Phishing Filter That Trusts Both a Model and a Red Flag Checklist

A full-stack email scanner that pairs TF-IDF logistic regression with frontend heuristics, then wraps the whole thing in a cyber-terminal interface that explains risk instead of just naming it.

8 min read • View on GitHub • More from Ishita197397

A cyber-terminal desk scene where one email splits into two paths. One path feeds a model box that extracts features and returns a classification, while the other path underlines suspicious phrases and assembles a visible threat meter. The image explains the repo’s core idea: phishing detection as two coordinated judgments, not one black-box answer.
The repo’s real trick is not detection alone. It is the fusion of a statistical classifier and a rule-based suspicion layer into one triage screen.
Key Takeaways

Most phishing tools stop at a score. This one adds a second opinion, and that changes the product. In email-phishing-detector, the backend predicts, but the frontend also interprets, so the user gets something closer to an analyst’s readout than a bare model result.

The real product is a two-layer judgment call

The architecture is simple in the best way. One path runs the classifier. The other path runs a visible heuristic pass that looks for red flags and turns them into a separate threat score. That split matters because phishing review is rarely just about whether an email is malicious. It is about why it looks suspicious right now.

A single email is judged twice. The model gives one answer, while the heuristic layer explains the risk in human terms.

Why the frontend does more than decorate

The frontend’s analysis layer is the surprise. According to the repo’s `generateMockAnalysis` logic, it scans for signals such as urgency, credential requests, and social engineering patterns, then assigns points to a threat score. That is not fake theater. It is a lightweight interpretation layer that makes the verdict legible.

const hasUrgency = /urgent|immediately|verify now/i.test(text);
const hasCredentialRequest = /password|login|account verification/i.test(text);
const hasSocialEngineering = /click here|limited time|suspended/i.test(text);

let threatScore = 0;
if (hasUrgency) threatScore += 20;
if (hasCredentialRequest) threatScore += 25;
if (hasSocialEngineering) threatScore += 30;
A close-up of a hand placing signal cards into a scoring tray. Each card represents urgency, credential requests, suspicious links, or social engineering cues. A sealed envelope labeled ML prediction sits beside them, showing that the final score is assembled from both rules and the model, not from the model alone.
The threat score is a fusion of weak signals. That is what makes the interface feel more like triage than automation.

The ML pipeline is classic on purpose

Under the hood, the model stack is intentionally familiar: regex cleanup, TF-IDF vectorization, and logistic regression. The vectorizer caps vocabulary at 5,000 features, which keeps the model lightweight and avoids letting rare noise dominate the decision. In a repo like this, that restraint is a feature. The point is to provide a dependable baseline that the UI can explain, not to win a benchmark.

ApproachStrengthWeaknessWhat email-phishing-detector does differently
Plain TF-IDF + logistic regression spam detectorFast and easy to trainHard to explain to a userAdds a visible heuristic layer that interprets risk
Black-box phishing demoCan look impressiveUsers cannot see why an email was flaggedShows cues like urgency and credential requests in the UI
Rule-only phishing checklistEasy to inspectMisses statistical patternsUses ML for classification and rules for triage

A terminal UI turns classification into triage

The cyber-terminal styling is not cosmetic fluff. It changes the user’s job from “trust this prediction” to “review this assessment.” Sender, subject, content, and links all sit inside one analyst-style workspace, which makes the product feel built for fast review under uncertainty. That framing is the reason the hybrid scoring system lands. It looks like a tool for decision-making, not a demo for model inference.

What the repo reveals about practical security tools

The best security tools often mix statistical prediction, hand-authored heuristics, and a presentation layer that makes uncertainty obvious. This repository does exactly that. It is small, but it captures a real pattern: the model is only half the product. The other half is how you tell a person what to do with the result.