Chetu248/Candidate_Ranking_System: How a Resume Screener Became a Heuristic Compiler

A LightGBM pipeline turns 100,000 candidates into a ranked shortlist by encoding recruiter taste, feature engineering, and explanation strings into one opinionated system.

8 min read • View on GitHub • More from Chetu248

A wide editorial illustration of a paper funnel that accepts hundreds of resume sheets and releases only a few through a narrow opening. The funnel is built from visible filters that suggest geography, experience, product background, domain depth, and platform trust signals. It explains that the system is selective by design, not neutral by default.
This repository does not try to rank everyone fairly. It tries to compress a very specific hiring philosophy into a machine that can sort at scale.
Key Takeaways

The system does not search for talent. It searches for a profile.

The most important thing about this repo is not the model. It is the target. The pipeline is tuned for a sharply defined Senior AI Engineer profile: India-based, roughly 5 to 9 years of experience, product-first, and deep in AI/ML work. That makes it less like a universal recruiter and more like a machine for enforcing taste.

That choice changes how you should read everything else. The 71 engineered features are not trying to capture the totality of a person. They are trying to make a preferred hiring pattern machine-readable, then rank a large pool against it.

A close-up editorial illustration of one candidate dossier assembled from modular cards. On one side are raw attributes such as title, years of experience, geography, niche skills, and platform signals. On the other side, those signals are stitched into a compact reasoning line that looks assembled rather than generated. It explains how structured features become a human-readable justification.
The surprising output is not just a score. It is a score plus a sentence that tries to justify why the candidate landed where they did.

The ranking engine explains itself

The CSV output is where the project becomes unusual. Instead of leaving the reader with a raw score, the system emits a reasoning field that reads like a compact justification. In the research notes, a top candidate is described as having a strong embedding and retrieval stack, specific experience fit, and a core AI/ML title. That is not just reporting. It is narrating the ranking.

candidate_id,score,reasoning
CAND_0002025,0.9916,"Core AI/ML engineering title; 5.9 yrs exp; 70 months deep AI/ML career; strong embedding & retrieval stack (5 skills)"
CAND_0001844,0.9841,"Product-first background; experience fit; India-based; LLM/RAG/MLOps signals aligned with ideal profile"

That matters because it gives the recruiter something closer to a review trail than a black-box score. The explanation is still synthetic, and it still reflects the system’s values, but it makes those values visible.

The pipeline is a translation layer. Raw attributes become features, features become a tier, and the tier becomes a sentence a human can skim.

Inside the notebook, the rubric becomes a model

The notebook reads like a compressed machine-learning competition submission. It loads JSONL records, engineers a large feature set, assigns synthetic tiers, trains a LightGBM multiclass classifier, and writes out the final shortlist. The project description says the system handles 100,000 candidates, uses 71 engineered features, and ranks them with sub-millisecond inference per candidate.

The important part is not that it learns from a mystical ground truth. It learns from a codified rubric. When labels are generated from a strict hiring philosophy, a tree model can reproduce that philosophy with striking consistency.

# Conceptual pipeline
records = load_jsonl("candidates.jsonl")
features = engineer_features(records)  # 71 signals
labels = assign_tiers(features)        # synthetic rubric tiers
model = LGBMClassifier(num_leaves=..., n_estimators=300)
model.fit(features, labels)
ranked = model.predict_proba(features)
explanations = build_reasoning_strings(features, ranked)

That is why the reported scores look almost too clean. A model cannot be more faithful to the rubric than the rubric itself. If the rubric is highly specific, the model can appear almost perfect.

Why LightGBM is the right kind of blunt instrument

This is a tabular problem with engineered signals, not a language problem. LightGBM fits that shape well. It handles mixed feature types, trains quickly, and makes inference cheap enough to score a large candidate pool in one pass.

A deep semantic system might sound more modern, but it would be less direct here. This repo wants to rank on explicit criteria such as experience fit, title relevance, and domain depth. A tree ensemble is a better fit for that kind of opinionated spreadsheet logic.

ApproachCore signalStrengthWeaknessWhat this repo does differently
TF-IDF or cosine similarityWord overlap between resume and job descriptionSimple and transparentMisses structured hiring preferencesUses engineered signals instead of text resemblance
Embedding-based rankingSemantic closeness in vector spaceCaptures broader meaningCan blur precise rubric choicesTurns preference into explicit features and tiers
LLM hiring agentPrompted natural-language judgmentFlexible and conversationalHard to audit and reproduceProduces deterministic rank plus reason strings
This repoLightGBM over handcrafted signalsFast, repeatable, explainable enoughBakes in the designer’s biasMakes the rubric the product

The other practical win is speed. Once the feature table exists, scoring 100,000 rows with a tree model is cheap. That makes the system feel like a screening engine rather than a research demo.

The project is more revealing than neutral screeners

What makes the repo editorially interesting is that it does not pretend to be neutral. Many screening tools hide preference inside a similarity score or a prompt. This one declares its preferences in the structure of the pipeline.

That makes it more useful in one narrow sense and more controversial in a broader one. It is useful because the hiring team can see exactly what the machine is optimizing for. It is controversial because the machine is optimizing for a very specific kind of candidate, not for fairness, diversity of path, or general ability.

Screening styleWhat it optimizesHow visible the bias isBest use case
General resume matcherText similarityLowBroad initial triage
LLM reviewerNarrative judgmentMediumSmall batch review with context
This repoCodified recruiter tasteHighNarrow search for a specific senior profile

What this system gets right, and what it leaves out

The strongest thing about the project is its clarity. It is fast, repeatable, and easy to reason about once you accept its assumptions. The explanation strings also give humans a usable bridge back from the model to the ranking.

The weakness is the same thing. A geographically and career-path specific rubric can be effective for a targeted search, but it also hard-codes exclusions. The repo is a clean example of a larger truth in hiring tech: software reflects the people who designed the rulebook.