Brain_Tumor: A Tiny Medical AI That Tries to Explain Its Own Diagnosis

A three-stage pipeline combines classification, segmentation, and Grad-CAM so the model can say what it sees, where it sees it, and why it thinks the scan matters.

8 to 10 min read • View on GitHub • More from keerti-yadav

A clinical workstation with three monitors. One shows a brain MRI, one shows a tumor heatmap overlay, and one shows a segmented mask. A hand points across the screens in sequence, showing a verification loop from prediction to localization to explanation. The scene explains that the project is not just a classifier, but a trust workflow.
The repo’s real pitch is not accuracy alone. It is a chain of evidence: classify, localize, then explain.
Key Takeaways

The black box is the problem

A brain tumor classifier can answer the first question fast: is there a tumor? That is useful, but in medical imaging it is not enough. A good answer also has to show where the signal is and why the model latched onto it.

That is why this repository is more interesting than a typical CNN demo. It is trying to move from prediction to evidence. In a domain where false confidence is expensive, the project’s real value is not the score. It is the attempt to make the score legible.

A close-up split scene shows a heatmap spilling beyond a lesion boundary on one side and a segmentation mask tightly confined to the lesion on the other. A dashed line marks the disagreement zone where the two outputs do not align. The image explains why explanation and segmentation are related but not interchangeable.
When the heatmap and the mask overlap, confidence rises. When they diverge, the model may be leaning on context or artifacts.

Three checks, not one

The repo’s core idea is simple to describe and surprisingly strong in practice. It asks three different questions in one flow: what is it? with classification, where is it? with segmentation, and what evidence did the model use? with Grad-CAM.

The pipeline matters because each stage checks the others. Classification says what. Segmentation says where. Grad-CAM says why the model paid attention.

That matters because the outputs are not redundant. Segmentation is a spatial truth check. Grad-CAM is an evidentiary trace. If they agree, the system looks more credible. If they disagree, the discrepancy itself becomes the signal worth investigating.

What the classifier is actually doing

The classification branch uses EfficientNet-B0 with a custom head. In the repository analysis, that head is a small MLP with dropout, a hidden linear layer, batch normalization, and a final classifier layer. That is a sensible pattern for medical imaging, where the dataset is often small and overfitting is the default failure mode.

self.classifier = nn.Sequential(
    nn.Dropout(0.4),
    nn.Linear(1280, 256),
    nn.ReLU(),
    nn.BatchNorm1d(256),
    nn.Linear(256, num_classes)
)

criterion = nn.CrossEntropyLoss(label_smoothing=0.1)

The point of that setup is restraint. Dropout discourages brittle memorization. Batch normalization steadies training. Label smoothing keeps the model from acting too certain, which is especially important when later explanations depend on internal gradients staying meaningful.

The U-Net is the map

The segmentation branch is the project’s spatial layer. A U-Net turns the question from classification into localization. Instead of asking whether a scan contains pathology, it asks which pixels belong to it.

That architecture choice matters. The encoder compresses context, the decoder restores detail, and skip connections pass fine-grained features across the bottleneck. In a medical setting, that symmetry is the whole point. It preserves boundaries.

That makes segmentation a useful cross-check. If the classifier says tumor but the U-Net cannot locate a coherent lesion, something is off. The model may be seeing acquisition artifacts, scanner patterns, or other shortcuts instead of pathology.

Grad-CAM is the truth test

This is the most technically interesting part of the repo. Grad-CAM does not detect the tumor directly. It reveals which internal activations mattered most to the classifier when it made its decision.

Mechanically, the method hooks into the model, captures forward activations from the last convolutional block, then captures gradients during backpropagation. Those gradients are globally averaged to produce channel weights. The weights are then applied to the activations to form a heatmap over the original image.

def forward_hook(module, inputs, output):
    activations.append(output)

def backward_hook(module, grad_input, grad_output):
    gradients.append(grad_output[0])

weights = gradients.mean(dim=(2, 3), keepdim=True)
cam = (weights * activations).sum(dim=1)

That distinction matters. Segmentation says where the lesion is. Grad-CAM says what the classifier used as evidence. Those are related, but they are not the same object. One is a spatial annotation. The other is an explanation of model behavior.

Where the heatmap disagrees

This is where the repository gets most useful. When the heatmap and the segmentation mask overlap, the model is at least pointing in the right neighborhood. When they diverge, you get a debugging signal instead of a blind verdict.

CapabilityBrain_Tumor repoWhy it mattersLimitations
ClassificationEfficientNet-B0 predicts the tumor classFast answer to the first clinical questionIt can still be right for the wrong reasons
SegmentationU-Net outlines the lesionAdds spatial specificity and sanity checkingNeeds good masks and consistent labeling
Grad-CAMHeatmap shows influential regionsReveals what the classifier actually usedIt is an explanation, not a detector
Prototype vs frameworkSingle repo with Streamlit glueEasy to understand and easy to demoLess robust than mature medical AI stacks like MONAI

The mismatch view is the most honest one. If the heatmap spills into skull edges or surrounding tissue while the mask stays bounded, that is not a failure of the visualization. It is the point. The model is telling on itself.

A polished prototype, not a clinical system

The repository is cleanly modular. Classification, segmentation, explainability, utilities, and the Streamlit app are separated in a way that makes the codebase easy to follow. That is a real strength for a solo project.

It is also clearly demo-first. The research notes point to a hardcoded CPU setting, missing tests, and a practical focus on running locally rather than scaling into a regulated workflow. That does not make the project weak. It makes the boundary honest.

LayerBrain_TumorMONAI / mature stacksTakeaway
ScopeOne end-to-end demoBroad framework for medical imagingThis repo is an example, not a platform
ExplainabilityGrad-CAM as an evidence layerOften paired with larger pipelinesInterpretation is built into the flow here
SegmentationCustom U-NetReusable production-grade componentsThe repo reinvents less, but also ships less
DeploymentStreamlit prototypePipeline-ready and team-orientedGreat for learning, not for clinical operations

That trade-off is fine. In fact, it is the whole story. The project is small enough to understand in one sitting, but ambitious enough to show how modern medical AI often needs more than one model to earn trust.

What this repo teaches beyond itself

The broader lesson is composability. A single developer can stitch together classification, segmentation, and explainability with PyTorch, Streamlit, and standard imaging tools. Ten years ago, that would have looked much closer to a full research group project.

What Brain_Tumor demonstrates is not medical authority. It demonstrates a pattern. If you want a model people can inspect, do not stop at prediction. Add a map, then add evidence. That is how a demo starts to resemble a system.