DurgSetu-AI: The Computer Vision System That Treats Forts Like Living Infrastructure

How a Django and PyTorch pipeline aligns mismatched photos, ignores vegetation noise, and lets humans verify the damage AI thinks it sees.

8 min read • View on GitHub • More from mitpatil07

A weathered fort wall with two offset photo plates being aligned by drafting compasses and registration marks. The overlap reveals a faint crack only when the images are brought into the same frame, explaining why change detection starts with alignment rather than detection alone.
Before the model can look for damage, it has to make two imperfect photos comparable.

DurgSetu-AI aims to provide a comprehensive AI-powered platform for analyzing and understanding the historical significance and architectural features of forts in Maharashtra.

Mithlesh Patil, Author/Maintainer · DurgSetu-AI GitHub Repository README
Key Takeaways

Why fort monitoring is a computer vision problem

Historic forts do not change like lab samples. The same wall is photographed from different angles, under different light, with vines, shadows, and handheld camera drift in the way. DurgSetu-AI starts from that reality, which is why its core problem is not classification. It is making two bad photos tell the truth together.

That is also why the project feels more practical than a demo. The repository is organized around a monitoring workflow for archaeological teams, not a one-off model showcase. The DurgSetu-AI repository combines image analysis, crowdsourced reporting, AI-generated summaries, and verification hooks in one pipeline.

Hedcut-style portrait of Mithlesh Patil based on the verified GitHub avatar at https://avatars.githubusercontent.com/u/127524455?v=4. The portrait should preserve his likeness while keeping the background clean and white, reinforcing that the project is person-led and practical rather than anonymous infrastructure.

The pipeline that makes two bad photos comparable

The detection engine in backend/home/structural_detector.py does not jump straight to a damage label. It first tries to solve the geometry problem, then the lighting problem, then the noise problem, then the texture problem. That order matters because every later step depends on the earlier one being honest.

The system is a chain of cleanup steps, not a single magic detector. Each stage strips away a different source of false signal before the result reaches a human.

# Conceptual flow from the repo's detection pipeline
photo_a, photo_b = load_images()
photo_a, photo_b = align_with_akaze_homography(photo_a, photo_b)
photo_a = apply_clahe(photo_a)
photo_b = apply_clahe(photo_b)
photo_a = mask_vegetation(photo_a)
photo_b = mask_vegetation(photo_b)
feat_a = resnet50_layer2(photo_a)
feat_b = resnet50_layer2(photo_b)
change_score = compare(feat_a, feat_b, metric='ssim_plus_distance')
result = route_to_human_review(change_score)
A close-up stone surface divided into three zones. On the left, moss and leaf clutter are swept away. In the center, texture bands are extracted as layered crosshatching. On the right, a thin crack emerges as an isolated line while a hand with a verification stamp waits outside the frame. This explains how the pipeline separates environmental noise from structural change.
The smartest part of the pipeline is not the detector. It is the filtering that makes the detector worth trusting.

Why layer 2 beats layer 4

The repo makes a deliberate choice in its CNN stack. It uses ResNet50 features from layer2 instead of pushing all the way to the deepest layers. That is not an accident. Deep layers are better at object semantics, but heritage damage often lives in small edges, mortar loss, surface erosion, and texture discontinuities.

ChoiceWhat it is good atWhat it misses
ResNet50 layer2Fine texture, cracks, erosion, surface irregularityHigher-level object meaning
ResNet50 layer4Semantics and coarse structureSmall visual defects and subtle surface change
Plain thresholdingFast and simpleLighting drift, shadows, and messy backgrounds
Human inspection aloneContext and expert judgmentScale, repeatability, and audit trail

That tradeoff is the point. The project is not trying to identify a fort, a tower, or a wall. It is trying to preserve the tiny visual cues that tell an expert something is getting worse.

The system does not trust itself

This is where DurgSetu-AI gets more interesting than a typical detection app. The data model includes fields such as is_false_positive and verified_by, which makes the system admit that shadows can look like cracks and moss can look like damage. Instead of hiding uncertainty, it routes uncertainty into review.

Automation firstHuman-in-the-loop
Assumes the model's output should ship on its ownAssumes the model needs expert confirmation when stakes are real
Optimizes for one-shot predictionOptimizes for auditability and correction
Treats false positives as noiseTreats false positives as part of the workflow
Feels fast in a demoFeels safer in a conservation setting

That is a better fit for heritage work than a fully autonomous claim. Preservation is not a leaderboard problem. It is a decision problem, and DurgSetu-AI leaves room for disagreement where it matters.

From detection to report to email

Once the system has a structural signal, it does not stop at an alert badge. The backend sends metrics such as risk scores and change counts into a report-generation flow that uses Llama 3.1, ReportLab, and email delivery. In other words, the model is not the product. The report is.

That matters because it closes the loop. A conservation workflow needs something a department can file, forward, and review later. The repository is opinionated about that last mile, which is where many CV projects disappear.

How DurgSetu-AI compares to generic CV stacks

SystemStrengthLimit
DurgSetu-AIPurpose-built fort monitoring with verification and reportingNarrow domain and early-stage maturity
Raw OpenCVFlexible image processing primitivesNo conservation workflow or human review layer
Generic PyTorch vision projectFast model experimentationOften stops at inference
QGIS or similar mapping toolsStrong spatial documentation and contextNot built for texture-level damage detection

That comparison is the clearest way to place the repo. It is not competing with OpenCV or QGIS on their own terms. It is trying to connect them into a conservation-specific operating system for a very specific kind of public infrastructure.

What this repo gets right, and what it still leaves open

The architecture is thoughtful for a solo-led project. The codebase is modular, the model loading is handled carefully, and the deployability scripts suggest someone has already thought past the notebook stage. The open question is validation, because a heritage tool only becomes real when people on the ground trust it enough to use it repeatedly.

That makes DurgSetu-AI feel like an early system with the right instincts. It knows that old stone is messy, photos are inconsistent, and certainty should be earned. That is a serious design posture, and it is why the repo stands out.