Prahar: The Open-Source AI Sentry Built to Catch Lies, Faces, and Gunfire

A modular defense intelligence framework that fuses RSS news, deepfake detection, acoustic reconnaissance, and a structured military knowledge base into one triage pipeline.

8 min read • View on GitHub • More from code-with-khushi26

A wide watchtower-like control room with multiple sensor screens feeding into one central judgment core. One screen shows a scrolling news feed, another a waveform, another a blurred face framed in a video still, and a fourth shows defense assets pinned on a map. The scene explains that Prahar is designed to combine several weak signals into a single triage decision.
Prahar treats defense intelligence as a fused judgment stack, not a single model verdict.
Key Takeaways

Prahar is not trying to be a general-purpose AI assistant. It is trying to be a digital sentry for noisy, adversarial environments, where text can be manipulated, faces can be forged, and sound can be the first clue that something is wrong.

That matters because defense intelligence is rarely one clean signal. A suspicious post, a doctored clip, and a distant gunshot can all point to the same event, but only if a system knows how to compare them instead of treating each one in isolation.

A Sentry, Not a Dashboard

Prahar’s value starts with its posture. It does not just display data. It watches multiple streams, scores them in different ways, and then asks a harder question: what does the combination of those signals mean in a defense context?

That is why the repo feels less like a demo and more like a triage stack. Text gets filtered for misinformation. Video gets checked for visual integrity. Audio gets converted into acoustic classes. The knowledge base then anchors those outputs against known defense entities and assets.

Prahar’s pipeline is a layered triage system. Every input gets its own test, then the results converge into one grounded judgment.

ApproachInputsOutputReasoning style
PraharText, video, audio, knowledge baseDefense-oriented verdictCross-checks multiple weak signals
Generic moderation stackMostly text or imagesPolicy flag or pass/failSingle-task classification
Traditional news reviewHuman reading and verificationEditorial judgmentManual, slower, and context-heavy

Why Multimodal Is the Whole Point

The architectural trick here is not novelty for its own sake. It is that each modality covers a different failure mode. Text catches narrative manipulation. Video catches synthetic faces and altered imagery. Audio catches events that never make it into the feed until much later.

In a defense setting, that separation matters. A model that only reads headlines can be fooled by a clean-looking lie. A model that only inspects frames can miss the surrounding story. Prahar’s bet is that coordination across modalities is more reliable than a single heroic classifier.

The Text Engine Starts the Triage

M1 is the first filter, and it is more interesting than a plain sentiment or fake-news detector. The repo uses DistilBERT for binary real-or-fake classification, BART for defense-specific zero-shot labels, and named entity recognition to pull out people, places, and organizations.

That means the output is not just a score. It is a structured assessment. A suspicious article can be tagged, entities can be extracted, and the result can be passed downstream as context instead of noise. In practice, that is a much better shape for a defense workflow.

def analyze_text(article):
    real_fake = distilbert_classifier(article)
    label = bart_zero_shot(article, labels=[
        'verified defense news from official sources',
        'unconfirmed or suspicious defense report'
    ])
    entities = ner_extractor(article)
    return {
        'real_fake': real_fake,
        'label': label,
        'entities': entities
    }

Deepfakes Are a Frame-by-Frame Problem

M2 handles the visual side, and its best move is refusing to trust a single still image. The pipeline isolates faces with Haar cascades, then scores them with EfficientNet-B0, then samples frames over time and averages the results.

A close-up pipeline shows a strip of video frames moving through face detection, then into a row of sampled frames with score stamps, and finally into a smoothed confidence meter. The image explains why Prahar treats deepfake detection as a temporal process rather than a single-frame judgment.
Prahar does not bet on one frame. It samples, scores, and averages before it decides.

That temporal averaging is the anti-hallucination trick. A single frame can be misleading, especially in compressed or unstable video. A rolling score forces the system to care about consistency, which is a better defense against isolated artifacts and adversarial manipulation.

Sound Becomes Signal

M3 is the most defense-native module in the repo. Raw audio gets turned into a mel-spectrogram, then a compact CNN classifies the result into categories such as gunshot, explosion, helicopter, siren, and ambient noise.

That conversion matters because it turns a messy waveform into an image-like representation the model can read consistently. It is a small move with a big effect: sound stops being an anecdote and becomes a structured signal that can be triaged alongside text and video.

def classify_audio(path):
    y, sr = librosa.load(path, sr=22050)
    mel = librosa.feature.melspectrogram(y=y, sr=sr, n_mels=128)
    log_mel = librosa.power_to_db(mel, ref=np.max)
    prediction = audio_cnn(log_mel)
    return prediction

The Knowledge Base Is the Ground Truth Anchor

M5 changes the system from detection to context. The SQLite layer and the defense knowledge base do not exist just to store records. They let Prahar cross-reference findings against known assets and specifications, which makes the result more actionable.

That is the difference between saying something looks suspicious and saying why it matters. Grounding is what turns a model output into a defense workflow. Without it, you have classifiers. With it, you have an analytic system.

LayerJobWhy it matters
DetectionFind suspicious text, frames, or audioFinds a potential event
GroundingCheck the result against defense knowledgeAdds context and reduces drift
SynthesisCombine all signals into one judgmentTurns outputs into triage

What Prahar Is Really Competing With

Prahar is not competing with code security tools, endpoint platforms, or generic moderation systems. Those products solve different problems. Some are built for developer workflows. Some are built for policy enforcement. Prahar is built for defense intelligence synthesis.

That makes the comparison unusually clear. A single-modality classifier can be good at one task. A conventional security product can be excellent inside its lane. Prahar’s ambition is broader and more brittle: coordinate weak signals from different channels without pretending they are the same thing.

The Prototype Tells You the Ambition

The repo is modular and sharp, but it still reads like a prototype. That is not a criticism. It is a clue. The code shows a strong concept and a clear system map, but it also shows the kind of details that usually get hardened later, like cleanup paths and deployment discipline.

That tension is the real story. Prahar has the shape of a serious defense triage system, and it already knows what kind of intelligence it wants to produce. The remaining question is whether the prototype matures into something robust enough for the environments it is trying to serve.