Real-Time-Malpractice-Detection-in-classroom: Real-Time Malpractice Detection in Classroom: How a Webcam Becomes an Evidence Machine

A lightweight Flask and OpenCV system does more than spot inattentiveness. It logs behavior, captures proof, and turns a simple face detector into a classroom audit trail.

8 min read • View on GitHub • More from divyasree04-git

A classroom webcam watches a monitor that is filled with timestamped screenshots, alert banners, an incident log, and a small dashboard graph. The scene explains that the system is not only detecting behavior, but preserving a record of what it believed happened.
The repo’s key move is not detection alone. It turns a fleeting visual signal into evidence that can be reviewed later.
Key Takeaways

Most classroom monitoring demos stop at the box and the beep. This one keeps going. It turns a brief event, a face leaving frame, into a record that can be reviewed later, which changes the project from a detector into a forensic system.

That is the sharpest thing about divyasree04-git/Real-Time-Malpractice-Detection-in-classroom. The repo is not trying to be a giant model zoo. It is trying to make a simple monitoring loop durable enough to matter.

The real product is not detection, it is proof

The repo’s central design choice is evidence-first monitoring. It watches for inattention, but it also captures screenshots, writes event data, and keeps a history of what the system saw. That gives the project a different feel from most computer vision demos, which often end at a warning overlay.

Here, the output matters as much as the inference. If a student is flagged, the system does not just shout. It leaves a trail.

A close-up of four weighted gears feeding a central gauge labeled attention score. The gears represent face, eyes, proximity, and pose, and the image explains how a small set of interpretable signals becomes a single decision.
The attention score is a weighted blend of simple signals. That makes the system easy to inspect, even if it is not as powerful as a heavier model.

Why a face detector is enough for the first pass

The monitoring loop is intentionally modest. A webcam stream goes into OpenCV. The code watches for a face, starts a timer when the face disappears, and triggers a warning once that absence lasts long enough. In the research notes, that threshold is three seconds.

That sounds basic, and that is the point. A lot of useful product behavior starts as state logic, not as a giant model. If no face is seen, the system does not need to infer a whole mental state. It only needs to decide whether the absence is long enough to matter.

if len(faces) == 0:
    if inattention_start_time is None:
        inattention_start_time = time.time()
    elif time.time() - inattention_start_time > 3.0:
        trigger_warning()
        save_screenshot()
        log_incident()
else:
    inattention_start_time = None

This is the system in miniature. A handful of signals become one score, then one threshold decision, then one evidence record.

The score is a stack of heuristics

The most interesting code is not a deep model. It is the scoring function. According to the research notes, `calculate_advanced_attention_score` combines face detection, eye detection, proximity, and pose into a weighted total. Face detection contributes 30 percent, eyes 40 percent, proximity 20 percent, and pose 10 percent.

That structure matters. It makes the system legible. You can see why a score changed, not just that it changed. The tradeoff is obvious too. A transparent heuristic system is easier to understand, but it is also easier to break with odd lighting, camera angles, or a student who simply looks different from the default assumption.

SignalWeightWhat it tells the systemTypical failure mode
Face detection30%Someone is present in frameOcclusion, side profile, bad lighting
Eye detection40%Eyes are visible and likely oriented forwardGlasses, blur, closed eyes, low resolution
Proximity20%The face is large enough to feel close to the cameraCamera position varies, laptop distance shifts
Pose offset10%Head angle looks centered enough to count as attentiveNormal movement can look like avoidance

From alerts to audit trails

This is where the repository becomes more than a classroom demo. The research notes describe an `AttentionLogger` class with SQLite tables for sessions, attention events, and look-away incidents. That means the system is not only reacting in real time. It is storing structured telemetry for later review.

That design makes the project feel like a black box recorder. A warning is only the beginning. The interesting part is that the app can later answer questions like when the incident started, how long it lasted, and what the system believed was happening at the time.

LayerWhat it storesWhy it matters
sessionsStart and end markers for a runDefines the boundaries of one exam or class
attention_eventsTimed telemetry snapshotsCaptures how the score moved over time
look_away_incidentsThreshold-crossing eventsTurns a moment into a reviewable case

Why this feels lighter than the usual CV stack

The project sits in a different class from MediaPipe-heavy or YOLO-based monitoring systems. It is lighter, simpler, and easier to run on modest hardware. The research notes also point out that the pose logic is a kind of poor-man’s estimation, using face position and basic heuristics instead of a larger inference stack.

That makes it easier to deploy, but not magically better. It trades model sophistication for runtime simplicity and interpretability. For this use case, that is the interesting bargain.

DimensionThis repoHeavier CV stacks
Input signalBasic webcam framesWebcam frames plus richer feature extraction
Model complexityLowHigher
Hardware demandLow to moderateOften higher
InterpretabilityHighLower
Setup complexityLowerHigher
Forensic outputBuilt inUsually needs extra work
Low-end laptop fitGoodLess certain

The ethical pressure point

The same thing that makes the project practical also makes it uncomfortable. When the rule is effectively “no face for three seconds,” the system is making a judgment about attention from a very small signal. That can be useful in a narrow monitoring context, but it can also be unfair if read as proof of bad intent.

The virtue of this repo is that it does not hide that tension. Its assumptions are visible in the code. That is better than pretending the system is neutral. It invites the harder question, which is whether a tool this legible is acceptable precisely because its limits are so easy to see.

What this repository teaches

This is a compact case study in how prototypes become products. A simple webcam script becomes a stateful tracker. The tracker grows a web UI. The UI writes to a database. The database feeds analytics. Then the whole thing stops being a demo and starts looking like infrastructure.

The larger lesson is not about cheating detection. It is about product design under constraint. If you need a cheap, understandable system, you often do not need more model. You need better packaging around a small signal so that the signal can survive review.