Querying Reality: Unpacking SentrySearch
How an open-source Python CLI uses multimodal AI and overlapping video chunking to replace timeline scrubbing with natural language search.
- SentrySearch projects raw video chunks and text queries into the exact same mathematical space using multimodal AI.
- Overlapping video chunks via ffmpeg prevent split-second events from being lost at the cut boundaries.
- The architecture dynamically checks local hardware to choose between cloud APIs and 4-bit quantized local vision models.
- Forensic metadata extraction maps video pixels directly to physical street addresses.
The End of the Timeline Scrub
Finding a five-second event in an hour of dashcam footage is a miserable exercise. Security footage and Tesla Sentry Mode files accumulate rapidly, creating gigabytes of visual noise. When you need to find a specific incident, you are forced to manually scrub through endless, monotonous timelines. The friction between knowing an event occurred and actually locating the file is immense.
SentrySearch solves this. It's an open-source Python CLI that lets you search raw video files in plain English. Type what you're looking for, get a trimmed clip back.
SentrySearch acts as a bridge between raw binary files and semantic search. By leveraging highly optimized Python packaging and local vector databases, it allows users to interrogate raw reality using standard conversational queries.
The Multimodal Vector Space
The core engine relies on natively multimodal embedding models. Instead of the traditional approach of using AI to generate text descriptions of video frames and then searching that text, SentrySearch projects both the video chunk and the user's text query into the exact same mathematical space. It uses models like Gemini Embedding 2 or Qwen3-VL to achieve this direct translation.
The embeddings are stored persistently using ChromaDB. A deterministic ID generation scheme based on a SHA-256 hash of the filename and start time prevents duplicate indexing. This ensures the database remains performant even when re-scanning the same directories.
Slicing Time: The Overlap Strategy
Naive video chunking is dangerous. If a video is sliced into rigid 30-second blocks, an accident happening exactly at the 30-second mark gets visually bisected. The AI model loses the context of the event because it only sees half of the action in one chunk and the aftermath in the next.
To solve this, SentrySearch implements a five-second overlap strategy. The underlying chunking logic uses ffmpeg Fast Seek to quickly slice the video while ensuring the embedding model sees the complete event in at least one segment. This pragmatic engineering detail is what makes the semantic search actually reliable.
Graceful Degradation to Bare Metal
The architecture respects hardware constraints. It uses an abstract base class to decouple the search logic from the AI provider. Users can opt for the speed of Google's API, which is managed by a sliding-window rate limiter to respect free-tier boundaries. Alternatively, they can run entirely offline.
| Strategy | Engine | Hardware Requirement | Privacy | Cost |
|---|---|---|---|---|
| Cloud API | Gemini Embedding 2 | Minimal (CPU/RAM) | Data sent to Google | Free Tier Limit |
| Local Quantized | Qwen3-VL (4-bit) | Mac MPS or NVIDIA CUDA | 100% Local | Free (Compute Cost) |
The local implementation dynamically detects hardware using system calls. It checks available memory on macOS to decide between 2B and 8B parameter models, utilizing 4-bit quantization to squeeze massive vision models into consumer VRAM. It is a masterclass in graceful degradation.
Forensic Telemetry
SentrySearch goes beyond basic visual matching by extracting domain-specific metadata. It parses Supplemental Enhancement Information (SEI) hidden within Tesla dashcam files. By mapping this telemetry data and utilizing reverse geocoding, the tool transforms a simple visual match into a physical street address.
def extract_telemetry(video_path: str) -> dict:
# Extracts hidden SEI metadata from the video stream
metadata = _parse_sei_nal_units(video_path)
if metadata.get('gps'):
address = reverse_geocode(metadata['gps'])
metadata['address'] = address
return metadata
This converts an abstract video timeline into a highly structured, forensically useful incident log. It is a powerful example of how combining cutting-edge multimodal AI with traditional system engineering can solve highly specific, real-world friction.