Packet_analyzer: The C++ DPI Engine That Turns Five-Tuple Hashing Into a Fast Path

A deep look at how this repository parses raw packets, tracks state per flow, extracts TLS SNI from encrypted traffic, and keeps every connection pinned to one worker thread.

9 min read • View on GitHub • More from perryvegehan

A packet sorting system routes stacked envelopes through a hashing gate into separate worker lanes, while one lane keeps both directions of a flow together at a single inspection desk. The visual explains that the engine treats deep packet inspection as a routing problem first and a parsing problem second.
The big idea is not just inspection. It is keeping each flow on one worker so state never has to fight the rest of the machine.
Key Takeaways

Why DPI Becomes a Flow-Distribution Problem

Deep packet inspection is only simple when packets are treated as independent. They are not. A flow remembers what came before, and memory is where concurrency gets expensive.

That is the clever move in Packet_analyzer: it does not just inspect packets, it decides where each packet should live. The unit of sharding is the connection, identified by the five-tuple of source IP, destination IP, source port, destination port, and protocol.

This diagram shows why consistent hashing matters. Keep both directions of a conversation on one worker, and the connection table stays local instead of becoming a shared bottleneck.

struct FiveTuple {
    uint32_t src_ip;
    uint32_t dst_ip;
    uint16_t src_port;
    uint16_t dst_port;
    uint8_t protocol;

    FiveTuple reverse() const {
        return {dst_ip, src_ip, dst_port, src_port, protocol};
    }
};

The Engine’s Core Shape: Reader, Load Balancer, Fast Path

The architecture is legible because it is strict. One reader ingests packets, a load balancer assigns them, fast-path workers inspect them, and an output writer emits the traffic that survives policy.

LayerJobWhy it matters
ReaderPull raw packets from inputKeeps capture separate from inspection
Load balancerHash the five-tuple to a workerPreserves thread affinity and flow state
Fast pathParse, classify, and apply rulesDoes the expensive work without central contention
WriterSerialize allowed trafficMakes the pipeline end in one predictable place

This is not generic task scheduling. It is a network pipeline built around locality. Once a flow lands on one worker, the worker can keep a private connection tracker and make decisions without asking a global table for permission every time.

Five-Tuple Hashing Is the Secret Sauce

The five-tuple is enough to define a conversation, which makes it enough to define ownership. If packet A and packet B belong to the same flow, they should not wander across threads looking for the same state.

That is why FiveTuple::reverse() matters. Bidirectional traffic should collapse into one tracking record, not split into two half-flows that disagree about what stage the connection is in.

Flow key = hash(src_ip, dst_ip, src_port, dst_port, protocol)
Reverse key = hash(dst_ip, src_ip, dst_port, src_port, protocol)
Same connection state, same worker thread

The payoff is simplicity under load. Once the shard choice is deterministic, the rest of the engine can stay stateful without paying the usual cost of a shared connection table.

A close-up reveals a TLS ClientHello as a nested mechanical shell inside Ethernet, IP, and TCP layers. A single domain tab protrudes from the innermost layer, showing how SNI can identify a destination without decrypting the payload.
The deep in DPI is a layered reveal. Even when the payload stays encrypted, the handshake can still leak useful metadata.

How the Engine Sees Through Encryption

Deep Packet Inspection (DPI) is a technology used to examine the contents of network packets as they pass through a checkpoint. Unlike simple firewalls that only look at packet headers (source/destination IP), DPI looks inside the packet payload.

Pratyush Narain, Project Creator · DPI Engine README

The useful part is not that encryption is defeated. It is that encryption still leaves structure behind. TLS SNI gives the engine enough signal to recognize a domain class, which is often enough for policy, analytics, or coarse filtering.

That is a practical DPI lesson. You do not always need full decryption to make a useful decision. Sometimes the handshake metadata is the whole game.

The Parser: Raw Bytes, Nested Headers, Real Consequences

Under the hood, the parser has one job: turn a byte buffer into protocol structs fast enough that the rest of the pipeline can keep moving. That means Ethernet, IPv4, TCP, and higher-layer fields are handled manually rather than hidden behind a thick abstraction.

// Pseudocode shape of the parser flow
EthernetHeader* eth = reinterpret_cast<EthernetHeader*>(buffer);
IPv4Header* ip = reinterpret_cast<IPv4Header*>(buffer + ethernet_len);
TCPHeader* tcp = reinterpret_cast<TCPHeader*>(buffer + ethernet_len + ip_len);

// Then inspect payload for TLS or HTTP signals

This kind of code is unforgiving, but that is the point. If you want speed and control, you usually accept more protocol knowledge at the edge.

Rules, Wildcards, and Shared Locks

The policy layer is built for many readers and few writers. That is why a shared mutex makes sense: workers can check rules concurrently, while updates stay rare and controlled.

Policy styleWhat it looks likeTrade-off
Exact matchBlock one domain or IPSimple and precise
Wildcard matchBlock a family like *.facebook.comBroader coverage, more careful rule checks
Shared-read accessMany workers consult the same rulesFast reads, serialized updates

The interesting part is the combination. Wildcards expand the policy surface, while thread-safe reads keep the fast path from slowing down every time a packet asks a question.

Where This Sits in the Open-Source Landscape

Packet_analyzer is not trying to be Wireshark, Suricata, or Scapy. It occupies a narrower and more teachable category: a transparent C++ implementation of the ideas behind high-speed DPI.

ProjectLanguageThreading modelBest for
Packet_analyzerC++17Reader plus worker pipelineLearning flow-affine DPI design
WiresharkCCapture plus GUI analysisInteractive protocol investigation
SuricataC and RustProduction multi-threaded IDSHigh-volume detection and prevention
ScapyPythonMostly single-threadedFlexible packet crafting and exploration

That category matters. Production tools hide complexity because they have to. Educational engines expose it because the reader needs to see the trade-offs.

Why This Repo Feels Educational and Serious at Once

The maturity signals are real. CMake support, Windows setup notes, atomics, thread-safe queues, and multiple entry points all point to a project that understands systems boundaries, not just demo code.

At the same time, the code stays readable enough to teach from. That combination is rare. A lot of networking repos are either toy examples or opaque production systems. This one sits in the useful middle.