The Graph-Powered Attacker: Inside VishNet

How an open-source vishing simulator uses real-time voice cloning and Neo4j to autonomously map human vulnerabilities.

8 min read • View on GitHub • More from NishithP2004

A 1920s telephone switchboard operator with mechanical joints weaving a spiderweb connecting filing cabinets. This represents the synthesis of old telephony and automated data graphing.
VishNet connects raw telephony to structured data mining.
Key Takeaways

The Weaponization of Empathy

Traditional phishing relies on static scripts and robotic delivery. VishNet takes a different approach by treating human emotions as an attack vector. The system orchestrates fluid conversations using Markdown-based identity templates stored in the personas/ directory. These templates contain explicit instructions for psychological triggers like flattery and urgency.

To bypass human skepticism, the agent relies on ElevenLabs v3 tags. The AI injects calculated markers like [sigh] and [laughs] directly into the text stream. The result is a synthetic voice that breathes, hesitates, and chuckles at the right moments.

Building the Victim Graph

The most surprising architectural choice in VishNet is its integration with Neo4j. The system does not merely record audio files. It treats the telephone call as a real-time data mining operation. Using the Model Context Protocol (MCP), a dedicated LLM processes the live transcription to extract Personally Identifiable Information (PII).

As the victim speaks, the transcription agent extracts entities like names, addresses, and account numbers. It then uses MCP to instantly execute Cypher queries. This wires the extracted data into a live relationship graph, automatically mapping out the victim's digital footprint before the call even ends.

The Voice-to-Graph Pipeline continuously extracts entities from live audio and maps them into a Neo4j database.

Orchestrating the Attack

The core of the system lives in agent/index.js. This orchestration hub manages the lifecycle of the attack using a dual-agent architecture. It toggles between an IBM BeeAI framework agent for normal reasoning and a custom LangGraph agent for highly targeted impersonation.

When a call connects, Twilio opens a bidirectional WebSocket to the Fastify server. This continuous stream allows the LLM to listen and speak simultaneously. The session manager maintains state across the call, ensuring the agent remembers earlier details and adapts its strategy dynamically.

Isolating the Digital Puppet

To clone a voice, you need clean source audio. VishNet handles this via its utils.js pipeline. By utilizing ffmpeg, the system separates the stereo recording from Twilio into distinct caller and recipient tracks. This channel separation is vital for accurate transcription.

Once the target's voice is isolated, the system can feed that clean audio directly into the ElevenLabs API. In minutes, the platform gains the ability to speak in the victim's exact voice, opening the door for secondary attacks against the victim's extended network.

A close-up of a telephone receiver split down the middle. The left is a normal plastic speaker, the right dissolves into a lattice of data nodes. This illustrates the transformation of voice into structured graph data.
Channel separation strips raw audio into discrete data streams for immediate cloning.
FeatureLegacy Phishing SimulatorsVishNet
Logic EngineStatic Decision TreesLLM-driven LangGraph
Data CaptureFlat CSV LogsReal-time Neo4j Knowledge Graph
Voice GenerationPre-recorded WAV filesElevenLabs cloning with expressive TTS tags
InteractionUnidirectional promptsBidirectional Twilio WebSocket streams