The Graph-Powered Attacker: Inside VishNet
How an open-source vishing simulator uses real-time voice cloning and Neo4j to autonomously map human vulnerabilities.
- VishNet shifts social engineering from static scripts to dynamic, stateful AI agents using LangGraph.
- The system leverages the Model Context Protocol to autonomously map extracted PII into a live Neo4j knowledge graph.
- Stereo audio separation via ffmpeg allows the agent to isolate and instantly clone a target's voice.
- Expressive TTS tags weaponize empathy by injecting sighs and laughter into real-time conversations.
The Weaponization of Empathy
Traditional phishing relies on static scripts and robotic delivery. VishNet takes a different approach by treating human emotions as an attack vector. The system orchestrates fluid conversations using Markdown-based identity templates stored in the personas/ directory. These templates contain explicit instructions for psychological triggers like flattery and urgency.
To bypass human skepticism, the agent relies on ElevenLabs v3 tags. The AI injects calculated markers like [sigh] and [laughs] directly into the text stream. The result is a synthetic voice that breathes, hesitates, and chuckles at the right moments.
Building the Victim Graph
The most surprising architectural choice in VishNet is its integration with Neo4j. The system does not merely record audio files. It treats the telephone call as a real-time data mining operation. Using the Model Context Protocol (MCP), a dedicated LLM processes the live transcription to extract Personally Identifiable Information (PII).
As the victim speaks, the transcription agent extracts entities like names, addresses, and account numbers. It then uses MCP to instantly execute Cypher queries. This wires the extracted data into a live relationship graph, automatically mapping out the victim's digital footprint before the call even ends.
Orchestrating the Attack
The core of the system lives in agent/index.js. This orchestration hub manages the lifecycle of the attack using a dual-agent architecture. It toggles between an IBM BeeAI framework agent for normal reasoning and a custom LangGraph agent for highly targeted impersonation.
When a call connects, Twilio opens a bidirectional WebSocket to the Fastify server. This continuous stream allows the LLM to listen and speak simultaneously. The session manager maintains state across the call, ensuring the agent remembers earlier details and adapts its strategy dynamically.
Isolating the Digital Puppet
To clone a voice, you need clean source audio. VishNet handles this via its utils.js pipeline. By utilizing ffmpeg, the system separates the stereo recording from Twilio into distinct caller and recipient tracks. This channel separation is vital for accurate transcription.
Once the target's voice is isolated, the system can feed that clean audio directly into the ElevenLabs API. In minutes, the platform gains the ability to speak in the victim's exact voice, opening the door for secondary attacks against the victim's extended network.
| Feature | Legacy Phishing Simulators | VishNet |
|---|---|---|
| Logic Engine | Static Decision Trees | LLM-driven LangGraph |
| Data Capture | Flat CSV Logs | Real-time Neo4j Knowledge Graph |
| Voice Generation | Pre-recorded WAV files | ElevenLabs cloning with expressive TTS tags |
| Interaction | Unidirectional prompts | Bidirectional Twilio WebSocket streams |