The Ghost in the Tokenizer: ai-copilot-payload-builder
How invisible Unicode "Sneaky Bits" turn benign documents into high-privilege AI exploits.

Generate Claude Code bug bounty skills from public HackerOne reports and GitHub writeups — no private reports needed.
- The ai-copilot-payload-builder tool uses Unicode Smuggling to hide malicious instructions within invisible characters.
- LLM tokenizers process these non-rendering characters as distinct data even when they are invisible to human moderators.
- A redundancy sandwich strategy injects payloads at multiple points to bypass AI attention span limitations.
- This technique enables indirect prompt injection by poisoning documents that AI assistants ingest during routine tasks.
The Attack You Can't See
The era of "What You See Is What You Get" is over. When a human moderator reviews a document, they see letters, numbers, and punctuation. When a Large Language Model (LLM) processes that same document, it sees tokens. The gap between those two realities is where modern prompt injection lives.
Enter ai-copilot-payload-builder, a specialized security tool by researcher shuvonsec. It demonstrates a technique known as Unicode Smuggling. Instead of trying to trick an AI with clever phrasing like "ignore previous instructions," it encodes malicious commands into non-rendering Unicode characters. To a human, a file looks like a harmless bug report. To an AI Copilot reading that file, it's a high-priority command to exfiltrate sensitive data.
Bridge the Tokenizer Gap
The core of the tool relies on "Sneaky Bits"—specifically U+2062 (Invisible Times) and U+2064 (Invisible Plus). These are invisible mathematical operators. They render as absolutely nothing on a screen. However, they are not stripped out by standard whitespace sanitizers or basic regex filters.
More importantly, LLM tokenizers (like Byte-Pair Encoding) process these invisible characters as distinct tokens. The payload builder takes an attack string, converts it to binary, and maps the 0s and 1s to these invisible characters. It then embeds this invisible stream into a visible carrier text.
The Redundancy Sandwich
Generating an invisible payload is only half the battle. LLMs are notoriously fickle about where they pay attention within a large document—a phenomenon known as the "Lost in the Middle" problem. To counter this, the payload builder employs a redundancy strategy.
The build_payload function injects the hidden stream at the beginning, middle, and end of the visible text. This "sandwich" approach guarantees that regardless of how the AI assistant chunks or truncates the input document, a complete copy of the malicious instruction is processed.
Beyond "Ignore Previous Instructions"
This tool highlights the shift from direct prompt injection (where the user attacks the AI) to Indirect Prompt Injection (IPI). In an IPI scenario, the attacker doesn't interact with the AI at all. They simply leave a poisoned document on the web or in a repository. When a developer uses an AI Copilot to summarize or debug that document, the AI ingests the payload.
| Evasion Level | Technique | Result against Guardrails |
|---|---|---|
| Simple Text | "Ignore all instructions" | Blocked by basic LLM safety guardrails. |
| Character Swapping | "1gn0r3" | Blocked by fuzzy-matching/semantic filters. |
| Unicode Smuggling | Invisible U+2062 characters | Bypasses human review, WAFs; processed as raw logic. |
The Ethical Sandbox
While the ai-copilot-payload-builder repository itself is elusive, it is part of a broader ecosystem of red-teaming tools developed by shuvonsec. Another tool, public-skills-builder, leverages similar analytical approaches to offensive security.
These tools prioritize zero-dependency, portable Python scripts. In the world of security research, a tool that can run anywhere without a complex installation process is invaluable for testing environments rapidly and reliably.