`ai-copilot-probe`: The Burp Suite for AI copilots

A tiny Python tool that fingerprints what an embedded assistant can see, what it can do, and whether its context layer leaks data across permission boundaries.

7 min read • View on GitHub • More from shuvonsec

A forensic workbench holds a locked AI copilot box, five diagnostic probes, and a thin key-like tool reaching into a hidden drawer. The scene explains that the core question is access, not chat behavior.
This is a reconnaissance tool disguised as a simple script. It tests the boundary between what the assistant can answer and what it is allowed to know.
Key Takeaways

The real attack surface is context, not chat

The easy mistake is to read this repo as another prompt-injection toy. It is not. `ai-copilot-probe` asks a more useful question: can an embedded assistant reach data it should not be able to summarize?

That shift matters because a copilot is rarely just a chatbot. It sits on top of real application permissions, real records, and real retrieval layers. If the model can see a ticket, report, or account through its context pipe, the vulnerability is not language. It is authorization.

A request payload is being assembled by hand while a hidden report slips into view behind a half-open panel. The image explains how a small context selector can turn a normal chat call into an authorization test.
The payload matters because it decides which resource the assistant may summarize. That is the same question security engineers already ask of any API.

Why this tool exists at all

The repo is built by shuvonsec as a lightweight utility for security researchers, not as a polished platform. That is the right shape for the job. In a bug bounty session, portability beats ceremony.

One file, one dependency, no framework. The project is flat on purpose, which makes it easy to drop into a scratch VM, a Kali box, or a quick investigation against a target API. The README frames the tool as fingerprinting, and that is the right mental model. First learn what exists, then learn what it can see.

The five-phase fingerprinting loop

The smartest thing in the repo is its sequence. It does not jump straight to the juicy payload. It moves in steps: check connectivity, verify identity with `/me`, enumerate reachable resources, test a normal chat call, then probe context with a named resource. That order is the whole story.

The repo's method is a recon ladder: connect, identify, enumerate, chat, then test whether a named context stays inside the user's permission boundary.

That final step is the point. `context_ids` turns a chat endpoint into a permission test. If the assistant will summarize a resource named by the caller, the real question is whether the backend checked that the caller was allowed to reach it in the first place.

Inside the Python engine

Under the hood, the implementation stays deliberately small. `requests.Session()` carries auth headers across probes, so the tool can keep a stable identity while it moves from one endpoint to another. It supports both Basic Auth and bearer tokens, which makes it useful against messy real-world APIs.

The chat method is the most interesting part. If the first response includes a `completion_id`, the client switches into polling mode until the job is `completed` or `failed`. That handles the common split between synchronous and asynchronous AI APIs without forcing the user to care which style the target uses.

Resource enumeration is equally pragmatic. The probe looks for JSON:API-style structures and extracts what it can from whatever the backend exposes. The goal is not perfect parsing. The goal is fast reconnaissance, enough structure to tell you which objects are in play before you try the context probe.

Why this is not just another AI red-team tool

Broad testing frameworks are useful, but they aim wider. Garak asks whether a model can be pushed into unsafe behavior. promptfoo asks whether prompts and evals behave as expected. `ai-copilot-probe` is narrower, and that narrowness is the feature. It is focused on the seam between model behavior and API permissioning.

ToolPrimary questionScopeStrengthWeaknessBest fit
GarakCan the model be induced into unsafe behavior?Broad LLM red teamingDeep probe coverageNot built for app-specific reconGeneral model assessment
promptfooDo prompts, outputs, and evals behave as expected?Prompt and workflow testingRepeatable evaluationsLess about backend permissionsRegression testing
Manual Burp-style testingWhat leaks under real auth flows?Ad hoc API reconnaissanceMaximum flexibilitySlow and inconsistentOne-off investigations
ai-copilot-probeWhat can the copilot reach, and does context obey permissions?Embedded copilots and backing APIsPurpose-built context probingNarrower than broad suitesCopilot authorization audits

That is why the repo feels closer to classic security tooling than to AI novelty. It enumerates, fingerprints, and tests boundaries. The object is not to make the assistant say something wild. The object is to see whether the assistant can touch data it should never see.

What the repo suggests about the next security layer

The broader lesson is simple. AI copilots are turning into data-access layers, not just chat surfaces. Once that happens, security tooling has to follow the boundary, not the model. Prompt safety is not enough if the context layer is porous.

That is the real value of `ai-copilot-probe`. It reframes AI security as authorization work. If a copilot can be steered into summarizing the wrong record, the bug is not in the wording. It is in the access path that made the record reachable at all.