bug-bounty-installer: claude-bug-bounty: The bug bounty harness that makes Claude ask for proof

A Claude Code wrapper with agents, commands, memory, and proxy visibility, built to move from recon to report without skipping the evidence.

8 min read • View on GitHub • More from shuvonsec

A wide editorial illustration of a security workbench viewed from slightly above. Raw recon notes and proxy traces enter from the left, while a neat report packet exits on the right. In the middle, drawers, switches, and checkpoints show how the harness forces findings through evidence before they move forward.
The project is less about a bot and more about a control room, where every result has to earn the next step.
Key Takeaways

The model is not the point

Most bug bounty automation tries to make tool output cheaper. This repo tries to make judgment less fragile. It wraps Claude Code in a harness that only moves forward when there is enough evidence to justify the next step.

That is the real shift. The model is not being trusted to improvise its way from recon to report. It is being constrained to behave like an operator with a checklist, memory, and checkpoints.

Why bug bounty work needed a harness

Bug bounty work is a bad fit for a single chat window. You are juggling recon, live traffic, target history, validation, and writeup quality at the same time. A good scanner can find signals. A good harness keeps those signals from evaporating the moment the conversation changes.

That filtering matters because most false positives are not technical failures. They are process failures. The evidence never got separated from the noise.

A close-up black-ink scene of a glass inspection chamber receiving a chaotic stream of proxy traffic and scan fragments. A fine metal sieve filters out noisy pieces while a clipboard catches only the surviving evidence. The image explains why live traffic and validation matter more than raw output.
Burp history and scan output become useful only after the harness separates signal from noise.

What lives inside claude-bug-bounty

The documented surface is a small control stack: 13 commands, 7 agents, more than 25 tools, plus Burp MCP and HackerOne MCP integrations. The installer drops commands and skills into `~/.claude/`, which makes the workflow feel native inside Claude Code instead of bolted on.

The important part is not the count. It is the separation of concerns. Recon, validation, memory, and reporting each have a place to live.

The loop: recon, test, validate, report

The harness becomes interesting when you follow a finding through the loop. Recon generates candidates, proxy visibility adds live context, validation decides whether the signal deserves more work, and memory stores the result so the next pass starts smarter.

The workflow is not linear. It is a gated loop that can branch backward when the evidence is thin and pause when human review is needed.

A weak lead does not become a report just because the model is confident. It gets rejected, retested, or parked until the evidence improves. That discipline is the difference between a chatty assistant and a usable operator.

Autopilot exists, but it still runs inside checkpoints, not around them. The design promise is not full autonomy. It is repeatable judgment with fewer dead ends.


Why this beats a setup script, and where it doesn't

A setup script provisions tools. `claude-bug-bounty` coordinates them. That sounds like a small distinction until you care about the handoff between discovery, proof, and writeup quality.

A split black-ink illustration showing two different bug bounty setups. On the left, a classic installer spills tools into a messy pile around a laptop. On the right, the harness arranges the same tools into drawers and checkpoints around a disciplined desk. The image contrasts provisioning with orchestration.
Provisioning gets you a toolkit. Orchestration turns that toolkit into a workflow.
FeatureTraditional installerclaude-bug-bountySemi-autonomous agents
Core jobInstalls a toolkit.Coordinates research and testing.Coordinates broader multi-agent work.
Human roleRun tools manually after setup.Steer, review, and approve checkpoints.Set goals, tune policies, and handle review.
Evidence gateNone.Built in at validation steps.Often present but looser or custom.
MemoryNone beyond shell state.Persistent target memory.Varies by system.
Tool visibilityLow once setup ends.High through MCP and commands.Depends on the framework.
OutputA working environment.Findings and submission-ready reports.Usually task-specific outputs.
Failure modeUseful machine with no judgment.Slower than pure automation if evidence is thin.More autonomy, but more drift and complexity.

Traditional installers are useful if you need a machine ready to scan. Semi-autonomous agent systems are useful if you want broader autonomy. This repo sits in the middle: opinionated enough to enforce evidence, but close enough to Claude Code that a human can still steer every important turn.

The tradeoff is obvious. If the operator wants magic, this will feel strict. If the operator wants fewer bad leads, strictness is the feature.


What the repo gets right, and what stays hard

The strongest design choice is proxy visibility. Reading Burp history inside the same workflow where the model is reasoning keeps the loop grounded in actual traffic, not just tool summaries. Persistent memory is the second win, because bug bounty work is only useful when lessons survive the next target.

What stays hard is the part no harness can fake: deciding when a lead deserves more time. The repo can reduce false positives and organize attention, but it cannot turn weak evidence into a real finding. That last mile still belongs to the researcher.

That is why the project feels bigger than automation. It is a control plane for judgment.