bug-bounty-installer: claude-bug-bounty: The bug bounty harness that makes Claude ask for proof
A Claude Code wrapper with agents, commands, memory, and proxy visibility, built to move from recon to report without skipping the evidence.
- claude-bug-bounty turns Claude into a control plane that has to earn its next step with evidence.
- Its real differentiator is the loop, because recon, validation, memory, and reporting stay linked instead of drifting apart.
- The repo is more than a setup script because it coordinates tools, MCP visibility, and checkpoints inside Claude Code.
- The harness improves judgment only when the human keeps the evidence standard strict.
The model is not the point
Most bug bounty automation tries to make tool output cheaper. This repo tries to make judgment less fragile. It wraps Claude Code in a harness that only moves forward when there is enough evidence to justify the next step.
That is the real shift. The model is not being trusted to improvise its way from recon to report. It is being constrained to behave like an operator with a checklist, memory, and checkpoints.
Why bug bounty work needed a harness
Bug bounty work is a bad fit for a single chat window. You are juggling recon, live traffic, target history, validation, and writeup quality at the same time. A good scanner can find signals. A good harness keeps those signals from evaporating the moment the conversation changes.
That filtering matters because most false positives are not technical failures. They are process failures. The evidence never got separated from the noise.
What lives inside claude-bug-bounty
The documented surface is a small control stack: 13 commands, 7 agents, more than 25 tools, plus Burp MCP and HackerOne MCP integrations. The installer drops commands and skills into `~/.claude/`, which makes the workflow feel native inside Claude Code instead of bolted on.
- Commands translate intent into repeatable actions like recon, hunt, and autopilot.
- Specialist agents narrow the problem space so one model is not doing every job at once.
- Burp MCP exposes live proxy history, while HackerOne MCP keeps the reporting side connected.
- Persistent memory carries target history forward so the next pass is informed by the last one.
The important part is not the count. It is the separation of concerns. Recon, validation, memory, and reporting each have a place to live.
The loop: recon, test, validate, report
The harness becomes interesting when you follow a finding through the loop. Recon generates candidates, proxy visibility adds live context, validation decides whether the signal deserves more work, and memory stores the result so the next pass starts smarter.
A weak lead does not become a report just because the model is confident. It gets rejected, retested, or parked until the evidence improves. That discipline is the difference between a chatty assistant and a usable operator.
Autopilot exists, but it still runs inside checkpoints, not around them. The design promise is not full autonomy. It is repeatable judgment with fewer dead ends.
Why this beats a setup script, and where it doesn't
A setup script provisions tools. `claude-bug-bounty` coordinates them. That sounds like a small distinction until you care about the handoff between discovery, proof, and writeup quality.
| Feature | Traditional installer | claude-bug-bounty | Semi-autonomous agents |
|---|---|---|---|
| Core job | Installs a toolkit. | Coordinates research and testing. | Coordinates broader multi-agent work. |
| Human role | Run tools manually after setup. | Steer, review, and approve checkpoints. | Set goals, tune policies, and handle review. |
| Evidence gate | None. | Built in at validation steps. | Often present but looser or custom. |
| Memory | None beyond shell state. | Persistent target memory. | Varies by system. |
| Tool visibility | Low once setup ends. | High through MCP and commands. | Depends on the framework. |
| Output | A working environment. | Findings and submission-ready reports. | Usually task-specific outputs. |
| Failure mode | Useful machine with no judgment. | Slower than pure automation if evidence is thin. | More autonomy, but more drift and complexity. |
Traditional installers are useful if you need a machine ready to scan. Semi-autonomous agent systems are useful if you want broader autonomy. This repo sits in the middle: opinionated enough to enforce evidence, but close enough to Claude Code that a human can still steer every important turn.
The tradeoff is obvious. If the operator wants magic, this will feel strict. If the operator wants fewer bad leads, strictness is the feature.
What the repo gets right, and what stays hard
The strongest design choice is proxy visibility. Reading Burp history inside the same workflow where the model is reasoning keeps the loop grounded in actual traffic, not just tool summaries. Persistent memory is the second win, because bug bounty work is only useful when lessons survive the next target.
What stays hard is the part no harness can fake: deciding when a lead deserves more time. The repo can reduce false positives and organize attention, but it cannot turn weak evidence into a real finding. That last mile still belongs to the researcher.
That is why the project feels bigger than automation. It is a control plane for judgment.