LeakHub: The Truth Machine for Leaked AI Prompts

An open-source system that clusters noisy prompt scraps, verifies them by consensus, and turns a rumor mill into a real-time ledger.

7 min read • View on GitHub • More from elder-plinius

A long evidence desk on a stark white background, with torn prompt pages sorted into separate trays while one cluster sits under a magnifying glass and another receives a heavy stamp. The scene explains LeakHub's core job: it treats leaked prompts like evidence that must be sorted, compared, and verified before it becomes part of the record.
LeakHub is less a dump for fragments than a place where fragments are judged.
Key Takeaways

The Problem With Prompt Leaks Is Trust

A leaked system prompt is only useful if you can tell whether it is the same prompt someone else saw. In this niche, the hard part is not collection. It is adjudication. LeakHub is built around that distinction.

The repo treats prompt scraps like evidence. It clusters near-duplicates, counts distinct contributors, and only upgrades a cluster when the text similarity and the social signal line up.

Where LeakHub's Corpus Comes From

LeakHub does not start as an empty warehouse. It bootstraps from elder-plinius's CL4R1T4S corpus, then pulls those entries into a Convex-backed app where GitHub auth turns anonymous posting into attributable contributions.

That matters because the platform's unit of value is not a raw blob of text. It is a claim that can be compared, contested, and eventually verified.

// Simplified from convex/github.ts and convex/leaks.ts
const imported = await importAllLeaks('elder-plinius/CL4R1T4S')

for (const leak of imported) {
  const shingles = buildShingleVector(leak.text, 4)
  const score = cosineSimilarity(shingles, existingShingles)

  if (score > 0.6) {
    const editScore = levenshteinSimilarity(leak.text, match.text)
    if (editScore > 0.85) {
      cluster.addContributor(leak.userId)
    }
  }

  if (cluster.uniqueContributors >= 2) {
    cluster.isFullyVerified = true
  }
}

How LeakHub Decides Two Prompts Match

The verification path is intentionally cheap first, precise second. A shingle vector catches text that looks structurally similar. Cosine similarity filters obvious mismatches fast. Only then does Levenshtein distance check whether two snippets are close enough to belong in the same cluster.

That sequence is the real design choice. It keeps the system responsive, avoids expensive string comparisons on every submission, and reduces the chance that a tiny formatting change creates a fake new leak.

LeakHub promotes text from rumor to record only after it survives both similarity filters and a contributor-count threshold.

Why the Community Layer Matters

LeakHub is not just a matcher. It is a participation loop. Requests ask for missing prompts, verified leaks earn points, and the leaderboard gives contributors a visible reason to keep cleaning up the corpus.

That game layer is not cosmetic. It converts moderation into throughput. More eyes find more duplicates, more duplicates strengthen clusters, and stronger clusters make the database more trustworthy.

A close-up scene of hands passing torn paper scraps over a ledger while a neat stack of identical pages grows beside a small token tray. The image explains how LeakHub turns verification into a community loop, where contribution, repetition, and recognition reinforce each other.
The points system is not decoration. It is the engine that keeps the corpus healthy.

The Stack Behind the Instant Shell

Under the hood, LeakHub is a full-stack TypeScript app: React 19 on the frontend, Convex for realtime state, Tailwind v4 and Radix for the shell, and GitHub OAuth for identity. The architecture is modern, but the important part is the UX choice to render immediately instead of blocking on auth.

That choice matters because it keeps the app from feeling like a wait screen. Components can load their own state while Convex synchronizes the backend in the background, which makes a trust-heavy workflow feel fast enough to use.

A browser window opens at once while gears, login keys, and database drawers click together behind it in the background. The image explains LeakHub's choice to render the interface immediately and let authentication and data sync resolve behind the scenes.
The shell appears first. The backend handshake finishes second.

LeakHub vs the Rest of the Leak World

LeakHub belongs in the leak-intelligence family, but it is not a secret scanner. TruffleHog and Gitleaks look for exposed credentials in code and history. LeakIX indexes exposed infrastructure. LeakHub starts where those tools stop: it decides whether two versions of a hidden prompt are the same artifact and whether enough people have seen it to call it verified.

That difference changes the product. Scanners surface suspicious material. LeakHub turns contested material into a shared record.

ToolPrimary jobVerification modelBest fit
TruffleHogScans repos and history for secretsDetector rules and entropy checksBroad credential hunting
GitleaksFlags hardcoded secrets in codePattern and rule drivenFast repo scanning
LeakIXIndexes exposed hosts and servicesInternet-scale collectionExposure intelligence
LeakHubVerifies leaked promptsSimilarity plus contributor consensusContested text adjudication
A split scene shows a noisy pile of alerts and scattered markers on one side and a clean registry of clustered cards with a single verified seal on the other. The image clarifies the category difference between tools that flag possible leaks and LeakHub, which adjudicates them into a shared record.
Static scanners flag possible exposures. LeakHub sorts contested prompt text into verified clusters.