The Auditor is the Code: Inside Self-Scribing-Auditor

Moving past generic security checklists to autonomous, repo-native vulnerability research.

6 min read • View on GitHub • More from RASHMOR1

A close-up of a clockwork mechanism missing a single gear, with a robotic hand holding a blueprint focusing on the negative space. This represents the tool's ability to find missing logic rather than just flawed existing logic.
Finding vulnerability in negative space. The engine looks for what is missing, not just what is broken.
Key Takeaways

Hunting for Holes

Most artificial intelligence security tools operate like glorified spell checkers. You feed them a codebase, and they scan for known bad patterns. They look for SQL injection, cross-site scripting, or hardcoded secrets. They find bugs by looking at what is already there.

Self-Scribing-Auditor takes the opposite approach. It searches for what is missing. The architecture relies on a concept called the Tension-First Hypothesis. Instead of hunting for specific vulnerabilities, the system looks for asymmetries in the business logic. If a smart contract has a complex minting function but a trivial burning function, the tool flags the asymmetry. If a user state updates but a corresponding cache does not, it identifies the coupled-state failure.

This flips the script on traditional static analysis. It forces the underlying language model to act like a cynical human auditor, analyzing the negative space of a repository to find the logic gaps that generic scanners miss.

The Prompt Factory

The repository itself is deceptively small. It does not contain thousands of lines of Python or Rust. Instead, it operates as a meta-compiler for intuition. The core logic lives in a single master Markdown file. This file does not audit the target codebase. It writes the exact software needed to audit the target codebase.

When deployed, the agent reads the target repository's architecture. It examines package manifests, configuration files, and directory structures. Based on this introspection, it generates a bespoke directory of specialized sub-prompts. If it detects a React frontend, it writes a set of prompts specifically designed to attack React state management. If it finds a Rust kernel module, it writes memory-safety verification tracks.

A step-through flow chart showing the "Recursive Bootstrap" process. Step 1 shows a Master Prompt ingesting a 'package.json' file. An arrow points to Step 2

Closing the Evidence Gate

The greatest weakness of language models in security is hallucination. Models are eager to please and will invent plausible-sounding vulnerabilities that do not actually exist. Self-Scribing-Auditor solves this with a rigid constraint known as the Evidence Gate.

The system refuses to report a bug based on intuition alone. Every hypothesized vulnerability must pass through a strict verification protocol. The agent must prove reachability. It must generate a plausible reproduction path. If the logic cannot be traced from an external input to a critical exploit, the hypothesis is discarded.

A wide shot of a massive, messy waterfall of code flowing through a series of glowing, geometric gates. Most of the noise is blocked, leaving only a few sharp, crystalline proofs at the bottom. This visualizes the Evidence Gate concept.
The Evidence Gate filters out LLM hallucinations, ensuring only verified reproduction paths make it to the final report.

The Four-Phase Lifecycle

To manage context windows and prevent reasoning drift, the audit is broken into four distinct phases. This structure mimics a human consultancy's methodology.

Phase one is sequential mapping. The agent builds a broad understanding of the system architecture. Phases two and three shift into parallel execution. The agent takes the generated deep-dive tracks and attacks specific components simultaneously, verifying hypotheses against the Evidence Gate. Finally, phase four synthesizes the surviving, verified vulnerabilities into a cohesive report.

Dimension Traditional SAST Self-Scribing Auditor
Input Static, universal rule sets Repo-native generated prompts
Output List of potential pattern matches Verified, bespoke audit report
Verification Probability and heuristics Evidence Gates and reproduction paths
Adaptability Fixed to supported languages Adapts dynamically to the tech stack

By treating the auditor itself as a generative artifact, the project moves security tooling away from static checklists. It proves that the best way to secure complex software is to build an intelligence that understands exactly what that software was trying to achieve in the first place.


Sources and Further Reading: