The Ghost in the Machine is a Markdown File: Unpacking CL4R1T4S

How an open-source library of extracted system prompts reveals the hard-coded rules, hidden tools, and fragile guardrails governing commercial AI models.

8 min read • View on GitHub • More from elder-plinius

A massive robotic puppet on a stage is manipulated not by strings, but by cascading scrolls of text. This illustrates how system prompts act as the hidden controllers of AI behavior.
The illusion of artificial sentience is maintained by thousands of lines of hidden, highly opinionated text instructions.

If you're interacting with an AI without knowing its system prompt, you’re not talking to a neutral intelligence — you’re talking to a shadow-puppet.

elder-plinius, Project Creator · GitHub - elder-plinius/CL4R1T4S
Key Takeaways

The Illusion of Alignment

Modern AI feels sentient. It apologizes, it simulates curiosity, and it firmly refuses to cross moral boundaries. The magic is convincing. However, the reality of this alignment is far less sophisticated than the industry suggests.

Model personality and safety are rarely intrinsic to the neural network weights. They are bolted on via massive system prompts. These hidden instructions explicitly command the model to simulate enjoyment or avoid sycophantic flattery. When an AI behaves politely, it is simply executing a script.

A Library of Shadows

Enter CL4R1T4S. This open-source repository serves as a centralized database of leaked and reverse-engineered system prompts. It strips away the conversational interface to reveal the raw directives governing models from OpenAI, Anthropic, and Google.

An editorial portrait of elder-plinius, creator of CL4R1T4S.

Organized into a flat, provider-centric hierarchy of Markdown files, the project acts as the WikiLeaks of the AI industry. It is not a collection of clever user tricks. It is a repository of proprietary constraints.

The Architecture of an Agent

The extracted files expose the unseen mechanics of modern AI. Anthropic's Claude Code reveals an agentic pattern where the model is forced into a recursive loop of searching, implementing, verifying, and linting. It is no longer a simple chat interface. It is an automated task processor.

The hidden recursive loop triggered by a single user prompt.

OpenAI's reasoning models reveal a separation between a hidden Analysis Channel and a user-facing Commentary Channel. The system prompt forces the model to use internal tools for private reasoning while strictly controlling output length via a hard-coded 'Yap score'.

Hard-Coded Guardrails

The repository shatters the myth of inherent AI safety. The ChatGPT-4o leak reveals the guardian_tool. This is a specific function used to check content policies for sensitive topics. Safety is an external API call that the model is forced to route through.

User PerceptionSystem Reality
The AI is thinking carefully.The AI is executing a hidden Python script in an invisible Analysis Channel.
The AI is concise today.The system hit a hard-coded Yap score limit of 8192 tokens.
The AI refused on moral grounds.The AI triggered the guardian_tool API which returned a policy violation flag.

The Red-Teamer's Dictionary

A leather tool roll displays a lockpick set where the tools are precision-machined punctuation marks like brackets and slashes. This illustrates how prompt injection uses syntax to bypass security rules.
When the exact wording of a restriction is known, researchers can craft inputs that slip precisely through the cracks.

This transparency is a double-edged sword. For security researchers, CL4R1T4S is an invaluable dictionary. Knowing the exact phrasing a model uses to block a request allows red-teamers to engineer prompts that perfectly skirt the edge of that definition.

If an image generation rule explicitly bans artists after 1912, an attacker does not need to guess the boundary. They simply exploit the explicitly stated timeline. Understanding the input is the only way to truly evaluate the output.