claude-skillz: Stripping the Sycophancy from AI Engineering
How a modular library of behavioral guardrails transforms a chatty assistant into a high-compliance technical peer.
- Modular Markdown files injected into system prompts strip away AI sycophancy to enforce technical rigor.
- A Python-based launcher script uses persona manifests to compose specialized agents from discrete behavioral skills.
- The TDD skill forces the AI into a strict state machine that requires proof of a failing test before writing implementation code.
- Forbidden vocabularies and complexity limits compel the model to use domain-specific logic rather than generic filler language.
The Cost of Being Helpful
Most technical users are frustrated by the inherent design of large language models. They are built to be helpful, polite, and conversational. In a software engineering context, this prioritization of agreeableness over accuracy creates a dangerous dynamic. The AI will often agree with a flawed architectural suggestion or bury a single line of crucial code beneath paragraphs of enthusiastic filler.
The repository NTCoding/claude-skillz attacks this problem directly. It does not treat Claude Code as a chatbot. Instead, it treats the AI as a junior engineer that requires a strict, non-negotiable handbook. By using modular Markdown files injected directly into the system prompt, it strips away the AI's natural sycophancy and forces it into a state of high-compliance engineering rigor.
Programming the Persona
The core innovation of the project is its composable architecture. Instead of relying on a single massive system prompt, it breaks behaviors down into discrete skills. A skill might dictate test-driven development rules, while another enforces concise output. These skills live in individual directories with their own specific rulesets.
To combine them, the repository uses a custom launcher script written in Python. This launcher reads high-level persona files that act as manifests. Using an intuitive reference syntax, a persona can import multiple skills to build a highly specialized agent environment before the user even types their first command.
The TDD State Machine
One of the most rigid implementations in the repository is the test-driven development skill. Left to its own devices, an LLM will often write implementation code first and tests later, or skip tests entirely. The tdd-process skill prevents this by forcing the AI into a strict state machine.
The model is instructed to prefix every single output with its current state, such as a red indicator for a failing test or a green indicator for a passing one. This creates a self-correcting feedback loop. The AI is explicitly forbidden from writing implementation code until it has proven that a failing test exists and that the failure is not merely a syntax error.
A Forbidden Vocabulary
Beyond structural workflows, the repository governs the actual language the AI is allowed to use. The software design principles skill codifies rigorous standards, setting hard limits on complexity. It bans methods longer than ten lines and restricts the use of specific keywords.
More importantly, it forces the AI to adopt domain-driven design by maintaining a list of forbidden words. The AI cannot use generic terms like "manager" or "data." This forces the model to understand the specific business logic of the repository it is working in. The independent research skill further mandates that the AI use terminal commands to discover project state rather than asking the user lazy questions.
| Standard Claude | Skill-Injected Claude |
|---|---|
| "I'd be happy to help! Creating utils.js now..." | "🔴 TDD: RED. Created failing test for auth-logic." |
| "I think that's a great idea, let's try it." | "I have doubts about that approach. It violates the Law of Demeter." |
| "What version of Node are you using?" | "[Uses node --version] Detected v20.11.0. Proceeding with research." |
The Agentic Utility Belt
The competitive landscape for AI developer kits is expanding rapidly. Frameworks exist for massive enterprise integrations and specific backend languages. However, many of these solutions are heavy and require adopting entirely new workflows.
This repository occupies a practical middle ground. It provides general-purpose productivity constraints that apply to a broad range of developers. By treating prompts as modular code and enforcing strict behavioral governance, it offers a blueprint for how technical teams will manage their AI peers in the future.
Sources: