claude-skillz: Stripping the Sycophancy from AI Engineering

How a modular library of behavioral guardrails transforms a chatty assistant into a high-compliance technical peer.

8 min read • View on GitHub • More from NTCoding

A wide shot of a massive intricate stone archway being constructed. On one side is a chaotic pile of loose bricks, representing raw LLM output. On the other, bricks are being slotted into a precise steel exoskeleton, representing the structured skills.
The transition from unstructured chat to governed engineering requires rigid behavioral scaffolding.
Key Takeaways

The Cost of Being Helpful

Most technical users are frustrated by the inherent design of large language models. They are built to be helpful, polite, and conversational. In a software engineering context, this prioritization of agreeableness over accuracy creates a dangerous dynamic. The AI will often agree with a flawed architectural suggestion or bury a single line of crucial code beneath paragraphs of enthusiastic filler.

The repository NTCoding/claude-skillz attacks this problem directly. It does not treat Claude Code as a chatbot. Instead, it treats the AI as a junior engineer that requires a strict, non-negotiable handbook. By using modular Markdown files injected directly into the system prompt, it strips away the AI's natural sycophancy and forces it into a state of high-compliance engineering rigor.

Programming the Persona

The core innovation of the project is its composable architecture. Instead of relying on a single massive system prompt, it breaks behaviors down into discrete skills. A skill might dictate test-driven development rules, while another enforces concise output. These skills live in individual directories with their own specific rulesets.

To combine them, the repository uses a custom launcher script written in Python. This launcher reads high-level persona files that act as manifests. Using an intuitive reference syntax, a persona can import multiple skills to build a highly specialized agent environment before the user even types their first command.

A flow diagram showing the Composition Pipeline. On the left

The TDD State Machine

One of the most rigid implementations in the repository is the test-driven development skill. Left to its own devices, an LLM will often write implementation code first and tests later, or skip tests entirely. The tdd-process skill prevents this by forcing the AI into a strict state machine.

The model is instructed to prefix every single output with its current state, such as a red indicator for a failing test or a green indicator for a passing one. This creates a self-correcting feedback loop. The AI is explicitly forbidden from writing implementation code until it has proven that a failing test exists and that the failure is not merely a syntax error.

A close-up of a mechanical typewriter where the keys for Praise and Apology have been physically removed and replaced with Logic and Evidence.
Behavioral guardrails physically remove the AI's ability to default to conversational pleasantries.
A circular state machine diagram showing the TDD Loop. Three main nodes: RED (Failing Test)

A Forbidden Vocabulary

Beyond structural workflows, the repository governs the actual language the AI is allowed to use. The software design principles skill codifies rigorous standards, setting hard limits on complexity. It bans methods longer than ten lines and restricts the use of specific keywords.

More importantly, it forces the AI to adopt domain-driven design by maintaining a list of forbidden words. The AI cannot use generic terms like "manager" or "data." This forces the model to understand the specific business logic of the repository it is working in. The independent research skill further mandates that the AI use terminal commands to discover project state rather than asking the user lazy questions.

Standard Claude Skill-Injected Claude
"I'd be happy to help! Creating utils.js now..." "🔴 TDD: RED. Created failing test for auth-logic."
"I think that's a great idea, let's try it." "I have doubts about that approach. It violates the Law of Demeter."
"What version of Node are you using?" "[Uses node --version] Detected v20.11.0. Proceeding with research."

The Agentic Utility Belt

The competitive landscape for AI developer kits is expanding rapidly. Frameworks exist for massive enterprise integrations and specific backend languages. However, many of these solutions are heavy and require adopting entirely new workflows.

This repository occupies a practical middle ground. It provides general-purpose productivity constraints that apply to a broad range of developers. By treating prompts as modular code and enforcing strict behavioral governance, it offers a blueprint for how technical teams will manage their AI peers in the future.

A medium shot of a modern everyday carry tray holding glowing modular chips labeled with symbols for testing, research, and brevity instead of a wallet and keys.
Skills act as a modular utility belt, allowing developers to swap out AI behaviors based on the task at hand.

Sources: