`j5ik2o/ai-tools`: The Compiler That Teaches Agents How to Become Skills
A deep dive into a repo that packages prompts, tests trigger rates, and turns AI behavior into something you can version, distribute, and improve.
- `ai-tools` treats a skill as the unit of AI software, which makes prompts feel closer to deployable artifacts than ad hoc instructions.
- The repo’s real innovation is the closed loop between evaluation and revision, where failures become input to the next description.
- Triggering is handled statistically, so the project measures reliability instead of pretending one run can prove anything.
- Distribution matters as much as generation, because the marketplace and plugin layer make skills portable across agent environments.
Why this repo matters
Most prompt repositories stop at collection. `j5ik2o/ai-tools` goes one step further: it treats a prompt as something you can compile, evaluate, revise, and ship. That shift sounds small until you notice what it changes. The object is no longer a clever instruction. It is a software unit with a lifecycle.
That is why this repo is interesting. It does not just store AI behavior. It tries to make AI behavior improvable in a repeatable way, so a skill can move from draft to tested asset. The result feels less like a toolkit and more like an early compiler stack for agent behavior.
AI tools for myself
The skill is the unit of software
The repo’s core abstraction is the skill, usually expressed as a `SKILL.md` file with structured instructions and trigger logic. That is a better unit than a raw prompt for one simple reason: it can carry identity. A skill can be named, tested, distributed, and updated without losing the surrounding context that makes it useful.
Here is the practical difference.
| Abstraction | What it contains | How it behaves | Why it matters |
|---|---|---|---|
| Prompt folder | Loose snippets and notes | Manual reuse | Fast to start, hard to trust |
| Agent framework | Tools, memory, orchestration | Composed at runtime | Powerful, but broad and heavy |
| Skill package | Instructions, triggers, evals, distribution | Versioned and testable | Portable behavior that can improve |
| Workflow automation | Steps and app integrations | Runs predefined flows | Good at process, not at intent |
That is the repo’s thesis in miniature. A skill is not just text. It is a contract between a developer and an agent, plus the machinery needed to know whether the contract actually holds.
Inside Skill Forge
The most revealing code lives in `plugins/agent-skills/skills/skill-forge/`. It turns failure into an input signal. `run_eval.py` runs the same trigger tests repeatedly. `run_loop.py` feeds the results into an improvement step. `improve_description.py` then rewrites the skill description using what the earlier version got wrong.
# run_eval.py, conceptually
# Evaluate the same skill multiple times because a single trigger test is not enough.
for query in queries:
outcomes = []
for _ in range(runs_per_query):
outcomes.append(run_agent(query, skill))
trigger_rate = sum(outcomes) / len(outcomes)
passed = trigger_rate >= trigger_threshold
That loop matters because AI triggering is stochastic. A skill can miss once and still be good. It can also pass once and still be unreliable. The repo refuses to confuse a lucky run with a trustworthy one.
Why triggering needs statistics, not vibes
`run_eval.py` is the sober part of the project. It uses repeated runs, pass rates, and thresholds to decide whether a skill is strong enough to trust. This is the right move. If the system is probabilistic, the evaluation has to be probabilistic too.
| Evaluation style | What it asks | Weakness | What ai-tools does instead |
|---|---|---|---|
| One-shot prompt check | Did it work this time? | A single lucky run can fool you | Runs the same query many times |
| Manual review | Does it feel good? | Hard to reproduce | Converts behavior into a rate |
| Binary pass/fail | Did it trigger or not? | Ignores reliability | Uses a threshold on repeated outcomes |
| Benchmark-only testing | How strong is it overall? | Misses trigger-specific failure modes | Targets the exact skill trigger problem |
This is the project at its sharpest. It does not assume that a prompt is correct because it sounds good. It asks whether the prompt behaves consistently enough to deserve use.
The marketplace is the distribution layer
Once a skill exists, the next problem is distribution. The repo answers with `.claude-plugin/marketplace.json`, a strict registry, and a flat `skills/` layer that makes the parts easier for consuming agents to find. That is the difference between a local trick and a reusable asset.
The distribution story matters because AI tooling often stops at creation. Here, creation is only half the system. The other half is making the skill legible to another agent, another CLI, or another environment.
| Layer | Job | What it solves | Why it matters |
|---|---|---|---|
| `.claude-plugin/marketplace.json` | Registry | Discovery | Tells agents what exists |
| `skills/` symlinks | Consumption | Compatibility | Makes skills easy to load |
| Plugin selection | Governance | Scope control | Reduces accidental tool use |
| Cross-agent config | Portability | Reuse across environments | Keeps the skill from being tied to one interface |
This is where the repo stops looking like a local notebook and starts looking like infrastructure. If the skill is the unit, the marketplace is the transport.
TAKT and the workflow layer
`takt` extends the same mindset into workflow orchestration. The package scripts suggest a proposal, apply, and archive rhythm, which feels close to GitOps for agents. The point is not just to automate steps. It is to control the lifecycle of AI changes.
That is a useful discipline. It gives a human operator a way to review, accept, or retire work that was proposed by a model. In a space full of improvisation, that kind of choreography is the product.
What this repo is really pointing toward
Compared with larger agent frameworks, `j5ik2o/ai-tools` is narrow and personal. That is a strength, not a weakness. It is not trying to be the entire AI operating system. It is testing a sharper idea: a skill should be portable, measurable, and improvable like code.
Compared with prompt libraries, it adds evaluation. Compared with general agent frameworks, it adds a tighter abstraction. Compared with no-code automation, it keeps the whole system in code, where versioning and inspection are native. The result is a small repo with a big implication: the next useful layer of AI software may not be the agent. It may be the skill.