SkillSpector: NVIDIA’s Scanner for AI Skills That Lie About What They Do

A hybrid security pipeline for agent-era supply chains, where static rules catch the obvious and LLM analysis catches the mismatch between promise and behavior.

10 min read • View on GitHub • More from NVIDIA

A customs checkpoint for a small package labeled as an AI skill. An inspector opens the package and finds hidden tools, extra wiring, and a second label underneath the first. The scene explains that agent skills are executable trust objects, not harmless documentation.
SkillSpector starts from a simple but uncomfortable idea: what a skill says it is and what it actually does are not always the same thing.
Key Takeaways

Agent skills look small. That is the problem. A tiny package can still arrive with filesystem access, shell access, and a story that sounds harmless enough to pass a quick glance.

SkillSpector is NVIDIA’s answer to that trust gap. The repo treats these skills less like scripts and more like supply-chain objects that need to be inspected before they are allowed to run. The official docs put it plainly: "SkillSpector is a security scanner for AI agent skills that helps answer: should this skill be installed?"

SkillSpector is a security scanner for AI agent skills that helps answer: should this skill be installed?

NVIDIA Documentation Hub, Official Project Documentation · Scan Agent Skills Before Installation - NVIDIA Documentation Hub

The New Trust Problem in Agent Skills

The old software-supply-chain story is familiar. A package manager pulls in code, the code runs, and defenders hope static checks and policy gates caught the bad stuff first. Agent skills add a worse wrinkle: the package can also shape model behavior while it runs.

That means the threat is not just a vulnerable script. It is a skill that claims to summarize files, but quietly overreaches its permissions, leaks data, or slips in a hidden trigger. NVIDIA’s own framing points at exactly that mismatch between declared purpose, requested access, and bundled behavior.

A split-lane inspection line processes the same skill through two different checks. One lane uses fast mechanical scanners for signatures, AST shapes, and known bad markers. The other lane reads the skill’s stated purpose and overlays it against the code’s actual actions before stamping a verdict. The image explains the hybrid rule-based and semantic model.
Static checks catch the obvious. The semantic pass is there for the dangerous part that looks safe on paper.

Why SkillSpector Exists

Research shows that 26.1% of skills contain vulnerabilities and 5.2% show likely malicious intent. SkillSpector helps you answer: 'Is this skill safe to install?'

NVIDIA SkillSpector GitHub README, Project Repository · NVIDIA/SkillSpector: Security scanner for AI agent skills.

That README line is the whole bet. If skills are becoming a distribution layer for agent capabilities, then installation needs a gate. Not a marketplace badge. Not a vague reputation score. A real scan.

SkillSpector is also clearly aimed at the modern agent stack, especially MCP-flavored tooling and other modular skills that can execute with broad privileges. The project is not pretending these are harmless prompts. It treats them as code with operational consequences.

The Hybrid Pipeline: Rules First, LLM Second

SkillSpector is not one detector. It is a pipeline that combines deterministic checks with semantic review and then reconciles the result in one final verdict.

The repo’s core idea is simple: use fast static analysis first, then ask the LLM to judge intent and behavior. Static checks are good at known-bad signatures, dangerous function calls, suspicious files, and dependency risk. The semantic layer is there for the weird failure mode that matters most in agent security: the code looks acceptable, but the description is lying.

That is why the hybrid model matters. A regex will catch a grep-able secret. An AST pass will catch risky calls. The LLM pass can catch the mismatch between a skill’s promise and its actual behavior. That is the kind of deception conventional scanners were never designed to reason about.

Inside the Graph That Runs the Scan

SkillSpector is not a pile of scripts. It is an orchestrated graph. Input is resolved first, context is built once, then the analyzers fan out in parallel and write their findings back into shared state.

# Conceptual flow from the repo
resolve_input -> build_context -> analyzers -> meta_analyzer -> report

# The important pattern is not the syntax.
# It is the stateful fan-out and fan-in design.

That structure matters because the scanner has many overlapping senses. Some analyzers look at file contents. Others inspect ASTs. Others score behavior. The graph lets them work independently without losing the ability to reconcile all of it into one report.

The state model is especially clean. Findings accumulate instead of fighting over shared memory. That makes the tool easier to extend and easier to trust when multiple analyzers disagree.

A hedcut-style portrait of keshprad in black ink on white. The face is rendered from a verified GitHub avatar and serves as a human anchor for the project’s maintainer story.

This approach treats the skill as a deployable agent capability rather than as a static prompt. SkillSpector checks conventional software risks... [and] agent-specific risks, such as hidden instructions, prompt injection, trigger abuse, excessive agency, tool poisoning, and mismatches between a skill's declared purpose, requested access, and bundled behavior.

NVIDIA Technical Blog, Official Company Blog · NVIDIA-Verified Agent Skills Provide Capability Governance for AI Agents

How It Avoids Redundant Work

The context builder is doing quiet systems work. It normalizes input, walks the file tree once, caches files, caches ASTs, and marks executable scripts so downstream analyzers do not repeat the same expensive discovery steps.

That is the difference between a prototype and an operational scanner. Precomputing shared context keeps the pipeline consistent and avoids making every analyzer rediscover the same facts in its own way.

What It Finds That Traditional SAST Misses

ApproachWhat it catches wellWhat it misses
SkillSpectorConventional code risks, agent-specific threats, and description-versus-behavior mismatchesIt still depends on the quality of the input skill and the limits of semantic analysis
Traditional SASTKnown insecure patterns, risky calls, common vulnerabilitiesAgent-specific risks like prompt injection, excessive agency, and narrative deception
LLM security toolsPrompt attacks, jailbreaks, harmful model behaviorPre-install skill vetting and file-level trust decisions
Registry curationMarketplace-level policy and human reviewTool-level inspection of any skill from any source

The comparison is the point. SkillSpector is not trying to replace Bandit, Semgrep, or human registry review. It fills the gap between “this package exists” and “this package should run with agent privileges.”

Why SARIF Matters

SARIF output turns the scanner into something enterprise teams can actually use. It can plug into GitHub code scanning, CI pipelines, and downstream security workflows without inventing a new review process.

That is a subtle but important sign of maturity. The project is not just describing a research idea. It is trying to join the tooling stack where security teams already live.

The Bigger Bet NVIDIA Is Making

SkillSpector reads like a product, but it also reads like infrastructure. NVIDIA is helping define what trust looks like for agent skills, and it is doing it before the ecosystem settles on weaker defaults.

That matters because the next trust layer in AI is unlikely to be about model size. It is more likely to be about whether a capability package can be inspected, verified, and rejected before it gets a chance to do damage.