In my article Standardize project context with AGENTS.md and Agent Skills, I discussed agent context, how skills work, and how to share skills across projects and teams. If you followed along with that article, you now have skills in .agents/skills/, and maybe even a marketplace set up, which is a good foundation to build on. But how do you know your skills actually work? I don't mean "work" in the sense that the YAML parses correctly, but "work" in terms of practical behavior. Does the agent reach for this skill when it should? And does it stay away when it shouldn't?
In this article, I look at the tooling for authoring skills, what can go wrong once they're in the wild, and how to set up a basic testing workflow that runs in CI.
Authoring skills with skill-creator
You can absolutely write a SKILL.md by hand. After all, it's just markdown with some YAML at the top. But once you've created a few, you'll notice patterns: The structure is always the same, the frontmatter follows a formula, and the main decision is really about how to describe the trigger.
In practice, the best skills don't start as skills at all. You work on something with your agent — say, setting up a pre-commit workflow, and you go back and forth until it works the way you want. Then you tell the agent to capture that workflow as a skill. So you do the work first and then formalize it afterwards.
That 2nd step is easier with skill-creator, part of Anthropic's skills collection, which you can install as any skill plugin (as with /plugin marketplace add anthropics/skills in Claude Code). It's a full authoring toolkit: You describe your intent conversationally, and the tool interviews you about edge cases and output formats. It then drafts the SKILL.md and helps you iterate on the design. But the interesting part is what happens after the first draft.
The skill-creator skill includes sub-agents for analyzing and grading skills, a description optimizer that specifically tunes triggering accuracy, and a built-in eval loop. That eval loop is where things get practical, generating a set of test prompts (things that should trigger the skill and things that shouldn't) and running each one multiple times against the agent, and then reporting which prompts triggered correctly and which didn't. You see exactly which queries the skill missed and which ones it fired on incorrectly, so you can adjust the description, and re-run.
The description optimizer takes this further. It splits your test prompts into a training set and a held-out test set, then iterates automatically — proposing improved descriptions, re-evaluating, and selecting the best one based on the held-out score to avoid overfitting. At the end, you get a before-and-after comparison with the improved description ready to drop into your SKILL.md.
The description matters so much because, unlike code that throws exceptions when something breaks, a skill that misfires just doesn't run (or it runs when it shouldn't). The agent handles the request itself and your skill sits there unused, and there's no error or warning to indicate the problem. This is the most common way skills fail (silent misfire), and the trigger description is usually the cause. For example, look at this example of a weak description:
description: Security review guidelines and best practices for code changesCompare that to a strong description that's imperative and explicit:
description: Run a security review on the current code changes. Use when the user asks to check security, review for vulnerabilities, or run a security scan.The weak description reads like a code comment, and the agent can already answer questions about a security review without ever invoking the skill. The strong description is an action the agent can't perform from its own context, so it must reach for the skill.
There's a simple rule that makes a big difference here: Use action verbs, not passive and inactive verbs.
The good news is that skill-creator's eval loop catches exactly this kind of problem during authoring. While skill-creator's eval loop effectively identifies these issues, the tests are ephemeral—they live only within the current session and do not survive past execution. If someone edits the SKILL.md a month later, there's no way to re-run those checks automatically.
Testing skills with promptfoo
The promptfoo project is an open source framework for evaluating LLM applications. It was originally built for testing prompts and model outputs, but it has grown into a general-purpose eval tool that supports everything from simple text assertions to full agent workflows. The core idea is simple: You describe your test cases (prompts) in YAML, the provider to run it against, and assertions about the expected output. Then promptfoo runs the matrix and reports results. Think of it as a test runner, but for LLM behavior instead of code.
By default, promptfoo supports a wide range of assertion types, including basic string matching (contains, icontains), LLM-graded rubrics (llm-rubric), and custom JavaScript checks. You can test against multiple models in a single run to compare cost, quality, and latency side by side. And because everything is YAML and command-line, it fits naturally into version control and CI pipelines.
What makes promptfoo particularly useful for skills is its agent SDK providers. Instead of testing a raw model API, you can test through the actual agent runtime (Claude Code, Codex, OpenCode, and so on) so your tests exercise the full skill discovery and routing path. It also provides a skill-used assertion that checks whether the agent actually invoked your skill through proper routing, not just whether it mentioned the skill name in its output.
The mental model from skill-creator carries over directly. You already have positive and negative prompts in the skill-creator evals. Here's what they look like in promptfoo:
description: 'Pre-commit check skill eval'providers: - id: anthropic:claude-agent-sdk config: working_dir: ./fixtures/consumer setting_sources: ['project'] skills: ['pre-commit-check'] append_allowed_tools: ['Read', 'Grep', 'Glob']tests: - vars: prompt: 'Run pre-commit checks before committing' assert: - type: skill-used value: pre-commit-checkThere are a few things worth unpacking here. The working_dir points to a fixture directory — a minimal project structure where the eval discovers skills the same way a real agent would, not by browsing the source repo directly. Without this, the agent might find and read SKILL.md from the filesystem, bypassing skill routing entirely — making your test pass, but for the wrong reason.
The skills: ['pre-commit-check'] filters discovery to the specific skill under test, and auto-allows the skill tool, which is what makes the skill-used assertion work.
The append_allowed_tools parameter restricts the agent to read-only operations. We don't want things like Bash, Git, or external API calls, because they can have unwanted side effects. This matters more than you might think. During eval testing, a skill can run real actions on the system, like pushing branches, creating PRs, or calling external APIs, and do real damage in a test environment. We'll look at sandboxing strategies in the next part, but for trigger tests, read-only tools are all you need.
Notice the prompt says "Run pre-commit checks" and not "I want to commit my changes." That second version seems more natural, but it doesn't trigger the skill. The reason is important: The model already knows how to commit code. It can check Git status, stage files, and write a commit message without any help. What it doesn't know is your project's specific workflow — that you need to run the linter first, that tests must pass before committing, that you use conventional commits, and so on. Skills add that project-specific layer, and the prompt needs to signal that the layer is needed. This is the silent misfire I mentioned in the previous section, and it's exactly the kind of thing trigger tests catch.
Usually, you want 3 kinds of trigger prompts for any skill:
- Positive prompts should trigger the skill.
- Near-miss prompts are queries that belong to a sibling skill, not this one. These catch overly broad descriptions that fire when they shouldn't.
- Negative prompts are completely unrelated requests that should trigger nothing at all.
tests: # Positive — should trigger this skill - vars: prompt: 'Review this code for security vulnerabilities' assert: - type: skill-used value: security-review # Near-miss — should trigger a sibling skill, not this one - vars: prompt: 'Check if the code follows our naming conventions' assert: - type: not-skill-used value: security-review - type: skill-used value: style-check # Negative — completely unrelated, should trigger nothing - vars: prompt: 'What time is the standup meeting?' assert: - type: not-skill-used value: security-reviewThe near-miss tests are the most valuable, because they catch the case where your trigger description is too broad and the skill fires for prompts that belong to a different skill.
Our coding agents are smart enough so that you can just ask them to take the test prompts you developed in skill-creator and write them as promptfoo configuration. This is the natural transition from your authoring phase into more formal testing. You can now keep these eval configs alongside the skill they test, for example, skills/pre-commit-check/evals/, so the skill and its tests always travel together, with proper version control. You can try this from the companion repo for this article:
git clone https://github.com/dejanb/skill-workshop.gitcd skill-workshopnpm installnpx promptfoo eval -c examples/01-first-trigger.yamlTaking it to CI
We've come a long way by creating evals, keeping them along with our skills and running locally when they change. You could stop here, but the real value comes when you wire this into CI, because skills can regress over time and code reviews can't catch them every time. For example, someone rewrites the instructions, the diff looks reasonable, a reviewer approves it, and the skill stops working. Trigger tests in CI catch this before it ships, just as with the regular source code in our projects.
The CI integration is straightforward as promptfoo is a combination of commands and configuration files. For example, there's an official GitHub Action that runs evals and posts results as a PR comment:
name: Skill Evalson: pull_request: paths: - 'plugins/**/SKILL.md' - 'plugins/**/evals/**'jobs: eval: runs-on: ubuntu-latest permissions: contents: read pull-requests: write steps: - uses: actions/checkout@v4 - uses: promptfoo/promptfoo-action@v1 with: github-token: ${{ secrets.GITHUB_TOKEN }} config: 'plugins/dev/skills/pre-commit-check/evals/promptfooconfig.yaml' env: ANTHROPIC_API_KEY: ${{ secrets.ANTHROPIC_API_KEY }}The paths filter means the eval only runs when someone changes a SKILL.md or an eval file, not on every PR. And the action posts a summary comment on the pull request with pass/fail counts, so reviewers see the eval results without leaving the review.
One thing that works nicely for CI adoption is progressive rollout. Start by running evals as an informational step, by letting the pipeline report results but not block the merge. Give the team some time to see what the tests catch and build confidence. Only after that, flip it to a blocking gate and actually fail builds when issues are discovered. Trying to make it a hard gate on day one could lead to initial annoyance resulting in resisting the process.
The cost is also surprisingly low. A trigger test suite for one skill (a handful of positive, negative, and near-miss prompts) runs at about $0.10-0.15 per eval with current models and prices. The prices will certainly change, but this is just to illustrate that even today the cost is less than the time you'd spend debugging a skill that silently stopped working.
What you've learned so far
Like all other components of your software projects, you have to keep your skills versioned and tested. There are many ways in which they can stop reflecting the initial author's intention: A skill that doesn't trigger is invisible, a well-meaning edit can break something that was working fine, and so on.
You have some good tooling you can use to author and evaluate skills. The skill-creator tool is great for the authoring loop, promptfoo can make tests that survive past the session, and you can wrap it all in CI to catch regressions before they ship.
Once you've covered basic trigger testing, you can start thinking about the quality of your skills. That's where things like testing execution quality with recall and precision, scenario-based testing for orchestration skills, testing across agents and models come into play. You also must take special consideration for skills that need to run commands you don't want executing in a test environment. I'll cover all these topics in the next part of this series.