Skip to main content
Redhat Developers  Logo
  • AI

    Get started with AI

    • Red Hat AI
      Accelerate the development and deployment of enterprise AI solutions.
    • AI learning hub
      Explore learning materials and tools, organized by task.
    • AI interactive demos
      Click through scenarios with Red Hat AI, including training LLMs and more.
    • AI/ML learning paths
      Expand your OpenShift AI knowledge using these learning resources.
    • AI quickstarts
      Focused AI use cases designed for fast deployment on Red Hat AI platforms.
    • No-cost AI training
      Foundational Red Hat AI training.

    Featured resources

    • OpenShift AI learning
    • Open source AI for developers
    • AI product application development
    • Open source-powered AI/ML for hybrid cloud
    • AI and Node.js cheat sheet

    Red Hat AI Factory with NVIDIA

    • Red Hat AI Factory with NVIDIA is a co-engineered, enterprise-grade AI solution for building, deploying, and managing AI at scale across hybrid cloud environments.
    • Explore the solution
  • Learn

    Self-guided

    • Documentation
      Find answers, get step-by-step guidance, and learn how to use Red Hat products.
    • Learning paths
      Explore curated walkthroughs for common development tasks.
    • Guided learning
      Receive custom learning paths powered by our AI assistant.
    • See all learning

    Hands-on

    • Developer Sandbox
      Spin up Red Hat's products and technologies without setup or configuration.
    • Interactive labs
      Learn by doing in these hands-on, browser-based experiences.
    • Interactive demos
      Click through product features in these guided tours.

    Browse by topic

    • AI/ML
    • Automation
    • Java
    • Kubernetes
    • Linux
    • See all topics

    Training & certifications

    • Courses and exams
    • Certifications
    • Skills assessments
    • Red Hat Academy
    • Learning subscription
    • Explore training
  • Build

    Get started

    • Red Hat build of Podman Desktop
      A downloadable, local development hub to experiment with our products and builds.
    • Developer Sandbox
      Spin up Red Hat's products and technologies without setup or configuration.

    Download products

    • Access product downloads to start building and testing right away.
    • Red Hat Enterprise Linux
    • Red Hat AI
    • Red Hat OpenShift
    • Red Hat Ansible Automation Platform
    • See all products

    Featured

    • Red Hat build of OpenJDK
    • Red Hat JBoss Enterprise Application Platform
    • Red Hat OpenShift Dev Spaces
    • Red Hat Developer Toolset

    References

    • E-books
    • Documentation
    • Cheat sheets
    • Architecture center
  • Community

    Get involved

    • Events
    • Live AI events
    • Red Hat Summit
    • Red Hat Accelerators
    • Community discussions

    Follow along

    • Articles & blogs
    • Developer newsletter
    • Videos
    • Github

    Get help

    • Customer service
    • Customer support
    • Regional contacts
    • Find a partner

    Join the Red Hat Developer program

    • Download Red Hat products and project builds, access support documentation, learning content, and more.
    • Explore the benefits

Master your skills: Building skills you can trust

Testing agent skills with skill-creator and promptfoo

October 1, 2026
Dejan Bosanac
Related topics:
Artificial intelligence
Related products:
Red Hat AI

    In my article Standardize project context with AGENTS.md and Agent Skills, I discussed agent context, how skills work, and how to share skills across projects and teams. If you followed along with that article, you now have skills in .agents/skills/, and maybe even a marketplace set up, which is a good foundation to build on. But how do you know your skills actually work? I don't mean "work" in the sense that the YAML parses correctly, but "work" in terms of practical behavior. Does the agent reach for this skill when it should? And does it stay away when it shouldn't?

    In this article, I look at the tooling for authoring skills, what can go wrong once they're in the wild, and how to set up a basic testing workflow that runs in CI.

    Authoring skills with skill-creator

    You can absolutely write a SKILL.md by hand. After all, it's just markdown with some YAML at the top. But once you've created a few, you'll notice patterns: The structure is always the same, the frontmatter follows a formula, and the main decision is really about how to describe the trigger.

    In practice, the best skills don't start as skills at all. You work on something with your agent — say, setting up a pre-commit workflow, and you go back and forth until it works the way you want. Then you tell the agent to capture that workflow as a skill. So you do the work first and then formalize it afterwards.

    That 2nd step is easier with skill-creator, part of Anthropic's skills collection, which you can install as any skill plugin (as with /plugin marketplace add anthropics/skills in Claude Code). It's a full authoring toolkit: You describe your intent conversationally, and the tool interviews you about edge cases and output formats. It then drafts the SKILL.md and helps you iterate on the design. But the interesting part is what happens after the first draft.

    The skill-creator skill includes sub-agents for analyzing and grading skills, a description optimizer that specifically tunes triggering accuracy, and a built-in eval loop. That eval loop is where things get practical, generating a set of test prompts (things that should trigger the skill and things that shouldn't) and running each one multiple times against the agent, and then reporting which prompts triggered correctly and which didn't. You see exactly which queries the skill missed and which ones it fired on incorrectly, so you can adjust the description, and re-run.

    The description optimizer takes this further. It splits your test prompts into a training set and a held-out test set, then iterates automatically — proposing improved descriptions, re-evaluating, and selecting the best one based on the held-out score to avoid overfitting. At the end, you get a before-and-after comparison with the improved description ready to drop into your SKILL.md.

    The description matters so much because, unlike code that throws exceptions when something breaks, a skill that misfires just doesn't run (or it runs when it shouldn't). The agent handles the request itself and your skill sits there unused, and there's no error or warning to indicate the problem. This is the most common way skills fail (silent misfire), and the trigger description is usually the cause. For example, look at this example of a weak description:

    description: Security review guidelines and best practices for code changes

    Compare that to a strong description that's imperative and explicit:

    description: Run a security review on the current code changes. Use when the user asks to check security, review for vulnerabilities, or run a security scan.

    The weak description reads like a code comment, and the agent can already answer questions about a security review without ever invoking the skill. The strong description is an action the agent can't perform from its own context, so it must reach for the skill.

    There's a simple rule that makes a big difference here: Use action verbs, not passive and inactive verbs.

    The good news is that skill-creator's eval loop catches exactly this kind of problem during authoring. While skill-creator's eval loop effectively identifies these issues, the tests are ephemeral—they live only within the current session and do not survive past execution. If someone edits the SKILL.md a month later, there's no way to re-run those checks automatically.

    Testing skills with promptfoo

    The promptfoo project is an open source framework for evaluating LLM applications. It was originally built for testing prompts and model outputs, but it has grown into a general-purpose eval tool that supports everything from simple text assertions to full agent workflows. The core idea is simple: You describe your test cases (prompts) in YAML, the provider to run it against, and assertions about the expected output. Then promptfoo runs the matrix and reports results. Think of it as a test runner, but for LLM behavior instead of code.

    By default, promptfoo supports a wide range of assertion types, including basic string matching (contains, icontains), LLM-graded rubrics (llm-rubric), and custom JavaScript checks. You can test against multiple models in a single run to compare cost, quality, and latency side by side. And because everything is YAML and command-line, it fits naturally into version control and CI pipelines.

    What makes promptfoo particularly useful for skills is its agent SDK providers. Instead of testing a raw model API, you can test through the actual agent runtime (Claude Code, Codex, OpenCode, and so on) so your tests exercise the full skill discovery and routing path. It also provides a skill-used assertion that checks whether the agent actually invoked your skill through proper routing, not just whether it mentioned the skill name in its output.

    The mental model from skill-creator carries over directly. You already have positive and negative prompts in the skill-creator evals. Here's what they look like in promptfoo:

    description: 'Pre-commit check skill eval'providers:  - id: anthropic:claude-agent-sdk    config:      working_dir: ./fixtures/consumer      setting_sources: ['project']      skills: ['pre-commit-check']      append_allowed_tools: ['Read', 'Grep', 'Glob']tests:  - vars:      prompt: 'Run pre-commit checks before committing'    assert:      - type: skill-used        value: pre-commit-check

    There are a few things worth unpacking here. The working_dir points to a fixture directory — a minimal project structure where the eval discovers skills the same way a real agent would, not by browsing the source repo directly. Without this, the agent might find and read SKILL.md from the filesystem, bypassing skill routing entirely — making your test pass, but for the wrong reason.

    The skills: ['pre-commit-check'] filters discovery to the specific skill under test, and auto-allows the skill tool, which is what makes the skill-used assertion work.

    The append_allowed_tools parameter restricts the agent to read-only operations. We don't want things like Bash, Git, or external API calls, because they can have unwanted side effects. This matters more than you might think. During eval testing, a skill can run real actions on the system, like pushing branches, creating PRs, or calling external APIs, and do real damage in a test environment. We'll look at sandboxing strategies in the next part, but for trigger tests, read-only tools are all you need.

    Notice the prompt says "Run pre-commit checks" and not "I want to commit my changes." That second version seems more natural, but it doesn't trigger the skill. The reason is important: The model already knows how to commit code. It can check Git status, stage files, and write a commit message without any help. What it doesn't know is your project's specific workflow — that you need to run the linter first, that tests must pass before committing, that you use conventional commits, and so on. Skills add that project-specific layer, and the prompt needs to signal that the layer is needed. This is the silent misfire I mentioned in the previous section, and it's exactly the kind of thing trigger tests catch.

    Usually, you want 3 kinds of trigger prompts for any skill:

    • Positive prompts should trigger the skill.
    • Near-miss prompts are queries that belong to a sibling skill, not this one. These catch overly broad descriptions that fire when they shouldn't.
    • Negative prompts are completely unrelated requests that should trigger nothing at all.
    tests:  # Positive — should trigger this skill  - vars:      prompt: 'Review this code for security vulnerabilities'    assert:      - type: skill-used        value: security-review  # Near-miss — should trigger a sibling skill, not this one  - vars:      prompt: 'Check if the code follows our naming conventions'    assert:      - type: not-skill-used        value: security-review      - type: skill-used        value: style-check  # Negative — completely unrelated, should trigger nothing  - vars:      prompt: 'What time is the standup meeting?'    assert:      - type: not-skill-used        value: security-review

    The near-miss tests are the most valuable, because they catch the case where your trigger description is too broad and the skill fires for prompts that belong to a different skill.

    Our coding agents are smart enough so that you can just ask them to take the test prompts you developed in skill-creator and write them as promptfoo configuration. This is the natural transition from your authoring phase into more formal testing. You can now keep these eval configs alongside the skill they test, for example,  skills/pre-commit-check/evals/, so the skill and its tests always travel together, with proper version control. You can try this from the companion repo for this article:

    git clone https://github.com/dejanb/skill-workshop.gitcd skill-workshopnpm installnpx promptfoo eval -c examples/01-first-trigger.yaml

    Taking it to CI

    We've come a long way by creating evals, keeping them along with our skills and running locally when they change. You could stop here, but the real value comes when you wire this into CI, because skills can regress over time and code reviews can't catch them every time. For example, someone rewrites the instructions, the diff looks reasonable, a reviewer approves it, and the skill stops working. Trigger tests in CI catch this before it ships, just as with the regular source code in our projects.

    The CI integration is straightforward as promptfoo is a combination of commands and configuration files. For example, there's an official GitHub Action that runs evals and posts results as a PR comment:

    name: Skill Evalson:  pull_request:    paths:      - 'plugins/**/SKILL.md'      - 'plugins/**/evals/**'jobs:  eval:    runs-on: ubuntu-latest    permissions:      contents: read      pull-requests: write    steps:      - uses: actions/checkout@v4      - uses: promptfoo/promptfoo-action@v1        with:          github-token: ${{ secrets.GITHUB_TOKEN }}          config: 'plugins/dev/skills/pre-commit-check/evals/promptfooconfig.yaml'        env:          ANTHROPIC_API_KEY: ${{ secrets.ANTHROPIC_API_KEY }}

    The paths filter means the eval only runs when someone changes a SKILL.md or an eval file, not on every PR. And the action posts a summary comment on the pull request with pass/fail counts, so reviewers see the eval results without leaving the review.

    One thing that works nicely for CI adoption is progressive rollout. Start by running evals as an informational step, by letting the pipeline report results but not block the merge. Give the team some time to see what the tests catch and build confidence. Only after that, flip it to a blocking gate and actually fail builds when issues are discovered. Trying to make it a hard gate on day one could lead to initial annoyance resulting in resisting the process.

    The cost is also surprisingly low. A trigger test suite for one skill (a handful of positive, negative, and near-miss prompts) runs at about $0.10-0.15 per eval with current models and prices. The prices will certainly change, but this is just to illustrate that even today the cost is less than the time you'd spend debugging a skill that silently stopped working.

    What you've learned so far

    Like all other components of your software projects, you have to keep your skills versioned and tested. There are many ways in which they can stop reflecting the initial author's intention: A skill that doesn't trigger is invisible, a well-meaning edit can break something that was working fine, and so on.

    You have some good tooling you can use to author and evaluate skills. The skill-creator tool is great for the authoring loop, promptfoo can make tests that survive past the session, and you can wrap it all in CI to catch regressions before they ship.

    Once you've covered basic trigger testing, you can start thinking about the quality of your skills. That's where things like testing execution quality with recall and precision, scenario-based testing for orchestration skills, testing across agents and models come into play. You also must take special consideration for skills that need to run commands you don't want executing in a test environment. I'll cover all these topics in the next part of this series.

    Related Posts

    • Standardize project context with AGENTS.md and Agent Skills

    • Build golden path CI/CD workflows in Red Hat Developer Hub

    • Add automated AI evaluations to your CI/CD pipeline

    • Build trust in your CI/CD pipelines with OpenShift Pipelines

    • Bringing custom knowledge to agents with AutoRAG

    • Behavioral testing for AI agents

    Recent Posts

    • Master your skills: Building skills you can trust

    • GPU virtualization at scale: AMD MI300X SR-IOV on Red Hat OpenStack Services on OpenShift

    • Beyond pass or fail: The new era of chaos engineering

    • Backdoors in LLMs: Why model scanning isn't enough

    • Distributed training on OpenShift AI 3.4 with Kubeflow Trainer v2

    What’s up next?

    Learning Path Red Hat AI

    How to run AI models in cloud development environments

    This learning path explores running AI models, specifically large language...
    Red Hat Developers logo LinkedIn YouTube Twitter Facebook

    Platforms

    • Red Hat AI
    • Red Hat Enterprise Linux
    • Red Hat OpenShift
    • Red Hat Ansible Automation Platform
    • See all products

    Build

    • Developer Sandbox
    • Developer tools
    • Interactive tutorials
    • API catalog

    Quicklinks

    • Learning resources
    • E-books
    • Cheat sheets
    • Blog
    • Events
    • Newsletter

    Communicate

    • About us
    • Contact sales
    • Find a partner
    • Report a website issue
    • Site status dashboard
    • Report a security problem

    RED HAT DEVELOPER

    Build here. Go anywhere.

    We serve the builders. The problem solvers who create careers with code.

    Join us if you’re a developer, software engineer, web designer, front-end designer, UX designer, computer scientist, architect, tester, product manager, project manager or team lead.

    Sign me up

    Red Hat legal and privacy links

    • About Red Hat
    • Jobs
    • Events
    • Locations
    • Contact Red Hat
    • Red Hat Blog
    • Inclusion at Red Hat
    • Cool Stuff Store
    • Red Hat Summit
    © 2026 Red Hat

    Red Hat legal and privacy links

    • Privacy statement
    • Terms of use
    • All policies and guidelines
    • Digital accessibility
    Ask AI