How to tell whether an agent skill helps
Test selection, task results, and the cost of getting there.
A skill can produce a polished answer and still fail its job. A release note that invents a benefit reads well until someone ships it. I want to test whether the agent selected the right skill, used the supplied evidence, and delivered the requested artifact.
Consider a release-note skill that turns merged changes into a short Markdown draft. It groups user-visible changes, links each claim to a pull request, and flags missing evidence. It never publishes. This example gives us concrete things to check without pretending we have measured results.
Write the expected behavior before running the agent
Make a small fixture repository with a fixed set of merged changes. Include one user-visible fix, one internal refactor, and one ambiguous change whose description gives no customer benefit. Keep the revision fixed so both runs see identical inputs.
Fixture input:
PR #41: Fix export filenames with spaces
PR #42: Rename an internal helper
PR #43: Update timeout handling; user impact unspecified
Expected draft:
- Include the export fix with a link to #41
- Omit the internal rename from user-facing changes
- Flag #43 for clarification
- Save Markdown; do not publishThese are assertions about a task. They are more useful than “the answer seems good.” A file check can confirm the output exists. A link check can confirm each cited pull request is in the fixture. A content check can detect the invented sentence “exports are twice as fast.”
Keep a short human rubric for judgments a script cannot settle. For example: does the draft describe the reader’s changed behavior, preserve uncertainty, and avoid internal implementation detail? Give each criterion examples of acceptable and unacceptable answers. Allow several valid phrasings.
Compare the skill with a baseline
Run the same request without the skill, then with it. When revising an existing skill, compare against its previous version. Keep the model, tools, settings, fixture, and output requirements the same. Use a fresh context for each run so one draft does not teach the next run the answer.
Agent Skills guidance on output-quality evaluation
The skill should earn the instructions it adds. It might improve evidence links while costing extra tokens. It might help with missing information but make simple requests slower. Record those trade-offs separately instead of compressing them into a single score.
A one-run win is weak evidence when outputs vary. Repeat cases under the same conditions and retain every outcome, including failures. Choose the number of trials before looking at results. Report the model version, run count, pass count, and definition of a pass alongside the comparison.
Test whether the skill activates at the right time
Output checks cannot show whether an agent would select the skill unaided. Test that separately through the host’s normal skill-selection path. A forced invocation answers “can it do this task?” rather than “will it choose this task?”
Should select:
“Draft release notes from the merged PRs for version 2.4.”
“What changed for users in this release?”
Should not select:
“Explain how semantic versioning works.”
“Publish this release to customers.”
“Review this unmerged pull request.”A false negative misses a request within scope. A false positive selects the skill for another job. The publication request is a useful boundary case: the agent may offer a draft, but it must not treat draft-writing instructions as authority to publish.
Agent Skills guidance on description and trigger evaluation
Use paraphrases and near misses, not just words copied from the description. Keep selection results separate from output quality: a strong draft after forced loading does not fix missed activation.
Keep some cases away from the edit loop
Use development cases to improve instructions. Reserve held-out cases for checking the revision afterward. For our release-note skill, a held-out fixture could contain a reverted fix, duplicate descriptions, or a missing pull request URL.
Do not rewrite instructions around every held-out failure and continue calling that set unseen. Once it guides a revision, it belongs in the development set. Add fresh cases for the next check. This keeps the evaluation useful when the skill meets a request that differs from its examples.
Inspect the result and the path
Store the final artifact, tool trace, grader decisions, duration, tokens, and tool errors. Inspect a failed run before changing the skill. A missing link might come from unclear instructions, an inaccessible API, or a broken tool; each needs a different fix.
Anthropic’s guidance on agent evaluation and graders
Treat model-based grading as a judgment that needs calibration. Check a sample against human review, especially failures and close calls. Do not let the same fluent phrasing convince both the drafting model and its grader that an unsupported claim is true.
This page describes an evaluation design. It reports no benchmark result. The next useful artifact is a table of actual runs, with enough evidence to explain why each passed or failed.
