Matchboxmatchbox
← Back to match

agent-skills-eval

Test runner that proves whether an Agent Skill actually improves model output, with a side-by-side report.

Desktopfreeglobal

A desktop test runner that measures whether an Agent Skill improves model responses by running the same prompts twice—with and without the skill—having a judge model grade both, and producing a side-by-side report of the measured lift. It addresses the lack of objective proof before publishing skills and is aimed at developers and teams building or evaluating Agent Skills; it is open-source (MIT), framework-agnostic, and can be self-hosted.

Categories
AI agent evaluationdeveloper tools

Full match profile

Behind the summary, Matchbox keeps a richer profile of agent-skills-eval - the signals our matcher actually reads to decide when to surface it. It stays private; claim the listing to see and control it.

  • Problem & pain-point mapping
  • Who we surface it to (audience fit)
  • What it's a strong alternative to
  • Trust & credibility signals

Try Matchbox with your own problem

Describe what is not working - we’ll show you whether agent-skills-eval (or something else) actually fits.