← Back to match
agent-skills-eval
Test runner that proves whether an Agent Skill actually improves model output, with a side-by-side report.
Desktopfreeglobal
A desktop test runner that measures whether an Agent Skill improves model responses by running the same prompts twice—with and without the skill—having a judge model grade both, and producing a side-by-side report of the measured lift. It addresses the lack of objective proof before publishing skills and is aimed at developers and teams building or evaluating Agent Skills; it is open-source (MIT), framework-agnostic, and can be self-hosted.
Categories
AI agent evaluationdeveloper tools

