← Back to match
Caliper
Evaluation harness that measures whether an AI agent skill actually works, across Claude Code, Codex, Pi, Hermes.
Desktopfreeglobal
This is an evaluation harness for AI agent skills that runs against Claude Code, Codex, Pi, or Hermes. Users write a specification describing what a working skill looks like, run it repeatedly to get a trackable success rate, and can remove a skill and re-run the same tasks to prove it is actually contributing. It suits developers building and maintaining AI agent skills who want to catch regressions.
Categories
developer-toolstestingai-agents
Something wrong with this listing — dead link, not a real product, wrong info?

