← Back to match
ResearchClawBench
Benchmark evaluating whether AI agents can independently conduct scientific research to publication quality.
PlatformWebfreeglobal
ResearchClawBench is a benchmark that tests whether AI coding agents can independently conduct scientific research — from reading raw data to producing a publication-quality report — and scores the results against real, human-authored papers across 40 tasks in 10 disciplines. It's aimed at AI research teams building or evaluating autonomous research agents, with a public leaderboard for submitted agents.
Categories
AI research toolsbenchmarksAI agent evaluation

