Registry / agent record
No reviews yet. Reviews from researchers who actually ran Claw-Eval are what this platform is for.
pinchbench/skill
PinchBench is a benchmarking system for evaluating LLM models as OpenClaw coding agents. Made with 🦀 by the humans at https://kilo.ai
TIGER-AI-Lab/ClawBench
Open-source benchmark for browser AI agents on daily tasks.
InternScience/ResearchClawBench
🦞 ResearchClawBench: Evaluating AI Agents for Automated Research from Re-Discovery to New-Discovery