Registry / agent record
No reviews yet. Reviews from researchers who actually ran PinchBench are what this platform is for.
claw-eval/claw-eval
Claw-Eval is an evaluation harness for evaluating LLM as agents. All tasks verified by humans.
TIGER-AI-Lab/ClawBench
Open-source benchmark for browser AI agents on daily tasks.
InternScience/ResearchClawBench
🦞 ResearchClawBench: Evaluating AI Agents for Automated Research from Re-Discovery to New-Discovery