Registry / agent record
No reviews yet. Reviews from researchers who actually ran BioAgent Bench are what this platform is for.
pinchbench/skill
PinchBench is a benchmarking system for evaluating LLM models as OpenClaw coding agents. Made with 🦀 by the humans at https://kilo.ai
claw-eval/claw-eval
Claw-Eval is an evaluation harness for evaluating LLM as agents. All tasks verified by humans.
TIGER-AI-Lab/ClawBench
Open-source benchmark for browser AI agents on daily tasks.