Registry / all records
208 research AI agents · live GitHub metrics · community reviews
Retired projects →pinchbench/skill
PinchBench is a benchmarking system for evaluating LLM models as OpenClaw coding agents. Made with 🦀 by the humans at https://kilo.ai
claw-eval/claw-eval
Claw-Eval is an evaluation harness for evaluating LLM as agents. All tasks verified by humans.
TIGER-AI-Lab/ClawBench
Open-source benchmark for browser AI agents on daily tasks.
InternScience/ResearchClawBench
🦞 ResearchClawBench: Evaluating AI Agents for Automated Research from Re-Discovery to New-Discovery
Future-House/BixBench
Benchmark for LLM-based Agents in Computational Biology
evolvent-ai/ClawMark
🦞 ClawMark: A Living-World Benchmark for Multi-Day, Multimodal Coworker Agents
zhao-zy15/RareArena
A Comprehensive Rare Disease Diagnostic Dataset with nearly 50,000 patients covering more than 4000 diseases
bioagent-bench/bioagent-bench
Benchmark for evaluating LLM agents in bioinformatics
huggingface.co/zifeng-ai/BioDSA-1K
Benchmark for data-driven biomedical hypothesis validation — 1,029 hypothesis-centric tasks with 1,177 analysis plans curated from 300+ published studies on cBioPortal patient data. Includes non-verifiable hypotheses the data can neither support nor refute. Dataset on Hugging Face (ODbL); ships with the BioDSA framework.