14 research AI agents in benchmarks, tracked with live GitHub metrics, companion-paper metadata and Verified Run reviews from researchers who ran them. The most-starred records right now: PinchBench, ClawBench, Claw-Eval.
pinchbench/skill
PinchBench is a benchmarking system for evaluating LLM models as OpenClaw coding agents. Made with 🦀 by the humans at https://kilo.ai
TIGER-AI-Lab/ClawBench
Open-source benchmark for browser AI agents on daily tasks.
claw-eval/claw-eval
Claw-Eval is an evaluation harness for evaluating LLM as agents. All tasks verified by humans.
InternScience/ResearchClawBench
🦞 ResearchClawBench: Evaluating AI Agents for Automated Research from Re-Discovery to New-Discovery
ktwu01/benchmark-radar
Track 14,810+ AI benchmark, eval, dataset, and data-quality records from 37 public sources, with linked evidence and daily updates.
ArcInstitute/cell-eval
Comprehensive suite for evaluating perturbation prediction models
Future-House/BixBench
Benchmark for LLM-based Agents in Computational Biology
evolvent-ai/ClawMark
🦞 ClawMark: A Living-World Benchmark for Multi-Day, Multimodal Coworker Agents
zhao-zy15/RareArena
A Comprehensive Rare Disease Diagnostic Dataset with nearly 50,000 patients covering more than 4000 diseases
bioagent-bench/bioagent-bench
Benchmark for evaluating LLM agents in bioinformatics
scaleapi/DrugDiscoveryBench
Opensource repository containing task and image data for DrugDiscoveryBench
Insilico-org/longeclaw
Digital Longevity Research
doi.org/10.1016/j.cell.2026.08.004
Cell paper — Arc Institute community benchmark of zero-shot perturbation-response generalization across cellular contexts.
huggingface.co/zifeng-ai/BioDSA-1K
Benchmark for data-driven biomedical hypothesis validation — 1,029 hypothesis-centric tasks with 1,177 analysis plans curated from 300+ published studies on cBioPortal patient data. Includes non-verifiable hypotheses the data can neither support nor refute. Dataset on Hugging Face (ODbL); ships with the BioDSA framework.
Star counts on cards are live GitHub metrics; full records include papers, licenses and review evidence.