No reviews yet. Reviews from researchers who actually ran AgentBench are what this platform is for.
SWE-bench/SWE-bench
SWE-bench: Can Language Models Resolve Real-world Github Issues?
xlang-ai/OSWorld
[NeurIPS 2024] OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
ArcInstitute/cell-eval
Comprehensive suite for evaluating perturbation prediction models