No reviews yet. Reviews from researchers who actually ran MCP Atlas are what this platform is for.
Future-House/BixBench
Benchmark for LLM-based Agents in Computational Biology
evolvent-ai/ClawMark
🦞 ClawMark: A Living-World Benchmark for Multi-Day, Multimodal Coworker Agents
THUDM/AgentBench
A Comprehensive Benchmark to Evaluate LLMs as Agents (ICLR'24)