Registry / benchmarks
37 benchmarks · 15 tracked with score observations · 144 source documents · 2520 observations
tracking since 2026-10-02 · last ingest 2026-10-02
Each benchmark carries five evidence layers, modeled after Benchmark Radar: identity and artifacts, the document registry that reports scores, score observations partitioned by source — never merged into cross-source rankings — and aggregate counts. Observations accumulate on every ingest, so each page doubles as a long-term tracking record. Score observations and source documents located via Benchmark Radar (benchmark-radar.org; code MIT, content CC BY-NC-SA 4.0, non-commercial use with attribution). Every row cites its primary source document; all descriptions are AgentX's own. Sources are kept as separate partitions and never merged into cross-source rankings.
37 of 37 benchmarks
| Benchmark | Observed | Standings | Documents | Released | Repo |
|---|---|---|---|---|---|
| GPQA Knowledge & QA | 851 794 subjects | 0.95 | 33 | Nov 2023 | idavidrein/gpqa |
| Humanity's Last Exam Knowledge & QA | 708 677 subjects | 0.55 | 24 | Jan 2025 | centerforaisafety/hle |
| SciCode Research & Discovery | 600 594 subjects | 0.60 | 4 | Jul 2024 | scicode-bench/SciCode |
| Tau2-Bench Tool Use | 89 41 subjects | 0.99 | 9 | Jun 2025 | sierra-research/tau2-bench |
| Terminal-Bench 2.0 Coding & SWE | 71 63 subjects | 0.83 | 22 | May 2025 | harbor-framework/terminal-bench-2 |
| OSWorld Computer Use | 48 47 subjects | 0.86 | 8 | Apr 2024 | xlang-ai/OSWorld |
| Toolathlon Tool Use | 47 42 subjects | 0.76 | 11 | Nov 2025 | hkust-nlp/Toolathlon |
| MCP Atlas Tool Use | 41 36 subjects | 0.88 | 9 | Nov 2025 | scaleapi/mcp-atlas |
| SuperGPQA Knowledge & QA | 38 37 subjects | 0.74 | 5 | Feb 2025 | SuperGPQA/SuperGPQA |
| Claw-Eval Computer Use | 14 14 subjects | 0.81 | 1 | – | claw-eval/claw-eval |
| PinchBench General Agents | 6 6 subjects | 0.90 | 1 | – | pinchbench/skill |
| PaperBench Research & Discovery | 3 3 subjects | 0.93 | 2 | Mar 2025 | openai/frontier-evals |
| MLE-bench Research & Discovery | 2 2 subjects | 0.64 | 3 | Oct 2024 | openai/mle-bench |
| BixBench Bio & Medicine | 1 1 subjects | 0.81 | 1 | Jan 2025 | Future-House/BixBench |
| ResearchClawBench Research & Discovery | 1 1 subjects | 0.17 | 2 | – | InternScience/ResearchClawBench |
| ActiveSciBench Research & Discovery | – 0 subjects | – | 1 | May 2026 | scientific-discovery/LLM-AutoSciLab |
| AgentBench General Agentsnot yet tracked | – 0 subjects | – | – | Jul 2023 | THUDM/AgentBench |
| AutoMedBench Bio & Medicine | – 0 subjects | – | 1 | Jun 2026 | AutoMedBench/AutoMedBench |
| BioAgent Bench Bio & Medicinenot yet tracked | – 0 subjects | – | – | Dec 2025 | bioagent-bench/bioagent-bench |
| BioDSA-1K Bio & Medicinenot yet tracked | – 0 subjects | – | – | Apr 2025 | huggingface.co/zifeng-ai/BioDSA-1K |
| ClawBench Computer Usenot yet tracked | – 0 subjects | – | – | – | TIGER-AI-Lab/ClawBench |
| ClawMark General Agentsnot yet tracked | – 0 subjects | – | – | – | evolvent-ai/ClawMark |
| DrugDiscoveryBench Chemistry & Drugsnot yet tracked | – 0 subjects | – | – | Aug 2026 | scaleapi/DrugDiscoveryBench |
| GENEB Bio & Medicine | – 0 subjects | – | 1 | Jun 2026 | darlednik/GENEB |
| Harbor Coding & SWEHarnessnot yet tracked | – 0 subjects | – | – | Aug 2026 | harbor-framework/harbor |
| HealthAgentBench Bio & Medicine | – 0 subjects | – | 1 | Jun 2026 | microsoft/HealthAgentBench |
| HeurekaBench Bio & Medicinenot yet tracked | – 0 subjects | – | – | Dec 2025 | mlbio-epfl/HeurekaBench |
| Holistic Agent Leaderboard General Agentsnot yet tracked | – 0 subjects | – | – | Sep 2025 | princeton-pli/hal-harness |
| HybridDeepResearch General Agents | – 0 subjects | – | 1 | Sep 2026 | Snowflake-AI-Research/HybridDeepResearch |
| Longevity Claw Research & Discoverynot yet tracked | – 0 subjects | – | – | Aug 2026 | Insilico-org/longeclaw |
| MedAgents-Bench Bio & Medicine | – 0 subjects | – | 1 | Mar 2025 | gersteinlab/MedicalAgentsBench |
| PatientAgentBench Bio & Medicine | – 0 subjects | – | 1 | Jul 2026 | amazon-science/PatientAgentBench |
| RE-Bench Research & Discovery | – 0 subjects | – | 1 | Nov 2024 | METR/RE-Bench |
| RareArena Bio & Medicinenot yet tracked | – 0 subjects | – | – | Jan 2026 | zhao-zy15/RareArena |
| SWE-bench Coding & SWEnot yet tracked | – 0 subjects | – | – | Sep 2023 | SWE-bench/SWE-bench |
| ScienceAgentBench Coding & SWEnot yet tracked | – 0 subjects | – | – | Sep 2024 | OSU-NLP-Group/ScienceAgentBench |
| deep_research_bench Research & Discovery | – 0 subjects | – | 1 | Jun 2025 | Ayanami0730/deep_research_bench |
“Observed” counts every recorded score observation with its source document — the tracking history, not a leaderboard. Standings show the best current observation per subject and instrument; the full history lives on each benchmark page and in the public API.