No reviews yet. Reviews from researchers who actually ran Humanity's Last Exam are what this platform is for.
bioagent-bench/bioagent-bench
Benchmark for evaluating LLM agents in bioinformatics
Future-House/BixBench
Benchmark for LLM-based Agents in Computational Biology
darlednik/GENEB
GENEB: ICML 2026 benchmark for genomic foundation models across 100 tasks and 13 functional categories.