Loading…
No reviews yet. Reviews from researchers who actually ran deep_research_bench are what this platform is for.
huggingface.co/zifeng-ai/BioDSA-1K
Benchmark for data-driven biomedical hypothesis validation — 1,029 hypothesis-centric tasks with 1,177 analysis plans curated from 300+ published studies on cBioPortal patient data. Includes non-verifiable hypotheses the data can neither support nor refute. Dataset on Hugging Face (ODbL); ships with the BioDSA framework.
InternScience/ResearchClawBench
🦞 ResearchClawBench: Evaluating AI Agents for Automated Research from Re-Discovery to New-Discovery