No reviews yet. Reviews from researchers who actually ran HybridDeepResearch are what this platform is for.
openai/mle-bench
MLE-bench is a benchmark for measuring how well AI agents perform at machine learning engineering
ktwu01/benchmark-radar
Track 20,710+ AI benchmark, eval, dataset, and data-quality records from 37 public sources, with linked evidence and daily updates.
SuperGPQA/SuperGPQA
No description yet — run the GitHub refresh to fetch one.