MLE-bench is a benchmark for measuring how well AI agents perform at machine learning engineering
Overview
MLE-bench is a benchmarks AI agent developed by openai with 1.8k GitHub stars, written in Python, currently Stable.
MLE-bench is a benchmark for measuring how well AI agents perform at machine learning engineering
GitHub metrics refresh daily — last fetched today.