No reviews yet. Reviews from researchers who actually ran PaperBench are what this platform is for.
harbor-framework/harbor
Framework for evaluating and improving agents
scicode-bench/SciCode
A benchmark that challenges language models to code solutions for scientific problems
claw-eval/claw-eval
Claw-Eval is an evaluation harness for evaluating LLM as agents. All tasks verified by humans.