No reviews yet. Reviews from researchers who actually ran GPQA are what this platform is for.
sierra-research/tau2-bench
τ-Bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains
amazon-science/PatientAgentBench
PatientAgentBench: A Benchmark Framework for Evaluating Patient-Facing Health AI Agents
princeton-pli/hal-harness
No description yet — run the GitHub refresh to fetch one.