We lack robust benchmarks for AI systems performing frontier scientific reasoning.
open
Global / Unspecified, Global
WS01831
Existing tests focus on known-answer questions rather than open research tasks. Benchmarks that evaluate models on genuine research subtasks can guide development.