Terminal-Bench-Science: Evaluating AI agents on scientific research workflows

(terminal-bench-science.ai)

13 points | by matt_d an hour ago ago

2 comments

  • akshay_akula 3 minutes ago

    Evals on actual research workflows is the right direction, most agent benches are toy tasks.

  • rubslopes 43 minutes ago

    I'm glad to see that GPT Sol beats Opus at least in Mathematical Sciences, because that's my need right now, and I much prefer GPT's prose style.