13 points | by matt_d an hour ago ago
2 comments
Evals on actual research workflows is the right direction, most agent benches are toy tasks.
I'm glad to see that GPT Sol beats Opus at least in Mathematical Sciences, because that's my need right now, and I much prefer GPT's prose style.
Evals on actual research workflows is the right direction, most agent benches are toy tasks.
I'm glad to see that GPT Sol beats Opus at least in Mathematical Sciences, because that's my need right now, and I much prefer GPT's prose style.