4 comments

  • sathish316 17 minutes ago

    What does it mean when Fable 5 is 1st place and Opus 5 is 3rd place, while Claude code is 7th place? Which model and effort is used for Claude in 7th place, compared to 1st and 3rd?

  • spullara an hour ago

    They are all different problems for the different languages. I was hoping this was a benchmark that attempted to see which languages were more efficient to use with which models.

  • dia80 an hour ago

    Why test Fable high effort vs Sol medium? Especially when Sol comes out 4-5x cheaper in their tests at those effort levels.

    • cbg0 an hour ago

      I think you just answered your own question.

      Edit: In DeepSWE Sol High scores the same as Fable High for ~1/3 of the cost.