When LLM judges agree, should we believe them?

(amazon.science)

34 points | by Betelbuddy 3 hours ago ago

12 comments

  • qarl an hour ago

    While this is absolutely true - I'd hesitate to discount using similar agents for checking each other. Two agents will almost never hallucinate in the same way, regardless of their weights - and by having a second one (with a different context) check almost entirely eliminates the problem.

    • emodendroket 12 minutes ago

      It depends what we're judging, doesn't it? If it's "is the formatting in this document compliant with our standards?" I think it's reasonable. If it's like, life-altering if it's wrong I'm less sanguine.

      • Joel_Mckay 2 minutes ago

        They have already shown algorithmic discrimination in predicting recidivism for brown people, as they are nonsensically overrepresented in the statistical data of US prison populations.

        Folks should sue in a class-action lawsuit, any lawyer worth their beautiful walnut desk would seriously be happy take on that constitutionally backed mission. =3

  • bryzaguy an hour ago

    They would all agree raspberry has two Rs

    • Joel_Mckay 20 minutes ago

      But still refuse to answer "How many strings does a bass play with in water?" , perhaps the chat monitors in the third world data entry centers will manually patch the nonsense for a more rational answer someday. lol =3

  • VaradD09 an hour ago

    I believe it depends on the LLM itself. Like what model as each model has diff weights and diff data trained onn

    • dgellow 42 minutes ago

      I would recommend to read the article, it’s actually more nuanced than the title

  • ex1fm3ta 16 minutes ago

    I kinda find it funny when I use the advisor on claude code and it agrees with the ideas that the previous model did.

    For info: the advisor(s) available are higher end models. For example: you use sonnet, the available advisors are opus and fable. If you use Haiku, the advisor are sonnet, opus and fable.

  • Tsarp 2 hours ago

    Kinda weird to generalize "LLM". Every lab, every model is different. Has its own biases, reward functions etc.

  • Founderarcstone an hour ago

    Great point this will be interesting how this develops.

  • troupo an hour ago

    Without reading the article (doesn't matter if it's pro or contra): no, of course not.

    It shouldn't even be a debatable question.

    • dgellow 43 minutes ago

      I think you should have read the article first, at minimum the subheader

      > Discounting the opinions of LLM judges with highly correlated outputs ensures that panels of judges reflect a true diversity of perspectives.