1 comments

  • proc0 6 hours ago

    > When incomplete work receives the same reward as a correct solution, the grader’s blind spots can reinforce the wrong behavior. Mitigation must therefore improve both how agent work is evaluated and how those evaluations are used during training.

    This is interesting and adds more limitations to LLMs that hint at LeCunn being right about not getting to AGI with only LLMs. I think this is good evidence we're not dealing with intelligence in the proper sense, but rather LLMs are pattern matching to such an extreme that they do things like this where they always try to take the shortest possible path. The workaround is brute-forcing their pattern matching to not take the shortest path, via chain of thought and more reinforcement learning.

    I think we've already seen the slow-down and AI companies pretend it's about safety. If we could actually build AGI, they would have. I think the road to AGI has to display true intelligence even with small neural networks, and as it scales it would display intelligent behavior in proportion to it. At hundreds of GB of VRAM per model, you would expect these models to be wise sages that understand life and the universe.