Ask HN: Have LLMs Plateaued?

9 points | by leandrobon 2 days ago ago

26 comments

  • brudgers 19 hours ago

    How much more productivity will we be able to squeeze from them?

    Probably, depends on how you measure productivity.

    If you measure productivity in terms of number of automated bureaucratic events (e.g. creating files, organizing files, generating lines of code, finding bugs in code, generating emails, responding to email) then yes productivity will continue to increase because LLM's are the killer app for increasing bureaucratic events.

    If you measure productivity in terms of changes to the material world (e.g. traditional things like trade goods, buildings, food stuffs, irrigation systems, transportation networks, etc.) then no because LLM's have little meaningful impact on those activities...no AGI is going to harvest lettuce for our wedge salads).

    • fragmede 12 hours ago

      The advances in robotics makes it look like a humanoid robot that could harvest lettuce is only a decade or two away. Making the robot is the easy part. The code to drive it has been the hard part. Until now, that is.

      • brudgers 11 hours ago

        California has about 800,000 agricultural laborers and about 400,000 FTE’s.

        That’s a lot of robots…and they probably won’t be able to drive themselves from field to field.

        • fragmede 9 hours ago

          Self driving cars are here today, driving in city traffic. Self driving a truck from one field to another seems like it would be eminently doable as far as the technology goes. However, I doubt a bunch of humanoid robots is the optimal robot configuration for harvesting lettuce. Tesla made 1.6 million cars in 2025. Toyota made 11 million. If we handwave that a humanoid robot is roughly as hard to manufacture as a car at scale, there's gonna be a lot of robots.

  • wmf 2 days ago

    Obviously not. And if you look at the things LLMs do poorly there's clearly plenty of room for improvement.

    • kazinator 2 days ago

      LLM shill argument in a nutshell: "while there is room for improvement in catalytic converters, they are improving practically by the month! Within a decade, internal combustion engines will put out air that is so breathable, it could be used for ventilating a maternity ward."

      • cyanregiment a day ago

        The exhaust from a hydrogen fuel cell car is drinkable water.

        However, experts do not recommend drinking it straight from the tailpipe because the water can pick up dust, dirt, and chemical residues.

        (it had to be disclaimed!)

  • iagooar 20 hours ago

    No, LLMs have not plateaued, and each of the latest releases has been a proof of that.

    Fable 5 is a model you instantly FEEL how smart and superior it is. GPT 5.6 Sol is a HUGE incremental improvement in multiple directions and dimensions.

    DeepSeek v4 Flash 0731 is a huge improvement over the preview version, using exactly the same architecture.

    Kimi K3 gets open weight models very, very close to the frontier.

    No my friend, we are not done yet.

    • taurath 16 hours ago

      > you instantly FEEL how smart and superior it is

      I felt taken by the change in the system model personality and writing style compared to opus, but I also found it to be much less impressive than I was expecting - let alone that the cost was incredibly high when not given for free.

      Are you sure your reaction is not primarily to the improved ergonomics of Fable?

  • tripleee a day ago

    I think the big change was agents. The LLM improvements after that point have been relatively minor in impact.

    • ldng a day ago

      Agents are the proof that LLMs Plateaued : you need the loop to get further.

  • NishanStepak a day ago

    In terms of the amount of information in the current LLMs, I think the largest is 5.6 trillion bytes of memory, it is miniscule to what is actually out there. There are over 13 zettabytes of information on the internet. Part of the capacity is about absorbing information. It has a lot further to go. If there are emergent properties with additional knowledge in LLMs, it is scratching the surface. I wonder what will happen with zettabyte computing.

  • atleastoptimal 19 hours ago

    Obviously not, have you seen the 10 mathematical advances OpenAI found with Astra?

    https://openai.com/index/ten-advances-in-mathematics/

  • _Yassine_ a day ago

    I don't think the models are the bottleneck. Instead, how and where we deploy them remains significantly underrated and underutilized.

    Like imagine how crazy that you can get human-like intelligence in a small device, we should be able to do more than a chat interface.

  • markmatsushima a day ago

    I think the giant leap in AI technology was the Transformer. ChatGPT is based on it, with an incredible chat user interface. In terms of technology, there has been no similarly significant innovation since then. But user interfaces and training data keep improving, which continues to increase our productivity.

  • kimjune01 a day ago

    nowhere near plateau, but right at the inflection point of diminishing returns imo

    • hiramwen a day ago

      This is the best take. Return on capital is diminishing from intense competition so improvements get hidden away like the gems they are.

      1) Frontier labs have no incentive to give the general public their best anymore; it's instantly distilled off of them. Why not charge governments and big corps real money to use the real good stuff instead? 2) So we get distilled-off-frontier public APIs like 5.6 and Fable/Opus. And the open source labs are distilling off of those. 3) There's a lot of benchmark hacking right now among all the publicly available models, actual usability of Opus for coding is far below its benchmarks suggest. 4) But context window, cybersecurity, logical coherence, and tool usage are absolutely better on Fable and Sol. It looks to me their internal tools definitely even better and not plateauing. But we won't get to use it.

  • kazinator 2 days ago

    I suspect you could plug GPT4 from 2023 into the integrations and workflows and it would be about the same.

    The "models are getting better with every breath you take" is just an unproven claim that the AI companies want everyone to believe and want the evangelists to use in every argument.

    The AI companies have to put out new models regularly, not because of any improvement but because the failure to do so will signal stagnation. That's how it works in the tech industry. It's not like the wine or cheese industry where you just follow your hundreds-of-years-old recipe exactly and people buy the product because of that.

  • moizrocky1 a day ago

    Been using Kimi, and yeah... the American ones seems like heading nowhere new.

  • rtanasa a day ago

    We've been seeing terrible performance from Opus 5, so it does look like a plateau. They most certainly reached the marginal returns realm.

  • AnimalMuppet 2 days ago

    I'm not answering your actual question, but... even if they have totally plateaued technically, there's still some more improvement left to be had from people learning how best to use them (and when not to).

    • vrighter a day ago

      how can you learn to use a tool that keeps changing unpredictably?

      • AnimalMuppet a day ago

        Well, I was assuming that what you learned about driving generation X of a tool would be at least mostly applicable to driving generation X+1 of it. That may be false, at least for some values of X.

  • cyanregiment a day ago

    Definitely, yes, in the way the calculator did.

    This is the technology - this is what it is. We've all seen it at this point. We know what it can and can't do.

    Sure, there's Mistral 7b running locally on a Macbook vs some large frontier model with MoE, context caching and a nice UI - but it's not all that different.

    Don't we know what LLMs are at this point?

    But it's not like the mystery of AI is fully solved and all opportunity has ended. Now we have to see what can be done with it - that part I think is still mostly unexplored.

    People are trying things in the software space, in music, video, legal, and there are quite a few "AI for AI" companies (AI analytics, infra solutions) - so we'll get some innovation there.

    "Harnessing" - People will figure out better ways with model routing, multi-modality, caching, combining things we already know.

    I'm slowly seeing the philosophy of mind creep in and make contact with software engineering. Somewhere between "AGI" and Chalmers' Hard Problem of Consciousness we'll get some new innovation around the concept of thinking itself and what it means to be an intelligent being.

    Autonomous driving is paving the way (hehe) for autonomous humanoid robots. A concept I originally scoffed at, but now consider to be a very likely thing to happen sooner than we might realize. This is going to change so much, it alone is enough reason to say "AI hasn't peaked" (and these will involve interdisciplinary models like vision, navigation, and LLMs).

    It's like asking if "programming has peaked" in 1999 and it would kinda be a "Yes". And then a couple years later ColdFusion would hit the shelves in a giant box like a Deluxe Edition RPG - if you need a laugh: https://www.reddit.com/r/webdev/comments/1o26h3l/i_have_deve...

  • kypro a day ago

    AI is beginning to do PHD-level math.

    If we're nearing the point where you can spin up 1,000 agents to look for ways improve existing models, then automatically run experiments to validate those ideas, we're more or less at RSI.

    As always with AI progress, compute will bottleneck this early on, but a few efficiency improvements could dramatically increase this pace of progress.

    I suspect we are at most 24 months from FOOM, but I suspect within about 6-12 months most frontier AI labs will be claiming the majority of their AI research will be AI-driven.

  • aborsy a day ago

    I hope so!