The AI Productivity Gap

(bjorg.bjornroche.com)

47 points | by kiyanwang 3 hours ago ago

41 comments

  • matthorse an hour ago

    Writing code is a small part of everyday's job of a software engineer. The article's table reflects this fairly well.

    AI compresses implementation time for an individual engineer, but architecture decisions, design reviews, integration, testing, deployment, and production validation remain largely serial activities. If code generation speeds up by 5x while those bottlenecks don't, you've mostly increased the team's work queue rather than its throughput.

    With the current capabilities, models still need constant babysitting and course correction. An engineer who lacks the skills to guide them can end up creating more work for the rest of the team. AI makes it easy to generate code faster than you can understand it, and that cost is paid during code review, debugging, and maintenance by colleagues, whose confidence in that engineer's skills may be affected by his use of AI.

    What looks like a productivity gain for one engineer can become a productivity loss for the team as a whole.

    • swiftcoder 41 minutes ago

      > you've mostly increased the team's work queue rather than its throughput

      Amdahl’s law remains unbeaten

    • joe_fishfish 40 minutes ago

      Good summary. Theory of constraints in action.

    • jorisw an hour ago

      Perfect summary of what's going on today

  • PostOnce an hour ago

    Pre AI and Post AI code review hours are both 0.75 in this made up example. I find that implausible.

    Even with the same amount of code, AI code is less trustworthy* and requires more attention... but we know it won't be the same amount, it will be more. This means it will take longer to review, or there will be unforeseen consequences of not spending that extra time.

    *meaning no human eyes have looked at it and said "this doesn't make sense", or "this is cheating", or "this doesn't meet requirements", and won't be caught until code review if at all.

    • yoz-y 32 minutes ago

      To me the biggest gotcha with AI code is that the bugs are not “normal”.

      When reviewing human code I focus on specific parts because I know that there are parts where a person will just not make a bug (unless very junior).

      AI on the other hand, will not do an off-by-one mistake, but it will happily just delete perfectly working code for no obvious reason. Or monkey patch a dependency because it missed a config flag. Or generally fail in a very novel and creative way.

      The effort it takes to review AI code is much greater. And this is in a code base I am deeply familiar with.

      Imo the future lies in a solid core programs with powerful plugin frameworks that expect all plugins to be code that was never read.

    • rightbyte an hour ago

      The hard part is that LLM code looks like there is some sort of flow. It is like a nice statistical smooth flow. It looks very convincing at a glance. No one would write code like that and not know what they are doing comments self assured and all.

      • ACCount37 34 minutes ago

        LLMs are incredibly good at replicating common, coarse statistical features - which is what backs "looks very convincing at a glance".

        If it's a general signal that's easy for you to recognize at a glance, it's a signal that's natural and easy for an LLM to replicate.

        They're much worse at making the underlying structure work. Not incapable at all, especially not the modern LLMs. Frontier models kick ass. But it's true that an LLM denies you a lot of the classic "tell at a glance" by its very nature.

  • laszlojamf an hour ago

    What I have noticed in my own work that a lot of the time that used to be for coding is now just waiting. I have three agents working on three different features in parallel, and I'll go back and forth with all of them, correcting things and steering etc, but then I find myself with three busy agents and nothing to myself except stare at the screen while they code away. There is a mental budget for me where I can't have more than those three running at the same time and still keep track so what I end up doing is just scrolling HN...

    • dgellow an hour ago

      I stopped using coding agents after more than one and a half year of active use, it really started to become way too boring, and I’m t a point where I just hate having to babysit them and for the 200th time make it understand what the actual goal is… and to be honest, going back to writing code by hand without assistance is really hard at first you continuously have that little voice telling you how simple that would be with an agent. Then after a little bit you’re back to being productive, but I still get that voice in my mind. I’m wondering if that’s how addiction feels (way lighter of course).

      That whole experience of going deep for a while into LLM coding, then trying to leave it behind made me pretty pessimistic about the future of our profession. We are creating a whole industry of people delegating their ability to work to a software stack currently controlled by basically 2 companies (that both have very sketchy financials). Doesn’t feel healthy

      • mawadev an hour ago

        I'd be really interested to see all the software that is written by agents. Whenever I touch agents or ai I can't get much use out of them. My understanding is the value when I think aloud with them/treat them as a better google search, but thats about it. Except one off web stuff, that is a pretty neat use case.

        But lets be real, anything moderately complex that is out of the domain of publicly available sample code is hit or miss compared to the time invested running the loop. I'd much rather invest the time in myself.

        What a lot of people don't talk about is the inherent security nightmare of trusting ai agents and the sheer data exfiltration happening behind the scenes.

        • dgellow 31 minutes ago

          For what it’s worth, I’ve made really good, state of the art software in my areas of interest using LLMs. So I do believe you can produce really good software using them. But it’s domains where I have decade of experience.

          But even with that result I don’t think it’s something we should bet the whole industry on, and something I personally don’t feel comfortable relying upon

      • himata4113 an hour ago

        Neither of these points feel true anymore.

        Models are very much predictable these days (except anthropic models). The real issue stems from letting them work on their own for far too long. Also we are not controlled by 2 companies anymore as kimi k3, deepseek flash (and soon pro) as the ultra-cheap variants, glm 5.2 especially is a direct replacement for opus 4.8.

        Models will only get better and cheaper I wouldn't feel too pessimistic and wouldn't feel too bad on relying on them to accelerate work and free up mental space from menial tasks.

        As a personal side-note I never let my agents do architectual design I only use them for implementing. I always found the actual coding part of programming extremely boring and coming up with designs, experimenting and testing the fun part.

        • pydry 42 minutes ago

          I always found that if you are good enough at whittling down boilerplate that coding becomes something akin to pure architecture.

          I find that mediocre programmers and LLMs are bad at both. They're helpful if you want to shit out some repetitive boilerplate or perform a complex search of some kind but otherwise you're better off without.

          • himata4113 13 minutes ago

            ehh, they're pretty good at automated performance research and bug fixes, especially when spanned across hundreds of them.

    • hahahaa 27 minutes ago

      You need loops so they run longer and use less of your context and brain power. Then (and this is where WFH is a super power) do stuff like walk, daydream, come up with killer ideas like a Madman episode laying on the office couch.

    • dwedge an hour ago

      This has been my observation too. Because I'm chatting it feels like I'm not working, so any output can be "productive" in that context but I'm hyper aware of all the negative time here. Correcting, pushing it back to the prompt, reminding it that it doesn't have full context so do what I told you not what you think, and then verifying it and correcting it (always) seems to take longer than just doing the work myself

      • hahahaa 24 minutes ago

        You should spend time as much as 50% on harness engineering for you and team to stop that and make the AI steer better.

    • pjmlp 38 minutes ago

      Which is why, when I can get away with it, I only use AI for improved code completion, or generating the initial boilerplate.

    • pop3zxcv an hour ago

      I find myself in the same situation, baby sitting AI agent, monitoring them. It's like l've become a coordinator.

  • Arshad-Talpur 7 minutes ago

    The role of developer is becoming more of a system thinker than a syntax writer, previously load balance was about what to do in a given timeline, now its more of what not to do , doing more is actually increasing technical debt.

  • raver1975 9 minutes ago

    It's like you just made up those numbers and then developed your thesis around that.

    • steelkilt 6 minutes ago

      That’s exactly what was done.

  • crnkofe 28 minutes ago

    It doesn't really matter how much more productive a developer is if all other roles at the company don't follow suit. Before a developer picks up something to work on a series of roles had to set their eyes on work to be done. Project/product leads, tech leads, business people stamping and deciding on priorities. Then there's all the work that happens after a developer finishes work which tends to be manual as well. Review, QA, education, ops changes, marketing material, education articles, webcasts, showcasing features to end users and lets not forget end users actually making good use of the amazing new features shipped and likely many more largely sequential processes depending on company size and product/project type.

    There's no real way to get to a 10x developer nowadays. Even if a company somehow achieved the magic productivity increase in all employees you still need a 10x consumer to gulp it all down.

  • vb-8448 an hour ago

    Based on personal observation, a lot of productivity has been thrown out of the window with unneeded refactoring, rewrites and "what-if" scenarios that the AI agent will spot.

    • dude250711 an hour ago

      "... unneeded refactoring, rewrites and "what-if" scenarios ..."

      Like so many senior developers I have encountered. That stuff is good for CV.

      • vb-8448 43 minutes ago

        It's on a completely different level now.

  • qsort an hour ago

    I've had similar conversations with a client recently while discussing estimates for a large project. Senior leadership has a mental model where AI makes everything X% faster, but that's very wrong. Some things get sped up by an insane amount and basically go to zero, some others not so much. Entirely new tasks emerge, such as directing agents to provide them the context they need, setting loops, etc.

    It's a very O-ring problem.

  • sandeepkd 37 minutes ago

    > There’s no doubt that AI has already improved the productivity of engineering teams

    Thought it might be an interesting read, however have up just after reading the first line.

    For the context, code had always been a copy-paste exercise, big part of it was understanding and differentiating between the different choices. Along with it people were growing as engineering practitioner's too. Human learning still needs to happen if they are expected to fix the code when LLM gives up.

    LLMs are quite useful tool in themselves, however the hype has unfortunately polarized the population.

  • adithyassekhar 33 minutes ago

    I don’t think human review is worth it for LLM generated code. We design abstractions and all around how humans think. LLMs writes code that is better understood by machines. If you are all in on LLMs, by all means, read the code figure out what it means. But trying to enforce a human flow to its logic is flawed and will be overwritten the next time.

    • alienbaby 8 minutes ago

      But LLM's will trash your established architecture and patterns if you let it run free over an established codebase for any length of time. In our experience anyways, even with careful guidance and rules to follow. It nearly always does something a bit odd.

      And it's test cases can sometimes leave a lot to be desired.

    • furyofantares 28 minutes ago

      LLMs definitely write worse code for LLM consumption than humans can. In my experience your claim can't be further from the truth. I can get much further with an LLM starting from a great codebase than I can starting from a vibed codebase.

      And I can do even better than that if I design the codebase specifically with LLM coding in mind, making choices that make it hard or impossible for the LLM to make certain categories of error it tends to make, and make it easier for the LLM to observe the results.

  • ghenna an hour ago

    Based on personal experience on a specific project, that 1.5 hours with AI let me accomplish work planned for a man-week in the pre-AI era. So it’s much more than 3x.

  • dwedge an hour ago

    > There’s no doubt that AI has already improved the productivity of engineering teams, and will only get better in the coming years.

    They lost me by begging the question in the very first sentence.

  • mawadev an hour ago

    "We can disagree about the specific numbers here, but if you think this is wildly off, you’ve probably never been a senior developer"

    I'd even say the productivity gap is even smaller, if not negative in some areas...

    • dwedge an hour ago

      I think juniors and fake seniors are a lot more productive because they were never really able to measure their productivity so spamming LoC and trusting AI output makes sense to them.

  • ectoloph an hour ago

    AI is a force multiplier.

    It multiplies both good and bad decisions. Both mine and it's 'own'.

    I can get some things done 10x faster and it might even catch mistakes or help me solve something difficult.

    But if I am being lazy or complacent then it bites me that much harder.

  • fearnot 39 minutes ago

    If a senior developper spends the same amount of time debugging, code reviewing, setting up CI/CD, documenting and doing admin work, he has not been using AI right.

  • hahahaa 30 minutes ago

    This assumes you are arranging deckchairs and not leaving the cruise ship for say, a boeing 747.

    One example, let's say there is a side bet that makes everyone 10x more productive with a success rate of 1%

    It takes 2 hrs to make the bet wit agent orchestration.

    10 people can get this done in their spare time freed up by AI in 5 weeks.

    Bet cashes in and you are much faster at everything.

    It won't feel faster. Because the brain probably scores emotionally in roadblocks cleared per hour.

    Back when you got a single punch card loaded in a day it felt like a fucking win.

    The other factor is you get paid the same and there is more disruption and competition and job insecurity.

    But objectively value gets shipped faster using AI.

    Just not much if you go the faster horses route with AI. You need the cars. (Or planes!)

  • tempfile 11 minutes ago

    > Reading and Debugging 1.5 1.0 > Code Reviews 0.75 0.75

    Since these numbers are made up, I may as well throw my personal anecdote in the ring. I find reading and reviewing far harder with coworkers who are using AI. Tickets contain about 5x as much meaningless junk as they used to, and testing notes - while far more thorough - are often now multiple pages in length. Reviews also contain much more code, people try to do more drive-by fixes because the models can generate those fixes so quickly, and people understand the code they're submitting far less clearly because the model is able to generate fixes they simply couldn't previously.

    I feel less productive than I was a year ago, and I don't see my team shipping more features than they were previously. But everyone reports that they're far more productive. I don't get it.

  • 6stringmerc 14 minutes ago

    > Sometimes I actually find AI makes non-coding work go slower…For now, let’s assume AI only helps.

    Haha fuck you dude no I’m not going to assume it only helps when your preceding paragraph gives a concrete example of HOW IT MAKES WORK LESS EFFICIENT.

    In turn, I’m not “assuming” this guy is delusional and “grasping at straws” I’m deducing it from his poorly constructed, self-defeating, fictional argument in favor of his assertion.