Why DuckDB 2.0 is faster

(motherduck.com)

95 points | by tosh 4 hours ago ago

22 comments

  • scythmic_waves 2 hours ago

    I also love the visualizations but I'm getting heavy LLM vibes from the prose:

    > One setting drives this,...

    > The cost is now about the rows you actually touch, not rounds times table size.

    Etc.

    I get the brain scramblies [1] from trying to parse this writing style at work so I hate to see it elsewhere. Apologies if I'm wrong. But if I'm not then OP don't use an LLM to write for you. It's hazardous to your reader's health [2].

    [1]: https://www.youtube.com/watch?v=ipUJq-odt5Q

    [2]: https://discourse.haskell.org/t/how-to-keep-enjoying-program...

    • kristianp an hour ago

      I noticed some AI tells, but found overall the article not too bad. It did seem to waffle at times though.

      > Storing a lake as thousands of 1 MB Parquet files is a bad practice anyway, and 2.0 does not rescue it.

      The "does not rescue it". No human would write like that.

      > I'll explain what that means on a table you already know.

      No I don't already know that table.

      Also

      > and claims 40x on graph reachability

      Is really hard to parse.

      The section on recursive CTEs wasn't well written and didn't explain how the optimisation was done. This article explains how the recursive CTEs were improved https://duckdb.org/2026/08/25/how-duckdb-runs-recursive-ctes...

      • vlovich123 29 minutes ago

        > The "does not rescue it". No human would write like that.

        This is what I don’t understand. Supposedly LLMs are trained on human text. Why do they come up with such unrealistic prose? Is it intentional because the companies want the tells to be obvious?

    • majormajor an hour ago

      I've lost count recently of how many times Claude has given me something in this style that I can't understand, and then I ask it to rework parts of it, and then it tells me that the original things it claimed weren't actually quite right anyway.

      I'm starting to read it as a sign of low LLM effort not just low human effort. It seems most common when one few-sentence prompt leads it to generate 4+ paragraphs (and the longer the output, the worse the odds). Prompting to dig into each resulting paragraph one by one, to make them readable, makes it do higher-effort deep dives.

    • sagarm 2 hours ago

      It really is unreadable. I guess you're supposed to skim an AI summary. Too bad you'd never see the visualizations that way.

    • gbalduzzi an hour ago

      I may be in the minority but I didn't get LLM vibes from this. Not enough to be bothered by it, at least

    • Eridrus an hour ago

      I have tried out the alpha releases of 2.0 and I got slopcoded vibes from it.

    • hobofan an hour ago

      > It's hazardous to your reader's health [2].

      Just because some guy on some forum said that doesn't make it true. That's not how you establish facts regarding health claims.

      • theendisney 4 minutes ago

        The entire internet is like a wikipedia talk page now. If you say it is true often enough it becomes self evident.

        Not that i question that reading mostly ai slop for long enough makes you feel dead inside.

    • gchamonlive 2 hours ago

        I get the brain scramblies [1] from trying to parse this writing style at work so I hate to see it elsewhere
      
      This sounds to me a lot like mass hysteria, people reading other people's behaviour online and reproducing it unconsciously.

      More and more really important and useful information will arrive like this for us to consume. There's no way around. So this is a disservice for newcomers that could come and go unscathed but instead is crippled by these kinds of comments that brings nothing of substance to the table and has the potential to make them hate something they otherwise wouldn't even notice.

      Also, from https://news.ycombinator.com/newsguidelines.html:

        Please don't post shallow dismissals, especially of other people's work. A good critical comment teaches us something. 
      
        Please don't complain about tangential annoyances—e.g. article or website formats, name collisions, or back-button breakage. They're too common to be interesting.
      • majormajor an hour ago

        Look at the specific AI habits called out in this comment: https://news.ycombinator.com/item?id=50037220

        This sort of writing decreases readability. People say "just have your own LLM rewrite it" like you don't lose value when you go from prompt -> slop you didn't review enough to clean this crap up in -> someone else's prompt to change the style -> finally someone reads it.

        Do you know what happens a double-digit percentage of the time when I ask Claude to rewrite some shit that it gives me like this? It says things like: "I overstated this, I rechecked and actually..." or "this claim doesn't hold up, actually [this other thing is true]..."

        So it's a sign that the claims in the post likely weren't vetted very hard.

        So if you aren't proofreading I'm gonna be skeptical. And saying "deal with it" doesn't rescue it. Does it?

        And then there's the reflexive "you must just be an ideological hater." No. I'm someone who uses the tools in a domain where quality matters enough that I have to dig into the quality of the tool output and spot the tells for when it's low output, so that I can deliver shit that works reliably and consistently.

        • gchamonlive an hour ago

          This is all absurd discussion. If you disagree with the article style, there is a flag button there.

            So it's a sign that the claims in the post likely weren't vetted very hard.
          
          Our time is better used submitting something useful instead of debating meaningless guesswork in the comments.
      • johnfn an hour ago

        Curious how you see low-effort writing as a "shallow dismissal" or "tangential"

        • gchamonlive an hour ago

          Everything is low-effort writing if your judgement is ideologically charged. Just read the damn post, it's very good, there's nothing low effort, but you need to read it as one single piece, not dissect it looking for hints of LLMisms, killing the article in the process.

          Also if you think LLMisms so bad it's spam, you also don't get a free pass. From the guidelines:

            If a story is spam or off-topic, flag it. Don't feed egregious comments by replying; flag them instead. If you flag, please don't also comment that you did.
          • johnfn an hour ago

            Who said I was "ideologically charged"? I like AI and I use it every day. I simultaneously hate AI writing. It is impossible to parse. If you can read it fine, good on you, but you seem to fundamentally misunderstand that other people perceive the world differently than you do. No one is "dissecting" this article under a microscope, it's obvious.

            • gchamonlive 32 minutes ago

              Understanding isn't a prerequisite for following the community guidelines. Everyone here is guilty of this, me included to be feeding this pointless discussion, and it's a bit disappointing from the mods not to terminate this thread. Either that or kill the main post. Keeping both up is contradictory.

    • augment_me 2 hours ago

      My manager was previously whining about something similar when it comes to generated reports and I just pulled down all of his writing and made the LLM write in his style, he is really content now. Do the same using your own writing and you wont have to be offended

      • gonzalohm 28 minutes ago

        So now we need a decoder to be able to read articles. It's true that it's painful to read and it's good to be called out for it

  • stacktraceyo 3 hours ago

    Great visualization. Side note their new c++ extension api is also gonna be faster from the perspective of development / distribution of those extensions

  • jiggawatts an hour ago

    I wish more database engines used a Task-based design like Umbra / CedarDB.

    Most of the DB engines out there still seem to use a "n-threads" style parallelism with exchange operations and poor async I/O management.

    DuckDB is improving on this front, but in some sense is catching up to R&D (and implementation!) that is now decades old.

    A "rhetorical challenge" I like to give software developers working on systems like this is the following: If I gave you a computer with 1,024 cores and matching network and storage bandwidth -- but with significant latency -- could you keep a system like this 100% utilised with one query?

    The answer for almost all software is "no".

    For example, SQL Server tops out at 64 hardware threads for any one query: https://learn.microsoft.com/en-us/sql/database-engine/config...

    GPU codes are starting to get there, but CPU codes are way behind on this frontier of computer science.

    It's not just databases! Can you (de)compress a file in parallel? Verify its hash in parallel? Upload/download from storage with CPU and I/O task parallelism? Can you overlap all of these operation so nothing is ever waiting on anything else it doesn't have to?

    This matters! I ran some tests with bioinformatics codes and found that most got stuck in tar pits. Many could not scale to modern SSDs with millions of IOPS or modern networking with hundreds of gigabits of throughput, no matter how many CPU cores were thrown at them.

    PS: AMD's Zen 6 era EPYC 9006 processors will have 512 cores and 1,024 threads per two-socket system, so this is not hypothetical: https://www.amd.com/en/products/processors/server/epyc/9006-...

  • larodi 2 hours ago

    ...because Atlas and Fable took turns to go back and forth through it (source code and runtime) and track suspected bottlenecks.

    • gwerbin 2 hours ago

      Totally valid technique in my opinion. Even with last year's technology, LLMs were really good at synthesis of small details, which, combined with their encyclopedic knowledge of everything ever written down about computers and programming, and their infinite capacity to run adhoc experiments and build out test infrastructure, made them very good at debugging and performance optimization.