Europe's Bitter Lesson

(fabianlindfors.se)

3 points | by fabianlindfors 5 hours ago ago

12 comments

  • ben_w 5 hours ago

    It remains unclear to me if anyone funding new models will actually benefit economically from having funded these models.

    I think is possible that these models can *culturally*(/memetically) benefit the creator via e.g. being trained to propagandise for the values of the originator, but it is unclear to me if this hypothesis has even been properly tested.

    However, even absent benefit from creating new models, there are still good reasons to support the economic changes that would allow such models to be created in Europe:

    We need a more unified, EU-wide, approach to investment both private and public; we need to upgrade the power grids even just because they're old and a poor fit for existing supplies and demands let alone forecast new demands and new supplies; and we need local chip factories to protect against supply chain problems in the event of conflict involving existing suppliers.

    I expect the AI-training investment bubble to burst before the EU gets any of those things, those things are good regardless.

    • fabianlindfors 5 hours ago

      Yes, it probably won't be clear until we have the benefit of some hindsight.

      One thing I'm inclined to believe though is that the investment being made into compute and research today may not benefit the investors, but it will be net-positive on a longer time scale for the wider populace. A comparison to the internet and the dot-com bubble seems fitting.

      And from that perspective I'd also be worried about the bubble bursting before Europe even gets started, as what remains after a bubble would still largely serve the US and China.

      • eigenspace 4 hours ago

        I think one problem with that comparison is that due to how energy intensive these datacentres are, they are very fast depreciating assets, unlike e.g. the telecom infrastructure that was (over-)built during the dot-com bubble, or the railway tracks from the era of Railway Mania.

        The only reason a GPU-based compute-heavy datacentre built 5 years ago is still in operation today is the component shortage.

        Without the current component shortages, those datacentres would just be so much less energy efficient per flop that it would literally be cheaper to just scrap the thing and rebuild it.

        _________

        This is not to say that we shouldn't build datacentres, but just that I wouldn't be so sure that the datacentres being built today won't be stranded assets long before they pay themselves off.

        • fabianlindfors 4 hours ago

          I agree, the datacenters and the GPUs inside them might be the least valuable thing to come out of the boom. The research in chips, architectures, and training, and the supporting infrastructure in fabs and energy grids still seems highly valuable. Europe will still see very little of that, but at least the Chinese labs still publish some of their research.

  • eigenspace 5 hours ago

    Regarding

    > Mistral seems to have given up on reaching the frontier and focuses instead on niches like moderation, OCR, and robot navigation.

    That's not actually true. They at least claim they have a new attempt at a large model they plan to release near the end of this summer.

    Their last 'large' model was a flop that needed to be replaced by their Medium 3.5, but they have not yet given up on large models, even if they're also making smaller more specialized models for now (because that's somewhere they actually do rather well).

    • fabianlindfors 5 hours ago

      I considered mentioning that in the footnote but my own skepticism stopped me. I'm now thinking I need to see it to believe it in terms of Mistral training a frontier model.

      That's not to say I don't hope they will. It would be incredible if they could do that despite their limited funding and the wide swath of projects they are working on. I do think that lack of focus is going to be a big problem.

      • eigenspace 4 hours ago

        Fair enough, though I think it's rather unlikely that they'd be lying about their work on a big model, even if one might be reasonably skeptical if the model will be any good.

        I think it's rather important for them that they at least keep attempting to build frontier models, because if they give up on that, then governments will no longer see them as a strateigic asset and instead just see them as a regular business.

        I think that for better or worse, Mistral wants the buy-in from governments to scale up, and even if they have faltering outcomes, it's better to have that, and be able to say "look, this is what we tried, we failed due to lack of compute. We need more money for more compute."

        • fabianlindfors 4 hours ago

          Oh definitely, not saying they are lying about working on a big model, but as you say I'm skeptical it will be any good.

          I hope they will get that boost, but I'm afraid that they haven't built credibility so far to convince governments to buy in at the level needed. Their slip away from the frontier might even have made governments more cautious on investing in AI development, which is the opposite of what we need right now.

  • weezing 4 hours ago

    I don't really want my tax money to be wasted on training someone's else clankers.

    • fabianlindfors 4 hours ago

      My point is that it wouldn't be someone else's "clankers". The lab would be majority owned by the populace through their democratically elected government.

      • weezing 41 minutes ago

        Unless we all get sellable shares then it's not owned by us. I don't really want my "democratically elected" government of thieves to actually own any clanker related thing bought with my money.

        • fabianlindfors 32 minutes ago

          It’s worked out pretty well for the Norwegians!