U.S. Department of Energy Launches the Genesis Open Models Initiative

(genesisopenmodels.anl.gov)

80 points | by moelf 3 hours ago ago

32 comments

  • firasd an hour ago

    Just realized that there are basically no American open models right now ever since the Llama series was abandoned. Basically Gemma and GPT-OSS I guess?

    Ah but Mira Murati's new Inkling is Apache 2.0

    But it makes sense that if you're a university researcher you are thinking about what's a model that will be open weight and developed over the long term and doesn't raise 'Chyna' concerns in Washington DC

    • ipsum2 an hour ago

      There's a bunch of American open models. Inkling, Nemotron, Trinity come to mind, but I'm sure there's others.

      • embedding-shape 31 minutes ago

        Laguna S 2.1 is really great too, in the "preview" release they've done so far at least. Still pending some reasoning-looping, but besides that, it's a really strong model to run within 96GB VRAM with the NVFP4 variants, and it's really good at coding (specifically).

        • behnamoh 15 minutes ago

          No it doesn't follow instructions and is substantially slower than ds4.

          • jauntywundrkind 8 minutes ago

            Like glm-5.x I think it has enormous self introspection that it often trips up on, but that this self reflection is actually a superpower, that enables incredibly good output. And from very small models.

            If you watch it think, which you can, unlike American closed models, you can steer it. You can provide a a massive rocket ship stratospheric boost to help it orient itself. You have no self correction, there is no multiplayer in American proprietary models.

            Sure it's great having super powerful mystic oracles that have the "right" answers. But I love respect & revere the open thinking. No it's not automous. But it is brilliant. And it considers. A lot. Deeply. It chases. That to me is the most human of models, even as it falls far astray.

            You should help it. You can. Unlike these vicious dark surfaces which yield and tell you nothing.

        • kadoban 17 minutes ago

          Yeah I think it got bad press because the chat templates (or something?) were messed up on first release, but I've been using a quant of it and it's a powerhouse, better than qwen 3.6 27b for local on a 3090, which is saying a lot.

      • firasd an hour ago

        Just looked into some Nemotron stats

        Looks like on <https://arena.ai> agent arena (grouped by lab) Nvidia is 15/15 (much worse than Thinky and Mistral) and on text arena it's 18/27

        On <https://openrouter.ai/models?order=most-popular> I definitely see usage though (probably mostly cause Nemotron 3 Ultra is free) the grouped order is DeepSeek, Tencent, Xiaomi, OpenAI, Z.ai, Nvidia

        • coder543 33 minutes ago

          I think glancing at a random snapshot from today misses all the context. Nemotron 3 is far more significant than you're giving it credit for.

          At this point, Nemotron 3 is really an 8 month old model series. That's when Nemotron 3 Nano was released, and the Nemotron 3 Super/Ultra models this year are obviously based on that recipe, mostly just bigger with a few tweaks here and there. Against today's models, no, not that interesting. Each of the Nemotron 3 models were briefly competitive when they launched, but never exceptional, and less competitive with each scale up. The fact that it took so long for Nemotron 3 Ultra to launch really hampered its competitiveness.

          The Nemotron 3 series is extremely open about training recipes and training data, far more open than most open weight models, and that is valuable.

          Before Nemotron 3, Nvidia had never released a single LLM that I would consider interesting at all, so Nemotron 3 was a big step up. The closest thing was Mistral NeMo, but a significant part of the credit there goes to the Mistral team, not Nvidia.

          Given how much Nemotron 3 improved, I'm curious to see if Nemotron 4 will take them to a leading edge level instead of just briefly competitive.

          (Nvidia released a Nemotron 3 and a Nemotron 4 like 3 years ago... this year's Nemotron 3 is entirely unrelated. Nvidia's naming schemes leave a little bit to be desired.)

      • written-beyond an hour ago

        Don't forget IBM

    • loeg an hour ago

      I would not be shocked if another open model eventually shakes out of Facebook (based on Zuckerberg's public remarks).

      • solomatov an hour ago

        Which remarks? Could you share a link?

    • wmf an hour ago

      Also Nemotron and Arcee.

    • mistrial9 39 minutes ago

      review of AllenAI Olmo research team and commitment to OSS -- AI2 complete transparency including training data, code, intermediate checkpoints, and detailed logs for reproducibility and scientific rigor.

    • connorbrinton 31 minutes ago

      Laguna S 2.1 is another fairly impressive-for-the-size American open model

  • andsoitis 31 minutes ago

    Does Europe have an equivalent program?

    • behnamoh 14 minutes ago

      No, Europe is monitoring the situation and condemning humanitarian crisis overseas that they contributed to a century ago.

  • an0malous 18 minutes ago

    Do all these models have any significant architectural differences or training data sources? What are the factors going into the diversity of their performance?

    • ux266478 12 minutes ago

      The article posted is basically entirely about that.

  • Smith42 an hour ago

    What would the selected participants get from this? Looks like there is no offer of funding?

  • datlife 16 minutes ago

    This is refreshing considering all the FUD (mostly from 1 frontier lab) happening around Open weight models.

  • andsoitis 31 minutes ago

    I wonder why it took so long.

    • dmix 18 minutes ago

      Mostly because it's generally a bad idea for government to try to compete with a brand new tech industry with hundreds of billions in private capital developing commercial models. If the American private industry does actually wash out vs Chinese open models there might be talent available for them to put money into, so maybe they are just preparing for that scenario in the meantime.

      • MangoCoffee 12 minutes ago

        The American attitude is generally to let private companies build up a new industry so it can create jobs and pay taxes. However, in the LLM race, the Chinese open weight playbook pretty much killed that. China has basically commoditized LLMs. Chinese models are good enough, so the race has come down to who can offer the cheapest tokens.

        • dmix 4 minutes ago

          In the 1980s databases were dominated by private companies like Oracle and IBM and over time we ended up with MySQL/Postgres.

          If open models end up being the main endgame I don't see why American companies won't just end up funding open models, just like databases and languages and other fundamental tech they build off of. There will always be cutting-edge commercial variations for the important stuff and open models for general business and public work, while Apple and Google will eventually have models built into phones, etc.

          Ultimately we have no idea what the future will look like. It's fair for the US government to experiment early on with open models, but at this point we're pretty far from a situation where the US needs to rush, just because China created some billion dollar funds to start competing with American industry. Chinese business will eventually need to pay those bills as well.

  • Thegn an hour ago

    “Gomi” is the Japanese word for garbage. Gotta wonder if someone has a sense of humor…

    • greggsy 23 minutes ago

      The Australian Liberal Party (basically our version of conservative republicans) proposed the National Energy Guarantee policy in 2017, which inevitably failed due to the media and public’s relative literacy and tendency to turn policy names into acronyms.

  • yewenjie 2 hours ago

    I couldn't find any details about size or training data for the model.

    • robotbikes an hour ago

      It looks like they're taking applications for training data (due August 14th), so I think it's safe to say this is just an announcement of intent and a call for involvement vs. something that is readily available. Seems almost quaint in comparison to the strategy of sucking up every piece of data you can find anywhere on the Internet and feeding it to your LLM but I suspect their intent is to be more careful in what they train their model on.

  • riffic 20 minutes ago

    stewards of the nuclear weapons biz. they'll do great here.

  • thegreatpeter an hour ago

    Pretty cool I’ll take it. Thanks!

  • shenenee 26 minutes ago

    Genesis is skynet

  • placedrock 30 minutes ago

    Modeling with my life as data.