Open-weight promise seems nice, but 1) I've seen people commenting that these hopes amounted to nothing for some previous releases (no idea which promises were made though); and 2) if the backbone is released as Dev, what will be missing? I can't easily tell from the post.
dev variants are usually cfg distilled which means that directly finetuning isn’t as effective. In the past, for the flux2 klein models,they released base versions that are not distilled. So it will probably be a while before you can fully take advantage of the open weights.
The term "world model" as it was once used in model-based RL can now apparently refer to anything as silly as linear regression. Then again, the RL folks probably borrowed the term from behavioral scientists before them. It's probably best to simply accept this :/
A similar thing happened to "object oriented" which has been misused by philosophers and visual artists alike.
In the 1990s "Object-Oriented Ontology" was introduced by Graham Harman [1]. The name was borrowed from computing, but its meaning has little or nothing to do with Simula or Smalltalk.
It would be interesting to see some time-series sentiment analysis of HN. Subjectively it feels like there's much negativity on HN these days, I hope I'm wrong.
> time-series sentiment analysis [...] much negativity on HN these days, I hope I'm wrong
Depending on the members, there certainly is. Put aside the "dismissers", those who have a habit or a hormonal reliance to cast a "meh". Those who objectively assess according to the input that the development of facts provide may bend their "apparent mood" accordingly. This may be more evident here because in brighter times we may be more inclined to post and submit about more idle intellectual beauty ("complications in ancient clocks"), and in darker times it makes sense that we are more focused on the problems.
It's possible it's because some people care about the consequences of what we're doing as a society beyond the mindless satiation of curiosity. Nothing wrong with curiosity in my books by the way, but I do think isolating it the way we've done in the technical fields is a dangerous and irresponsible attitude.
As I understand it HN is a community for hackers to discuss interesting and curious topics.
Pessimistic opinions on the labor market, views about society and politics that border on the dystopian, as well as exaggerated concerns about datacenter environmental impacts, all seem like a poor fit for this website.
I am very excited for this. 2.3 was excellent and I've seen clips from people who got early access along with reading their reflections on it and I think this will be the new SOTA for home use.
A very beautiful place on earth: one friend who lives there enjoys road biking in the surrounding mountains. Another friend does cross-country skiing in his lunch break (in winter times).
You are completely ignorant. This is one of the best places to live and work on earth. Nonetheless, they also have positions in San Francisco for what is worth.
Flux 2 Dev Klein has practically been the best you could use on most commercial hardware so I really hope Flux 3 has a comparable updated open-weights model to it. if not it'd be a great loss to most hobbyists.
> It jointly learns from images, videos, and audio within a unified architecture, because what it needs to learn is not any one of these elements in isolation.
I'm confused, videos contain images and audio ...?
That's most likely a disagreement on terms. In the media world, video is only the moving images, not audio. This is separate from images, that are meant to be still images.
It's more a comment about the feature detection I think; all image, video and audio input contribute to the same weights/activations that can produce image, video and audio output.
I have a feeling open-weight models ought to be outperforming proprietary ones by now, but that still hasn’t happened. So far, Nano Banana and GPT-2 Image seem to be the best in class, and Flux still isn’t crossing that quality bar.
i don't get why they are investing money on image/video gen. All generations i have see looked blurry, lacking in fine details and missing the artistic touch (lacks meaning? lifeless?)
Open-weight plans are near the bottom (Launch section):
- Video and audio generation and editing through APIs and private weight access. (“FLUX 3 Video”)
- Action prediction through selected research and commercial partners, beginning with mimic robotics (“FLUX-mimic and FLUX 3 Action”)
- Image synthesis and editing through APIs and private weight access. (“FLUX 3 Image”)
- Open-weight access to a multimodal backbone, for content creation (video, audio and image) and action prediction. (“FLUX 3 Dev”)
imagine spending nine figures training a model to learn that the sound has to match the impact. my 8-month-old figured that out by dropping a spoon on the floor twice.
Well, unified multimodal intelligence is the only way we will get to The Terminator, which seems to be the goal now of Silicon Valley and every Nation State with a military budget, so have at.
Well, probably not, but I am not engaging in socialization. I am using it to help me decide among competing statistical modeling approaches, or other aspects of my research, including coding, neuroimaging pipelines, and an assortment of other thorny issues that come with longitudinal/developmental neuroscience.
Sorry because pointing this is a bit tired by now, but reading already the first two paragraph thete is this unmistakable stench of LLM slop writing. Immediately disengaged.
The fact that people keep developing this technology shows that the true problem is not that machines are likely to become intelligent, but that people have already become machines - unthinking and without any care to the future whatsoever.
"Horseless slop trained on centuries of coachmakers' and blacksmiths' work. It'll never replace a real horse."
- Hacker News commenter in 1889 criticizing the automobile
I hope the open-weight versions will be SOTA.
> Over the next few weeks and months, we will make the following capabilities available
> Open-weight access to a multimodal backbone, for content creation (video, audio and image) and action prediction. (“FLUX 3 Dev”)
> We will also release more technical details on the underlying approach.
Open-weight promise seems nice, but 1) I've seen people commenting that these hopes amounted to nothing for some previous releases (no idea which promises were made though); and 2) if the backbone is released as Dev, what will be missing? I can't easily tell from the post.
dev variants are usually cfg distilled which means that directly finetuning isn’t as effective. In the past, for the flux2 klein models,they released base versions that are not distilled. So it will probably be a while before you can fully take advantage of the open weights.
I run some of the 2.3 models locally so I'm not sure where you saw that they didn't follow their words?
- Showed close to zero examples of people.
- Frivolous use of the term World Model.
- Claims 20 seconds of video, shows only jumpcuts.
Coming soon!
> Frivolous use of the term World Model
The term "world model" as it was once used in model-based RL can now apparently refer to anything as silly as linear regression. Then again, the RL folks probably borrowed the term from behavioral scientists before them. It's probably best to simply accept this :/
A similar thing happened to "object oriented" which has been misused by philosophers and visual artists alike.
> probably borrowed the term from behavioral scientists
It's from epistemology. It is not there to refer to the subjective but to the objective.
> A similar thing happened to "object oriented" which has been misused by philosophers and visual artists alike
For instance?
In the 1990s "Object-Oriented Ontology" was introduced by Graham Harman [1]. The name was borrowed from computing, but its meaning has little or nothing to do with Simula or Smalltalk.
[1] https://en.wikipedia.org/wiki/Graham_Harman
This is very interesting especially in the terms of what I think some call a "rabbit hole",
but we could be curious on how and why you saw misuse.
Many video examples on /r/stablediffusion
And they are stunning
Yeah, very weird launch. I was like... where's the videos?
It's incredible how negative and dismissive the comments in here are while here I am thinking the model actually looks impressively capable.
But then again I heard the downers have always been the first to leave their dung comments here so let's see...
It would be interesting to see some time-series sentiment analysis of HN. Subjectively it feels like there's much negativity on HN these days, I hope I'm wrong.
Bunch of software engineers are worried about their future employment prospects.
> time-series sentiment analysis [...] much negativity on HN these days, I hope I'm wrong
Depending on the members, there certainly is. Put aside the "dismissers", those who have a habit or a hormonal reliance to cast a "meh". Those who objectively assess according to the input that the development of facts provide may bend their "apparent mood" accordingly. This may be more evident here because in brighter times we may be more inclined to post and submit about more idle intellectual beauty ("complications in ancient clocks"), and in darker times it makes sense that we are more focused on the problems.
Advertisers love these sorts of reactions.
I see nothing here but them TELLING us how great it is. Not showing us.
> Not showing us
Have you watched the 46s video full screen on a monitor and not marvelled at the incredible 4K detail of the FPV motorcycle racing clip?
none of this is exciting compared to holodeck in star trek
It's possible it's because some people care about the consequences of what we're doing as a society beyond the mindless satiation of curiosity. Nothing wrong with curiosity in my books by the way, but I do think isolating it the way we've done in the technical fields is a dangerous and irresponsible attitude.
As I understand it HN is a community for hackers to discuss interesting and curious topics.
Pessimistic opinions on the labor market, views about society and politics that border on the dystopian, as well as exaggerated concerns about datacenter environmental impacts, all seem like a poor fit for this website.
I am very excited for this. 2.3 was excellent and I've seen clips from people who got early access along with reading their reflections on it and I think this will be the new SOTA for home use.
These people are hiring in... Freiburg im Breisgau? Wonder how hiring is working out for them there.
A very beautiful place on earth: one friend who lives there enjoys road biking in the surrounding mountains. Another friend does cross-country skiing in his lunch break (in winter times).
You are completely ignorant. This is one of the best places to live and work on earth. Nonetheless, they also have positions in San Francisco for what is worth.
there's a reason people call it Flyburg im Nicegau.
first AI thing coming from Europe that gives high hopes
I thought the clips were real footage until they were dancing in a flooded room.
> a model must learn a representation of the world: […] and how events sound
I honestly hope they put an unrealistic amount of wilhelm scream into the learning process, just for fun.
Flux 2 Dev Klein has practically been the best you could use on most commercial hardware so I really hope Flux 3 has a comparable updated open-weights model to it. if not it'd be a great loss to most hobbyists.
Show a person laying in the grass first.
I'm not sure many people expect a positive release of BFL anymore. They botched their releases so hard IMO. They'll find a way to fuck it up.
I wonder if this also creates people with huge heads and short necks like Flux Klein does.
> It jointly learns from images, videos, and audio within a unified architecture, because what it needs to learn is not any one of these elements in isolation.
I'm confused, videos contain images and audio ...?
That's most likely a disagreement on terms. In the media world, video is only the moving images, not audio. This is separate from images, that are meant to be still images.
It's more a comment about the feature detection I think; all image, video and audio input contribute to the same weights/activations that can produce image, video and audio output.
I have a feeling open-weight models ought to be outperforming proprietary ones by now, but that still hasn’t happened. So far, Nano Banana and GPT-2 Image seem to be the best in class, and Flux still isn’t crossing that quality bar.
i don't get why they are investing money on image/video gen. All generations i have see looked blurry, lacking in fine details and missing the artistic touch (lacks meaning? lifeless?)
> why they are investing money
Maybe they presume that after a series of "good enough to some" they may be getting near the Real Thing?
Lots of words about multi-modal but then this:
> our mission to develop real-world visual intelligence
Visual is mono-modal, isn't it?
its doing video, audio, images and motion. I think that counts as multimodal.
Is this really the value-add comment you’re going with?
Says the one who posts this comment?
Open-weight plans are near the bottom (Launch section):
imagine spending nine figures training a model to learn that the sound has to match the impact. my 8-month-old figured that out by dropping a spoon on the floor twice.
"One-shot learning" is still part of the discipline, actively studied (definitely in the past and surely in the present).
Well, unified multimodal intelligence is the only way we will get to The Terminator, which seems to be the goal now of Silicon Valley and every Nation State with a military budget, so have at.
I wish you were being hyperbolic but I know better.
I don't even know anymore, to be honest. I talk to AIs more than I do with humans, these days...So who knows
That doesn’t sound healthy
Well, probably not, but I am not engaging in socialization. I am using it to help me decide among competing statistical modeling approaches, or other aspects of my research, including coding, neuroimaging pipelines, and an assortment of other thorny issues that come with longitudinal/developmental neuroscience.
Sorry because pointing this is a bit tired by now, but reading already the first two paragraph thete is this unmistakable stench of LLM slop writing. Immediately disengaged.
Thanks for the heads-up
Correct reaction
The fact that people keep developing this technology shows that the true problem is not that machines are likely to become intelligent, but that people have already become machines - unthinking and without any care to the future whatsoever.
Show it, don't just say it. The argument would be?
AI slop trained on copyrighted content.
"Horseless slop trained on centuries of coachmakers' and blacksmiths' work. It'll never replace a real horse." - Hacker News commenter in 1889 criticizing the automobile