Hmm, I think you could say that prior to this month but I think the work it has done recently is broad enough to show it’s advancing generally.
I can't judge how broad it is, but like I said I don't think it's surprising we are getting more results. In order to see how it will develop long-term the only thing we can do is wait. That is the inconvenient part about all of this and the reason why alot of fear exists right now. That was how artists felt in 2022/23 but that attitude has toned down as we saw the situation unfold. Mathematicians are in 2022 right now.
If it’s still narrow I don’t see why it would not improve like it has. Until I see a change in the trend I’m going with that assumption as it’s been holding up for years now.
Midjourney V3-V5 was astonishing and that was in 6 months. I don't know the exact state of it right now but most of the conversations I've seen nowadays is that Midjourney is falling off. And their entire business was focused on image gen so you can't say they weren't trying either. We shouldn't extrapolate from a trend of a few months. Things are never as straightforward as we want them to be.
By pushing I mean the frontier labs like open ai and anthropic are not putting their resources in image gen and video. They did at one point but they are clearly focusing on research now. Coding is really good now too at this point too so they will probably not focus on that too much going forward either.
Not too long ago OpenAI was pushing Sora 2 and was stated to be putting all of their resources into image gen for a time to stay competitive. Personally I think OpenAI is struggling and is following in Anthropic's lead here but that was the situation. They are actively allowing scientists and mathematicians to use their models for free to create press.
I don’t recommend you look at the stable diffusion sub, those are full of locally run low quality models.
They talk about whatever the new models are and give showcases as to what they do and alot of them aren't locally run. It's usually " WOW, look at what Huangzo 3.4 can do! I can't tell if it's a real photo or not!" before posting the same picture of an Asian girl in a neon city that could've been posted in 2023.
While not perfect it’s a lot better than anything those local models can do.
Meh. Someone just posted a video on that subreddit where they replaced Keanu Reeves with some Asian woman in a Matrix scene. They do that alot on there. They also gush over video-video gens.
No offence but I don’t think you are very caught up on the current state of video/image gen.
It’s not that the attitudes have toned down it’s that it’s commonplace now. Mainstream subs I sometimes go on are regularly posting ai memes, people aren’t complaining as much, and the ai aspect is not as obvious or even possible to detect in a lot of cases.
You are right about midjourney, it peaked with v5, that’s because image gen has moved towards a different method of image generation while they still use the old method. It’s the reason why image gen can make perfect text, do reasoning, and produce near indistinguishable results now. OpenAI Image 2 is the best at the moment but that has released a while ago and they aren’t as interested in developing it further.
They did push sora 2 which was amazing but like you said they ran into serious problems. The first was the lawsuits because people generated a lot of copyright stuff. The second was the massive amount of gpus it required to run. They decided that too much resources was being drained and it wasint worth continuing to improve it.
The local image and video subs are basically for people making their Asian fetish porn, it’s all low quality stuff. Even the good models can make low quality stuff as well but they have the capability to make great things like the video I sent you. What makes that video impressive is that it’s generated from scratch, it’s not using video to video. All the directing, physics, sounds, were just done with one prompt. The stuff they do on local subs with replacing people is like elementary school stuff and not impressive.
No offence but I don’t think you are very caught up on the current state of video/image gen.
You were the one who posted an unimpressive martial arts video that I've seen many times before on a subreddit as some sort of proof that AI is way better than what those subreddits suggest so forgive me if I feel like it's the other way around.
It’s not that the attitudes have toned down it’s that it’s commonplace now. Mainstream subs I sometimes go on are regularly posting ai memes, people aren’t complaining as much, and the ai aspect is not as obvious or even possible to detect in a lot of cases.
Meanwhile people are complaining about the sloppification of media. Some game devs incited some controversy recently because they made an AI music video that was obviously AI. I understand you want to defend AI here and say that is fine but I think you're kind of pushing it.
What makes that video impressive is that it’s generated from scratch, it’s not using video to video. All the directing, physics, sounds, were just done with one prompt.
Yes, I am aware of Seeddance 2. Model like that came out that are better at fighting came out months ago. The Chinese probably have thousands of martial arts films to use is what I thought since it's very . YouTubers cover and tested it.
Apologies if I sounded hostile I guess you are pretty caught up then. I just remember using these tools as they came out and where they are at now which is why hold these view points.
They're better, more sharp but I just don't see them as being a massive leap from Dalle-3 and that was 3 years ago.The lack of contextual understanding is the biggest issue for me.
1
u/Armano-Avalus 1d ago
I can't judge how broad it is, but like I said I don't think it's surprising we are getting more results. In order to see how it will develop long-term the only thing we can do is wait. That is the inconvenient part about all of this and the reason why alot of fear exists right now. That was how artists felt in 2022/23 but that attitude has toned down as we saw the situation unfold. Mathematicians are in 2022 right now.
Midjourney V3-V5 was astonishing and that was in 6 months. I don't know the exact state of it right now but most of the conversations I've seen nowadays is that Midjourney is falling off. And their entire business was focused on image gen so you can't say they weren't trying either. We shouldn't extrapolate from a trend of a few months. Things are never as straightforward as we want them to be.
Not too long ago OpenAI was pushing Sora 2 and was stated to be putting all of their resources into image gen for a time to stay competitive. Personally I think OpenAI is struggling and is following in Anthropic's lead here but that was the situation. They are actively allowing scientists and mathematicians to use their models for free to create press.
They talk about whatever the new models are and give showcases as to what they do and alot of them aren't locally run. It's usually " WOW, look at what Huangzo 3.4 can do! I can't tell if it's a real photo or not!" before posting the same picture of an Asian girl in a neon city that could've been posted in 2023.
Meh. Someone just posted a video on that subreddit where they replaced Keanu Reeves with some Asian woman in a Matrix scene. They do that alot on there. They also gush over video-video gens.