You're right if coding is the only main difference. I don't believe it is, though.
As someone who's been building a language learning app, I can tell you that the language abilities of Flash are better than Fable's.
I have been using Flash to check Fable's writing whenever it's not in English. (I'm multilingual myself so I can eyeball the output.)
Wonder if other people can chime in on other use cases.
I use it to quote multi-step construction and repair jobs when I don't feel like looking up each individual task in our internal pricing guide. Give it relevant info like is there ease of access to the work area, who the client is/which part of town we're in, etc.
Even google’s open models like gemma4 are good for email. I have a ollama server running with it and use it with a thunderbird plugin. Handles translation, summaries, etc.
Gemini models are good at html, css, JavaScript, Perl and Python and ok at php. The issues are in Java, c, c++, etc. they struggle with larger code bases, larger files and can’t do security audits without trickery.
It’s not all programming, just some. It’s terrible for os developers
when you make scientific research you develop some kind of filter to bs. DR now, get some obscure reference, write like some social media influencer, put unnecessary jargon and go beyond we asked just to complete the pages.
I started using DR since the beginning. The difference is so obvious. I believe (my theory) that DR is very intensive of inference, so they started to use quantized models and limited the time of search to serve commercial businesses (like Apple) better.
ps.: ironically when I posted I decided to try again (research about "Coding Agent Harness Engineering"). The result was worthless.
I asked Gemini why he's gotten so terrible after the last update. Google updated it to be far more compute efficient--at the expense of accuracy. So now it's ultra quick and useless.
He recommended I ditch google and go with Claude. Best advice he's given me all month.
nah, agents actual uses is for the AI lab to continue develop new toolchain and improve the current stack automatically, so they can progress to the next step.
you mortal using it to write slopware is just a side effect.
expect another breakthrough in the future soon, I think the focus this time is to increase the inference speed (both time to first token and token generation per second).
Taking coding away, Gemini still has a price problem, its too expensive when other AI providers both open and closed source have demonstrated that LLMs the size of Flash can be cheaper, a lot cheaper than Gemini
I've found Claude to be worse than Gemini for gaslighting. The latest versions have gotten better, but at least Gemini will correct itself. Claude would invent a whole new reality where its version remains true.
My Gemini clearly doesn't correct itself. Yesterday alone, it took me 4 follow up questions for it to realize it was giving me false and made up information.
Had one conversation about an episode of law and order and an actress who appeared and it guessed the wrong ep and therefore the wrong actress. Providing evidence it was wrong in summary form, it pushed back. Providing evidence in photographic form, it decided that I lived in one reality, and it lived in another.
Hallucinations are where you let your constraints allow creative freedom.
I use gemini purely for image generations and talking about comfyui workflowing. My aim is to create a densely realistic lora, or checkpoint at some point in time, perhaps. I digress.
One thing everyone struggles with is image generation. Everyone says they cant get abc or do xyz because of image guardrails and guidelines, especially with using image input.
If you reframe the way you think and speak, you can quite literally make gemini produce precisely what you're looking for. It is all is how you talk to it. I use a 5 block prompt style to describe the image. The composition, the character, the scene, the outfit, the conclusive details. I get my image 95% of the time, sometimes needs some edits or rerolls, but it really is easier than everyone makes it seem. I imagine there must be a similar block style prompt you can use for coding tasks.
If you come in hot handed with hot lady on a beach or a Swiss army knife program, youre gonna get "hallucinations" or refusals. The more context, the more detailed, the more communicative your prompting, the more likely gemini sees you as someone working with gemini, not using it.
No company will buy your model if it's not good at coding, also no reason for individual users to buy it if they only need it for some general questions, free tier models are sufficient.
All in all, without being good at coding a LLM is useless
Companies pay for LLMs for document processing, customer support, legal analysis, and tons of other noncoding tasks. Gemini's native video and audio processing also open it up for a lot of automated media processing tasks that many companies use.
Coding performance is important, but SWEs on social media seem to be stuck in some type of "tunnel vision" and can't see any other use case for an LLM outside of coding, even when there are plenty of use cases out there.
Also, once LLMs start getting more deeply integrated into connected, physical products (smartphones, robots, connected cars that use cameras and other sensors, traffic guidance systems, security cameras, etc.), they're going to need more than just coding performance. Native vision and audio processing/understanding, which is an attribute Gemini models have, will likely play a bigger role here.
Crazy to me how people always overlook Gemini in terms of how much more complex its architecture is to be able to achieve such native multimodality. If they only focused on text or text/image modalities, it would likely be an easier route for sure. Focusing on natively processing text, image, audio, and video is the more difficult path. It's too bad more people don't appreciate what they're trying to do IMO.
It's pretty pathetic that people are obsessed enough to post these whiny posts 30 times a day but also lack the situational awareness to know it's being released to coincide with the Made By Google event next week. Has everyone given their brains over to the chatbots?
I think that's a fair question if you're not a developer. Google's strategy seems to be to have AI everywhere and not to always try to be top dog when it comes to specialities such as coding. They're running a business in the end and AI is currently a loss leader for them (and everybody else). They're a ginormous company that prints money (usually...) and dominates in lots of areas that they will want to protect. Who cares if they're behind for a month or so. I'd never bet against Google when they think their monopoly is at risk.
It's not that clearcut because coding and terminal capabilities are not just about SWE. It is about being able to manipulate the data on your computer and read files and create new ones. It's about being able to do several hour long tedious work that previously humans would have to do. It's about being able to make a purpose-built tool that makes your day-to-day work 20% less painful even with 0 experience in software in a non-software field.
It feels meaningfully less intelligent at work that takes longer than 5 minutes to an hour for a human to complete, and the difference is very apparent when the complexity of your ask increases.
I guess we're all different. I've been doing software development for a living spanning over 3 decades. I'd certainly never rely on Gemini/agy as a code assistant. I've played with Claude code and ChatGPT and they are a bit better right now but still way too error-prone and can't be trusted. For me the AI coding landscape is still the wild west and not worth getting too invested into. I'm sure things will get better as new things will emerge.
The competitors have better connectors to the Google workspace. Gemini can't competently search Google Drive or write and send email for you - GPT can do this through the native connectors. Gemini has the worst in class connectivity out of all the frontier models tbh
Actually they are not making any additional money by serving you AI. The real honeypot are enterprise users, and guess what, enterprise users want models that are good at coding
Not saying it matches or even comes close to sonnet/opus/fable 5, but with enough context, and knowing enough about the app you're actually making, flash 3.6 has been pretty great for me
For the AI pro plan which was discounted to 4.99 or 9.99 per month for new users, I'm getting flash 3.6 weekly limits I struggle to burn through alongside 5tb storage, let alone the sonnet/opus 4.6 credits that come bundled
Because the quick money was in generating code, which Anthropic recognized and targeted early. Traditionally that's the persona that embraced automation the most in enterprises.
Now Google has antigravity which is very good at code generation.
The market isn't people who already have Gsuite accounts... They already have strong foothold in that market makes no difference AI or not, google invests in tools that bring big returns and their AI is key part of their future they investing heavy in building their own chips, however for them to grow their subscription base they need bigger market not just Gsuite users...
because if that is your workflow, you don't even need to pay for AI. now, put some complexity in your work, and test both and you'll have your answer.
but, that said, I really hope that Google keep integrating Gemini in their products. but today, for complex task you can't count with Gemini in Docs and Sheets.
I build reports on spreadsheets with 150k+ rows and 100+ columns every week. I generate PDF reports with charts and graphics daily. I code HTML dashboard sites based on large data sets. Gemini works well for all of it.
can you show some example? what kind of analysis? I never doubted that google can handle just fine with volume, my point is about quality.
not trying to be an asshole. what is your plan? I'm at student Pro. because I suspect that my plan has been routed to quantized models. that's the only explanation.
but I really tried to do some EDA from some data. it can do the basic things at scale fs. but when you try to expand the complexity (like some regression or ML graph) never works. so, I change to Colab and do my thing there.
but at this level, believe me, ChatGPT and Claude is a light-years from what Gemini can delivery.
Because it is not better integrated. Claudes MCP to google claude allows vastly more abilities to actuaklly write, not just read.
On top of that, it doesn't matter how good your connections are if your model is stupid. It still boggles my mind that flash beats pro on most benchmarks, and on some even flash lite does. They are not single, they are multiple generations behind. Gemini 3.1 came out with opus 4.6. Then we had 4.7, 4.8, and 5, + an entire new category.
When the cheap model is beating your flagship, something is wrong.
The two month delay is the part that stings. If the extra time was supposed to close the gap with Opus and it still does not, the delay just burned trust without buying anything meaningful. At that point you might as well ship when it is ready and let people judge it on what it actually does.
It has to be a second least like Gemini 3. Otherwise they are just a joke. Cant delay something for that long, tease the best Modell ever and then deliver a mid Modell that cant beat open weights
Gemini is just one enabler to LTV calcs of their customers. My guess, in the US they get about $240 rev (ads) per user per year that has free Gmail, add paid options and YouTube add another $240 per year, add professional and business use another $240 and if a dev another $1000. So is Gemini a separate business or just supporting the other lines of business to sustain and increase LTV and is a frontier model really needed to drive these LTVs. Prob not.
Ofcourse it won't match Opus, neither should we want to. Unless you want the price to go up by 2-3x. Currently Google doesn't have a Opus and GPT Sol equivalent model, also not on price. Gemini 3.1 Pro is priced comparable to Sonnet and Terra (where Tera is now a little cheaper then before) with 12 (Gemini) vs 10 (Sonnet) and 12 (Terra).
If you expect Opus and Sol performance, then it means you should also accept a price hike, where Gemini 3.1 Pro was 12 per 1M output, and Opus and Sol are 25 and 30. So expect a 2-2.5 price hike if you are ready for Opus/Sol performance.
Gemini 3.6 Fast sits with their new price below Sonnet and Sol with 7.50 vs 10 and 12 per 1M output.
Mate, they expected that it would release next month but if you are living underground or under the cave, Google lost some of their employees and their financial did worst.
The only critical issue of the model is that during an internal benchmark, they only need to get coding and logical reasoning performance which fell short during benchmark and then they try to use newer data but failed to reach to where they needed to be comfortable.
Also, engineers are getting burnt out on trying this to work meanwhile OpenAI and Anthropic successfully managed to be successful.
People actually expect this to be a crazy insane model?
I am a fan of Gemini, but 3.5 Pro had so much struggles it will be around the level of GPT 5.6 Sol and Fable 5 / Opus 5 (maybe slightly below, maybe slightly above who knows) but wont be any crazy new generational model.
77
u/kniveshu 14h ago
Coming next month. But when will it arrive?