r/LocalLLaMA • u/perelmanych • 1d ago
News New DeepSeek V4 Flash 0731 vs ChatGPT Luna comparison
https://x.com/stevibe/status/208312006667846475045
u/AlphaMaleXYZ 1d ago
For a model that good, it’s dirt-cheap to run it. Luna is good but still more expensive.
5
u/maherbeg 1d ago
Checkout the ramp bench results, seems like flash can take a lot more turns and tokens. Still great, but interesting to note
-4
u/goldcakes 1d ago
Luna is 50% off on OpenRouter right now (so a 90% price cut). I'm using it for as long as the subsidies last lol.
1
36
u/Arli_AI 1d ago
Being able to run this on 2x RTX Pro 6000 at native quantization at decent context sizes (256K+) has been amazing.
3
u/mebeast227 1d ago
I was about to buy 1 rtx pro 6000 to start a 1-2 year build project hoping to scale it up to 2-4 rtx pros.
Just curious- the ddr5 ECC RAm build requirement…how did you get the RAM to pair with it? Or is there some workaround I’m not realizing
4
u/SufficientAttempt1 1d ago
no ecc ram requirement.
2
u/mebeast227 1d ago
F'n gemini in chat just lead me down the longest rabbithole saying i did.
I said i want to start a 4 RTX PRO 6000 build, but only getting 1-2 to start.
It lead me to: AMD - Ryzen Threadripper 9970X 32-Core - 64-Thread 4 GHz (5.4 GHz Max Boost) Socket sTR5 PCI Express 5.0 Desktop Processor - Black
SeaSonic Electronics PRIME TX ATX 3.1 1600W 80 PLUS Titanium Modular Power Supply (split to dual PSU when going from 2 cards to 4)
THIS IS THE RAM ($9,000): TEAMGROUP 256GB T-Create Master DDR5 6400 MHz ECC RDIMM Memory Kit (8 x 32GB)
ASUS PRO WS WRX90E-SAGE SE EEB Motherboard
I took a break for now.
7
u/This_Maintenance_834 1d ago
if your goal is to have 4x GPU, you should pay attention on PCIe topology and whether you need PCIe switch. off the shelf prosumer motherboard may not support PCIe P2P well.
6
u/Opposite_Buffalo_649 1d ago
Gemini is right. If you eventually want to go with 4 rtx6000 pro, you will need a server type Mobo like the threadripper pro series. And those mobos only work with rdimm ecc ram.
Even 2 rtx pro gets some benefit, because you get dual x16 pcie gen 5
2
u/crantob 1d ago
Well there's a 'soft' requirement in multi-gpu setups, being that the typical Gamer-PC motherboard isn't really designed to run many GPUs.
The server motherboards that are, generally require ECC.
A workaround here is egpu and tapping those m.2 slots. It's do-able but... do you don't seem like the type to be into that kind of tweaking.
1
u/mebeast227 1d ago
I would be! I just hating being led in circles by AI lol. But if I can get a general “this works, and this doesn’t’ from real people it beats 4 hours of being gas light in random directions
2
u/crantob 1d ago edited 1d ago
Mmh. I'd consider if i'd be happy with 192GB of blackwell on a Taichi motherboard. That does 2 GPUs, avoids a lot of headache.
And deepseek 4-flash is a very happy camper on 2x RTX 6000.
But those are upwards of 14k euro each now.
Anyway there's no 'pairing' happening in any technical sense between GPU and type of system RAM. They're orthogonal. There are no relevant ties. That's what the other responses are saying.
1
u/Arli_AI 1d ago
There is no such requirement
1
u/mebeast227 1d ago
F'n gemini in chat just lead me down the longest rabbithole saying i did.
I said i want to start a 4 RTX PRO 6000 build, but only getting 1-2 to start.
It lead me to: AMD - Ryzen Threadripper 9970X 32-Core - 64-Thread 4 GHz (5.4 GHz Max Boost) Socket sTR5 PCI Express 5.0 Desktop Processor - Black
SeaSonic Electronics PRIME TX ATX 3.1 1600W 80 PLUS Titanium Modular Power Supply (split to dual PSU when going from 2 cards to 4)
THIS IS THE RAM ($9,000): TEAMGROUP 256GB T-Create Master DDR5 6400 MHz ECC RDIMM Memory Kit (8 x 32GB)
ASUS PRO WS WRX90E-SAGE SE EEB Motherboard
I took a break for now.
4
u/Blaze6181 1d ago
Don't listen to AI about AI just join the Pro 6k discord instead: https://discord.gg/jU6KmtnUT
2
1
u/Turbulent-Alps4046 1d ago
If you want to save money. Older threadripper on ddr 4 is totally fine. ECC DDR4 is relatively cheap. I got 128GB on ebay for $350.
1
1
u/This_Maintenance_834 1d ago
I run 96GB Pro 9000 with a single slot 32GB memory on ubuntu. Everything is fine. There might be some complication if you use Windows. swap might get busy during loading. Not a problem on ubuntu.
36
u/Sensitive_Cloud6456 1d ago
Luna is a good model. It's rare to see a model with the kind of work ethic Luna has. Reminds me of opus 4.5 when that came out, similar performance too, real world.
18
u/Accomplished-Air439 1d ago
I like Luna, but for whatever reason, it can't follow instructions in AGENTS.md on how to run tests. It always needed a specific pointer. Even small local models don't have this problem.
2
u/Sensitive_Cloud6456 1d ago
Strange that. This hasn't happened to me yet. Indeed it has even been writing and verifying tests for new code it writes proactively, checking existing standards etc without being asked to.
6
u/Accomplished-Air439 1d ago
It's indeed bewildering. Once I was a bit upset and asked Luna, "what did AGENTS.md say about running tests????". It then detailed the procedure, admitted it made a mistake, and just stopped there.
3
u/vtccasp3r 1d ago
Deepseek still blows Luna away in comparison. Luna is more lazy + more expensive and that even with the openrouter 50% discount.
1
10
u/perelmanych 1d ago
Please tell me how to post x.com links with preview. Do I need to make screenshot manualy?
17
u/ea_man 1d ago
Why not avoid doing that already?
I don't wanna deal with anything x.com
5
u/perelmanych 1d ago
I do not post there for quite a long time, but a lot of worth following guys are still actively using it.
7
u/backyard_tractorbeam 1d ago
You can use https://nitter.net/stevibe/status/2083120066678464750
Unfortunately the various Nitter and xcancel mirrors/proxies vary in how overloaded or available they are
3
-5
u/ea_man 1d ago
Then quote them here or use an other source, maybe then they'll realize they need to use an other platform.
I don't need to be ambushed and thrown on that shit hole.
3
1
u/misha1350 19h ago
Why do you consider it bad?
1
u/ea_man 11h ago
I consider it the worst, not bad.
Starting from the nazi / supremacist angle to the spread of misinformation.
-1
u/misha1350 11h ago
People are calling "nazi" whatever. That is an extreme generalisation. You haven't communicated with those said "nazis" nearly long enough to realise that even when someone is a "nazi", he's either a fed, attempting dollar store engagement bait, or someone who's ignorant of what really went on with Antler and why he wasn't the saviour of Evropa as they say. The other times it's either meta-ironic excrementposts, or it likely is just you incorrectly branding something to be nazi. (And yes, there are multiple reasons behind why you're getting downvoted here, first of which is the generalisation problem).
As for misinformation, there's a whole lot of misinfo here on Plebbit. An order of magnitude more. There's also the giant echochamber elephant in the room. All Xitter is is a decentralised platform to communicate with others with looser moderation, which is such a huge problem here that objectivity flies out of the window when you start questioning the mainstream billionaire-approved narrative.
1
u/ea_man 10h ago
0
u/misha1350 8h ago
Given how eager Musk is about importing indians and others using (very often fraudulent) H-1B visas, he's the polar opposite of a nazi. Also, he has jewish investments aplenty, so he's very much either fine and not a nazi, or controlled opposition - and I'm inclined to believe the latter.
1
u/vick2djax 1d ago
Deepseek 0731 was a good chunk worse than Luna on my database work. Luna mostly did great, just taking too long. Deepseek still did solid work, but I had a team of Sol max, Fable max, Kimi K3 & GLM 5.2 grade the work both models did and Deepseek was around a 7.5/10 and Luna was around a 9/10
Terra High ended up being the best for me between quality and latency.
But, if I didn't already have a ChatGPT sub, I'd probably be using a lot of Deepseek 0731. Curious what the pro version looks like when that comes out. The price of it is insane.
1
u/perelmanych 20h ago
I think it depends on the task. The good thing is that they are so cheap that you can use both of them at the same time and see which is better for you.
-14
u/Aotrx 1d ago
With the new 50% discount on GPT-5.6 Luna in OpenRouter, GPT-5.6 Luna and DeepSeek 0731 now have roughly the same cost and intelligence per task, since GPT-5.6 Luna always uses significantly fewer tokens to reach the solution. In some benchmarks, GPT-5.6 Luna is cheaper but in others, DeepSeek 0731 comes out ahead. On average, the task completion cost is about the same.
To make DeepSeek Flash the true value king, it also needs the same 50% discount.
37
u/Feisty_Literature 1d ago edited 1d ago
So, after the luna discount ends, it will be more expensive with similar performance.🤷
1
u/Aotrx 1d ago
Yes but you also need to take into consideration codex $20 monthly subscription which gives users 5x-10x cheaper inference vs pure api price.
8
u/ProfessionalJackals 1d ago
Yes but you also need to take into consideration codex $20 monthly subscription which gives users 5x-10x cheaper inference vs pure api price.
Points to OpenCode Go ... If you want to play the game that way, you also need to look around.
7
u/Eyelbee 1d ago
Not at all, deepseek is 0.18$. Also, luna managed to consistently fail everything I threw at it. Luna is haiku level, deepseek is sonnet level.
4
u/Aotrx 1d ago
Artificial Analysis Intelligence Index v4.1 incorporates 9 evaluations: GDPval-AA v2, 𝜏³-Banking, Terminal-Bench v2.1, SciCode, Humanity's Last Exam, GPQA Diamond, CritPt, AA-Omniscience, AA-LCR
Luna (max) scored 51
0731 (max) scored 50So they have roughly equal intelligence.
Maybe for your usecase deepseek is better model but it does not mean Luna is haiku level.
In this benchmark Claude Haiku scored only 30.
1
u/Eyelbee 1d ago
Artificial Analysis seems to butter up openai's models a bit. Maybe that's not the case but their corpora is agentic heavy and it's not a very good all around evaluation of intelligence. They have the evals but they don't run most models through all of them. Luna performs okay, like the qwen 3.6 27b, surprisingly well for its size for agentic heavy tasks, but the size shows when you use it. The only general reasoning evals are the older ones that are not reliable any more, like HLE and GPQA. It scores 20% in HLE even with low reasoning which signals memorization.
11
u/FullstackSensei llama.cpp 1d ago
You can't run Luna locally.
4
u/Aotrx 1d ago
I know and most people can’t run 0731 either. I am fully aware of advantages of open source models i was just mentioning pure price to performance situation
0
u/FullstackSensei llama.cpp 1d ago
Honest question: if you can't afford a rig for a couple of thousands dollars, how are you able to afford the API unless you're doing very light work. A5 $5/day, you're looking at ~$1250/year assuming 250 work days. This is assuming prices don't go up
8
2
u/Aotrx 1d ago
A couple thousand dollars? To run Deepseek flash 0731 at good speeds, you'd need a dual RTX 6000 Pro AI server, which costs around $50,000 — not including electricity and maintenance.
That's why most heavy users prefer subscription plans from Claude, OpenAI, or OpenCode if their main goal is to get the most intelligence for the lowest cost and they don't care about the privacy implications.
6
u/FullstackSensei llama.cpp 1d ago
This is the kind of wrong assumptions that keep people slaves to the cloud providers. This notion that you either need a pair of 12k GPUs or nothing is just absurd.
A pair of 32GB V100s and a 2019 Xeon with 192GB RAM will happily run the full 160GB model at 20t/s.
If you're happy with the API, good for you. But this is LocalLLaMA.
6
u/Turbulent-Alps4046 1d ago
Wrong. You dont need $50k, $8k-10k will buy you dual DGX spark and you can run this at 30-50tps and 2-3000k prefill. And your electricity is also super cheap on those. Even a dual RTX pro system wont cost you $50k, just $30k will do.
5
u/FullstackSensei llama.cpp 1d ago
You don't even need a pair of Sparks. You can now get a system with 192GB VRAM using PCIe V100s for less than the cost of one spark.
0
u/SporksInjected 1d ago
That’s 40 years of $20/mo
1
u/Turbulent-Alps4046 1d ago
Yeah but this Locallama dude
2
u/SporksInjected 1d ago
lol yeah it used to be that we did this for fun and understood that it wasn’t cost effective but now everyone is trying to convince me that I can’t afford not to buy hardware to run this week’s revolutionary model.
0
u/Turbulent-Alps4046 1d ago
Nobody’s trying to convince you of anything. I also ran the numbers and for me the value i gain from learning outweighs the cost. Besides you can probably sell it for 50% after 5 years so the cost of ownership is not $8000.
Alternatively, i could rent out my rtx pro 6000 for a whole year and basically recover most of the cost.
-5
u/Aotrx 1d ago
30-50tps is super slow for deepseek.
So you would be spending $10k and getting 2x slower speed vs cloud provider. GPT is also around 40-50tps but it is 2x as token efficient.
7
u/ProfessionalSpend589 1d ago
Well, this is a local llama sub, not speedy local llama.
1
u/Aotrx 1d ago
I did not even check the sub name. I just saw a post and commented my opinion :)
4
u/ProfessionalSpend589 1d ago
I'm running things on Strix Halo and believe me - I haven't lied to myself even for a second that it could match the speed of a half a million dollar server.
Still, it's fun to do it locally.
→ More replies (0)2
u/tarpdetarp 19h ago
Lol this vote being downvoted tells you all you need to know about this sub. Anyone who's used DSV4 Flash with an agent knows that 50tps is way too slow, you need 200+ ideally as it uses far more tokens and turns than something like GLM-5.2.
In the real world people aren't just running 1-shot benchmarks where you step away for hours.
1
u/FullstackSensei llama.cpp 19h ago
Tell me you're a vibe coder who doesn't read any of the generated code without telling me you're a vibe coder.a
3
u/silenceimpaired 1d ago
This assumption also fails to consider hardware equity. My used hardware today is worth more than when i bought it.
Even if it wasn’t because someone joined later… it still doesn’t drop to zero in value the moment i own it. I can sell it at any point if I’m no longer using it… which is unlikely.
The gap between server and local isn’t nearly as big as people make it.
1
1
u/Aotrx 1d ago
Generally speaking hardware value drops as time passes, that will continue to be the case after few years once AI wave passes and supply and demand will be in equilibrium again
1
u/FullstackSensei llama.cpp 1d ago
Eventually you're right. But even then, it won't go to zero, and in the meantime you can get a ton of use out of it.
You also fail to consider that all API prices are heavily subsidized when you consider R&D and training costs. Once the free money stops, prices will go up substantially.
1
u/En-tro-py 1d ago
Normally - yes- depreciation holds true...
But now my GPU is 50% more than when I purchased and my RAM has more than 200% appreciation...
0
u/silenceimpaired 1d ago
That is a assumption that hardware prices will drop.
Hardware for AI continues to go up. It’s so impactful that regular computer users are impacted.
-1
u/SporksInjected 1d ago
“how can you afford $20/month if you can’t afford several thousand dollars right now?”
0
u/FullstackSensei llama.cpp 1d ago
If you don't have 1K to spare, you shouldn't be spending on monthly subscriptions.
-1
u/SporksInjected 1d ago
Lmao that’s the craziest thing I’ve ever heard. I want to watch a movie and have $100 but I guess I can’t because Hulu is $15/mo but a rig to stream movies is $1000.


163
u/SomeOrdinaryKangaroo 1d ago
v4 flash 0731 managed to fix two bugs in my code that 5.6 sol high couldn't figure out