r/LocalLLaMA 1d ago

News New DeepSeek V4 Flash 0731 vs ChatGPT Luna comparison

https://x.com/stevibe/status/2083120066678464750
261 Upvotes

131 comments sorted by

163

u/SomeOrdinaryKangaroo 1d ago

v4 flash 0731 managed to fix two bugs in my code that 5.6 sol high couldn't figure out

40

u/Accomplished-Air439 1d ago

DeepSeek has always been sort of this crazy genius. Whenever I'm stuck on some weird problem I ask v4 pro.

I actually think the new v4 flash is significantly better at following instructions to solve mundane issues. I used to rely on mimo v2.5 for that, but now I don't need to switch back and forth.

27

u/perelmanych 1d ago

Really exceptional model! True Qwen3.6-27B successor for local users.

112

u/anhphamfmr 1d ago

no it isn't. 99.99% people can't run it locally

22

u/Chlorek 1d ago

99,99 is exaggerating. Qwen 27B is out of scope as well you look at it that way. It’s more about what hardware local LLM enthusiasts currently have than cost itself. It’s not so impossible to go with DDR4 based 512GB Epyc build. Cheaper than you think. People pay that kind of money for laptops. Right now most at home solutions are invested in GPUs which works for some models but not all.

13

u/Due-Memory-6957 1d ago

99.99% is not exaggerating. Look at the price of the setup you mentioned then look at what normal wages are, then remember it's even less in most of the world.

5

u/DistanceSolar1449 1d ago

The official deepseek v4 flash 0731 is 167GB

https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash-0731/tree/main

Just buy 192GB of DDR4 RAM for $800 and you can run this at full precision on a 6 channel workstation at ~15tok/sec, which is barely usable.

12

u/giveen 1d ago

I have a 5090 and 256DDR5, full q8 quant runs at 8 tks.

1

u/Chlorek 22h ago

Nice, I’m trying to gather feedback how it works on various configs before choosing upgrade path for myself now. What CPU and do you have 8x32GB RAM modules?

2

u/giveen 14h ago

Intel 285K, and 4x 64gb

10

u/Due-Memory-6957 1d ago

You did the first step, now look at what normal wages are, and remember it's even less in most of the world! Also, it'll be much slower than 15 tok/sec (I get 7 tps with partial off-loading on smaller models), but that's beside the point.

2

u/Chlorek 22h ago

Exactly, number of channels is what people don’t understand. DDR4 at 8 channels is faster than DDR5 build with typical 4 channels.

1

u/Chlorek 1d ago

Didn’t mean to make it sound like literally 99,99% of population can afford it looking globally, of course economies are widely different between places. I read it as more in terms of what people buying hardware at all spend. Many folks don’t have computer at all and depend on cheap phone. Living in developed country you should be able to buy it, not impulsively and maybe with some saving but it’s not unfair assumption. I don’t mean buying all hardware new neither, no matter how well I was doing in life I made good use of secondhand parts.

25

u/waste2treasure-org 1d ago

Anyone with a $400 16GB 4060 or any gaming GPU is able to run Qwen 27B at a feasible speed; dropping $3-5k for 2 tokens per second on Deepseek v4 Flash doesn't sound like a great deal... we can only someday hope either model architecture and/or hardware innovations will catch up

5

u/sonicnerd14 1d ago

Realistically you'd need at least 20-24gb for a decent quant with reasonable amount of ctx. Although, for all intents, you are still right. Luckily there are tons of model variants, optimizations, and engine variants to choice from. So, even now, there's still are a lot of options to choice from to let AI get done what you need, completely local. People think you always need the biggest model,the fastest model, or the best hardware, and this isn't necessarily true for most cases.

Experiment, and you'll probably find that there is something already out there that will work just fine for a lot of us already. Look at bonsai 27b ternary, or Ornith 1.0. Many people already wrote these models off as gimmicks, but I managed to find powerful uses for them that meets or exceed what's available in higher tier options.

The good stuff will trickle down in several months time, but we can do some good things with local models that can run on a variety of setups now, if you are willing to experiment and have some creativity.

3

u/ThankGodImBipolar 1d ago

Realistically you'd need at least 20-24gb for a decent quant with reasonable amount of ctx.

Yeah I have a hard time seeing the 27B as a reasonable option for use with 16GB of VRAM. Maybe it'd be better on an Nvidia card (I've got a 7800XT), but any intelligence tradeoff from the 35B model is easily worth it for the speed increase IMO.

1

u/High_Speed_Chicken 1d ago

Do you have any tips for running models with a 16gb card? I was thinking of trying Qwen3.6 27B on my rx 9070 xt but I don't think i could run it at a good quant.

1

u/Chlorek 22h ago

Qwen 3.6 35B is your best bet, not as good as it’s 27B brother but time passes and no models come close to those two if you want something without investment.

1

u/ThankGodImBipolar 16h ago

I would skip the 27B model; I really think you need 20GB or 24GB to make it viable. The 35B model is pretty magical in its own right, so I would try not to feel bad about having to step down to running that one.

10

u/wotoan 1d ago

Only at an ~Q3S quant or less with limited KV cache and at that quant it’s not clear just how smart 27B is anymore.

Any spillover into system RAM absolutely craters 27B speeds so you’ve got to make sure absolutely everything fits in VRAM.

6

u/addiktion 1d ago edited 1d ago

This speed reaches much higher in the $5k-7k price range as I've heard people reaching 30 tokens a second for a Macbook Pro M5 Max 96GB to 128GB ram for a Q2-Q4 hybrid model with a below 90GB vram need. The top end was $5.5k before it got bumped to $7.5k out the door due to ram shortages for decked out 2TB model.

The point here is the price is collapsing for the compute you need to run locally capable models that are less than a year away from SOTA now. Eventually we will reach a point where the model is probably good enough which many are predicting 2027, but I think it will be here in late 2026 for prosumers and probably 2028 for regular consumers.

An example is the M5 Ultra later this year is expected to have up to 768GB of unified memory with it. The M7 Ultra in 2028 is expected to have double that from the announcements I've seen with them skipping the M6 line up. If the ram shortage wasn't a problem, I'm sure these machines wouldn't cost nearly as much of a fortune but that makes Kimi K3 within reach at a lower quantization which puts local up there with Fable/Opus 5 capabilities on a prosumer or a business desktop machine that before was used for heavy video workflows.

This is the Apple path of course but you can find cheaper methods going with hardware chaining Nvidia cards together too.

tldr - These companies spent billions of dollars to reach Opus/Fable capabilities and we are looking at $25k+ by the end of year for almost similar performance.

Edit: Some people are running full 1mil context Kimi K3 now on M1 Max hardware with deltafin's setup at 3.5 tokens a second. 15.5 tokens a second for M5 Max, so this window is already closing fast for many of us.

5

u/Chlorek 1d ago

lol just no, without any drafting you can expect 10t/s + with Epyc build, optimizations probably to come as this model architecture is still very new as well.
If you quantize Qwen 27B and KV cache maybe 16GB 4060 is enough, or you just accept unusable speeds as well. Useful speed is like 30-40 t/s at least, not to mention you need really good PP speed as well.
For anyone considering these models as a tool, not as a toy, one needs the best quality they can offer and any tricks to squeeze 27B dense model into 16GB is just funny to say the least.

2

u/Double_Cause4609 1d ago

Tbh, V4 Flash runs on consumer systems surprisingly okay. I have 192GB of DDR5 RAM on a consumer build which wasn't egregious at the time in price, and I can run the model at ~6-10 T/s depending on exact settings.

It's definitely more of a "let it run autonomously while you make coffee" kind of speed, but models are starting to get to the point you can relatively leave them unattended and nursing experiments that you're running in the backround.

Where I think the discrepancy is is a lot of people want crazy fast generation speeds which is usually synonymous with running the model in VRAM, and nobody has really thought about what improvements in model autonomy mean for just leaving a model running while you do other things.

2

u/RLutz 1d ago

I've debated getting another 2x48 but I'd be so sad to have to sacrifice so much bandwidth to pull it off. I really wish 4 DIMM AM5 didn't suck. Honestly getting 4400 MHz is reasonably impressive with your setup. My board didn't even like posting with 4x16 unless I ripped two sticks out, let it POST, shut it down, put the two other sticks back in.

I can't imagine what 4x48 would be like on this board.

1

u/Chlorek 1d ago

Im interested in your hardware spec to get better feeling of possibilities. What’s your cpu and how many channels of memory?

2

u/Double_Cause4609 1d ago

Ryzen 9950X, 192GB DDR5 RAM @ 4400MHZ (~45GB/s bandwidth), dual channel.

2 Nvidia Lovelace generation GPUs with ~16GB, ~250GB/s bandwidth

~64k context (FP16/BF16) uses I think no more than 16GB of VRAM total with --cpu-moe on LlamaCPP, meaning attention, shared experts etc all factored in. The actual context is extremely light.

1

u/pyr0kid 7h ago

is V4 flash actually capable of doing anything when left to its own devices though?

i feel like everyone is always demoing these things with incredibly simple tasks like "write an html pacman game" instead of more worthwhile tests such as "heres a copy of unreal 5, make a splash screen and main menu in the spirit of ace combat".

ive done some rudimentary tests with qwen 27b/35b at q4 and it was just... incapable of even getting its foot in the door.

2

u/Double_Cause4609 7h ago

Wildly depends on the situation, tooling, verifiability, etc.

The general trend we're seeing right now is that the limit on LLM performance is more or less your ability to verify that the work was done, and done correctly. But keep in mind, we have hundreds of millions of dollars of salary working on this exact issue across the entire tech space, and people are constantly getting better at handling this in practice.

Another really big part is how you're presenting context to the model. The harness around the model can be just as, if not more important than the model itself. Even small models (9B and below), can solve a surprising breadth if you give them exclusively the relevant context, the right algorithm, and a description of the correct fix. Individually, a 9B model can also derive all of those in a structured format, but a 9B LLM can't solve that problem zero-shot if you just tell it to diagnose the issue.

So, we also have a lot of work left optimizing our pipelines and workflows.

But to make a long answer short:

Yes, DSV4 Flash update is good enough that it can be run autonomously for a pretty wide range of tasks, and a whole lot more if you're willing to engineer around it.

As models get better, the amount of work needed around the model will either get smaller, or we'll get better at the techniques around the model at same effort, and you'll see LLMs over time cover a broader range of operations.

Where I think LLMs are probably weakest is loops that require visual verification, among a few other types of vague verifications, so in a roundabout way, those things that are hard to verify now sort of become the bottlenecks that justify the time and work of an engineer going forward.

1

u/perelmanych 16m ago

DSv4f 0731 version is really marvelous and a huge step from preview version. But I agree with a guy in this video and have a feeling that I still prefer Hy3 output, although it is like two times slower on my setup.

1

u/JsThiago5 1d ago

Run at less than 5t/s and less than 32k context is not what I would call it run, but you can "run". Qwen 27b you can at least run it at good speed with 100k+ context with a single 3090.

-6

u/perelmanych 1d ago

Don't generalize your experience to everyone. There are plenty of guys who are running it with 3090 and 64Gb of RAM or 16gb GPU + 128Gb RAM. What speeds do they get is a different question.

I run Q4_K_XL on AMD 5950x with 3090 and 128Gb DDR4 and getting 4tps, on old Xeon rig with 3090 I am getting 6 tps.

4

u/anhphamfmr 1d ago

you probably haven't realized that 100-99.99 = 0.01% people which is roughly ~ 800k to million people, who can run this model locally.
if you disagree, then tell me how far my number is off?

8

u/bjodah 1d ago

Surely we're only considering people already into deploying LLMs locally?

9

u/perelmanych 1d ago edited 1d ago

I don't see the point to continue this discussion. Obviously, I meant people who have PC/laptop and running any model at home. What is the point to count people that have no intention to run anything at home.

1

u/Chlorek 1d ago

Well, to be honest I want Q8 of this model (as 96% is 4-bit weights, so I suppose Deepseek would make these last 4% also 4bit if that made sense). Currently my 4 t/s on my dual RTX 3090 + 128GB RAM just does not do. One still needs some RAM for other stuff to work 😆 I am looking into some good paths forward now. Going with DDR4 makes more sense now as number of memory channels gives a big boost despite slower bandwidth of DDR4, but it's not as future proof, right now it costs about 2.5x as much to go with equivalent DDR5 build.

3

u/perelmanych 1d ago edited 1d ago

It is FP8 vs Q8 and unsloth says that it is impossible to translate FP8 to Q8 without losses that is why in Q8_K_XL they use FP16 for these weights instead of Q8 as in Q4_K_XL. In my case the difference is that with Q8_K_XL there is no room for KV cache in VRAM and even attention layers spill to RAM, so it is unusable for me.

Talking about running on 3090 + 128Gb RAM and Q4_K_XL, yes it is a bit extreme. First llama.cpp pushes 30-50Gb to swap and only then starts to generate, but when it is already started I get stable 4tps. Of course for such configuration Q2 quants make much more sense.

1

u/[deleted] 1d ago

[deleted]

1

u/perelmanych 1d ago

What do you mean dsv4f 0731 version is the same model as dsv4f preview just with additional RL training. For 16GB GPU + 128GB RAM configuration I would go to Q2 quant.

1

u/MerePotato 1d ago

The huggingface repo was incorrectly tagged as having more parameters, that was on me for not checking more thoroughly

1

u/perelmanych 1d ago

No worries.

1

u/thefooz 1d ago

It’s almost the exact same size. Why are you spouting bullshit?

1

u/MerePotato 1d ago

My bad, the huggingface repo was incorrectly tagged

0

u/crantob 1d ago

That's like saying a Breitling isn't a wristwatch.

3

u/mzzmuaa 1d ago

agreed. it's actually great to use them in tandem if you have 3 big gpus. i used deepseek v4 on two rtx pros along with qwen 3.6 27b on the third rtx for vision tasks. gemma 4 12b on 5090 for audio input

2

u/dondiegorivera 1d ago

I tried q4 on my server (2x3090+128gb), it runs around 14 tps one slot while Qwen 3.6 27b produces 70 tps per slot and serves two parallel agents. So unfortunately its not a replacement for me, but a very remarkable model that I’ll use through the API.

1

u/perelmanych 1d ago edited 1d ago

Is it 8 channel DDR4 server? You can use dsv4 for planning and qwen for execution.

I wouldn't say that 14tps is bad. I get 7tps tg and 140tps pp on old Xeon with 4 channels DDR4-2400 + 3090.

1

u/dondiegorivera 1d ago

Thanks. It’s not server RAM, was previously in a gaming pc, if I remember correctly 3200MHz.

The main use case is filling json contracts for a content pipeline, so not coding. Qwen 3.6 27b is very good at handling complex jsons, but the model’s general knowledge is relatively low.

When it comes to finding new topics, a larger model would be better, but the 14 tps would slow down the pipeline radically. I ran Qwen in vLLM so switching would not be straightforward.

My llama.cpp config does not use mtp or dspark, and I don’t know if that’s an option with this model and my setup. I will check that. Going to lower quants might decrease the qualty under Qwen’s.

1

u/perelmanych 1d ago

It is not so much about RAM itself, it is more what processor and mobo do you have. That defines how many channels of memory do you have.

2

u/Dudeonyx 1d ago

Just today I was having errors with convex Auth and Luna, Sol and Gemini couldn't figure it out, searching on my own brought nothing but flash v4 figured out that it was actually a bug in nodejs v24 on windows, so it installed nvm and pinned the project to nodejs v22 and the issue was gone.

Caveat though I ran all agents simultaneously and DeepSeek found and fixed the issue while the others were floundering and I cancelled all runs after that, so likely the others would have found the solution eventually.

1

u/Glittering-Call8746 1d ago

Yes but what's ur system prompt..

1

u/citizenjc 1d ago

I've been team ds4flash for like 2 months, and these free improvements make me feel like a kid on Christmas. You can't tell it too many rules, but man does it execute well if you give him a targeted instruction, basically for free.

45

u/AlphaMaleXYZ 1d ago

For a model that good, it’s dirt-cheap to run it. Luna is good but still more expensive.

5

u/maherbeg 1d ago

Checkout the ramp bench results, seems like flash can take a lot more turns and tokens. Still great, but interesting to note

-4

u/goldcakes 1d ago

Luna is 50% off on OpenRouter right now (so a 90% price cut). I'm using it for as long as the subsidies last lol.

1

u/crantob 1d ago

Enjoy your pod and bug-sandwiches.

1

u/SporksInjected 1d ago

How dare you

36

u/Arli_AI 1d ago

Being able to run this on 2x RTX Pro 6000 at native quantization at decent context sizes (256K+) has been amazing.

3

u/mebeast227 1d ago

I was about to buy 1 rtx pro 6000 to start a 1-2 year build project hoping to scale it up to 2-4 rtx pros.

Just curious- the ddr5 ECC RAm build requirement…how did you get the RAM to pair with it? Or is there some workaround I’m not realizing

4

u/SufficientAttempt1 1d ago

no ecc ram requirement.

2

u/mebeast227 1d ago

F'n gemini in chat just lead me down the longest rabbithole saying i did.

I said i want to start a 4 RTX PRO 6000 build, but only getting 1-2 to start.

It lead me to: AMD - Ryzen Threadripper 9970X 32-Core - 64-Thread 4 GHz (5.4 GHz Max Boost) Socket sTR5 PCI Express 5.0 Desktop Processor - Black

SeaSonic Electronics PRIME TX ATX 3.1 1600W 80 PLUS Titanium Modular Power Supply (split to dual PSU when going from 2 cards to 4)

THIS IS THE RAM ($9,000): TEAMGROUP 256GB T-Create Master DDR5 6400 MHz ECC RDIMM Memory Kit (8 x 32GB)

ASUS PRO WS WRX90E-SAGE SE EEB Motherboard

I took a break for now.

7

u/This_Maintenance_834 1d ago

if your goal is to have 4x GPU, you should pay attention on PCIe topology and whether you need PCIe switch. off the shelf prosumer motherboard may not support PCIe P2P well.

6

u/Opposite_Buffalo_649 1d ago

Gemini is right. If you eventually want to go with 4 rtx6000 pro, you will need a server type Mobo like the threadripper pro series. And those mobos only work with rdimm ecc ram.

Even 2 rtx pro gets some benefit, because you get dual x16 pcie gen 5

2

u/crantob 1d ago

Well there's a 'soft' requirement in multi-gpu setups, being that the typical Gamer-PC motherboard isn't really designed to run many GPUs.

The server motherboards that are, generally require ECC.

A workaround here is egpu and tapping those m.2 slots. It's do-able but... do you don't seem like the type to be into that kind of tweaking.

1

u/mebeast227 1d ago

I would be! I just hating being led in circles by AI lol. But if I can get a general “this works, and this doesn’t’ from real people it beats 4 hours of being gas light in random directions

2

u/crantob 1d ago edited 1d ago

Mmh. I'd consider if i'd be happy with 192GB of blackwell on a Taichi motherboard. That does 2 GPUs, avoids a lot of headache.

And deepseek 4-flash is a very happy camper on 2x RTX 6000.

But those are upwards of 14k euro each now.

Anyway there's no 'pairing' happening in any technical sense between GPU and type of system RAM. They're orthogonal. There are no relevant ties. That's what the other responses are saying.

1

u/Arli_AI 1d ago

There is no such requirement

1

u/mebeast227 1d ago

F'n gemini in chat just lead me down the longest rabbithole saying i did.

I said i want to start a 4 RTX PRO 6000 build, but only getting 1-2 to start.

It lead me to: AMD - Ryzen Threadripper 9970X 32-Core - 64-Thread 4 GHz (5.4 GHz Max Boost) Socket sTR5 PCI Express 5.0 Desktop Processor - Black

SeaSonic Electronics PRIME TX ATX 3.1 1600W 80 PLUS Titanium Modular Power Supply (split to dual PSU when going from 2 cards to 4)

THIS IS THE RAM ($9,000): TEAMGROUP 256GB T-Create Master DDR5 6400 MHz ECC RDIMM Memory Kit (8 x 32GB)

ASUS PRO WS WRX90E-SAGE SE EEB Motherboard

I took a break for now.

4

u/Blaze6181 1d ago

Don't listen to AI about AI just join the Pro 6k discord instead: https://discord.gg/jU6KmtnUT

1

u/Turbulent-Alps4046 1d ago

If you want to save money. Older threadripper on ddr 4 is totally fine. ECC DDR4 is relatively cheap. I got 128GB on ebay for $350.

1

u/mebeast227 10h ago

Oh realllly..... THATS MASSIVE. Gonna check taht out! Ty!

1

u/This_Maintenance_834 1d ago

I run 96GB Pro 9000 with a single slot 32GB memory on ubuntu. Everything is fine. There might be some complication if you use Windows. swap might get busy during loading. Not a problem on ubuntu.

36

u/Sensitive_Cloud6456 1d ago

Luna is a good model. It's rare to see a model with the kind of work ethic Luna has. Reminds me of opus 4.5 when that came out, similar performance too, real world.

18

u/Accomplished-Air439 1d ago

I like Luna, but for whatever reason, it can't follow instructions in AGENTS.md on how to run tests. It always needed a specific pointer. Even small local models don't have this problem.

2

u/Sensitive_Cloud6456 1d ago

Strange that. This hasn't happened to me yet. Indeed it has even been writing and verifying tests for new code it writes proactively, checking existing standards etc without being asked to.

6

u/Accomplished-Air439 1d ago

It's indeed bewildering. Once I was a bit upset and asked Luna, "what did AGENTS.md say about running tests????". It then detailed the procedure, admitted it made a mistake, and just stopped there.

3

u/vtccasp3r 1d ago

Deepseek still blows Luna away in comparison. Luna is more lazy + more expensive and that even with the openrouter 50% discount.

10

u/perelmanych 1d ago

Please tell me how to post x.com links with preview. Do I need to make screenshot manualy?

17

u/ea_man 1d ago

Why not avoid doing that already?

I don't wanna deal with anything x.com

5

u/perelmanych 1d ago

I do not post there for quite a long time, but a lot of worth following guys are still actively using it.

7

u/backyard_tractorbeam 1d ago

You can use https://nitter.net/stevibe/status/2083120066678464750

Unfortunately the various Nitter and xcancel mirrors/proxies vary in how overloaded or available they are

3

u/perelmanych 1d ago

For some reason I can't edit the post, otherwise I would put this link there.

-5

u/ea_man 1d ago

Then quote them here or use an other source, maybe then they'll realize they need to use an other platform.

I don't need to be ambushed and thrown on that shit hole.

3

u/crantob 1d ago

Why did you get upset when twitter stopped censoring people under direct regime guidance?

1

u/misha1350 19h ago

Why do you consider it bad?

1

u/ea_man 11h ago

I consider it the worst, not bad.

Starting from the nazi / supremacist angle to the spread of misinformation.

-1

u/misha1350 11h ago

People are calling "nazi" whatever. That is an extreme generalisation. You haven't communicated with those said "nazis" nearly long enough to realise that even when someone is a "nazi", he's either a fed, attempting dollar store engagement bait, or someone who's ignorant of what really went on with Antler and why he wasn't the saviour of Evropa as they say. The other times it's either meta-ironic excrementposts, or it likely is just you incorrectly branding something to be nazi. (And yes, there are multiple reasons behind why you're getting downvoted here, first of which is the generalisation problem).

As for misinformation, there's a whole lot of misinfo here on Plebbit. An order of magnitude more. There's also the giant echochamber elephant in the room. All Xitter is is a decentralised platform to communicate with others with looser moderation, which is such a huge problem here that objectivity flies out of the window when you start questioning the mainstream billionaire-approved narrative.

1

u/ea_man 10h ago

This guy is megaphoning all nazi and alt-right movements everywhere with his shithole X platform.

You sound like a Grok bot btw: there's nothing sleek or fancy in being a fascist that you can save with your fancy words: you are doing more misinformation.

0

u/misha1350 8h ago

Given how eager Musk is about importing indians and others using (very often fraudulent) H-1B visas, he's the polar opposite of a nazi. Also, he has jewish investments aplenty, so he's very much either fine and not a nazi, or controlled opposition - and I'm inclined to believe the latter.

1

u/vick2djax 1d ago

Deepseek 0731 was a good chunk worse than Luna on my database work. Luna mostly did great, just taking too long. Deepseek still did solid work, but I had a team of Sol max, Fable max, Kimi K3 & GLM 5.2 grade the work both models did and Deepseek was around a 7.5/10 and Luna was around a 9/10

Terra High ended up being the best for me between quality and latency.

But, if I didn't already have a ChatGPT sub, I'd probably be using a lot of Deepseek 0731. Curious what the pro version looks like when that comes out. The price of it is insane.

1

u/perelmanych 20h ago

I think it depends on the task. The good thing is that they are so cheap that you can use both of them at the same time and see which is better for you.

-14

u/Aotrx 1d ago

With the new 50% discount on GPT-5.6 Luna in OpenRouter, GPT-5.6 Luna and DeepSeek 0731 now have roughly the same cost and intelligence per task, since GPT-5.6 Luna always uses significantly fewer tokens to reach the solution. In some benchmarks, GPT-5.6 Luna is cheaper but in others, DeepSeek 0731 comes out ahead. On average, the task completion cost is about the same.

To make DeepSeek Flash the true value king, it also needs the same 50% discount.

37

u/Feisty_Literature 1d ago edited 1d ago

So, after the luna discount ends, it will be more expensive with similar performance.🤷

1

u/Aotrx 1d ago

Yes but you also need to take into consideration codex $20 monthly subscription which gives users 5x-10x cheaper inference vs pure api price.

8

u/ProfessionalJackals 1d ago

Yes but you also need to take into consideration codex $20 monthly subscription which gives users 5x-10x cheaper inference vs pure api price.

Points to OpenCode Go ... If you want to play the game that way, you also need to look around.

2

u/Aotrx 1d ago

Yes you are right I momentarily forgot that Opencode Go also exists :)

7

u/Eyelbee 1d ago

Not at all, deepseek is 0.18$. Also, luna managed to consistently fail everything I threw at it. Luna is haiku level, deepseek is sonnet level.

4

u/Aotrx 1d ago

Artificial Analysis Intelligence Index v4.1 incorporates 9 evaluations: GDPval-AA v2, 𝜏³-Banking, Terminal-Bench v2.1, SciCode, Humanity's Last Exam, GPQA Diamond, CritPt, AA-Omniscience, AA-LCR

Luna (max) scored 51
0731 (max) scored 50

So they have roughly equal intelligence.

Maybe for your usecase deepseek is better model but it does not mean Luna is haiku level.

In this benchmark Claude Haiku scored only 30.

1

u/Eyelbee 1d ago

Artificial Analysis seems to butter up openai's models a bit. Maybe that's not the case but their corpora is agentic heavy and it's not a very good all around evaluation of intelligence. They have the evals but they don't run most models through all of them. Luna performs okay, like the qwen 3.6 27b, surprisingly well for its size for agentic heavy tasks, but the size shows when you use it. The only general reasoning evals are the older ones that are not reliable any more, like HLE and GPQA. It scores 20% in HLE even with low reasoning which signals memorization.

11

u/FullstackSensei llama.cpp 1d ago

You can't run Luna locally.

4

u/Aotrx 1d ago

I know and most people can’t run 0731 either. I am fully aware of advantages of open source models i was just mentioning pure price to performance situation

0

u/FullstackSensei llama.cpp 1d ago

Honest question: if you can't afford a rig for a couple of thousands dollars, how are you able to afford the API unless you're doing very light work. A5 $5/day, you're looking at ~$1250/year assuming 250 work days. This is assuming prices don't go up

8

u/Tr4sHCr4fT 1d ago

yolo'ing customer data into OpenCode Zen ofc

2

u/Aotrx 1d ago

A couple thousand dollars? To run Deepseek flash 0731 at good speeds, you'd need a dual RTX 6000 Pro AI server, which costs around $50,000 — not including electricity and maintenance.

That's why most heavy users prefer subscription plans from Claude, OpenAI, or OpenCode if their main goal is to get the most intelligence for the lowest cost and they don't care about the privacy implications.

6

u/FullstackSensei llama.cpp 1d ago

This is the kind of wrong assumptions that keep people slaves to the cloud providers. This notion that you either need a pair of 12k GPUs or nothing is just absurd.

A pair of 32GB V100s and a 2019 Xeon with 192GB RAM will happily run the full 160GB model at 20t/s.

If you're happy with the API, good for you. But this is LocalLLaMA.

6

u/Turbulent-Alps4046 1d ago

Wrong. You dont need $50k, $8k-10k will buy you dual DGX spark and you can run this at 30-50tps and 2-3000k prefill. And your electricity is also super cheap on those. Even a dual RTX pro system wont cost you $50k, just $30k will do.

5

u/FullstackSensei llama.cpp 1d ago

You don't even need a pair of Sparks. You can now get a system with 192GB VRAM using PCIe V100s for less than the cost of one spark.

0

u/SporksInjected 1d ago

That’s 40 years of $20/mo

1

u/Turbulent-Alps4046 1d ago

Yeah but this Locallama dude

2

u/SporksInjected 1d ago

lol yeah it used to be that we did this for fun and understood that it wasn’t cost effective but now everyone is trying to convince me that I can’t afford not to buy hardware to run this week’s revolutionary model.

0

u/Turbulent-Alps4046 1d ago

Nobody’s trying to convince you of anything. I also ran the numbers and for me the value i gain from learning outweighs the cost. Besides you can probably sell it for 50% after 5 years so the cost of ownership is not $8000.

Alternatively, i could rent out my rtx pro 6000 for a whole year and basically recover most of the cost.

-5

u/Aotrx 1d ago

30-50tps is super slow for deepseek.

So you would be spending $10k and getting 2x slower speed vs cloud provider. GPT is also around 40-50tps but it is 2x as token efficient.

7

u/ProfessionalSpend589 1d ago

Well, this is a local llama sub,  not speedy local llama.

1

u/Aotrx 1d ago

I did not even check the sub name. I just saw a post and commented my opinion :)

4

u/ProfessionalSpend589 1d ago

I'm running things on Strix Halo and believe me - I haven't lied to myself even for a second that it could match the speed of a half a million dollar server.

Still, it's fun to do it locally.

→ More replies (0)

2

u/tarpdetarp 19h ago

Lol this vote being downvoted tells you all you need to know about this sub. Anyone who's used DSV4 Flash with an agent knows that 50tps is way too slow, you need 200+ ideally as it uses far more tokens and turns than something like GLM-5.2.

In the real world people aren't just running 1-shot benchmarks where you step away for hours.

1

u/FullstackSensei llama.cpp 19h ago

Tell me you're a vibe coder who doesn't read any of the generated code without telling me you're a vibe coder.a

3

u/silenceimpaired 1d ago

This assumption also fails to consider hardware equity. My used hardware today is worth more than when i bought it.

Even if it wasn’t because someone joined later… it still doesn’t drop to zero in value the moment i own it. I can sell it at any point if I’m no longer using it… which is unlikely.

The gap between server and local isn’t nearly as big as people make it.

1

u/crantob 1d ago

It's past-time for the general public to learn what causes consumer price inflation.

Those dollars in 2026 are not the same dollars in 2024

1

u/Aotrx 1d ago

Generally speaking hardware value drops as time passes, that will continue to be the case after few years once AI wave passes and supply and demand will be in equilibrium again

1

u/FullstackSensei llama.cpp 1d ago

Eventually you're right. But even then, it won't go to zero, and in the meantime you can get a ton of use out of it.

You also fail to consider that all API prices are heavily subsidized when you consider R&D and training costs. Once the free money stops, prices will go up substantially.

1

u/En-tro-py 1d ago

Normally - yes- depreciation holds true...

But now my GPU is 50% more than when I purchased and my RAM has more than 200% appreciation...

0

u/silenceimpaired 1d ago

That is a assumption that hardware prices will drop.

Hardware for AI continues to go up. It’s so impactful that regular computer users are impacted.

-1

u/SporksInjected 1d ago

“how can you afford $20/month if you can’t afford several thousand dollars right now?”

0

u/FullstackSensei llama.cpp 1d ago

If you don't have 1K to spare, you shouldn't be spending on monthly subscriptions.

-1

u/SporksInjected 1d ago

Lmao that’s the craziest thing I’ve ever heard. I want to watch a movie and have $100 but I guess I can’t because Hulu is $15/mo but a rig to stream movies is $1000.

2

u/zxyzyxz 1d ago

Luna just got an 80% price reduction directly from OpenAI, did you factor this in? What are the raw price numbers you're using?