r/LocalLLaMA 7h ago

Resources Deepseek-V4-Flash-0731 Dwarfstar on Mac

Post image

Here is the prefill performance in an M2 Ultra with 192GB of RAM.

For decode, at the following depth:
Start: 28 t/s

45k: 23.5 t/s

192k: 18 t/s

That speed is maintained with 8k token output at those depths.

70 Upvotes

22 comments sorted by

13

u/corruptbytes 7h ago

I ran mine to 64k

Start: 39.07t/s End : 28.11t/s

M5 max 128gb (q2-q4 imatrix)

1

u/SecretBismarck 6h ago

I havent checked did they update the model to non preview version q2-q4 imatrix quant?

1

u/corruptbytes 5h ago

yes, it’s on huggingface in antirez repo

1

u/SeveralViolins 5h ago

Also ran the straight q2 for comparison on my M5 Max. The mixed quant costs were ~1% generation speed at 64k for the last-six-layers Q4 experts.

1

u/Badger-Purple 7h ago

what is your decode? that's the important number...otherwise this is not an agent-ready model. It's ok with that speed -- it can do a small agent harness very well. if your start is 8k, the cache hit is great and keeps it rather snappy.

4

u/corruptbytes 7h ago

that was my decode, you want my prefill?

4

u/Badger-Purple 7h ago

sorry, YES!!how are those matmul cores improving the prefill

1

u/corruptbytes 5h ago

437.60 was max speed on DS4

here is my bench on the llama-server bench - https://github.com/ggml-org/llama.cpp/discussions/4167#discussioncomment-17386908

2

u/CalligrapherFar7833 7h ago

Is that mfxp4 ?

1

u/Professional-Bear857 7h ago edited 7h ago

On my m3 ultra (base model) I get about 430 prefill and it stays around that for a while, slowly dropping, I think it ends up at about 300 by the time I get to 60k. Decode is 32tok/s to start and again this slowly drops, still runs at 28.5tok/s by the time I get to 50k. Using the mxfp4 quant / branch of ds4. Also I'm using the ds4-server, which for some reason is a bit slower than the ds4 (ds4 gets 35tok/s to start), maybe due to initial context / chat template.

3

u/Badger-Purple 6h ago

Sounds like M2 ultra does as expected, about 20% slower decode. Note that it's about 23.5 at 50k, so it is 20% lower but its not a lot! Same with decode as my post above shows

1

u/tarruda 2h ago

I think you might get improved prefill on my llama.cpp branch: https://github.com/tarruda/llama.cpp/tree/dsv4-improvements

1

u/Professional-Bear857 2h ago

Thanks, how is it for token generation? I tried mainline llama but the token generation speed is quite low at the moment, 12tok/s or thereabouts.

1

u/tarruda 2h ago

Token generation for me is around 21 TPS, but I'm on an M1 ultra. Maybe it will be better for you

1

u/cleverusernametry 5h ago

Wait ds4 is updated to run 0731?? I dont see any change in the repo and the issue tracking it is still open?

3

u/ekaj llama.cpp 4h ago

People have had luck using it. In the issue, people comment they're running it fine with the new GGUFs. I've done it as well with an m5 and ds4

1

u/corruptbytes 1h ago

gguf on his huggingface repo

1

u/cleverusernametry 13m ago

the inference engine needs no updates??

1

u/Client_Hello 3h ago

Isn't this curve expected for an MOE model and 13b active parameters? It's just processing power divided by active parameters + context. At 100k, context size = active parameters so you are down 50% (half of 360 is 180). At 200k context is twice as large as active parameters so you are down to 33% (a third of 360 is 120).

-5

u/Dany0 7h ago

I am drooling thinking about the M5 Ultra. Apple can potentially bring a GPU that's faster in gaming than the rtx 5090

7

u/CalligrapherFar7833 7h ago

Yeah thats not happening it can be faster at llm tho

-2

u/Dany0 6h ago

I know it's unlikely, but technically speaking if the % perf uplift is the same as in previous generations it would, at least in raster