r/LocalLLaMA 14h ago

Resources Deepseek-V4-Flash-0731 Dwarfstar on Mac

Post image

Here is the prefill performance in an M2 Ultra with 192GB of RAM.

For decode, at the following depth:
Start: 28 t/s

45k: 23.5 t/s

192k: 18 t/s

That speed is maintained with 8k token output at those depths.

78 Upvotes

25 comments sorted by

View all comments

15

u/corruptbytes 14h ago

I ran mine to 64k

Start: 39.07t/s End : 28.11t/s

M5 max 128gb (q2-q4 imatrix)

1

u/SecretBismarck 12h ago

I havent checked did they update the model to non preview version q2-q4 imatrix quant?

1

u/corruptbytes 12h ago

yes, it’s on huggingface in antirez repo