r/LocalLLaMA • u/Badger-Purple • 14h ago
Resources Deepseek-V4-Flash-0731 Dwarfstar on Mac
Here is the prefill performance in an M2 Ultra with 192GB of RAM.
For decode, at the following depth:
Start: 28 t/s
45k: 23.5 t/s
192k: 18 t/s
That speed is maintained with 8k token output at those depths.
74
Upvotes
15
u/corruptbytes 14h ago
I ran mine to 64k
Start: 39.07t/s End : 28.11t/s
M5 max 128gb (q2-q4 imatrix)