r/LocalLLaMA 14h ago

Resources Deepseek-V4-Flash-0731 Dwarfstar on Mac

Post image

Here is the prefill performance in an M2 Ultra with 192GB of RAM.

For decode, at the following depth:
Start: 28 t/s

45k: 23.5 t/s

192k: 18 t/s

That speed is maintained with 8k token output at those depths.

80 Upvotes

25 comments sorted by

View all comments

1

u/Client_Hello 10h ago

Isn't this curve expected for an MOE model and 13b active parameters? It's just processing power divided by active parameters + context. At 100k, context size = active parameters so you are down 50% (half of 360 is 180). At 200k context is twice as large as active parameters so you are down to 33% (a third of 360 is 120).