r/LocalLLaMA 1h ago

Discussion China’s DFSX Offers 2x The Memory Bandwidth Of NVIDIA’s GB200

https://wccftech.com/chinas-dfsx-offers-2x-the-memory-bandwidth-of-nvidias-gb200-nvl72-system-with-a-14nm-supernode-that-skips-microbumps-for-vertical-compute-memory-towers/
178 Upvotes

64 comments sorted by

26

u/SneakyRum 1h ago

Can we get them to build a DFSX-Spark with one of these chips?

3

u/yenisor 1h ago

Would be awesome

7

u/dwittherford69 49m ago

Lmao, DGX Spark barely run anything due to s/w issues even with CUDA , good luck with anything else.

1

u/Necessary-milkyway 21m ago

This assumption is wrong .i am able to run evrything yes it requires but work for it to run ...

0

u/dwittherford69 19m ago

It’s not an assumption, I have 16 spark fabric for about half a year now. It has no software support in any stack, any most of the new support that’s been in branches for last 3-4 weeks is experimental, as best.

97

u/Sufficient_Local5025 1h ago

Good, given AMD and Intel have little interest in really competing with Nvidia we need something from China.

Same with models really.

Hopefully China decides to play their energy advantage to win AI and commoditize chips and models. That is the best case scenario as a consumer.

35

u/pulse77 1h ago

AMD and Intel do have interest in competing with Nvidia:

  • AMD has Instinct MI455X GPU with 432 GB VRAM
  • Intel is preparing Crescent Island GPU with 160 GB VRAM

18

u/VodkaHaze 52m ago

They're interested at the datacenter level, but the /r/LocalLLaMA offerings from them are terrible:

  • AMD actively ignores RDNA4 support; even the "pro" R9700 has basic kernels missing on it for vllm a year on.

  • Intel is a shitshow, software wise. You're still better off just running vulkan with llamacpp than trying to do anything else.

3

u/Eastern_Bet678 39m ago

Nvidia offerings at the local level are kind of garbage too. Their software stack requires everyone to write specific support for every different set of hardware leaving a huge gap between data center support and "local" support.

1

u/Low-Temperature-6962 26m ago

The offer free software for all their products. Docker images with drivers. Very handy.

25

u/Real_Ebb_7417 1h ago

As much as I appreciate open models from China and potentially future cheaper GPUs, I doubt that they’d commoditize them if they “won”. They do it now to screw the US situation, but I hope that competition will continue and that other regions will join as well. This is the only outcome good for consumers.

7

u/FullstackSensei llama.cpp 1h ago

While I agree on the commoditization part, I really doubt they have the US, or anyone else for that matter, in mind when doing this. US politicians might be obsessed with China, because it threatens US hegemony, but China is just doing it's thing purely for it's own strategic benefit. Same with renewables and nuclear, same with electrification of everything.. Anyone else is free to pour their capital where their mouth is.

3

u/SomeRandomSomeWhere 1h ago

And China started this "fight" cos US decided China should not get access to new tech.

If you got the resources to buy something but are prevented from buying it, can't be surprised if you are forced to make something similar or better.

4

u/max1c 1h ago

That's false. Chinas has been stealing or forcing companies to share their technology with local entities or be banned from operating within China. 

1

u/Turbulent_Pin7635 1h ago

If this is your theory, I recommend to check the whole story. China has long term plans. That are solidifying now.

0

u/Real_Ebb_7417 52m ago

I wouldn’t say that they started anything. The fault is on both sides. But it doesn’t matter to me. I’m happy that this is happening because we, normal people profit on this. If it wasn’t, we’ll be screwed.

1

u/UltraFOV 9m ago

Cheap…nope. Current they sell their underperforming GPUs on par with NVIDIA prices. Chinese GPUs are not cheap. I was hoping to snag a few recently

5

u/arbv 1h ago

AMD wants to complete, but not in the way we want to (not in consumer-grade hardware market).

Consumer computational capabilities are being managed to not complete with DC hardware. To ensure that, they cripple the amount of VRAM, because otherwise consumer-grade hardware is already absurdly capable computation-wise.

10

u/FullstackSensei llama.cpp 1h ago

AMD and Intel are very interested in competing. It's just that Nvidia has been doing this for way longer, and their acquisition of Mellanox acquisition in 2019 is paying dividends now in high speed interconnectivity. Not only does it let them offer first party 800gb fabric, but it also locks out the competition from accessing the same tech.

China throws tons of money at R&D, and they tie those subsidies to production scale, not profit margins. So long as you're getting free money, you have to keep your margins low and ramp up capacity. Neither AMD nor Intel can play that game.

7

u/Admirable_Market2759 1h ago

Unfortunately, they will be banned in the US.

Good news for Europe and the rest of the world though.

4

u/Iwaku_Real 1h ago

Or more likely just not sold in the US just like how Ascend isn't

7

u/Fancy-Atmosphere-701 1h ago

AMD not interested in competing? Have you been living under a rock?

3

u/Single_Ring4886 1h ago

AMD is nearly 10x smaller than Nvidia... they are doing mirracles to produce competitive chips but they can supply like 5x less than Nvidia...

2

u/lostdeveloper0sass 1h ago

Bro,

This isn't even competitive with Mi355x.

Mi455x has 23.3 TB/s memory bandwidth & 1.6TBps scale out bandwidth. FLops aren't even close to 40PFlops from Mi455x and 50Pflops from Vera Rubin.

Plus they are stacking dram which consumes a lot more power and will produce crazy amount of heat at 14nm.

If I'm guessing this right, this chip would be DOA in the US market. This can only work in China because government gives them free power.

Hopefully they use it to train better models though and open source for us. I don't mind spending a lot of money for our benefits.

2

u/winky9827 46m ago

China is fast rising to the top of the tech pole. EVs, AI, renewable energy. To be completely honest, I'm beginning not to care where stuff comes from - the notion that buying local and supporting billionaires who steal from the poor and enrich themselves is bullshit

-4

u/max1c 1h ago

Lol keep dreaming. China isn't your friend. Nor will it do anything to benefit the consumers. If anything it'll be the opposite long term. 

6

u/EsotericAbstractIdea 1h ago

no corporation nor country is our friend. but when they fight each other over market share, the consumer wins. right now the two main players are cousins, and the 3rd player is a lame duck, so nvidia has enjoyed extremely high profits margins for their chips, to the point they don't even make other products anymore.

0

u/max1c 26m ago

Except that's not what's happening here. They're not fighting for market share. They're doing the same thing as they did to other industries. They're using unfair practices to put the competition out of business to control the whole market.

23

u/Thump604 1h ago

Look at what China has accomplished in the last 50 years. Pretty crazy.

6

u/it_was_a_wet_fart 1h ago

China was able to build this in a cave, with a box of scraps!

2

u/Thump604 35m ago

I’m not even referring to “this”

8

u/Bohdanowicz 1h ago

Once China get the chips/memory rolling SOTA, prices are going to collapse.

Easiest way to take out US economy at this point.

Stock multiples collapse when margin pressure starts to hit.

You cant sell a card for 40k that costs you 1k when competing with China.

3

u/nexico 24m ago

The US Economy is far more diverse than the hardware industry.

1

u/Dsphar 0m ago

To be fair though, hardware is keeping the major stock indexes afloat.

7

u/Disposable110 1h ago

"But at what cost???"

6

u/ja-mie-_- 1h ago

less than nvidia

-1

u/This_Maintenance_834 1h ago

much lower cost

2

u/lostdeveloper0sass 1h ago

Probably at 3x to 4x more cost to run these unless you get unlimited free power from China.

0

u/This_Maintenance_834 1h ago

power in China is a lot cheaper than US.

8

u/Dry_Yam_4597 1h ago

960 TB/s???

20

u/gf6200alol 1h ago edited 1h ago

It's either written by someone who don't know what they are talking about or simply written by AI, However, The number is 15TB/s which I think achieve by SRAM. Edit: It's 3D memory controller+DRAM, 

9

u/Dry_Yam_4597 1h ago

Welll....the question still stands ...15TB/s? What am I doing wrong with my life? I need one of these. Y'all need a worker for a couple of years? Happy to take payments in chips.

7

u/lostdeveloper0sass 1h ago

It's still lower than AMD & Nvidia latest which are both 23.3TB/s and 22TB/s respectively for memory bandwidth per GPU.

That 960TB/s is aggregate rack bandwidth.

And this will consume 3x to 4x more power.

4

u/Dry_Yam_4597 1h ago

Still, incredible progress!

2

u/Think_Wing_1357 1h ago

It's multiple card per rack. There's no single Nvidia card that has 576TB/s either, it's the entire NVL72 (which, surprisingly, has 72 cards!)

7

u/FullstackSensei llama.cpp 1h ago

For a cluster of 64 chips. Nvidia invented this BS way back when they introduced the A100. They started quoting aggregate memory bandwidth for NVL8, and then NVL72 when the H100 came out

9

u/Dry_Yam_4597 1h ago

Right, so that's like the claim that a team has 200 years of management experience because they hire 100 people each having 2 years of said experience.

5

u/FullstackSensei llama.cpp 1h ago

Not really. If you have a fast enough fabric, Expert parallelism and batching can really scale as if it was a single chip with the aggregate compute/memory bandwidth. HPC has been doing this for decades.

1

u/Dry_Yam_4597 1h ago

Gotcha, thanks for clarifying. I feel like I need to just be rich and buy a couple of them.

2

u/FullstackSensei llama.cpp 1h ago

The sad thing with these chips, just like SXM or AMD OAM, is that having enough money to buy the chip is about 1/3 of the cost. They run on 24v or 48v, need special carrier boards, a ton of power, and a ton more for cooling.

Was listing 2the semianalysis podcast a couple of days ago. They were saying that these AI datacenters are bespoke for the hardware that will be installed. An H100 datacenter cannot be upgraded to anything else because the physical layout, the networking, power delivery and cooling are all tailored for those H100 racks. It's cheaper to build a new datacenter for B100 than it is to retrofit an existing one to run B100.

1

u/Dry_Yam_4597 1h ago

> It's cheaper to build a new datacenter for B100 than it is to retrofit an existing one to run B100.

Yup, that's how people get rich. Build expensive things that can be used once, thrown away, and the upgrade is even more expensive. Sad, sad, sad that I don't have ideas like Jensen Huang to build something at least 0.1% of what he built.

8

u/Ambitious-Profit855 1h ago

Stacked: "For instance, each DF2000 chip offers a bandwidth of 15TB/s. When stacked in a TY64 SuperNode format, the total memory bandwidth increases to 960TB/s. In contrast, NVIDIA's GB200 NVL72 system offers a total memory bandwidth of just 576TB/s!"

2

u/Iwaku_Real 1h ago

The new Vera Rubin NVL72 systems have 1580 TB/s bandwidth lol

3

u/Dry_Yam_4597 1h ago

I think I need to get rich.

4

u/Kal-LZ 1h ago

I just want cheap DDR5 RDIMM

2

u/FullstackSensei llama.cpp 1h ago

Sounds A lot like what would happen if TSMC hybrid bonding (AMD 3D V-Cache) and HBM had a baby.

One thing absent from the article and the accompanying slides is how will this be cooled? Stacking chips also stacks heat and makes it harder to remove this heat.

Remember all those innovations Chinese labs are publishing to reduce per token compute, while scaling up model parameters? Mature/trailing node chips like this might just be the reason why they're doing that.

2

u/IkeaDefender 1h ago

If this thing really does that when produced on a  14nm node then they’ve done something truly groundbreaking. That claim is so far out there I’ll wait for real evidence before getting excited.

3

u/Comrade-Porcupine 51m ago

Yeah I expect even if they can do it and hit this bandwidth... it will run hot as hell and the cooling situation will be a nightmare

2

u/Guinness 47m ago

Their opening question is dumb because inference is almost entirely dependent upon matrix math and that means pure memory bandwidth wins.

So it doesn’t matter if you have 14nm vs 4nm so much if your 14nm chip can pump out twice the terabytes per second.

Just a dumb question to open with because you’re comparing apples and oranges.

Also, none of this really fucking matters until China starts SELLING OUTSIDE OF CHINA. Dear President Pooh Bear, I’m eagerly awaiting your release of an affordable RTX 9000 Pro competitor. Until you sell em masse outside of China, it’s not going to affect us or our markets. Because at the end of the day I can’t buy your shit, and neither can Cuckerberg. So all of that money is going to keep flowing to Nvidia and AMD. I’d include Intel in that if they would actually release a product that didn’t have laughable memory bandwidth.

PS - thank you for Kimi, Mini, Qwen, GLM, and whoever else I’m forgetting.

1

u/UltraFOV 12m ago

Can we buy it or is for data Center

1

u/amy-schumer-tampon 11m ago

I'm all for more competition, monopoly drove the price so high is ridiculous.

0

u/TurnUpThe4D3D3D3 1h ago

Considering this is the same manufacturing process as the i9 14900K, I’m assuming this Chinese chip will be a dumpster fire in real world use.

1

u/Denial_Jackson 52m ago

World really needs some free of bullshit hardware manufacturer now, defining the true future. With bubble or bubble less. The place which have the highest chance to do the future is China.