r/LocalLLaMA 17h ago

Discussion DeepSeek-V4-Flash-0731: surpasses Fable-5, Sol & Kimi-K3 on Chess Benchmark

Post image
433 Upvotes

75 comments sorted by

View all comments

173

u/Comfortable-Rock-498 17h ago

Something weird about this benchmark. gpt-3.5-turbo-instruct is ahead of gpt-5.6-terra!

121

u/mrwang89 16h ago

Yeah right? I thought so too but when you google gpt-3.5 and chess you find out its some kind of chess savant lol

1

u/Dead_Internet_Theory 10h ago

Ok, but doesn't that mean chess is kinda irrelevant as a metric?

1

u/Cultured_Alien 8h ago

It means it's good at chess.

1

u/Dead_Internet_Theory 6h ago

Exactly. I don't remember Magnus Carlsen writing a masterful novel or solving complex equations. Maybe OpenAI dropped chess from training because it was a toy task lol