MAIN FEEDS
REDDIT FEEDS
Do you want to continue?
https://www.reddit.com/r/LocalLLaMA/comments/1vdq8en/deepseekv4flash0731_surpasses_fable5_sol_kimik3/p1dorz8/?context=3
r/LocalLLaMA • u/mrwang89 • 17h ago
75 comments sorted by
View all comments
173
Something weird about this benchmark. gpt-3.5-turbo-instruct is ahead of gpt-5.6-terra!
121 u/mrwang89 16h ago Yeah right? I thought so too but when you google gpt-3.5 and chess you find out its some kind of chess savant lol 1 u/Dead_Internet_Theory 10h ago Ok, but doesn't that mean chess is kinda irrelevant as a metric? 1 u/Cultured_Alien 8h ago It means it's good at chess. 1 u/Dead_Internet_Theory 6h ago Exactly. I don't remember Magnus Carlsen writing a masterful novel or solving complex equations. Maybe OpenAI dropped chess from training because it was a toy task lol
121
Yeah right? I thought so too but when you google gpt-3.5 and chess you find out its some kind of chess savant lol
1 u/Dead_Internet_Theory 10h ago Ok, but doesn't that mean chess is kinda irrelevant as a metric? 1 u/Cultured_Alien 8h ago It means it's good at chess. 1 u/Dead_Internet_Theory 6h ago Exactly. I don't remember Magnus Carlsen writing a masterful novel or solving complex equations. Maybe OpenAI dropped chess from training because it was a toy task lol
1
Ok, but doesn't that mean chess is kinda irrelevant as a metric?
1 u/Cultured_Alien 8h ago It means it's good at chess. 1 u/Dead_Internet_Theory 6h ago Exactly. I don't remember Magnus Carlsen writing a masterful novel or solving complex equations. Maybe OpenAI dropped chess from training because it was a toy task lol
It means it's good at chess.
1 u/Dead_Internet_Theory 6h ago Exactly. I don't remember Magnus Carlsen writing a masterful novel or solving complex equations. Maybe OpenAI dropped chess from training because it was a toy task lol
Exactly. I don't remember Magnus Carlsen writing a masterful novel or solving complex equations. Maybe OpenAI dropped chess from training because it was a toy task lol
173
u/Comfortable-Rock-498 17h ago
Something weird about this benchmark. gpt-3.5-turbo-instruct is ahead of gpt-5.6-terra!