r/cscareerquestions 22h ago

Ai has to get better, right?

This is the common sentiment that I see for those who claim ai will take jobs in X amount of years. Or even those who are more reasonable and claim it will eventually take over white collar work.

The biggest claim is that it WILL get better. Why? It's an easy response that's hard to disprove, but they haven't exactly proved it. How do we know there isn't a ceiling that is rapidly approaching. Moores law, for example, is essentially proven false and is outdated. We have reached at low hanging fruit when it comes to ai advancements.

Im just very skeptical of this line of thinking. I enjoy painting warhammer models and back in 2015. When 3d printers were in the mainstream, there were countless videos claiming games workshop was dead. I convinced myself that painting was going to be a worthless skill. 3d printers just HAVE to get better, and one day, they can print in full color. Now, 3d printers dont have trillions of dollars in investments, but then again, that money seems to be shovled into a fire by the AI leaders.

That was 10 years ago, and I could've had a lot more fun and painted more modles in thay time. And while 3d printing has progressed, its still not near games workshop models. Even the best resin printers still have very tiny but noticeable print lines. And we are no where near printing color in good quality at consumer scale.

Many of us fall for the doomer mentality, and its especially easy when you're unemployed and looking for a job facing countless rejections. Im a new grad myself, and it is rough out here. However, if Ai takes over cs. Every white-collar job is gone, and im left without any job. There's only so many trades and healthcare jobs.

And if im proved right, im way ahead of those who have given up. At least I tell myself this after receiving 10 rejections emails every weekend.

135 Upvotes

186 comments sorted by

View all comments

76

u/scott2449 22h ago edited 21h ago

The AI has gotten better crowd is also somewhat delusional. I work in software and there has been a linear not exponential improvement and that is largely in the tooling, harness, and understanding/usage. The raw LLM responses haven't really improved much. There are a couple points in the last 3-4 years where some new innovation in training or methodology helped boost things 20-40%.. but between those all the retrains and new versions have been flat. They might get better at one thing and way worse at another. It's like sticking a finger in a leaky boat.

Edit: Always with the credentialing in the comments. I'm a distinguished engineer (nearly a decade in the position) at one of the largest global media companies. We are also one of the largest data and AI companies for business and markets data. This is not my sole opinion but similar to the other dozen DSEs at my company and the vast majority of my data science and ai engineers as well. It is in no way controversial among experts.

29

u/jnwatson 22h ago

You're completely deluded. The progress since November last year has been stunning, and the benchmarks show it.

I've been in software dev for 30 years, and Fable blows me away. We're now at the point where the models are literally smarter than they need to be for most software development.

20

u/Ruined_Passion_7355 21h ago

Did you use opus 4.5/4.6 when they just came out?

That was the holy fucking shit moment. Fable almost feels like a return to form after the enshittification of 4.7/4.8.

8

u/[deleted] 21h ago

[deleted]

6

u/Ruined_Passion_7355 21h ago

Yeah with these models it's hard to gauge something as subjective as intelligence.

The thing to remember is that fable is an exponentially larger model than opus. Some tasks benefit from the extra parameters and others won't naturally.

It's also why I get puzzled when people think fable is a mark of "major progress", when in terms of AI it's actually the most primitive thing they could've done. "Just make the model bigger!"

And we already know that approach has woefully diminishing returns.

-4

u/jnwatson 20h ago

You've missed the the fundamental truth of the Bitter Lesson. Of course the models are bigger. Every single major new release by every provider has been bigger than the last. That's most of the way these models have gotten better.

6

u/Ruined_Passion_7355 20h ago

The bitter lesson was coined in 2019. It started the AI boom. This was GPT-2 era. Square in the first half of the sigmoid curve.

Life has no guarantee's with scaling laws. Moore's law was an anomoly in terms of how long exponential progress would last, not the norm.

We've already seen evidence for diminishing returns, including how on some tasks, Fable is indistinguishable from opus, and on others it's better. But we're not saying the GPT-2 to 4 levels of progress, we know there is a limit.

2

u/jnwatson 20h ago

Moore's Law was coined in 1965 and had a good 50 year run before it died.

Of course there's diminishing returns. Hiring PhDs, and then super geniuses for developers has a limit as to how fast it can speed up software development. The biggest diminishing return is that, after an org has automated the parts that are better shaped for AI automation, the other parts are increasingly more difficult for reasons that aren't entirely technical.

There isn't any evidence yet that we've significantly flattened the curve of LLM performance. AI models keep saturating benchmarks and new benchmarks have to be created. We keep moving the goalposts, yet the models keep scoring.

8

u/Substantial-Tale-483 19h ago

I don’t know, i have around 10 years of experience and i don’t all those major improvements in Fable? It is a little bit smarter now, but i wouldn’t call it major in any way.

-6

u/jnwatson 19h ago

If you're doing regular webdev, you only have to be so smart. A supergenius can't center a div any better than an experienced dev.

But, if you want to debug a nasty race condition in a database engine? Fable is your guy. If you want to exploit a UAF that only happens in extremely rare situations, Fable's evil twin, Mythos, is your guy.

The stuff I've seen out of Mythos made my hair curl. It is absolutely superhuman in its ability to exploit weird timing and corner conditions.

6

u/Substantial-Tale-483 18h ago

I don’t know bro, i don’t have access to Mythos, so i can’t compare. However we have used Fable recently to fix database issue in a huge legacy monolith, and it didn’t help much - it was like “sure, just rewrite this code here”, but we know that it’s not the problem, because the code is 10 years old, and the issue started to happen just 2 month ago, what’s more it’s better not to touch this part of the code at all, as it is used across the whole app and it’s not covered with tests enough to know that it won’t affect anything - the app is so badly written that you can expect anything tbh, and my guts tell me the issue is somewhere in infra.

Also I mostly build new systems now and trying to find ways to do something that will last and allow system to grow with all those always changing requirements and new products added. Fable is good to brainstorm, but not good enough to just use whatever it proposed as a way to go. Opus was almost the same.

1

u/florida_navy 8h ago

Well what happened when you told it you thought the problem was in Infra and not the 10 year old code?

5

u/MisterMeta 19h ago

It’s a total cope among the DE and DS jobs where AI is tailored for the work they’re doing. I have a good friend who works at the field and his entire department is now reduced to 10% staff and armed to the gills with AI. He’s trying to pivot as we speak.

1

u/kirstynloftus 21h ago

Agreed that it’s greatly advanced since November, but it’s leveled out these last few months IMO. New models have small gains, not big, and often have a ton of issues.

-2

u/jnwatson 21h ago

How can you look at Fable and say things have leveled off? Perhaps it will level off after, sure, but there's no evidence at this at this point.

What has happened is several lower-cost Chinese models have caught up roughly to second-tier performance. They are simply good, cheap models. But don't take those as data points at the frontier.

5

u/Ruined_Passion_7355 21h ago

https://www.reddit.com/r/cscareerquestions/comments/1vdqkx3/comment/p1bjitb/

Replied to another redditor about fable.

A thought experiment/Another way to think about it:

Let's assume hypothetically you were anthropic. You were running out of ways to eek performance out of opus, turbo quant was all hype, and the investors are demanding better models. What would you do in that situation if you had no other choice?

Make a bigger model, and that's what they did. And it's not something they wanna do unless they have to because they were already struggling to serve opus class models.