Eight Myths on Software Engineering and GenAI
queue.acm.org289 points by tchalla 2 days ago
289 points by tchalla 2 days ago
>On my visits to the Bay Area, I would ask AI researchers or interns why they are doing their current research or projects, when in a year or three agentic LLMs could probably do them;
This is such a weird point to make that doesn't become correct just because everyone makes it, all the time. Why clean the ocean if some magic future tech will clean them? Why save the world now if some benevolent AI is 'just around the corner' and will do it for us? And people have been making this point for years now, and it's not like my job got any easier. I just got more AI.
https://www.poetryfoundation.org/poems/51294/waiting-for-the...
And I say that as someone who uses Claude Code in complex environments almost hourly; I, as the human, still have to do the thinking as Claude still 'can't jump' [1] and I have seen no evidence that they (or similar AI, any time soon) will 'jump' like a human brain does.
> This is such a weird point to make
I think it is a great point to make, because if everyone really believed that AIs will do everything without human intervention in a handful of years, as the marketing repeats again and again (AGI, singularity, etc.) and have been saying for years... why then get bothered?
Because we DO know LLMs have their hallucinations, limitations, perform tasks not previously seen way worse than humans, etc. And it seems that, for now, there is not a good or magic solution to it, it is inherent limitations of the paradigm.
Yes, you can feed more and more and more (curated data) and eventually make AIs excellent at task X or Y, but then you spend your time specializing those engines. So the work does not really disappear, it just shifts and you make it more replicable for a bound set of problems.
Needless to say that at some point I prefer to learn (and combine with AIs, it is ok) than acritically getting inputs from something until I become totally useless.
Unless we have a paradigm for which a fully autonomous AI can do everything, this will just become improving our productivity in some ways, with all the in-between bottlenecks that it has.
It's a very silly point to make to AI researchers specifically. If they don't work on those projects, the AI won't advance and won't magically be able to replicate the work in "one to three years".
Can you imagine scenarios that would make it less silly? I will give an example:
- The AI researcher might be working for a lab or company with much less funds than the top dogs. Are they likely to discover something that is worth it before a bigger model becomes more capable?
Ask the researchers working on Deepseek. They seem to be doing pretty well for themselves.
Just because a AI will be able to do it in the future does not mean that us plebs will be allowed to have access to it. That alone is enough of a reason for smaller labs to keep going; having a seat at the table.
Why spend money and time making the new flagship model when a future flagship model can make you the flagship model?
Then why not stop researching and doing the definitve model that will solve every problem? Why some people are not doing it?
Bc they are aware of the marketing and limitations. If they did believe it, then they would switch area of research.
> you can feed more and more and more
I have a meta thought..
Hypothetically what happens once there is no more data to be fed to the system? Are we expecting AI to invent its own data and reach full cognition?
Currently we are feeding it the data that humans created but if we stop (i.e "why bother?") thinking that AI will do it all?
This is a well-known problem that has been an issue for years now. You can't use models to generate data for models because it leads to "model collapse" where it amplifies quirks in the generated data until it's all quirks. Here is a random university press release about it (grain of salt etc)
https://www.utoronto.ca/news/training-ai-machine-generated-t...
In practice you can do it a bit (generated data from a better / different model is fine, some generated data might be useful if there is non generated data etc.)
The AI researchers are not the ones making the marketing, much less believing in it.
Is this really true? At least one very headline AI researcher is pretty much an Anthropic spokesperson.
That they believe in the PR is itself PR, they're paid to be spokespeople as well as researchers.
I agree. But this is not what you see on the headlines and what money-incentivized stakeholders are saying.
> inherent limitations of the paradigm
This is such a weird point to make. We are currently ( only ) discovering that paradigm; we are not inventing anything. We found a bunch of laws that produce rather cool results but our paradigm is incomplete which leads more or less wordy or frame-rich weird stuff like hallucinations, singularity and so on ... it's childish, really and on that funny pseudo-profound, pseudo-intellectual, pseudo-spiritual ( personal opinion, if it gets you horny, you go, baby ) "universe consciousness unity, Rick James, bitch" level ...
Our bodies and minds need proper AI, not all the stuff we already outsource to middle and/or passionate men and women. Other species on the planet would certainly like to see us get augmented by AI so we can solve as many survivability issues as possible to keep as many ecosystems running long enough ... whatever that means but whether animals and plants are aware of chance and potential is another philosophical debate.
To individuals, software is a hammer and chisel, a knife, a brush and canvas, pen and paper, a reading help, and to a good amount of people it's a microscope and a fine scalpel.
To collectives, it's a tool to work on consensus and conventions, to share and gather.
It's baby steps for civilizations and it looks like our particular species is gonna get stuck in a puddle of our own monkey shit, with bottles of champagne in our hands and monkeys grinding up and down the few ivory towers in proximity.
> why then get bothered
Humans are on different levels. Most have decided that "nature realized the/a bug and wanted someone dead" or "their survival is a matter of chance" is not acceptable at all and some people decided that sabotage, poison, abuse, rape, murder are acceptable means to get chicken shit ...
The "paradigm" of life is far from explored/discovered, so we simply can't content ourselves with presumptions about inherent limitations of the LLM and AI paradigm for any other reason than to uncover ( not invent ) other parts of the paradigm.
We are happy with what AI can do for us but "AIs will do everything without human intervention" sounds weird because babies are born and the older they get and the less sabotaged ( vs influence, cultural manipulation ) they get to grow up, the more breadth and depth humans want to experience. For this they need to learn and use their hands & fingers. They need to feed body and mind to find what triggers what, and what excitement and curiosity are inherent and which can or need to be added/acquired/experienced extrinsically.
How many associations will we be able to make if AIs will do everything without human intervention?
No we don't know your 'points'
The Hallucinations are becoming less, significantly by now.
It also might be already were it is cheaper for one of the big few companies to spend millions and billions to teach the LLM / creating the training data necessary for an LLM to do something which it is not yet good enough due to the fact, that they sell this capability then to everyone who wants to use this capabilitiy.
We have not seen the end of Reinforcement Learning, which does need a lot less training data but more compute.
I'm 'vibing' on the side a handfull of small things, no LLM trained on particular what i'm asking to do. Its very capable of stringing together enough things so it can clearly follow handwavy things i tell it to do, analyse error messages, analysing screenshots etc. all by itself.
There is not a single real ceilling in sight, we only have clear barriers like compute but constant fast progress.
The field of mathematics went from 'useless' to 'you start better using it' to 'gamechanger' in how fast? 1 year after coding? less?
I want signes that we hit a real problem, instead I get cheaper tokens, Chinese models becoming very good as open models, new model updates from the others, mathematicans now saying how good it is etc.
Only half a year ago I had to babysit an LLM, now i tell it 1-3 sentences and it just goes and does it. And that stuff runs without compile errors etc.
If AI makes us 10% or 20% betteer, which is not that much, this alone will lead to companies reduing their expensive staff by 10-20%, which will has real impact on a job area. Some jobs are already hard to sell like cyber security and basic image tasks.
Hallucinations were low hanging fruit in some ways. As someone working on a large-ish complex-ish distributed system that has to be maintained and support customers, it's still very high value to have Claude in the mix, but the core problem of needing to monitor, advise, course correct, and make sure you don't end up with more code and complexity than you need is, at least in my experience, still roughly the same. The sharp edges are being filed off very rapidly, but the core experience of "make and maintain a large system" isn't advancing nearly as fast, IMO.
I'm waiting for the agentic ai platform layer.
We see AI factories going in this direction but there is no real 'the open source ai platform' thingy.
It needs connectors to integrate with k8s, hyperscalers etc. it needs to be able to have a basic router, a way of configuring expert agents and interaction options for the human in the loop.
There is for sure things we need to build or change, but it def feels like to me that it would immediadly fix a few things today.
I'm not disappointed that it doesn't advance as fast as it feels
> The Hallucinations are becoming less, significantly by now.
Yes? What is the mega-solid technique that is used for it? Armies of people using curated data and reviewing it by hand? That is exactly one of my points: shifting the work elsewhere for specialized tasks. More replicable, improved, but, it scales infinitely and is autonomous? Can you assert that?
I am not denying there is some use (a lot of uses!) for this, but this is more nuanced than just: oh, they will replace us. Not at all, that day, with the current technology, is not going to arrive. This is just a systematization, fitting and tweaking of human knowledge by curated data. It is not the one true superintelligence they are selling us. To begin with, they do not have a concept of truth, but of probabilistic truth. Only that poses already a very, very big problem for the path to perfection.
> We have not seen the end of Reinforcement Learning, which does need a lot less training data but more compute.
Noone said the opposite, but I would like to know at which cost and if it is feasible. We do not have even enough compute power for current technology.
> . Its very capable of stringing together enough things so it can clearly follow handwavy things i tell it to do, analyse error messages, analysing screenshots etc. all by itself.
I use it every day for these tasks and it works well BECAUSE I review the output and makes me go faster. It finds a lot of things I would have not found and it also hallucinates another handful of them, which confirms my point about AIs not being able to be fully autonomous in any future point in time unless tweaked exactly for the task, and even then, it can still miss judgement a human could have for edge cases. So I am not sure of how bad or good it can be compared to a human but I am pretty sure it cannot be more reliable than an expert in many situations.
> Chinese models becoming very good as open models
I think they will be better in the long term if they follow this path. Not absolutely better but when mixing with economics and the fact that no frontier model is totally reliable anyway... why pay a lot for something that needs human inspection anyway?
> There is not a single real ceilling in sight, we only have clear barriers like compute but constant fast progress.
The ceiling is the paradigm itself, as I mentioned above. There is not a single chance with current technology that something could become "generically knowledgeable" and "reliable" both at the same time. If it becomes generically knowledgeable and reliable, it is bc of data fed into it and curated and tweaked by humans. This is not an original idea from myself, there are armies of people doing this every day around the world, you can check. This is where a lot of improvement comes from. Can this be reused? Of course. It is a generic solution? No way.
> Only half a year ago I had to babysit an LLM, now i tell it 1-3 sentences and it just goes and does it. And that stuff runs without compile errors etc.
Yes, I also do one-off scripts like this and code snippets, even reviews and others. Now go design a full distributed system. Use agents if you want. We come back in six months and compare it to a system that was properly written and tested by humans and we can compare the quality on some grounds:
1. how long it takes to add new features?
2. which ones act more according to spec once added?
3. when adding features, which ones have more bugs?
4. in the face of an error, will the agent delete my whole AWS infra (count the money losses if possible also)?
5. will I understand (or need to understand, but I bet yes) this code at some point in the future?
You have to count all that money also, not just I vibe coded something and it seemed to work. With full systems things become super messy. Now add the human factor of requirements and back and forth (iterations can be admittedly faster with AI, especially prototypes, but that comes with other costs also)...Not easy at all.
> shifting the work elsewhere for specialized tasks. More replicable, improved, but, it scales infinitely and is autonomous? Can you assert that?
I would say yes and it will scale. It will either happen through central LLM just paying for it and scaling it up to everyone on the planet (literaly) OR by the agentic layer every big business is building into their systems.
You needed some human to use your tool optimized for their company? With agentic layer you no longer need this. And if you look at companies like Google, they were pushing this notion for ages already because they saw an adoption problem of more 'complex' tools and trying to make it simpler and easier. Now you can act from the other side too.
> We do not have even enough compute power for current technology.
Exactly. Right? There is no ceiling if its clear that we don't even have the hrdware. But the hardware is a bottleneck not a ceiling.
> The ceiling is the paradigm itself, as I mentioned above. There is not a single chance with current technology that something could become "generically knowledgeable" and "reliable" both at the same time.
It doesn't need to be perfect, it only needs to be better than the avg human. And the current LLMs are already better than aat least 1-2 people in my team.
> Now add the human factor of requirements and back and forth
Yeah for now. Grill me skill made it a lot easier. Harness engineering is also being worked on, agentic layer, ai factories etc.
And all of this can be copy and pasted. There is only one harness needed which becomes the expert security reviewer and tomorrow everyone can have it.
I'm still discusing progress with LLMs with people and still not everyone is using it or playing around with harnesses or developgn an agentic layer. We still have a lot of work to do to even see how good it will become while it already is really good.
People are already borred of AI today and making wrong decisions based on the current level of AI while i think we will see continues progress for years.
> I would say yes and it will scale
So you mean AI will be useful for any general job without lots of training for those jobs? How about new tasks? Tasks it has not been tweaked for. When I deviated from the average, and not really weird things, when programming, the output was way worse than average stuff. And this is an explicit target of AIs nowadays.
I think you are missing a lot of details here, honestly.
> Exactly. Right? There is no ceiling if its clear that we don't even have the hrdware. But the hardware is a bottleneck not a ceiling.
No, the hardware is a bottleneck, the paradigm as we know it is a ceiling unless you massively and continuously feed this system with average tasks (which is useful). Which is exactly the opposite of what singularity and AGI have been promising.
The systems we have now (unless the paradigm changes) will keep doing, essentially, fitting. No concept of truth and limited inference. That inference is based on already existing data, not on future data. In fact, there have been experiments about feeding output back to the input of LLMs and the degradation of the quality is very visible. If they are supposed to be so "intelligent", why it happens?
> It doesn't need to be perfect
I can agree that for lots of tasks it does not. But for others it is just not a tool good enough.
> Yeah for now. Grill me skill made it a lot easier. Harness engineering is also being worked on, agentic layer, ai factories etc.
I will not deny there could be progress, but nothing similar to "autonomous", "reliable", "super intelligence" or "singularity" with this paradigm.
In fact, often in my experience, this is a waste of tokens for subpar results that shift the technical debt elsewhere. I mean if you try to develop full systems by "vibe-code like" techniques. If you use them judiciously, you can accelerate your workflow, maybe 2x, but not much beyond that if you want to have something worth to be used. Note that here I am talking about the full thing: with testing, quality, maintenance concerns and everything together.
If you want to ship a sub-par thing that will go to the rubbish in a couple of months, then yes, you can do that. But that will fail commercially any way. Unless your job is convincing enough people that you can go 10x faster every time, deliver some sub-par thing, and find another customer, which, to me, would equal a scam.
> So you mean AI will be useful for any general job without lots of training for those jobs? How about new tasks?
I'm pretty sure we will solve this issue. Either already through World Models or another architecture.
It could also be, that we just need a 10 or 100 Trillion Parameter model to match so many generic ways of solving tasks and keeping the concept in the LLMs 'head' to solve it that it will just emerge with parameter size. Like with fable they said that it can chain together exploits which wouldn't work as standalone exploits.
What if the only real barrier is the depth of understanding of concepts and this is exactly what is getting solved with parameter count?
But look how young this field really is if you start counting it when it became relevant on mass. Its not 'just' an LLM which is changing the world, its machine learning overall. Robotics wouldn't be were it is today if its not for machine learning. Took humans time and energy to take the leap, to start learning what the status quo is and then actually doing more with it.
> the paradigm as we know it is a ceiling unless you massively and continuously feed this system with average tasks
While I do think its doable to achieve AGI in 5-15 years, even if it doesn't happen and it always means that people train an LLM or whatever, if you need 10 experts to teach this to an LLM OR every single senior has to teach this to their juniors every single time, the LLM will always win.
I'm now team lead for 10 years and every single year I teach them the same thing over and over and over again.
the craziest thing about this? If i wouldn't tell them what they are doing wrong, they wouldn't even know it.
Quality is already a very flexible term for a lot of people.
> I can agree that for lots of tasks it does not. But for others it is just not a tool good enough.
Yet.
> If you want to ship a sub-par thing that will go to the rubbish in a couple of months, then yes, you can do that. But that will fail commercially any way.
Now we come to the reality: I have seen so much garbage software its crazy. People using md5 as a password hash in 2024! No clue what coding best practices are, teams without code review, teams without a security expert not even knowing what crazy things they do day in day out.
Just a few month ago a team build an API for my team including a Swagger UI. Half of it didn't work. You pressed a button on the Swagger UI and a 500 returned.
And do'nt underestimate what it means that a lot of business people don't like software people. You know that fruit basket we get? and water and stuff? they don't do it because they like us they do it because thats what you have to do. if a Product Owner starts vibe coding with AI, he will have leadership convinved in no time, then it goes on production and it will run for waaaaaay longer than anyone would have guest.
Besides that there is plenty of software were complexity is less relevant or security is not that big of an issue.
> I'm pretty sure we will solve this issue. Either already through World Models or another architecture.
Please elaborate. How? With which technique? Currently the only path forward is to feed more data and tweak for specific situations (fitting, basically). How does that help in the general case or in new situations with current tecchnology (LLMs, concretely). Noatter how far you get, this is not a general or reliable solution. It van only simulate more generality or more reliability by training and tweaking. Nothing else. At least, with this paradigm.
This does not mean they will not be useful. What I challenge here is the AGI or singularity. We are far from that.
> I'm now team lead for 10 years and every single year I teach them the same thing over and over and over again.
I have been a lead and an architect also for years at different position. I think you miss how much tacit knowledge and judgement there is inside the brains of each of us that an LLM is not capable of. And if it is, then you have to dumo so much context that it is better to go do it yourself. There is a cost to that also actually. It is not just so "dry and technical" the knowledge. Maybe yes to learn Java patterns or C++ constructors or the like.
But not for "given this situation with all these specifics", which solution would you bet on? Probably the LLM will give you a shitty REST API that is not what u need at all.So u tell the AI. It gives u something else generati g 30-50% of "decorated code". Now it seems to workso you use it. Now you do this every day. Come back in 2 months. You generated a lot of fat.
Now you have a bug. You do not know even where to start. Thisis theprice to payfor speed, as usual: technical debt.
Now you tell me you put three agents to talk and burn 2000 usd in tokens. Great! Is the final solution better than what you would have achieved? Not sure at all.
TBH I am not into agents bc I do not trust a tool sniffing all my code and for copyright concerns. But I saw some and use a prompt with limited access and the best I can take out for my speed + control when coding is tech discussions to decide on it, error catching, test generation, one-off scripts... But never "make an app like this or that". If I ever do that (I did it a couple of times) is for scaffolding and later throw away 70%.
Namely, to see something that runs on screen quickly. But later you need to spend time yourself as usual. Not a bad thing, just that this is not what you deliver and need the work done. Iterations etc.
> Please elaborate. How? With which technique?
Reinforcement learning can just solve things even if they are new. It doesn't understand how a tool works? Give it a vm with the tool, a thousand agents and let it discover it automatically.
Use the thumbs up/down emoji + chat analysis when a customer is unhappy, feed that to a RL Loop.
The AI Researchers though work on World Models, grounding the AI and letting it simulate. It can do the simulation in parallel (unlimited) and choose what is best.
> I think you miss how much tacit knowledge and judgement there is inside the brains of each of us that an LLM is not capable of.
But thats my problem. Soooo many do not have this even as senior developers.
> Now you have a bug. You do not know even where to start. Thisis theprice to payfor speed, as usual: technical debt.
Yeah now i just ask the LLM to describe to me the bug. Works very well.
> Now you tell me you put three agents to talk and burn 2000 usd in tokens. Great! Is the final solution better than what you would have achieved? Not sure at all.
This is the thing. It only needs to make the team 10-30% better to compensate token budget with one work collegue. We have reached this level in my opinion already. Choosing a head count vs. choosing tokens.
But it becomes cheaper and easier and better. So you will not just be able to do ith with 3 agents but with 20, 50 or 100.
It will be better if your team is an avg team. It will be worse if you have a high profile team, for now. But man our industry has such a weird broad quality spectrum.
> TBH I am not into agents bc I do not trust a tool sniffing all my code and for copyright concerns
In worst case, my team always do code reviews, I do a code review on an ai instead of a human and adjust the harness or the infos the ai can access. I can actually work on making this workflow better and then i can clone it or spin it up for every single PR. For a human? I have to train them and they might leave.
But there are plenty of cases were code doesn't matter. Researchers write a lot of random shitty uggly code as long as it does what it does, it doesn't matter. I have scripts for small tasks, we have microservices which do one thing because it is a tech stack we only need for one use case.