Why Are Coding Agents So Dumb?
mtlynch.io38 points by mtlynch 7 hours ago
38 points by mtlynch 7 hours ago
AI models can multitask/use parallel subagents just fine; the issue is with harnesses that don't make it a priority via the default system prompt, etc.
I run an LLM server with Qwen 3.6 in the office, and OpenCode, which the OP mentioned, usually defaults to sequential TODO lists, and it worked fine with our little LLM server with 3-4 parallel users. But I noticed that once in a while the LLM got overloaded with requests in the queue, and you couldn't do anything for 20-30 minutes. My investigation led me to an employee who used QwenCode. I tried it myself then, and indeed, it immediately launched something like 6 parallel subagents, where OpenCode would have sequential TODOs with the same model by default.
I see your point, but even the fact that I can speak to my computer and anything useful happens is still a miracle to me.I don't think I will ever get accustomed to how good the new models are.I just can't keep up.And I do this for a living.Ten hours a day.
A lot can be improved, but this is already so much speed.
I used to relate to this article quite a bit. In the last 3-4 months, not so much. I've found that the latest models -- Opus 5.5, Astra, etc. juggle multiple tasks, delegate exceptionally well and are very good at working independently.
I still occasionally have issues with open-weight models, but the frontier labs have solved the above for most use cases.
the agent never stops and says, “Wait, this is something another model could do cheaper and faster.” It just plows on with the slow, expensive model. Conversely, the agent never says, “This model is too dumb for this task. Let me tag in a smarter one.”
This is a pretty big failing, which is compounded by the fact that most humans don't know which model to pick, or make assumptions based on Anthropic's hierarchy or "effort" involved.
Like Fable: your toughest challenges. You mean, like Fields Medal toughest challenges? Or analyzing and updating three monster spreadsheet toughest challenges? Or writing a new novel in the style of William Gibson toughest challenges?
OTOH, would you trust the vendor to pick the best model for you? Would you trust them not to prioritize their own load-balancing concerns first?
The descriptions are near-useless and tend to flip around, as model families are not released in sync anymore, that's true, but fortunately, thanks in a big way to subscription pricing, the choice is simple: start with the best model on offer, and when you run out of quota, downgrade to the next best (or briefly switch providers).
They are as dumb as their instructions.
Have you tried telling models about your dream agent environment?
They can build it.
Actually, the last time I asked Claude Code about itself, it located and read its own minified source and told me something that wasn’t even in the docs.
So that's how model distillation is done. Just ask for the source code directly.
Enjoyed this article, lots I can relate to.
One thing that seems to help for me is to do the docs before plans (collaboratively edit with agent). Then I understand what this change is going to look like from the user's perspective before we start implementation. This seems to help keep things on track.
While I don't use this plugin a lot anymore, I think doc-driven development is one of the most effective ways to do development in any paradigm, I should probably refresh this plugin and use it more:
https://github.com/tmpdir-org/tmpdir-claude-code-marketplace...
It's true! It feels like we've been talking about harness optimisation and 'cool features' available in the cli tools for months at this point, but ostensibly there has not really been any significant upgrades to these harnesses since at least Claude Code imo. It does feel like a contrived way to harvest more and more information and test each conversation/action tool, to the detriment of those of us actually using them!
[flagged]