I gave four coding agents $100 budget to build a PDF editor
blog.nielstron.de
[hidden]
[hidden]
My experience is agents really need cheap easily runnable third party ways to validate their work and stay on track. This usually means compilers and test suites but it can also mean visual checks (like all those vibecoded game ports flooding feeds right now). Of course, these things are useful for humans too. Tightening the feedback loop is always good. But for agents, the lack of feedback can be catastrophic. Better models can ward off catastrophe for longer but eventually without enough feedback, the errors not only accumulate but increase the chance of future errors.
To the points in the conclusion, humans still have the edge on dealing with vagueness and lack of clarity in goals. And we fill in the details of our goals in real time. In other words, the act of writing code feeds back into the other parts of building a program. Agents aren't like this, if they can see where they're going they can get there with remarkable tenacity but without it they'll get somewhere goal shaped fast. You can use agents to take big steps for you, but you're going to lose some fine touch.
nielstron[hidden]