Ephemeral Testing
lemire.me
[hidden]
The big challenge with designing an API is that the only way to be confident in the design is to build a bunch of different things on top of it. But why invest all that effort in an API that you don't think is ready yet?
With coding agents the cost of building those prototypes drops to almost nothing. I can exercise a proposed API design five different ways before I commit to the shape.
[4 comments hidden]
[2 comments hidden]
[hidden]
[hidden]
- We haven't really solved software engineering yet, it's hard to tell upfront whether a design adapts well to future needs. The community has a collection of heuristics which are sometimes at odds with each other.
- Therefore: write all designs, each simulated against a random tree of future needs.
- Pick whichever minimises the average future diff.
In other words, monte carlo simulation for software design. This scheme obviously assumes the price of code production goes even closer to zero.
[hidden]
[hidden]
I was designing a DSL to help make a colleague’s work easier, and I tested it by having a coding agent re-implement some of their notebooks using a draft of the DSL. The LLM output was a decent enough approximation of “typical” use, and it uncovered some warts I didn’t catch by testing it myself. As its designer, I simply wouldn’t have thought to try using it in some of the ways the LLM-generated code did.
[2 comments hidden]
[hidden]
IMO, It's the same idea with REPL-driven development in a functional language, why write unit tests if an interface is impossible to behave differently over time without changes?
Typically someone then argues, well, what if you change the implementation. To that person I will point out that the change will be developed in a REPL just like the original version, and thus be tested when it hits. Oh well...
To some, QA means a lot of compute and green outputs, to others, QA means spending quality hammock time before hitting the REPL :)
[5 comments hidden]
[4 comments hidden]
[3 comments hidden]
The actual biggest improvement is reducing the number of review rounds a PR has to go through. We were having too many rounds of drip-feeding new findings, because the AI reviewer is a subjective grader. Even if it found the same finding before, the next time it runs it might score it wildly differently. And it tends to want to score its findings across the whole grading curve, because it thinks that's more correct-looking. Our answer is to use previous rounds scores as anchors for the next round. Since we can't just use the same context and still have good performance, we needed to condense the score report into a small json we send forward.
[hidden]
[hidden]
[hidden]
[2 comments hidden]
Gets annoying pretty quickly
[hidden]
[hidden]
[hidden]
[hidden]
that gut check of 'does this feel good to use' is funny to hear as QA because most of the feedback I give that boils down to 'this sucks to use' nearly universally gets met with 'but the AC! SLO! SLA! it's designed this way on purpose! it's a feature not a bug' etc.
I mean, I'm glad SWEs are finally discovering that using the product of your own work is actually a net good but it's a little like seeing a bullet coffee evangelist author a blog post about how he's found the real secret to preventing congestive heart failure and it's to not drink butter regularly
[hidden]
[hidden]
[hidden]
Eh. I think you're just going to end up with slop, or sloppy recommendations?
My experience is that you can make different trade-offs for different reasons. I think even asking for the best answer to "improve the code, make better trade-offs".. even if you got a perfect response, there's no reason to think that it's the same set of trade-offs your actual use cases would benefit from.
[hidden]
Sigh.
Write a C or C++ API that works with pointers. Make it handle null pointers and errors elegantly so that the API user can safely chain calls and only check the final result. Claude decides it's better to be safe than sorry and peppers its code with intermediate nullptr and return value checks anyway.
kqr[hidden]
Richard Gabriel wrote something that has really stuck with me:
> Abstractions must be carefully and expertly designed, especially when reuse or compression is intended. However, because abstractions are designed in a particular context and for a particular purpose, it is hard to design them while anticipating all purposes and forgetting all purposes, which is the hallmark of the well-designed abstractions.
This is one of my favourite quotes on abstraction, because “anticipating all purposes and forgetting all purposes” is such a good summary of what goes into abstraction design.
A language model in an agentic harness cannot (yet) do this at the same level as a good software engineer, but the advantage they have is speed of token generation, so they can actually build the things the engineer tries to imagine, and verifying those is easier. Very cool!