The distance between can and should
Short thought on deciding what is worth building
May 15, 2025
philosophy
technology
design
Five years ago I wrote about how, as a kid, I imagined professional developers working in programs so sophisticated that they could just think about how the software should behave and the tool would fill in the blanks with the right code. I called myself naive, and I wrote that you have to write that code yourself and that it’s gonna be like that for a while.
Well, “a while” turned out to be about five years. Today you can describe what you want and a model will write a big part of the code for you, and the same goes for images, text and plenty of other stuff. So the part I thought was the hard one got much easier, and I keep running into the next one: what should I ask for?
A more capable tool gives me more possible answers, but it doesn’t tell me which of them deserves to exist. Aristotle already separated these two in Book VI of the Nicomachean Ethics: knowing how to make something is one kind of reasoning, and judging what is good to do is another. He lived in a very different society, but I think the split still holds. Being good at building doesn’t pick the thing worth building.
Software shows this nicely. I can measure whether a page loads, whether a model returns an answer, whether somebody clicks. Which of those numbers should matter is a decision that somebody has to make, and a system can get excellent at hitting a target that was badly chosen.
One example: a neighborhood needs a place where people can meet. You could build an app. You could also rent a room with reliable opening hours, or pay someone to keep an existing place welcoming. Only the app looks like a tech product, and that says nothing about which option solves the problem.
With generative tools it’s even easier to skip the question. I can generate fifty versions of something in an afternoon, and if nobody agreed on what a good outcome looks like, I end up with fifty polished versions of the wrong thing.
Richard Sutton argues in The Bitter Lesson that methods which scale with search and learning end up beating the ones where researchers build in their own knowledge of the problem. I agree with him, and I read it as a statement about means. If a system discovers a better way to maximize an engagement metric, I still have to defend the decision to use that metric.
My own checklist is short. The thing should give somebody an ability they didn’t have before, it should answer a need that somebody really has, and I should be able to defend it in front of the people who carry its costs. These can go against each other (a useful service can also make its users dependent on it), so no benchmark will settle it for me.
So before I ask a model for more options, I’m trying to get better at explaining why one of them is worth doing at all.