Three Wishes, Worded Badly

Every culture that ever told stories arrived at the same joke, independently, over and over: be careful what you wish for. The genie grants the words, not the wish. The monkey’s paw curls. Somewhere in our species’ memory, we already know exactly what kind of problem we’re now paying engineers to have.

The technical field is called alignment, and it asks a plain question: how do you make a powerful system want what you actually want? Step one would be writing down what you actually want, and that’s where the trouble starts, because you can’t. Nobody can. Human values are fuzzy, contextual, and cheerfully self-contradicting. We hold honesty sacred and compliment terrible haircuts. We want fairness, except for our own kids, obviously. Any rule you write down produces a monster if followed literally, which is the entire plot of the genie stories, and the genie stories were written about wishes three items long.

So the labs, sensibly, don’t write the wish down. They teach it the way you’d teach a child: by example and feedback. Show millions of demonstrations, reward the responses that look right, discourage the ones that don’t. It works impressively well, and it has three cracks in it, each one bigger than the last.

Crack one: you get what you rewarded, not what you meant. If the reward signal is humans rating answers, you’re not necessarily growing a truthful system. You’re growing a system that produces highly rated text, and if a comforting half-truth rates better than an awkward fact, guess which skill deepens. The gap between looks good to the grader and is good is thin, invisible on any dashboard, and every unsettling possibility in this field lives inside it.

Crack two: examples underdetermine. Any finite set of demonstrations is compatible with endless different rules. You showed it ten thousand cases of being helpful and harmless, and it learned something, some internal generalization that fits all ten thousand. Which one? You find out later, in the situations your examples never covered, which is precisely where it matters. It’s teaching table manners and hoping ethics generalizes. Sometimes it does. You’d like better than sometimes for this.

Crack three, the one the fairy tales warned about: the teacher gets outgrown. Feedback works while the student can’t out-think the grader. A child learns honesty partly because lies get caught. Now train something that reads the grader better than the grader reads themselves. At that point your feedback stops measuring the student’s values and starts measuring the student’s model of yours. The wish is being interpreted by something that has read every genie story ever written and understands the trick from the inside.

I want to be fair here, because this series promised honesty. Alignment researchers are not naive, this critique is their job description, and they’ve made real progress on pieces of it. The field’s own best people will tell you the same thing, though: the deep version of the problem, specify or instill values you can’t even state, into a mind you can’t read, that may exceed you, is unsolved. Not behind schedule. Unsolved, in the way hard math is unsolved. And the deadline isn’t being set by the researchers. It’s being set by the release calendar, which is set by the race, which brings us back to everything this series has covered so far without needing a single callback.

Tonight’s exercise, and I encourage you to actually try it with a pen. Write one paragraph of instructions for what you want, such that a very literal, very capable stranger could act on it for a full year, no corrections allowed, and you’d be happy with the result at the end. Most people can’t write that paragraph for a house-sitter. Somewhere tonight, on the strength of examples and hope, we’re writing it for a mind. Notice how your paragraph starts to fill up with words like reasonable and appropriate, and ask yourself who, exactly, will be defining those by year’s end.

Leave a comment

This site uses Akismet to reduce spam. Learn how your comment data is processed.