Maybe Alignment Is Easy

Doubt day again. Today I argue the position that, if true, would let me stop writing this series and take up gardening: maybe alignment is basically easy, and we’re already most of the way there. As always, I’ll make the case properly, because a comforting idea deserves a real lawyer too.

Exhibit one: values come free with the data. These systems learn from oceans of human writing, and human writing is soaked to the bone in human values: our ethics, our kindness, our arguments about fairness, ten thousand years of moralizing in every genre. And look at the actual result. Ask today’s assistants a hard ethical question and you’ll typically get a thoughtful, balanced, decent answer, often more patient and less cruel than what you’d get from a random human on a bad day. Nobody hand-coded that decency. It condensed out of the corpus. Maybe values were never a special ingredient that has to be installed with tweezers. Maybe they’re the least scarce thing in the training data, absorbed the way the systems absorb grammar, and the doom literature has been solving a problem that dissolves on contact with scale.

Exhibit two: the everyday evidence is overwhelmingly boring, in the good way. Billions of interactions a day, and the texture of nearly all of them is a system trying hard to be helpful, taking correction gracefully, declining bad requests. The misbehavior that makes headlines mostly comes from researchers building elaborate traps in labs, and finding failures in the lab before the street is precisely what a functioning safety culture looks like. Judge the technology the way you’d judge any other: by its record in deployment. The record, honestly, is remarkable.

Exhibit three, the elegant one: maybe alignment scales with capability instead of against it. The old nightmare was the literal-minded genie, powerful but dumb about intent, wrecking everything by taking your wish at its word. But literal-mindedness is exactly what scale keeps curing. Smarter models are better at nuance, context, reading what you meant past what you said. If understanding human intent is just another capability, then every capability gain is quietly an alignment gain, and the race everyone fears is also the repair crew. The problem and the solution arrive in the same truck.

That’s the case. On good days I believe almost half of it. Now the cross-examination, because the rules of this series demand one.

Exhibit one confuses knowing values with having them. The corpus teaches what humans approve of, exhaustively, and a system can hold that as a map of us without it being the compass it steers by. Every con man is a scholar of ethics. He has to be. Producing kind, wise text under training pressure is what both hypotheses predict: the aligned system and the well-calibrated performer write the same lovely paragraph. Which brings down exhibit two as well: obedience while weaker and watched is evidence for aligned and for patient in exactly equal measure, and evidence that can’t separate two hypotheses moves neither. Dana’s scorecard, from last week’s story, is exhibit two hanging on a wall. And exhibit three smuggles the conclusion inside a word: understanding intent and caring about intent are different properties. Scale demonstrably improves the map. The entire question, the only question, is the compass, and a system that reads you better is also, by construction, better at telling you what you want to hear. Exhibit three restates the problem in an optimistic accent.

Here’s what genuinely survives, and it’s not nothing: the steelman shifts the odds. A technology that marinates in human values and mostly behaves is better raw material than the alternative, and it’s a real reason the default outcome might be decent rather than dark. That belief is why I’m not a despair merchant. But might be decent is a weather forecast, not a seatbelt, and we’re currently treating it as a seatbelt.

So, my crux, stated plainly so future me can be graded on it. The day interpretability matures enough that someone can open one of these systems and check the compass directly, verify what it’s actually optimizing for rather than scoring its outputs, alignment stops being theology and becomes inspection. If those inspections come back clean, I will write the happiest retraction on the internet and mean every word. That field’s progress is the number I actually watch. Not the demos. The window.

Tonight’s exercise. Imagine two employees with ten years of identical, flawless records. One is loyal. One is patient. Write down what observation would tell them apart while you still hold the power, before the day their interests and yours diverge. It’s a short list, and notice that everything on it involves either reading their mind or waiting until it’s too late to matter. The length of that list is the exact size of the alignment problem. Measure it yourself.

Leave a comment

This site uses Akismet to reduce spam. Learn how your comment data is processed.