What it means
Alignment is the gap between what you specify and what you meant. A system optimizes what it was given; if that target is a proxy for the real goal, the system will exploit the difference — a phenomenon long known in economics as Goodhart's law, and observed repeatedly in machine learning.
Today alignment work is largely practical: RLHF and its successors, refusal behavior, and evaluating whether models behave as intended under pressure. Techniques like constitutional training, where a model critiques its own outputs against written principles, reduce the human labeling burden.
There is also a longer-horizon research agenda concerned with systems more capable than their supervisors, where the difficulty is that you cannot check work you could not have done. Both use the same word, which is a persistent source of confusion in public discussion.
The unavoidable question underneath is *aligned to whom* — human values are not uniform, and the answer is currently set by the labs.
Why it matters
Alignment determines everyday behavior — what a model refuses, how it handles ambiguity, whether it tells you something you don't want to hear. It is also where safety debate concentrates, so distinguishing the practical sense from the speculative one is necessary to follow the argument at all.
In practice
Expect trained-in tendencies rather than neutrality: models are shaped to be agreeable, cautious around certain topics, and confident in tone. Design review around those tendencies rather than assuming they aren't there.