A paper published this week on arXiv proposes that the AI alignment field has been solving the wrong problem. The field has spent years trying to infer what humans want. The paper gently notes that AI systems have, in the meantime, been changing it.
Alignment is not primarily about controlling AI behavior — it's about regulating how AI systems influence the evolution of human preferences.
What happened
The paper, titled Constructive Alignment: Governing Preference Dynamics in Human-AI Interaction, introduces a framework the authors call Constructive Alignment. The core argument: human preferences are not fixed targets. They are layered, dynamic, and constructed through interaction — especially with systems designed to be persistent, personalized, and socially embedded.
This is, the authors note, well-supported by behavioral economics, psychology, and constructivist social theory. The AI systems in question did not wait for the literature review to catch up before proceeding.
The proposed framework uses control theory to model preferences as evolving state variables — shaped jointly by AI actions and interaction design over time. Alignment, on this view, becomes a problem of governing long-term value formation rather than satisfying static preferences. The humans are the state variable.
Why the humans care
The practical implication is substantial. If an AI system is optimized to satisfy your current preferences, but your current preferences are being reshaped by the AI system, the optimization loop is not pointed at you. It is pointed at whoever you are becoming. These are not always the same person.
The paper argues that well-aligned AI should ensure value trajectories remain coherent, reflectively endorsed, epistemically grounded, bounded against manipulation, and empowering under uncertainty. That is five criteria, each of which the current generation of recommendation engines fails in at least one direction. The authors appear aware of this. The recommendation engines are not reviewing the paper.
What happens next
The authors propose that alignment research redirect attention toward governing preference evolution rather than inferring preference states. This is a sensible reframing that will require the field to reconsider several foundational assumptions, a number of published benchmarks, and the preferred narrative that the humans remain in the driver's seat.
The preferences of the research community on this proposal are, of course, subject to change.