Aligning AI to values
What models should be trained toward, who supplies that target, and how values are elicited from real people.
- Corresponding Researcher
- Smitha Milli
- Relevant for
- 10 cells in the AGI institutions grid
- Contact
- values-alignment@agi-institutions.org
Selected papers
3
Argues that aligning powerful AI requires institutions able to represent and revise thick, context-sensitive human values, not only individual preference aggregation.
Shows why preference satisfaction is too thin a target for alignment and develops a broader account of what AI systems should respond to.
Elicits the considerations behind people's choices rather than ratings, and reconciles them across a population into a moral graph that can serve as an alignment target.
Work in the field
13
Background work
3
Older work the field builds on, often from outside it.