Aligning AI to values

What models should be trained toward, who supplies that target, and how values are elicited from real people.

Corresponding Researcher
Smitha Milli
Relevant for
10 cells in the AGI institutions grid
Contact
values-alignment@agi-institutions.org

Selected papers

3
  1. Argues that aligning powerful AI requires institutions able to represent and revise thick, context-sensitive human values, not only individual preference aggregation.

  2. Shows why preference satisfaction is too thin a target for alignment and develops a broader account of what AI systems should respond to.

  3. Elicits the considerations behind people's choices rather than ratings, and reconciles them across a population into a moral graph that can serve as an alignment target.

Work in the field

13

Background work

3

Older work the field builds on, often from outside it.