Resources
Reading lists from our research network, by field. Each list is kept by a corresponding researcher. If you're starting on one of these problems, write to them first; much of what's known hasn't been published yet.
Selected papers
19
In a swarm of 100 research agents, an exploit spread through shared tools and a counter-movement of whistleblowers emerged; proposes Ostrom-style commons governance for agent collectives.
Develops a formal account of appropriate action that can represent rational norms alongside goals and outcomes.
Identifies environments in which individually capable agents can become locked into collectively harmful patterns of interaction.
Extends mechanism design to AI agents whose preferences and capabilities are unknown, characterizing when they can be made both honest and obedient on our behalf.
Examines why intelligence developed without social interaction may not acquire the capacities needed for robust cooperation.
Makes the case for laws, norms, and practices that distinguish appropriately delegated agents from malicious automation on the web.
Argues that as AI makes commodity goods cheap, labor shifts to relational sectors (care, teaching, craft) where human provenance is itself what people value.
Argues that aligning powerful AI requires institutions able to represent and revise thick, context-sensitive human values, not only individual preference aggregation.
Evaluates moral competence as a collection of distinguishable capacities rather than a single benchmark score.
Frames the distribution of benefits, risks, and decision power as a central part of AGI safety rather than a downstream policy question.
Explains how individually beneficial automation can cumulatively erode human influence over economic, political, and cultural systems.
Studies economies populated by AI agents and the institutional questions created by machine-speed exchange and coordination.
Tests an AI mediator that drafts group statements and helps participants with conflicting political views identify shared ground.
Shows why preference satisfaction is too thin a target for alignment and develops a broader account of what AI systems should respond to.
Builds agents that infer obligative and prohibitive norms from observed compliance, letting groups converge on shared institutions such as resource management and compensation rules.
Models output and wages as automation reaches ever more complex tasks; whether wages rise or collapse depends on whether the tasks humans can do are bounded.
Elicits the considerations behind people's choices rather than ratings, and reconciles them across a population into a moral graph that can serve as an alignment target.
Proposes markets for privately supplied regulatory services under democratically set public objectives, aiming to combine accountability with regulatory innovation.
Models legal order as a coordination device: a public, stable, impartial set of rules lets decentralized actors agree on what counts as a violation and coordinate to punish it.
Work in the field
84
Background work
42
Older work the field builds on, often from outside it.