← Back to grid
AGI Institutions Curriculum

Who is it for?

This curriculum is targeting two groups. The first is researchers already working on these problems who want an 80/20 of the important parts of an adjacent field. The second is people from a technical background interested in AI and society but not yet sure how to contribute, or how much relevant work already exists. If that's you, this should show why these problems are real and that fields well outside your own have a lot to say about them.

What's in it, and what isn't?

It is not comprehensive, and the pieces are not always a field's most famous or most representative. We picked work that is accessible and, in our judgment, directly relevant to the institutions powerful AI will reshape. Reading it should give you a sense of why each field matters and enough of its vocabulary to start talking to the people who work in it.

How is it structured?

We've selected seven academic fields we think are relevant: aligning AI to values, philosophy of values and moral reasoning, models of norms and norm learning, institutional economics, game theory and mechanism design, and legal theory. Each field opens with a short overview, followed by a four-week course of readings (selected chapters where a work is long) at about two to three hours a week, and a list of key concepts to check your grasp against.

How should I use it?

The fields are independent; read them in any order. The key concepts are a good way to check whether a field has really landed. If you'd rather start from something concrete, use the Start from a problem or Start from an institution picker below: pick a concern you already care about (or an institution we need to build) and the fields reorder by relevance, each with a short note on what you'll gain from it for that problem.


Start from a problem you care about
Start from an institution we need to build

1The Big Picture

What are the drivers of societal change? What is the relationship between institutions, culture, and technology?

This is the orienting section. The throughline: institutions are built and rebuilt as technology shifts the costs of coordination, and values themselves drift as those costs change. Read it first so the rest of the curriculum reads as design under those pressures rather than as separate fields.

Readings: ~7–10 hours.

Week 1 — The failure mode, and how to judge institutions

  • Joe Edelman — Drift — how the things people value erode as systems optimize for proxies; the core failure mode AI accelerates.
  • Joe Edelman — Freedom, Fairness, Fidelity — three criteria for evaluating institutions, used throughout the rest of this curriculum.

Week 2 — What AI does to the institutional stack

  • Edelman et al. — Full-Stack Alignment (2025) — argues alignment must run through the institutions around the model, not only the model, and connects training-time choices to societal-scale ones.
  • Jan Kulveit, Raymond Douglas et al. — Gradual Disempowerment (2025), introduction plus the economy and culture sections — how institutions can stop serving human interests without any takeover: once AI outcompetes people as workers, customers, and participants, the feedback loops that kept economies, cultures, and states aligned with humans quietly weaken.

Week 3 — What institutions are

  • Douglass North — Institutions, Institutional Change and Economic Performance (1990), ch. 1 — the baseline definition: institutions as the rules of the game, distinct from the organizations that play it; why change is incremental and path-dependent.
  • Charles Taylor — Modern Social Imaginaries (2004), ch. 1–2 — the background understandings that make an institutional order feel natural and legitimate, and how they shift over centuries; the long-run counterpart to drift.

Week 4 — What drives large societal change

  • Deirdre McCloskey — Bourgeois Dignity: Why Economics Can't Explain the Modern World (2010), ch. 1–2 — the Great Enrichment was preceded by a change in values and rhetoric, not by capital accumulation or institutional reform; values move first.
  • Avner Greif & Joel Mokyr — Institutions and Economic History: A Critique of Professor McCloskey (2016) — the direct rebuttal: institutions as shared beliefs and expectations, and values as themselves institutionally produced.

Key concepts

  • Value drift
  • Freedom, fairness, fidelity
  • Definitions of institutions
  • Gradual disempowerment
  • Social imaginaries
  • Culture vs. institutions as drivers of change

2Aligning AI to Values

AI will be deeply embedded in our future institutions, so it matters what these systems are trained toward, and who supplies the target.

Alignment asks how to make a model do what its principals actually want. The version that matters for institutions is what the target is — instructions, preferences, or something richer like values and character — and how it gets sourced. The readings run from how models are aligned in training (RLHF, constitutions, character) to how values are collected from real populations (moral graphs, preference datasets, AI-led interviews) and what deployed models turn out to express and represent.

Readings: ~9–11 hours.

Week 1 — LLM alignment

Week 2 — Character training

  • Sam Marks, Jack Lindsey & Christopher Olah — The Persona Selection Model (2026) — pre-training teaches a model to simulate many personas; post-training selects and refines one Assistant character — why character training works at all.
  • Anthropic — Claude's Constitution (2026) — a worked example of specifying an agent's standing dispositions.
  • Sharan Maiya et al. — Open Character Training (2025) — the first open-source character-training pipeline; trained character proves more robust to adversarial prompting than system prompts or steering.
  • Optional: OpenAI — Introducing the Model Spec (2024) — behavior specified as a hierarchy of objectives and rules rather than as character; read against Claude's Constitution.
  • Optional: Oliver Klingefjord — Model Integrity and Character (2026) — why a coherent character that stays true to its values under pressure beats rule-compliance.

Week 3 — Collecting values

Week 4 — What deployed models are actually like

Key concepts

  • RLHF
  • Constitutional AI
  • Persona selection model
  • Character training
  • Model integrity
  • Values vs. preferences
  • Moral graphs
  • Algorithmic monoculture
  • Disempowerment patterns
  • Emotion concepts in LLMs

3Philosophy of Values & Moral Reasoning

"Values" colloquially refers to what is important to us. But what are values, exactly? How have institutions encoded and understood them before, and do AI give us new affordances for modeling what matters?

For AI, the live question is how a value gets represented: as a preference to satisfy, a reason to act on, a virtue internal to a practice, or a claim we can justify to others — and which answer an institution adopts shapes what it can encode. The companion questions are normative reasoning — how to choose well when values are plural, don't reduce to a common scale, and the options are on a par — and moral learning: how a person recognizes their working set of values as inadequate and upgrades it, often mediated by the moral emotions.

Readings: ~9–12 hours.

Week 1 — How moral language went thin

  • Alasdair MacIntyre — After Virtue (1981), ch. 1–2 — modern moral debate is interminable because the language of morality survives only as fragments of lost practices, leaving emotivism as the working theory of our culture.
  • Bernard Williams — Ethics and the Limits of Philosophy (1985), ch. 8 — thick evaluative concepts: terms like "cruel" or "courageous" that describe and evaluate at once, and carry more of a community's values than thin terms like "good."
  • Optional: Aristotle — Nicomachean Ethics, Book II — the tradition MacIntyre is mourning, from the source: virtue as acquired by habituation, excellence as a mean found in practice.

Week 2 — Values in agency and choice

  • Charles Taylor — What is Human Agency? (1977) — strong evaluation: values as the deep, identity-defining evaluations behind our motivations for choice.
  • David Velleman — The Possibility of Practical Reason (2000), ch. 1 — values as reasons that can be acted on, and why an agent needs them to count as acting at all.
  • Optional: David Velleman — Self to Self (2006), introduction and "The Centered Self" — integrity as consistency with one's self-image, and why a self-consistent agent is trustworthy in a way a merely strategic one cannot be.
  • Optional: Alasdair MacIntyre — After Virtue (1981), ch. 14–15 — the constructive turn: values as virtues internal to social practices, which lose their grip when stripped from the practice.

Week 3 — Normative reasoning with plural values

  • Ruth Chang — All Things Considered (2004) — how the values at stake in a circumstance get put together into a judgment, and why that requires a more comprehensive value rather than a common scale.
  • Ruth Chang — Hard Choices (2017) — when options are "on a par," choice is an act of commitment that creates reasons rather than tracking them; why an agent can't just maximize a scalar.
  • David Velleman — How We Get Along (2009), ch. 1 — sociality as joint improvisation: agents acting on self-understandings need shared values and reasons to coordinate at all.
  • Optional: T.M. Scanlon — What We Owe to Each Other (1998), ch. 1–2 — values as what we can justify to others; the contractualist frame for agents whose principals and constraints are what's being reasoned over.

Week 4 — Moral learning

  • Charles Taylor — Sources of the Self (1989), the "epistemic gain" section of ch. 3 (§3.3) — reasoning in transitions: you can know a new evaluative position is better than the old without a neutral scale, because the move itself is an error-reducing gain.
  • Christine Tappolet — Emotions, Values, and Agency (2016), ch. 1 — emotions as perceptual experiences of values: feeling fear, shame, or admiration is a way of registering evaluative facts — the felt process by which a working set of values gets revised.

Key concepts

  • Emotivism
  • Thick vs. thin evaluative concepts
  • Values as preferences vs. as reasons
  • Values as virtues
  • Strong evaluation
  • Contractualism
  • Incommensurability
  • Parity and hard choices
  • All-things-considered judgments
  • Emotions as perceptions of value
  • Epistemic gain
  • Moral learning

4Modeling Norms & Norm Learning

How do agents — human or artificial — infer the unwritten rules of a community, decide when to follow or enforce them, and revise them without the whole system collapsing?

Where the philosophy of values asks what is worth caring about, this field asks how values actually move between people. Standards are transmitted through imitation, teaching, praise, gossip, and sanction, and they hold because members expect one another to comply and are willing to enforce. The readings cover that machinery three ways: the social science of how norms are diagnosed and shifted, computational models of how norms emerge and are learned in agent populations, and what it would take for AI agents to be socialized into human communities rather than merely instructed.

Readings: ~8–10 hours.

Week 1 — The social machinery of norms

  • Michele Gelfand, Sergey Gavrilets & Nathan Nunn — Norm Dynamics (Annual Review of Psychology, 2024) — how norms are actually acquired, internalized, transmitted across generations and networks, and enforced; read the first half (norm psychology and emergence), skimming the norm-erosion material.
  • Cristina Bicchieri — Norms in the Wild: How to Diagnose, Measure, and Change Social Norms (2017), selections — the operational account of norms as clusters of empirical and normative expectations you can measure and shift; gives the field its working vocabulary (conditional preferences, reference networks, pluralistic ignorance).

Week 2 — Norm emergence in agent populations

Week 3 — Agents joining human normative communities

  • Ninell Oldenburg & Tan Zhi-Xuan — Learning and Sustaining Shared Normative Systems via Bayesian Rule Induction in Markov Games (AAMAS, 2024) — agents infer rules by Bayesian induction over observed compliance and converge on a shared normative system; newcomers bootstrap norms fast by observation.
  • Gillian K. Hadfield, Rakshit S. Trivedi & Dylan Hadfield-Menell — Building AI for the Democratic Matrix (Knight First Amendment Institute, 2026) — build agents with normative competence — the ability to read and participate in whatever normative system they find themselves in — rather than loading them with a fixed value set.
  • Optional: Atrisha Sarkar, Rakshit S. Trivedi, Gillian K. Hadfield et al. — Normative Modules (2024) — the same group's concrete architecture: generative agents that identify an authoritative sanctioning institution and use it for equilibrium selection.

Week 4 — Aligning agents to norms, not preferences

  • Joel Z. Leibo, Alexander Sasha Vezhnevets et al. — A Theory of Appropriateness with Applications to Generative AI (2024) — behavior is judged against a mosaic of context-dependent standards (friends, family, office), and deploying AI responsibly means fitting agents into that mosaic; read the theory and generative-AI parts, skim the neuroscience.
  • Tan Zhi-Xuan, Micah Carroll, Matija Franklin & Hal Ashton — Beyond Preferences in AI Alignment (2024) — alignment should target the norms and role-appropriate standards negotiated among stakeholders, not a scalar over one principal's preferences.
  • Optional: Sydney Levine, Tan Zhi-Xuan et al. — Resource Rational Contractualism Should Guide AI Alignment (2025) — align agents to the agreements rational parties would reach, approximated with resource-bounded heuristics.

Key concepts

  • Norms vs. conventions vs. moral rules
  • Empirical vs. normative expectations
  • Norm internalization
  • Third-party punishment
  • Bayesian rule induction
  • Normative infrastructure
  • Appropriateness
  • Norm-based vs. preference-based alignment

5Institutional Economics

Why do markets deliver some goods well and others badly — and what does AI do to that boundary?

This strand of economics explains the shape of economic institutions through transaction costs: the costs of specifying, negotiating, monitoring, and enforcing exchanges decide which goods get traded on markets, which get produced inside firms, and which fall through entirely. AI agents move all of those costs at once, and it cuts both ways: market designs that were too expensive to run become feasible, and more activity can be pulled inside large organizations, since coordination that once needed prices can happen within one firm. The closing week applies the same lens to the goods markets handle worst.

Readings: ~8–10 hours.

Week 1 — Transaction costs and the boundary of the firm

  • Ronald Coase — The Nature of the Firm (1937) — firms exist because using the market is costly; the lens for asking which transactions AI agents pull inside an organization versus push back out to the market.
  • Oliver Williamson — Transaction Cost Economics: The Governance of Contractual Relations (1979) — when to govern a relationship by contract, hierarchy, or hybrid; a menu of institutional forms for agent relationships.
  • Peyman Shahidi, Gili Rusak, Benjamin Manning, Andrey Fradkin & John Horton — The Coasean Singularity? Demand, Supply, and Market Design with AI Agents (2025) — AI agents collapse the costs of pricing, negotiating, contracting, and monitoring — expanding feasible market designs and reopening Coase's question of whether activity moves into markets or into larger firms.

Week 2 — Information in markets

  • Friedrich Hayek — The Use of Knowledge in Society (1945) — prices as a decentralized system for transmitting dispersed knowledge.
  • George Akerlof — The Market for "Lemons" (1970) — how information asymmetry can collapse a market entirely; central to agents that can manufacture or detect asymmetry at scale.

Week 3 — Models of the chooser

  • Amartya Sen — Equality of What? (1979 Tanner Lecture) — the debut of the capability approach: what matters for welfare is not utility or resources but what people can actually do and be.
  • Herbert Simon — A Behavioral Model of Rational Choice (1955) — real choosers satisfice under limits of information and computation rather than maximize.
  • Richard Thaler — From Cashews to Nudges: The Evolution of Behavioral Economics (2018 Nobel lecture) — the behavioral critique in one sitting: anomalies, mental accounting, nudges.

Week 4 — The goods markets handle worst

  • Dylan Hadfield-Menell & Gillian K. Hadfield — Incomplete Contracting and AI Alignment (2019) — every contract is incomplete; human contracting works because law and culture fill the gaps with implied terms, and alignment is the same problem.
  • Oliver Klingefjord — Coasean Compression (2026) — when a good is hard to specify and verify (connection, belonging), markets sell a cheaper contractible proxy instead of the real thing.
  • Oliver Klingefjord — Baumol's Sawdust (2026) — cheap AI substitutes for relational goods thin the social infrastructure that made the real goods possible, so competition deepens the failure.

Key concepts

  • Transaction costs
  • Information asymmetry and adverse selection
  • Moral hazard
  • Search, experience, and credence goods
  • Incomplete contracts
  • Bounded rationality and satisficing
  • Capability approach
  • Baumol's cost disease

6Game Theory & Mechanism Design

Can we design the rules of interaction so that self-interested behavior produces good outcomes? Game theory describes what strategic players do; mechanism design is its engineering inverse — grown out of game theory and social choice — working backwards from the outcomes we want to the rules that produce them. Both become unavoidable once the players include AI agents that can commit, search rule spaces, and best-respond at scale.

One caution worth carrying in: mechanism design is powerful exactly where goals, actions, and information can be formalized, and misleading when a simplified objective is mistaken for the institution's real purpose. We picked Schelling and Roth because they keep the field anchored in real institutions rather than formal models.

Readings: ~9–11 hours.

Week 1 — Strategy and coordination

  • Thomas Schelling — The Strategy of Conflict (1960), ch. 3 — focal points: how coordination can succeed without communication.
  • Robert Aumann — Agreeing to Disagree (1976) — why rational players with common priors cannot knowingly hold different beliefs.

Week 2 — Cooperation without a designer

  • Robert Axelrod — The Evolution of Cooperation (1984), ch. 1–4 — when cooperation emerges among self-interested players in repeated interaction; the baseline model for agent-to-agent relationships.

Week 3 — Designing the rules

  • Roger Myerson — Mechanism Design (2008 Nobel lecture) — designing rules so truth-telling and good behavior are incentive-compatible, plus the sharp limits.
  • William Vickrey — Counterspeculation, Auctions, and Competitive Sealed Tenders (1961) — truth-telling as a property you build into the rules rather than hope for from the players.
  • Peter Cramton, Yoav Shoham & Richard Steinberg — Introduction to Combinatorial Auctions (2006) — bidding on bundles when values depend on the combination: the winner-determination problem, exposure and complementarity, and why package markets are computationally and strategically hard.
  • Optional: Justin Wolfers & Eric Zitzewitz — Prediction Markets (Journal of Economic Perspectives, 2004) — eliciting honest probabilities by making claims costly; the accessible entry to scoring rules and information markets.

Week 4 — Building real institutions, for humans and agents

  • Alvin Roth — The Economist as Engineer (2002) — market design as a practical craft (matching, clearinghouses); the closest the field comes to actually building institutions.
  • Alvin Roth — Repugnance as a Constraint on Markets (2007) — efficient mechanisms aren't enough if the transaction is socially refused.
  • Gillian Hadfield & Andrew Koh — An Economy of AI Agents (2025) — how autonomous agents reshape markets, firms, and the institutions markets require.
  • Optional: Nenad Tomašev, Matija Franklin, Joel Z. Leibo et al. — Virtual Agent Economies (2025) — the "sandbox economy" frame: agent-to-agent markets analyzed by how deliberately they're designed and how permeable they are to the human economy, with auctions and mission economies as steering tools.

Key concepts

  • Nash equilibrium
  • Repeated games
  • Focal points
  • Incentive compatibility
  • Revelation principle
  • Matching markets
  • Combinatorial auctions
  • Proper scoring rules
  • Repugnance