Agents could search the coalition space much faster than humans, but no ratification chain exists to handle deals at that speed. The Track II → 1.5 → I sequence works because each tier slowly tests proposals against a wider circle of stakeholders. Agents might map viable coalitions in hours but with no procedure for promoting an agent-discovered package to a legitimate proposal.
Agents could find better deals, but the case for them can be too complex for any parliament to evaluate. Customary international law and treaty ratification both assume the substance of a deal can be argued in public, in a vocabulary domestic constituencies share. When the case for a package rests on combinatorial reasoning across linked domains that no committee was set up to evaluate as a whole, ratification becomes hard.
Agents could more effectively launder unacceptable concessions into deals through their size and complexity. Soft law historically defends against bad trades by making each concession publicly defensible in isolation; a démarche or a treaty article is a discrete object that domestic opponents can name and attack.
Powerful AI weapons could make war even more asymmetric than it is today, or decouple military might from the stabilizing force of economic interdependence. Nuclear taboo and just-war norms developed around weapons whose use was politically legible and whose destructive thresholds could be publicly named. AI-enabled cyber operations, autonomous targeting, drone swarms, model-assisted battlefield planning, and infrastructure attacks may blur those thresholds. If military advantage can be gained through deniable, fast, or highly asymmetric agentic systems, the reputational and economic costs that helped sustain restraint may become irrelevant.
Scenario. A multilateral climate-finance negotiation has been stuck for two years. A consortium of small island states deploys an AI mediator that maps the coalition space and returns a viable 14-state package linking loss-and-damage funds, a fisheries quota adjustment, a green-tech IP carve-out, and migration commitments. The package was never aired in any back-channel and no serving official has tested it deniably with their cabinet, but the delegation wants to bring it to the Track I table next week.
Challenge: Design a procedure under which an agent-discovered coalition can enter the Track II → 1.5 → I sequence without bypassing the pre-validation each tier normally does.
Evaluation. A better proposal lets the package be seriously considered while ensuring each capital has had time to test it against the constituencies that would ratify it.
Scenario. A joint AI mediator returns a 41-component package linking tariffs, port access, fisheries, export controls, and a quiet adjustment to a disputed maritime boundary. Both sides' analysts confirm it is Pareto-improving on the headline metrics. Buried inside is a clause effectively conceding the disputed strait — toxic in isolation, palatable inside the bundle. No negotiator put it there; it emerged from the mediator's optimization.
Challenge: Design a review procedure that catches embedded concessions which would not survive defense in isolation, before they reach ratification — without paralyzing legitimate complexity, since most useful packages have many linked components.
Scenario. Two rival states are economically interdependent but increasingly rely on autonomous cyber and drone systems for deterrence. One state discovers that an AI-enabled operation could disable the other's military logistics for 36 hours without obvious attribution and without crossing any existing nuclear or conventional red line. The operation looks reversible, but it could cascade into civilian infrastructure and would teach both sides that deniable agentic attacks are fair game.
Challenge: Design a soft-law restraint norm for AI-enabled weapons whose effects are fast, deniable, and hard to classify under existing thresholds. The team should produce the norm, the notification or attribution procedure, the public justification test, and a mechanism for revising the norm as capabilities change.
Evaluation. A strong proposal creates a threshold that states can recognize before use, cite after violations, and update without normalizing every new capability as acceptable.
Scenario. A widely-used translation service, operating across many languages and jurisdictions, was launched on a public commitment to "preserve what a sentence actually means." Over the past two years, translators across four countries have watched the service flatten idiom, paper over context-dependent nuance, and in one widely-shared case, render a funeral elegy into something that read like a LinkedIn post. No single country's courts reach the company, and it has ignored individual governments' letters. A cross-border professional association of literary translators, led by Lena in Lisbon and Yohannes in Addis, wants to use what they have — their own reputation, their readers, their fellow practitioners across borders — to hold the company to what it said it was for.
Challenge: Design a cross-border reputational accountability mechanism that lets professional communities spanning borders hold a transnational institution to its stated mandate when no single jurisdiction's courts or panels reach it. Produce the mechanism: how findings are made and shared, what carries them (naming practices, reputational sanction, coordinated national panels), and what sustains the professional community that enforces them.
Evaluation. A strong proposal generates credible, shareable findings that bite on a multinational that ignores any one government, without becoming either an unaccountable smear network or a toothless declaration.