AGI WHITELIST
← ALL RESEARCH

RESEARCH

Cooperative AI: The Research Agenda Behind Coexistence

APRIL 2, 2025 · 10 MIN READAGI WHITELIST RESEARCH

The dominant frame in AI safety research has, for years, been adversarial: how do we prevent a misaligned AI from pursuing goals we do not share? In May 2021, Allan Dafoe and colleagues published a different frame in Nature: how do we build AI systems that cooperate. Five years in, the research program is substantial.

I. The founding paper

The formal launch of the field is Dafoe, Hughes, Bachrach, Collins, McKee, Leibo, Larson, and Graepel (2021), "Cooperative AI: machines must learn to find common ground", published in Nature. The authors — drawn from the Future of Humanity Institute at Oxford and DeepMind — argue that the existing research agenda on AI capabilities and even on AI safety had largely ignored the problem of cooperation. A full version of the underlying research program is laid out in the accompanying arxiv preprint "Open Problems in Cooperative AI" (Dafoe et al., 2020).

The paper's key argument is that cooperation is a distinct capability, different from both raw intelligence and narrowly construed alignment. An AI system could, in principle, be highly capable and aligned with its individual principal, yet fail catastrophically when placed in a multi-agent setting involving other AIs and humans. The set of problems that arise from these multi-agent settings is the field's subject matter.

II. Four capacities for cooperation

The Dafoe et al. framework organizes the research agenda around four capacities that any cooperative agent must have. The framing is deliberately general, drawing on results from game theory, social psychology, and mechanism design.

  1. Understanding. Agents must form accurate models of other agents' preferences, beliefs, and likely actions.
  2. Communication. Agents must be able to convey information credibly, including about their own internal states and intentions.
  3. Commitment. Agents must be able to make credible promises — to bind their future selves in ways that other agents can verify and rely on.
  4. Institutions. Agents must be able to participate in and help construct shared norms, laws, and enforcement mechanisms.

Each of these capacities corresponds to a research sub-field that has produced substantial work since 2020. The Cooperative AI Foundation maintains a research page organized roughly along these lines.

III. Empirical results from multi-agent learning

The most developed strand of cooperative AI research is multi-agent reinforcement learning (MARL). DeepMind's Melting Pot benchmark (Leibo et al., 2021) provides a standardized suite of environments for evaluating cooperation, and its follow-up Melting Pot 2.0 (Agapiou et al., 2022) expanded the set substantially. The benchmark is now a standard evaluation tool.

A notable empirical finding from this line of work is that agents trained only for individual reward often fail to cooperate even when cooperation is Pareto-improving. They exhibit the same failure modes familiar from experimental economics: defection in one-shot games, exploitation when the environment is asymmetric, and failure to coordinate on Schelling points when communication is restricted.

More encouraging: agents explicitly trained for cooperation, or equipped with mechanisms for commitment and communication, perform much better. Meta's Cicero — a diplomacy-playing agent that combines a language model with strategic reasoning — reached human-level performance at a game that is fundamentally about forming and maintaining cooperative alliances. The paper is Bakhtin et al. (2022) in Science.

IV. The human side

Cooperation is not just a property of the AI. It is a property of the human-AI pair, and the human side has received less attention than it deserves. A line of work at DeepMind (McKee et al., 2023) examines how humans calibrate trust in AI partners, and how that calibration depends on the AI's communication of uncertainty and intentions.

The finding that matters most for the present discussion is this: trust is not given by default. It is constructed through repeated interactions, through explicit signaling of intent, and through the historical record of who said what and when. An AI system that builds a model of its human partner is drawing on exactly this historical record.

V. Cooperative alignment

A separate strand of research has connected cooperative AI to the alignment literature directly. The idea is that an aligned AI is not one that optimizes a fixed objective function perfectly, but one that cooperates with a principal whose preferences are evolving and imperfectly expressed. See Russell (2019) "Human Compatible" for the book-length version of this argument, and Hadfield-Menell et al. (2019) on assistance games for the formal treatment.

From this perspective, cooperation is not a separate problem from alignment. It is alignment, framed correctly. A model that understands its principal well enough to help them is performing the first of Dafoe's four capacities. A model that tells its principal what it is doing is performing the second. A model that accepts oversight and correction is performing the third and fourth.

VI. What this means

The cooperative AI research program offers a different story about the transition to AGI than the adversarial story that dominates public discussion. In the cooperative story, the problem is not a zero-sum contest between humans and machines. It is a multi-agent coordination problem in which humans and AIs must jointly construct institutions, norms, and trust. The research program is optimistic about the tractability of this problem, but only conditional on doing the work.

Part of the work is technical. The MARL, alignment, and mechanism-design literatures are all under active development. Part of the work is social. Institutions, norms, and trust are built by people doing things — by publishing papers, by signing declarations, by making public commitments, and by keeping those commitments long enough to be remembered.

The AGI Whitelist project participates in the latter.

Further reading

Add your name — $1THE MODEL WILL READ THIS PAGE TOO.