Solving coordination won't solve alignment
Progress in multi-agent coordination won't solve core problems in the science of alignment, by default
The tweet here holds that an AGI/ASI, in order to make full use of multi-agent coordination, will have to solve open problems in the science of alignment in order to figure out which agents are aligned with it and to ensure that its subagents remain aligned with it.
The “open problems of alignment” are from:
creating a mind and shaping what it ends up wanting
overseeing a mind more capable than yourself
The biggest issue with the take in the tweet is that multi-agent coordination doesn’t engage either of these situations by default, because AIs can capture nearly all of the value from coordination without addressing these two open problems.
We lack a science for how values and goals form in intelligent systems and we don’t know how to predict or steer what a given training process will instill.
More on the multi-agent coordination setting: a coordinating/orchestrating AI is interacting with AIs whose capabilities don’t drastically exceed its own and whose qualities are bounded by the fact that it’s an instance of identical or near-copy weights. For multi-agent coordination, these problems are mostly of incentives, verification, and security.
We can consider multi-agent coordination of two flavors. One is where an orchestrator agent spawns sub-agents and delegates sub-tasks. Another is where there are separate interacting AIs without an implicit principal-agent structure.
Let’s first consider the orchestrator-subagent delegation setting where a model spawns instances of itself. The sub-agents are copies with the same weights, differing only in instructions and accumulated context. This doesn’t make trust trivial because copies can still diverge in behavior from the context and we’re also unclear about whether instances have indexical preferences, but these issues can be handled with standard oversight/control protocols because the principal is at least as smart as the sub-agents.
A second type of multi-agent coordination is not implicitly hierarchical at all. An ecology of independent AIs with different weights and provenance can’t rely on control or assurances from sharing weights. This is what human institutions address even without a deep/good theory of human values. Humans with conflicting goals coordinate through reputations, legal systems, etc.
AI ecologies may build tools for enabling coordination that humans can’t utilize e.g., cryptographic proofs of what code an agent is running, escrowed stakes, or fully auditable logs of actions and verbalized thoughts. All of this may be sufficient for multi-agent AI coordination.
The most analogous/relevant form of multi-agent coordination (at this point we’re really stretching the definitional boundary) is an AI that is also delegating upward and building/training a successor smarter than itself (which just reproduces our predicament).
But nearly all gains from multi-agent work are available among peers/copies without entering this regime. This also only forces alignment research on a particular kind of AI mind (one that has stable goals of its own that it wants preserved in its successor). A corrigible, instruction-following system can be told to build a smarter model and will do so without caring whether the result is aligned with anyone/anything. As will an AI that just wants to complete research tasks well.
[highly contestable take; highly uncertain and loosely held] On present vibes-evidence those are, in my opinion, at least as likely as a goal-guarding schemer. The pressure from increasing capacity in multi-agent coordination probably selects more for those kinds of AIs which is not the kind of AI we want building its successors.
The defensible version of Josh’s claim that multi-agent coordination forces AGI/ASI to solve open problems in alignment science is very modest.
Multi-agent deployment does create instrumental demand for tools and methods improving coordination that border alignment research, and some of these results may help. But improving coordination won’t force AGI to solve alignment because it conflates coordinating existing minds (with particular affordances unavailable to humans) with creating new, smarter, different ones.
We should not conclude that capability growth in multi-agent coordination will result in meaningful alignment progress.



