When we train multi-agent systems, one approach is to reward bridging — things that people who started far apart all find unexpectedly agreeable. When you make that the reward function, you nudge different people closer.
j 下一段next speechk 上一段previous speech