Towards cooperation in learning games
This piece is available as a PDF.
Suppose that several actors are going to deploy learning agents to act on their behalf. What principles should guide these actors in designing their agents, given that they may have competing goals? An appealing solution concept in this setting is welfare-optimal learning equilibrium. This means that the learning agents should constitute a Nash equilibrium whose payoff profile is optimal according to some measure of total welfare (welfare function). In this work, we construct a class of learning algorithms in this spirit called learning tit-for-tat (L-TFT). L-TFT algorithms maximize a welfare function according to a specified optimization schedule, and punish their counterpart when they detect that they are deviating from this plan. Because the policies of other agents are not in general fully observed, agents must infer whether their counterpart is following a cooperative learning algorithm. This requires us to develop new techniques for making inferences about counterpart learning algorithms. In two sequential social dilemmas, our L-TFT algorithms successfully cooperate in self-play while effectively avoiding exploitation by and punishing defecting learning algorithms.
Cite this
@online{clifton-towards-cooperation-in-learning-games,
title = {Towards cooperation in learning games},
author = {Jesse Clifton and Maxime Riché},
url = {https://longtermrisk.org/media/toward_cooperation_learning_games_oct_2020.pdf},
year = {2020},
date = {2020-10-01},
howpublished = {Working paper},
keywords = {},
pubstate = {published},
tppubtype = {online}
}