Theory of Mind for Multi-Agent Inverse Reinforcement Learning in Sequential Social Dilemmas
B.B. Kovács (TU Delft - Electrical Engineering, Mathematics and Computer Science)
L. Cavalcante Siebert – Mentor (TU Delft - Electrical Engineering, Mathematics and Computer Science)
A. Mone – Mentor (TU Delft - Electrical Engineering, Mathematics and Computer Science)
F.A. Oliehoek – Graduation committee member (TU Delft - Electrical Engineering, Mathematics and Computer Science)
More Info
expand_more
Other than for strictly personal use, it is not permitted to download, forward or distribute the text or part of it, without the consent of the author(s) and/or copyright holder(s), unless the work is under an open content license such as Creative Commons.
Abstract
To be useful, AI needs to act in accordance with human goals, particularly in situations that have trade-offs between cooperative and uncooperative behaviour, formalised as sequential social dilemmas (SSDs). A promising method for finding these human goals from observing humans is Multi-Agent Inverse Reinforcement Learning (MIRL), but current MIRL algorithms perform poorly in SSDs.
A specific feature that makes this difficult is that human behaviour in SSDs often depends on how we expect others to react to our actions, and change their beliefs about how we will act in the future. Humans do this by using `Theory of Mind' (ToM): each human has a model of how others might act towards them, which they update based on evidence.
Building on this, we propose the CUSToM (Change-aware Understanding of Social dilemmas with Theory of Mind) algorithm for MIRL, which treats agents as having a ToM that is explicitly aware of change in opponent behaviours.
We show experimentally that our algorithm is able to retrieve accurate reward functions even in SSD situations where current state of the art algorithms fail. This provides evidence for the benefit of ToM-based MIRL methods and shows that, to work well, these need to take into account the inherent changeability of opponent behaviours.