BK

B.B. Kovács

info

Please Note

2 records found

To be useful, AI needs to act in accordance with human goals, particularly in situations that have trade-offs between cooperative and uncooperative behaviour, formalised as sequential social dilemmas (SSDs). A promising method for finding these human goals from observing humans is Multi-Agent Inverse Reinforcement Learning (MIRL), but current MIRL algorithms perform poorly in SSDs.

A specific feature that makes this difficult is that human behaviour in SSDs often depends on how we expect others to react to our actions, and change their beliefs about how we will act in the future. Humans do this by using `Theory of Mind' (ToM): each human has a model of how others might act towards them, which they update based on evidence.

Building on this, we propose the CUSToM (Change-aware Understanding of Social dilemmas with Theory of Mind) algorithm for MIRL, which treats agents as having a ToM that is explicitly aware of change in opponent behaviours.

We show experimentally that our algorithm is able to retrieve accurate reward functions even in SSD situations where current state of the art algorithms fail. This provides evidence for the benefit of ToM-based MIRL methods and shows that, to work well, these need to take into account the inherent changeability of opponent behaviours. ...

Teacher-Student Curriculum Learning for Efficient Meta-Learning

We investigate whether a teacher-student curriculum learning approach using a teacher network with a simpler structure than the student network can achieve better results at meta-learning. The goal of meta-learning is to learn from a set of tasks, and then perform well on a new, structurally similar but unseen task with minimal retraining. Instead of sampling uniformly from all data to create the training batches, the curriculum-learning approach aims to create a sequence of mini-batches that enhances the training process, also known as a curriculum. During teacher-student curriculum learning a "teacher" network is trained in the standard manner, and then its outputs are used to order the training samples by difficulty and categorise them into mini-batches. This curriculum is then used to train the "student" network. Previous teacher-student models either had pre-trained more complex teachers, or teachers with the same structure as the student network. We investigate whether a teacher network with a simpler structure can also increase accuracy, while preserving computational resources. We find that using such a curriculum worsens performance compared to not using any curriculum at all. ...