BK
B.B. Kovács
info
Please Note
<p>This page displays the records of the person named above and is not linked to a unique person identifier. This record may need to be merged to a profile.</p>
2 records found
1
To be useful, AI needs to act in accordance with human goals, particularly in situations that have trade-offs between cooperative and uncooperative behaviour, formalised as sequential social dilemmas (SSDs). A promising method for finding these human goals from observing humans is Multi-Agent Inverse Reinforcement Learning (MIRL), but current MIRL algorithms perform poorly in SSDs.
A specific feature that makes this difficult is that human behaviour in SSDs often depends on how we expect others to react to our actions, and change their beliefs about how we will act in the future. Humans do this by using `Theory of Mind' (ToM): each human has a model of how others might act towards them, which they update based on evidence.
Building on this, we propose the CUSToM (Change-aware Understanding of Social dilemmas with Theory of Mind) algorithm for MIRL, which treats agents as having a ToM that is explicitly aware of change in opponent behaviours.
We show experimentally that our algorithm is able to retrieve accurate reward functions even in SSD situations where current state of the art algorithms fail. This provides evidence for the benefit of ToM-based MIRL methods and shows that, to work well, these need to take into account the inherent changeability of opponent behaviours. ...
A specific feature that makes this difficult is that human behaviour in SSDs often depends on how we expect others to react to our actions, and change their beliefs about how we will act in the future. Humans do this by using `Theory of Mind' (ToM): each human has a model of how others might act towards them, which they update based on evidence.
Building on this, we propose the CUSToM (Change-aware Understanding of Social dilemmas with Theory of Mind) algorithm for MIRL, which treats agents as having a ToM that is explicitly aware of change in opponent behaviours.
We show experimentally that our algorithm is able to retrieve accurate reward functions even in SSD situations where current state of the art algorithms fail. This provides evidence for the benefit of ToM-based MIRL methods and shows that, to work well, these need to take into account the inherent changeability of opponent behaviours. ...
To be useful, AI needs to act in accordance with human goals, particularly in situations that have trade-offs between cooperative and uncooperative behaviour, formalised as sequential social dilemmas (SSDs). A promising method for finding these human goals from observing humans is Multi-Agent Inverse Reinforcement Learning (MIRL), but current MIRL algorithms perform poorly in SSDs.
A specific feature that makes this difficult is that human behaviour in SSDs often depends on how we expect others to react to our actions, and change their beliefs about how we will act in the future. Humans do this by using `Theory of Mind' (ToM): each human has a model of how others might act towards them, which they update based on evidence.
Building on this, we propose the CUSToM (Change-aware Understanding of Social dilemmas with Theory of Mind) algorithm for MIRL, which treats agents as having a ToM that is explicitly aware of change in opponent behaviours.
We show experimentally that our algorithm is able to retrieve accurate reward functions even in SSD situations where current state of the art algorithms fail. This provides evidence for the benefit of ToM-based MIRL methods and shows that, to work well, these need to take into account the inherent changeability of opponent behaviours.
A specific feature that makes this difficult is that human behaviour in SSDs often depends on how we expect others to react to our actions, and change their beliefs about how we will act in the future. Humans do this by using `Theory of Mind' (ToM): each human has a model of how others might act towards them, which they update based on evidence.
Building on this, we propose the CUSToM (Change-aware Understanding of Social dilemmas with Theory of Mind) algorithm for MIRL, which treats agents as having a ToM that is explicitly aware of change in opponent behaviours.
We show experimentally that our algorithm is able to retrieve accurate reward functions even in SSD situations where current state of the art algorithms fail. This provides evidence for the benefit of ToM-based MIRL methods and shows that, to work well, these need to take into account the inherent changeability of opponent behaviours.
Teaching How to Learn to Learn
Teacher-Student Curriculum Learning for Efficient Meta-Learning
We investigate whether a teacher-student curriculum learning approach using a teacher network with a simpler structure than the student network can achieve better results at meta-learning. The goal of meta-learning is to learn from a set of tasks, and then perform well on a new, structurally similar but unseen task with minimal retraining. Instead of sampling uniformly from all data to create the training batches, the curriculum-learning approach aims to create a sequence of mini-batches that enhances the training process, also known as a curriculum. During teacher-student curriculum learning a "teacher" network is trained in the standard manner, and then its outputs are used to order the training samples by difficulty and categorise them into mini-batches. This curriculum is then used to train the "student" network. Previous teacher-student models either had pre-trained more complex teachers, or teachers with the same structure as the student network. We investigate whether a teacher network with a simpler structure can also increase accuracy, while preserving computational resources. We find that using such a curriculum worsens performance compared to not using any curriculum at all.
...
We investigate whether a teacher-student curriculum learning approach using a teacher network with a simpler structure than the student network can achieve better results at meta-learning. The goal of meta-learning is to learn from a set of tasks, and then perform well on a new, structurally similar but unseen task with minimal retraining. Instead of sampling uniformly from all data to create the training batches, the curriculum-learning approach aims to create a sequence of mini-batches that enhances the training process, also known as a curriculum. During teacher-student curriculum learning a "teacher" network is trained in the standard manner, and then its outputs are used to order the training samples by difficulty and categorise them into mini-batches. This curriculum is then used to train the "student" network. Previous teacher-student models either had pre-trained more complex teachers, or teachers with the same structure as the student network. We investigate whether a teacher network with a simpler structure can also increase accuracy, while preserving computational resources. We find that using such a curriculum worsens performance compared to not using any curriculum at all.