Circular Image

A. Mone

info

Please Note

2 records found

Diagnostics for MORL Policy Selection

Real-world decision-making often requires optimizing multiple competing objectives simultaneously. In reinforcement learning (RL), this is typically addressed by combining reward signals into a single scalar objective via a scalarization function, which can be fragile: small changes in the weights can induce drastically different policies. Multi-objective reinforcement learning (MORL) instead produces sets of policies that explicitly represent trade-offs between objectives. However, these policies are typically presented to the decision maker only through their value vectors, which can obscure substantial behavioral variation: policies that induce distinct trajectories may appear indistinguishable when evaluated solely by expected returns. We propose an exploratory diagnostic workflow that automatically highlights behavioral variation along the Pareto front that objective values alone do not reveal, providing both quantitative and visual tools to support policy inspection. We validate our approach on simple grid examples and scale it to continuous control benchmarks, demonstrating that it remains effective as problem complexity increases. ...
Inverse Reinforcement Learning (IRL) seeks to infer reward functions from expert demonstrations. When demonstrations come from multiple experts with different intentions, the problem becomes Multi-Intention IRL (MI-IRL). MI-IRL methods generally follow a rewardfirst paradigm: they ask which reward, if followed, could have generated each trajectory. This leads to trajectory similarity being based on reward likelihood rather than on the relationships between behaviors. In deep generative MI-IRL methods, this paradigm further couples behavior clustering and reward learning, making such methods dependent on prior knowledge of the true number of modes K∗. This limits adaptability to unseen behaviors, and restricts the analysis to learned rewards rather than relationships across behaviors. In contrast, we approach MI-IRL from a different standpoint. We propose Contrastive Multi-Intention IRL (CoMI-IRL), a transformer-based unsupervised behavior-first framework: it first learns a representation of trajectory dynamics that captures intrinsic behavioral similarity, then clusters trajectories in that space, and finally learns a reward for each cluster. Our experiments show that CoMI-IRL achieves higher clustering and reward learning performance when compared to Ess-InfoGAIL without needing K∗ or labels, while allowing for visual interpretation of behavior relationships and adaptation to unseen behavior without full retraining. ...