Robotic Skill Diversification via Active Mutation of Reward Functions in Reinforcement Learning during a Liquid Pouring Task

Conference Paper (2026)
Author(s)

Jannick Van Buuren (Student TU Delft)

Roberto Giglio (Politecnico di Milano)

Loris Roveda (Politecnico di Milano, Dalle Molle Institute for Artificial Intelligence (IDSIA))

Luka Peternel (TU Delft - Mechanical Engineering)

Research Group
Human-Robot Interaction
DOI related publication
https://doi.org/10.1109/AIM65483.2026.11658170 Final published version
More Info
expand_more
Publication Year
2026
Language
English
Research Group
Human-Robot Interaction
Publisher
IEEE
ISBN (electronic)
979-8-3195-3611-2
Event
2026 IEEE/ASME International Conference on Advanced Intelligent Mechatronics, AIM 2026 (2026-07-07 - 2026-07-10), Genova, Italy
Downloads counter
2
Reuse Rights

Other than for strictly personal use, it is not permitted to download, forward or distribute the text or part of it, without the consent of the author(s) and/or copyright holder(s), unless the work is under an open content license such as Creative Commons.

Abstract

This paper explores how deliberate mutations of reward function in reinforcement learning can produce diversified skill variations in robotic manipulation tasks, examined with a liquid pouring use case. To this end, we developed a new reward function mutation framework that is based on applying Gaussian noise to the weights of the different terms in the reward function. Inspired by the cost-benefit tradeoff model from human motor control, we designed the reward function with the following key terms: accuracy, time, and effort. The study was performed in a simulation environment created in NVIDIA Isaac Sim, and the setup included Franka Emika Panda robotic arm holding a glass with a liquid that needed to be poured into a container. The reinforcement learning algorithm was based on Proximal Policy Optimization. We systematically explored how different configurations of mutated weights in the rewards function would affect the learned policy. The resulting policies exhibit a wide range of behaviours: from variations in execution of the originally intended pouring task to novel skills useful for unexpected tasks, such as container rim cleaning, liquid mixing, and watering. This approach offers promising directions for robotic systems to perform diversified learning of specific tasks, while also potentially deriving meaningful skills for future tasks.

Files

– Personal use only – Dutch Copyright Act (Article 25fa)
warning

File under embargo until 25-02-2027