MP
M. Plooij
info
Please Note
<p>This page displays the records of the person named above and is not linked to a unique person identifier. This record may need to be merged to a profile.</p>
2 records found
1
Controlling the estimation bias in deep reinforcement learning problems with sparse rewards
Towards robust robotic object manipulation learning
Many recent robot learning problems, real and simulated, were addressed using deep reinforcement learning. The developed policies can deal with high-dimensional, continuous state and action spaces, and can also incorporate machine-generated or human demonstration data. A great number of them depend on state-action value estimates, especially the ones in the actor-critic framework. Deriving unbiased estimates for these values is still an open research question, mostly since the connection between accurate value estimates and system performance is not yet well-understood. This thesis work has three main research contributions. Firstly, it analyzes the connection between value estimates and performance for the TD3 algorithm. Secondly, it derives theoretical bounds for the true value function when dealing with environments where a reward is only given for successful completion of a task (sparse/binary reward). Lastly, a deliberate underestimation objective is added to the TD3 algorithm together with the theoretical bounds to improve system performance when using human demonstration data that only covers a specific part of the state and action space. All the algorithms are tested and evaluated using simulated robot manipulation tasks in the robosuite environment, where the robot is first trained on the demonstration data and then can gather more experiences in the simulation. Results show that the deliberate underestimation together with the value bounds enable the robot to learn from human demonstration, which was not possible for the standard TD3. Additionally, applying just the value bounds speeds up the learning process when using machine-generated datasets.
...
Many recent robot learning problems, real and simulated, were addressed using deep reinforcement learning. The developed policies can deal with high-dimensional, continuous state and action spaces, and can also incorporate machine-generated or human demonstration data. A great number of them depend on state-action value estimates, especially the ones in the actor-critic framework. Deriving unbiased estimates for these values is still an open research question, mostly since the connection between accurate value estimates and system performance is not yet well-understood. This thesis work has three main research contributions. Firstly, it analyzes the connection between value estimates and performance for the TD3 algorithm. Secondly, it derives theoretical bounds for the true value function when dealing with environments where a reward is only given for successful completion of a task (sparse/binary reward). Lastly, a deliberate underestimation objective is added to the TD3 algorithm together with the theoretical bounds to improve system performance when using human demonstration data that only covers a specific part of the state and action space. All the algorithms are tested and evaluated using simulated robot manipulation tasks in the robosuite environment, where the robot is first trained on the demonstration data and then can gather more experiences in the simulation. Results show that the deliberate underestimation together with the value bounds enable the robot to learn from human demonstration, which was not possible for the standard TD3. Additionally, applying just the value bounds speeds up the learning process when using machine-generated datasets.
Master thesis
(2022)
-
C.A. Langens, M. Wiertlewski, L. Willemet, M. Plooij, J. Kober, E. van der Kruk
Manipulating soft and fragile objects is a challenging task in robotic grasping. The key challenge for robotic grasping is to exert enough grip force to prevent slipping while being gentle enough to prevent damage to an object. Existing grippers used for processes like automatic harvesting of fruits, either apply excessive grip force leading to object damage or react to slip resulting in object release from the gripper. The aim of this study is to develop a grip force controller that uses tactile feedback to maintain a constant frictional safety margin over the minimum required grip force, called Safety Margin Control. Tactile sensors can provide information on friction, which is used to predict slip. An optical tactile sensor is modeled and used in simulations where Safety Margin Control regulates the grip force during interaction with various virtual objects. The deformation of the sensor’s soft viscoelastic membrane is described by local frictional behavior and used to estimate the safety margin. The desired safety margin is set to 30%, based on comparison to the way humans control grip force in their fingertips. The desired value can be tuned to favor release over damage and vice versa. Safety Margin Control is compared to two baseline controllers: React To Slip and Conservative Control. The performance is evaluated based on maximum pressure and total lateral displacement of the object relative to the sensor. Safety Margin Control results in a pressure decrease of 44% on average compared to Conservative Control, and no significant pressure change was observed compared to React To Slip. The total lateral displacement for Safety Margin Control is 0 mm, as opposed to 1.3 mm for React To Slip. Safety Margin Control provides a way forward for automated harvesting as the pressure exerted on an object can be reduced while no slip occurs.
...
Manipulating soft and fragile objects is a challenging task in robotic grasping. The key challenge for robotic grasping is to exert enough grip force to prevent slipping while being gentle enough to prevent damage to an object. Existing grippers used for processes like automatic harvesting of fruits, either apply excessive grip force leading to object damage or react to slip resulting in object release from the gripper. The aim of this study is to develop a grip force controller that uses tactile feedback to maintain a constant frictional safety margin over the minimum required grip force, called Safety Margin Control. Tactile sensors can provide information on friction, which is used to predict slip. An optical tactile sensor is modeled and used in simulations where Safety Margin Control regulates the grip force during interaction with various virtual objects. The deformation of the sensor’s soft viscoelastic membrane is described by local frictional behavior and used to estimate the safety margin. The desired safety margin is set to 30%, based on comparison to the way humans control grip force in their fingertips. The desired value can be tuned to favor release over damage and vice versa. Safety Margin Control is compared to two baseline controllers: React To Slip and Conservative Control. The performance is evaluated based on maximum pressure and total lateral displacement of the object relative to the sensor. Safety Margin Control results in a pressure decrease of 44% on average compared to Conservative Control, and no significant pressure change was observed compared to React To Slip. The total lateral displacement for Safety Margin Control is 0 mm, as opposed to 1.3 mm for React To Slip. Safety Margin Control provides a way forward for automated harvesting as the pressure exerted on an object can be reduced while no slip occurs.