D.J. Broekens
Please Note
54 records found
1
A Little Chit-Chat Goes a Long Way
Design and Evaluation of Task-and Person-Oriented Styles for Social Robots
Whereas the reception task is a promising application domain for social robots, knowledge is lacking about how to design the appropriate re-usable communication styles for a reception robot. This paper presents the use and evaluation of an iterative interaction-design (ID) method with which task- and person-oriented multi-modal communication styles have been designed for such a robot. First, we report on an evaluation study of the ID-method with Industrial Design students (N =13) who designed these two communication styles for a Pepper robot. This provided a set of distinct designs of the two styles, for which the differences in design parameters were in line with social science theory. The task-oriented style showed a more formal, shorter and less chatty communication. Second, we present findings from a Mechanical Turk study conducted to evaluate the perception of these style designs. Participants (N =301) were presented with videos showing the robot acting as a receptionist and were asked to rate their perception of the robot, the service experience and the orientation of the designs. Overall, the interaction with the robot was appreciated well. The robot with a person-oriented style was perceived to be more animate and likeable. Analysis showed that chit-chat was the main contributor to the perceived difference between the person-oriented and task-oriented styles. This is an important finding as it gives interaction designers a validated best-practice approach to make interaction style more or less personal.
Collecting Mementos
A Multimodal Dataset for Context-Sensitive Modeling of Affect and Memory Processing in Responses to Videos
In this article we introduce Mementos: the first multimodal corpus for computational modeling of affect and memory processing in response to video content. It was collected online via crowdsourcing and captures 1995 individual responses collected from 297 unique viewers responding to 42 different segments of music videos. Apart from webcam recordings of their upper-body behavior (totaling 2012 minutes) and self-reports of their emotional experience, it contains detailed descriptions of the occurrence and content of 989 personal memories triggered by the video content. Finally, the dataset includes self-report measures related to individual differences in participants' background and situation (Demographics, Personality, and Mood), thereby facilitating the exploration of important contextual factors in research using the dataset. We describe 1) the construction and contents of the corpus itself, 2) analyse the validity of its content by investigating biases and consistency with existing research on affect and memory processing, 3) review previously published work that demonstrates the usefulness of the multimodal data in the corpus for research on automated detection and prediction tasks, and 4) provide suggestions for how the dataset can be used in future research on modeling Video-Induced Emotions, Memory-Associated Affect, and Memory Evocation.
Intelligent tutoring systems need a model of learning goals for the personalization of educational content, tailoring of the learning path, progress monitoring, and adaptive feedback. This article presents such a model and corresponding interaction designs for the coaches and learners (respectively, a monitor-and-control dashboard and mobile app with supportive communications trough a virtual agent), all deployed and tested in a system for child diabetes self-management training. We developed a domain-independent upper ontology to structure learning goals and related concepts (such as achievements and tasks) and a domain ontology that specifies the knowledge base (for, in our case, diabetes self-management training). With this approach, we relate knowledge elements (e.g., skill) to educational tasks and to learners' knowledge development (e.g., achievements). The ontology was implemented in a multimodal tutoring system consisting of mobile educative games, a health diary, an embodied conversational agent (ECA), and a web application for authoring and monitoring. We show that our model provides a coherent and concise foundation for: 1) the formalization of learning in the diabetes self-management domain, but also for other domains such as mathematics; 2) personal goal setting and thereby personalization of the educational process including ECA's guidance; and 3) creating awareness of progress on the personal educational path. We found that a motivational tutoring system requires a rich set of learning activities and accompanying materials of which a subset is offered to the learner based on personal relevance. The implemented model proved to accommodate the personal agent-guided learning paths of children with diabetes, under different treatments from hospitals in Italy and the Netherlands.
The Role of Simulated Emotions in Reinforcement Learning
Insights from a Human-Robot Interaction Experiment.
Transparency of behavior is important for robots that work with humans. If such robots need to adapt to a variety of users and tasks, they need to learn to optimize their behavior, and Reinforcement Learning (RL) is a promising learning method for this purpose. However, the behavior generated by RL is not inherently transparent due to the exploration/exploitation tradeoff that is needed to optimize a policy.Emotions are -for humans- a natural way of communicating intent and situational appraisal. In this study, we implemented emotional expressions based on Temporal Differences as a means to increase the transparency of a robot's learning process. We analysed the effect on the human teacher's behavior and experience, and on the robot's learning result and learning process.A between-subject experiment with 61 participants and three robot conditions was performed: no emotions, simulated emotions, and simulated emotions with matching attribution. The learning task was one where a human teacher had to help a humanoid robot to learn the meaning of three colors.Our results demonstrate minimal differences between these three conditions. This means that for simple tasks, emotional expressions grounded in RL do not help nor hurt. We discuss our findings and propose three important criteria for interactive learning tasks when investigating the effect of emotional expressions grounded in RL. Such tasks need to be sufficiently complex, afford robot autonomy, and the emotion must be informative about how the user could influence the robot's actions.
Model-based Reinforcement Learning
A Survey
Sequential decision making, commonly formalized as Markov Decision Process (MDP) optimization, is an important challenge in artificial intelligence. Two key approaches to this problem are reinforcement learning (RL) and planning. This survey is an integration of both fields, better known as model-based reinforcement learning. Model-based RL has two main steps. First, we systematically cover approaches to dynamics model learning, including challenges like dealing with stochasticity, uncertainty, partial observability, and temporal abstraction. Second, we present a systematic categorization of planning-learning integration, including aspects like: where to start planning, what budgets to allocate to planning and real data collection, how to plan, and how to integrate planning in the learning and acting loop. After these two sections, we also discuss implicit model-based RL as an end-to-end alternative for model learning and planning, and we cover the potential benefits of model-based RL. Along the way, the survey also draws connections to several related RL fields, like hierarchical RL and transfer learning. Altogether, the survey presents a broad conceptual overview of the combination of planning and learning for MDP optimization.
Sequential decision making, commonly formalized as optimization of a Markov Decision Process, is a key challenge in artificial intelligence. Two successful approaches to MDP optimization are reinforcement learning and planning, which both largely have their own research communities. However, if both research fields solve the same problem, then we might be able to disentangle the common factors in their solution approaches. Therefore, this paper presents a unifying algorithmic framework for reinforcement learning and planning (FRAP), which identifies underlying dimensions on which MDP planning and learning algorithms have to decide. At the end of the paper, we compare a variety of well-known planning, model-free and model-based RL algorithms along these dimensions. Altogether, the framework may help provide deeper insight in the algorithmic design space of planning and reinforcement learning.
A Cloud-based Robot System for Long-term Interaction
Principles, Implementation, Lessons Learned
Making the transition to long-term interaction with social-robot systems has been identified as one of the main challenges in human-robot interaction. This article identifies four design principles to address this challenge and applies them in a real-world implementation: cloud-based robot control, a modular design, one common knowledge base for all applications, and hybrid artificial intelligence for decision making and reasoning. The control architecture for this robot includes a common Knowledge-base (ontologies), Data-base, "Hybrid Artificial Brain"(dialogue manager, action selection and explainable AI), Activities Centre (Timeline, Quiz, Break and Sort, Memory, Tip of the Day, ), Embodied Conversational Agent (ECA, i.e., robot and avatar), and Dashboards (for authoring and monitoring the interaction). Further, the ECA is integrated with an expandable set of (mobile) health applications. The resulting system is a Personal Assistant for a healthy Lifestyle (PAL), which supports diabetic children with self-management and educates them on health-related issues (48 children, aged 6-14, recruited via hospitals in the Netherlands and in Italy). It is capable of autonomous interaction "in the wild"for prolonged periods of time without the need for a "Wizard-of-Oz"(up until 6 months online). PAL is an exemplary system that provides personalised, stable and diverse, long-term human-robot interaction.
ComVis-Sail
Comparative Sailing Performance Visualization for Coaching
During training sessions, sailors rely on feedback provided by the coaches to reinforce their skills and improve their performance. Nowadays, the incorporation of sensors on the boats enables coaches to potentially provide more informed feedback to the sailors. A common exercise during practice sessions, consists of two boats of the same class, sailing side by side in a straight line with different boat handling techniques. Coaches try to understand which techniques are that make one boat go faster than the other. The analysis of the obtained data from the boats is challenging given its multi-dimensional, time-varying and spatial nature. At present, coaches only rely on aggregated statistics reducing the complexity of the data, hereby losing local and temporal information. We describe a new domain characterization and present a visualization design that allows coaches to analyse the data, structuring their analysis and explore the data from different perspectives. A central element of the tool is the glyph design to intuitively represent and aggregate multiple aspects of the sensor data. We have conducted multiple user studies with naive users, sailors and coaches to evaluate the design and potential of the overall tool. (Figure presented.).
Robots and virtual agents need to adapt existing and learn novel behavior to function autonomously in our society. Robot learning is often in interaction with or in the vicinity of humans. As a result the learning process needs to be transparent to humans. Reinforcement Learning (RL) has been used successfully for robot task learning. However, this learning process is often not transparent to the users. This results in a lack of understanding of what the robot is trying to do and why. The lack of transparency will directly impact robot learning. The expression of emotion is used by humans and other animals to signal information about the internal state of the individual in a language-independent, and even species-independent way, also during learning and exploration. In this article we argue that simulation and subsequent expression of emotion should be used to make the learning process of robots more transparent. We propose that the TDRL Theory of Emotion gives sufficient structure on how to develop such an emotionally expressive learning robot. Finally, we argue that next to such a generic model of RL-based emotion simulation we need personalized emotion interpretation for robots to better cope with individual expressive differences of users.
Empirical evidence suggests that the emotional meaning of facial behavior in isolation is often ambiguous in real-world conditions. While humans complement interpretations of others' faces with additional reasoning about context, automated approaches rarely display such context-sensitivity. Empirical findings indicate that the personal memories triggered by videos are crucial for predicting viewers' emotional response to such videos ?- in some cases, even more so than the video's audiovisual content. In this article, we explore the benefits of personal memories as context for facial behavior analysis. We conduct a series of multimodal machine learning experiments combining the automatic analysis of video-viewers' faces with that of two types of context information for affective predictions: \beginenumerate∗[label=(\arabic∗)] \item self-reported free-text descriptions of triggered memories and \item a video's audiovisual content \endenumerate∗. Our results demonstrate that both sources of context provide models with information about variation in viewers' affective responses that complement facial analysis and each other.
A Framework for Reinforcement Learning and Planning
Extended Abstract
Think Too Fast Nor Too Slow
The Computational Trade-off Between Planning And Reinforcement Learning
Virtual agents are increasingly being used for communication training such as public speaking-, job interviews-, as well as negotiation training. In these use-cases the agent is generally taking on the role of interviewer and its behaviour is altered according to the nonverbal cues of its human interlocutor. However, understanding how the agent's non-verbal cues influence human behaviour, perception or interactions outcomes is equally important. This contributes to appropriate behaviour generation in agents, but also to our understanding of the intricate interplay of non-verbal behaviours on human perception and interaction outcomes. This study focuses specifically on the perception of vocal dominance in human-agent negotiation. Earlier research showed that the perception of dominance influences decision making in the course of negotiations, as do concessions tactics. However, the effect of voice as well as the effect of the negotiator type in this regard have been so far under-explored. To close this gap, an online experiment was conducted, in which a total number of 121 participants negotiated with conversational agents displaying either low or high vocal dominance, and an individualistic or neutral concession tactic. The results showed that when taking into account the self-reported type of negotiator, significant differences caused by vocal dominance were evident in regard to the number of negotiation rounds and perceived fairness. The number of rounds was significantly higher for the competitive participant type negotiating with the low vocal dominance agent, and the perceived fairness was significantly lower with the collaborative participant type negotiating with the high vocal dominance agent.
Context in Human Emotion Perception for Automatic Affect Detection
A Survey of Audiovisual Databases
An important aspect of human emotion perception is the use of contextual information to understand others' feelings even in situations where their behavior is not very expressive or has an emotionally ambiguous meaning. For technology to successfully detect affect, it must mimic this human ability when analyzing audiovisual input. Databases upon which machine learning algorithms are trained should capture the context of social interactions as well as the behavior expressed in them. However, there is a lack of consensus about what constitutes relevant context in such databases. In this article, we make two contributions towards overcoming this challenge: (a) we identify two principal sources of context for emotion perceptions based on psychological theory, and (b) we provide an overview of how each of these has been considered in published databases covering social interactions. Our results show that a similar set of contextual features are present across the reviewed databases. Between all the different databases researchers seem to have taken into account a set of contextual features reflecting the sources of context seen in psychological theory. However, within individual databases, these features are not yet systematically varied. This is problematic because it prevents them from being used directly as resources for the modeling of context-sensitive affect detection. Based on our findings, we suggest improvements for the future development of affective databases.
Explanation of actions is important for transparency of-, and trust in the decisions of smart systems. Literature suggests that emotions and emotion words-in addition to beliefs and goals-are used in human explanations of behaviour. Furthermore, research in e-health support systems and human-robot interaction stresses the need for studying long-term interaction with users. However, state of the art explainable artificial intelligence for intelligent agents focuses mainly on explaining an agent's behaviour based on the underlying beliefs and goals in short-term experiments. In this paper, we report on a long-term experiment in which we tested the effect of cognitive, affective and lack of explanations on children's motivation to use an e-health support system. Children (aged 6-14) suffering from type 1 diabetes mellitus interacted with a virtual robot as part of the e-health system over a period of 2.5-3 months. Children alternated between the three conditions. Agent behaviours that were explained to the children included why 1) the agent asks a certain quiz question; 2) the agent provides a specific tip (a short instruction) about diabetes; or, 3) the agent provides a task suggestion, e.g., play a quiz, or, watch a video about diabetes. Their motivation was measured by counting how often children would follow the agent's suggestion, how often they would continue to play the quiz or ask for an additional tip, and how often they would request an explanation from the system. Surprisingly, children proved to follow task suggestions more often when no explanation was given, while other explanation effects did not appear. This is to our knowledge the first longterm study to report empirical evidence for an agent explanation effect, challenging the next studies to uncover the underlying mechanism.
Robots Expressing Dominance
Effects of Behaviours and Modulation
A mayor challenge in human-robot interaction and collaboration is the synthesis of non-verbal behaviour for the expression of social signals. Appropriate perception and expression of dominance (verticality) in non-verbal behaviour is essential for social interaction. In this paper, we present our work on algorithmic modulation of robot bodily movement to express varying degrees of dominance. We developed a parameter-based model for head tilt and body expansiveness. This model was applied to a variety of behaviours. These behaviours were evaluated by human observers in two different studies with respectively static pictures of key postures (N=772) and realtime gestures (N=31). Overall, specific behaviours proved to communicate different levels of dominance. Further, modulation of body expansiveness and head tilt robustly influenced perceived dominance independent of specific behaviours and observer viewing height and angle. The modulation did not influence perceived valence, but it did influence perceived arousal. Our study shows that dominance can be reliably expressed by both selection of specific behaviours and modulation of behaviours.
A mayor challenge in human-robot interaction is the synthesis of social signals through non-verbal behaviour expression. Appropriate perception and expression of dominance (verticality) is essential for social interaction. In this paper, we present our work on algorithmic modulation of robot bodily movement to control dominance expression. We developed a parameter-based model for body expansiveness. This model was applied to a variety of behaviours and evaluated by human observers in two different studies with respectively static postures (N=772) and gestures (N=31). Modulation of body expansiveness proved to robustly influence perceived dominance independent of behaviour and viewing angles.
The perception of warmth and competence in others influences social interaction and decision making. Virtual agents have been used in many domains including serious gaming and training. In this work we study the effect of warmth expressed in the behavior of a virtual agent on a human-agent negotiation. We design and conduct an experiment where participants negotiate with two versions of the same agent displaying varying levels of warmth. The results show that humans are more satisfied with the warm agent, are more willing to renegotiate with it, would recommend the agent more to their friends and had a better interaction experience, even though there is no difference in negotiation outcome (utility, agreement or rounds needed). While studies have shown effects of emotional displays on negotiation and collaboration, this is - to our knowledge - the first time that a clear effect of behavioral style is shown on the post-hoc appraisal of a human-agent collaboration, in our case a negotiation.