Research
Multimodal Robot Interaction (RoboCarnival)
A multimodal interaction concept across three robot platforms, voted Best Team by visiting high schoolers at RoboCarnival 2025.
- Role
- Workshop facilitator, observer & video editor
- Team
- Anna, Aseer, Emon, Fati, Jasmin (5-person team)
- Duration
- ~6 weeks, spring 2025
- Tools
- QT (voice), Misty (expression), Joy for All companion cat (touch)
The project
- 1
Course
Social Robots: Design, Research & Interaction, Tampere University
- 2
Research workshops
Visiting high schoolers test QT, Misty, and the cat (March 2025)
- 3
RoboCarnival
Public demo station, Hervanta campus (April 2025)
- 4
Best Team
Voted by the visiting students
A robot's face alone isn't enough to communicate an emotion, and people will trust a robot's touch faster than its words. That's the short version of what this project taught me. It's a university course project, not a flagship case study, but it's the clearest example in my portfolio of designing for a physical, multimodal interface, and it taught me things a screen-based project couldn't.
Our course, Social Robots: Design, Research and Interaction, was built around a framework called robot literacy: the idea that using and understanding robots is a skill, and that different people need different things to build that skill. The course split this into six areas, from basic awareness of what robots are and do, to programming them, to the ethics of trusting them. Each team picked one area to explore and build a public demo around.
My team of five was assigned interaction: how someone who has never met a robot before figures out how to engage with it. We turned that into a handful of research questions: how teenagers read a robot's non-verbal cues, how they think about robots as companions, what makes them willing to interact at all, and how modalities like touch, voice, and gesture change that willingness. We had three robots on loan to explore those questions with, one per channel.
My role
I want to be specific here, because this was a team project and I don't want to claim more than I did.
We ran several research workshops with visiting high school students over the course of the project. I facilitated one of these sessions myself: running the activities, keeping time, and making the calls on when to move on. For the others, I worked as an observer, staying quiet and watching what participants did and said when no one was managing them directly. Facilitating and observing gave me two different views of the same activities, and both fed into the same shared notes. After the workshops, the five of us went through those notes together, pulled out the patterns, and built our final presentation deck as a joint effort. None of that analysis or storytelling was mine alone.
One piece I did own individually: I shot and edited the short video we used for our pitch at RoboCarnival. It was a side task compared to the research, but it meant picking up a camera and an editing timeline on a tight deadline, which is its own kind of design constraint.
With roles sorted, here's what we were actually working with.
Three robots, three modalities
None of us wrote new behavior for these robots. QT, Misty, and the cat all came with existing capabilities, and programming them from scratch was actually a different team's project within the same course. Our job was to choose which of those existing capabilities to use, and to build a short, legible flow around them.
QT: voice
For QT, that meant loading in a set of spoken questions we wanted it to ask visitors, things like "where is the microphone?" or "would you find a robot like this creepy?" QT would ask, the visitor would answer out loud, and we'd watch how they responded to being addressed directly by a robot.
Misty: non-verbal expression
For Misty, it meant picking a handful of her preset facial expressions and sounds for an emotion-guessing game. An operator on our team would trigger an expression, visitors would write or draw what emotion they thought it was, and then we'd reveal the intended one.
The cat: touch
The cat, a Joy for All companion pet (the same product line used in a lot of eldercare and dementia research), needed no configuration at all. It reacted to touch: petting, holding, moving it. That simplicity turned out to matter a lot.
Testing modalities with real teenagers
One session in March is a good example of how these workshops worked: four groups of visiting high school students came through in one afternoon to test how these three channels actually landed. It wasn't a controlled study. Four groups, one afternoon. But the patterns showed up consistently enough across groups, and across the other sessions we ran, that I'd trust them as a real signal, not just a one-off reaction. Each group went through the same structure: a quick icebreaker about their attitudes toward robots, a round with Misty's emotion-guessing game, some free time with QT, and a closing discussion where they sketched ideas for how Misty's emotions could be made clearer.
A few things came out of this that shaped everything after.
The non-verbal channel was harder to read than we expected. Groups regularly mixed up Misty's expressions, mistaking her "party mode" for amazement, or unsure if a sound meant laughter or crying. Their own suggestions were telling: several students said Misty needed a mouth, not just eyes, because a face with fewer moving parts gives you less to read. One student made the connection to face masks during the pandemic, pointing out that people could still read emotion around a covered mouth, which meant the eyes alone should theoretically be enough. But on a simplified robot face, they weren't.
The touch channel needed no explanation at all. The cat was the most consistently popular thing in the room across all four groups. Students petted it without being told how, and the reactions were immediate and warm. One student said they couldn't "dehumanize" the cat enough to imagine using it for a chore like the dishes. Another compared it directly to their own pet.
The voice channel sat somewhere in between. Talking to QT was novel and often funny to them, but it also surfaced real questions: one student asked where the camera and microphone actually were, which opened into a genuine conversation about privacy that we hadn't planned for.
From notes to a public station
Once we pulled the patterns together across these workshops, the team sat down and turned them into a set of decisions for RoboCarnival. We used a simple impact-versus-feasibility exercise to sort through ideas, things like letting visitors redesign Misty's face by drawing, or running two robots through a short back-and-forth to show them interacting with each other. Most of those got parked. What survived was closer to what we already had: three robots, three modalities, side by side, with better facilitation built around each one.
The main change wasn't to the robots. It was to us. We built explicit talking points for the moments we knew would be confusing, like why Misty's face alone often wasn't enough, and we leaned harder into the parts that worked without any explanation, like just letting people pet the cat before we said anything at all.
RoboCarnival
RoboCarnival was the course's public exhibition, held in the Language Center lobby at Tampere University's Hervanta campus. Every team ran its own station side by side, each covering a different slice of robot literacy: ours was interaction, others covered programming, ethics, or trust. High school students from across Tampere came through in groups over the course of the day.
Getting ready for it meant packing our findings into something a stranger could absorb walking past a table in under a minute, plus the pitch video mentioned above, which was the other thing on our checklist that week.
At our table, QT, Misty, and the cat sat side by side, and we walked visitors through all three in a few minutes each. Having run the same activities several times already in our earlier workshops meant we had a sense of where people would get stuck and where they wouldn't need help at all, which made the day feel far less improvised than it could have.
Best team
At the end of the day, the visiting high school students voted on their favorite station, and ours was chosen. We were later given a certificate recognizing the team for the project, noting we'd been picked as the best team by the students who came through. It's a nice thing to have, and I want to be clear about what it is: a vote from a friendly, one-day audience of teenagers, not an industry award. I'm proud of it in that context.
You were chosen to be the best team by highschool students. Congratulations!!!
What I took from this
The biggest lesson was about redundancy across channels. Misty's face, on its own, wasn't a reliable way to communicate an emotion, however good the intention behind the design was. People needed more than one signal to be confident in what they were reading, and when we only gave them one, they guessed wrong more often than we expected going in. That's stuck with me. When I look at any interface now, robotic, voice, or otherwise, I ask what happens when the primary signal is ambiguous, and whether there's a second one backing it up.
Working with physical robots also changed how I think about failure. There's no console log to check when a robot doesn't respond the way you expect in front of a room of teenagers. You either have a backup plan ready, or the moment just falls flat. Our workshop script had built-in fallback activities for exactly this reason, and having them there took real pressure off the day.
Public demonstration taught me something different: how much you learn by watching someone react to your work in real time, with no chance to explain it first. A lot of my research experience since has involved trying to get that same kind of unscripted, first-contact reaction as early as possible, because it tells you things a script or a survey won't.
This is the project that first showed me how a signal I designed can be read completely differently by the person receiving it, and that gap between intention and interpretation is something I now check for by default in any multimodal or AI-driven interface I work on.