Robots are getting good at handing you a coffee, screwing in a bolt, or navigating a factory floor. But what happens when one of them fumbles the job — and then has to smooth things over with a human co-worker? A team of researchers has been probing exactly that awkward moment, and the results are a useful reality check on how far emotional AI can carry a machine.
The study, led by Seung Chan Hong as part of his undergraduate thesis at Monash University in Melbourne, Australia, was published on 18 May in IEEE Robotics and Automation Letters. The core idea: teach a collaborative robot to read human emotions not just from facial expressions, but from the full context of an interaction.
To do that, the team leaned on a vision language model (VLM) — think ChatGPT-style reasoning, but with the ability to interpret visual scenes. They used Gemini 2.5 as the backbone. The insight here is subtle but powerful: a furrowed brow doesn’t automatically mean anger. Someone drumming their fingers or pursing their lips might just be concentrating. Reading the whole scene, rather than freeze-framing a face, changes the interpretation entirely.
The numbers backed up the approach. When volunteers watched videos of robots handing over objects — sometimes successfully, sometimes not — and labeled the emotions on display, the VLM’s readings were compared against a conventional AI system built on standard facial analysis and object tracking. On a similarity scale from 0 (no match) to 1 (perfect match), the traditional system scored 0.77, while the VLM landed at 0.86.
The second experiment is where things get genuinely interesting. The team had 40 volunteers work with a robot running the VLM, then deliberately programmed the robot to make a mistake. The robot responded with one of two apologies: an emotionally adaptive one, tailored to the human’s reaction, or a canned, pre-scripted line.
People noticed the difference. A commanding 31 out of 40 participants preferred the emotionally adaptive apology over the boilerplate version. So far, so encouraging for the future of polite robots.
But here’s the catch. When surveyed, participants made it clear that a smooth apology counted for far less than the robot actually doing its job. After a robot failed a task, trust dropped — no matter how gracefully it said sorry. As Hong puts it, “A personalized apology acts as a social lubricant, but it cannot repair the trust lost by the robot failing its physical task.”
There’s a deeper limitation, too. The VLM matched third-party human observers well, but its accuracy fell off sharply when measured against people’s own self-reported feelings. “While the VLM is a good observer of outward social cues, it isn’t a mind reader,” Hong says.
The takeaway is refreshingly grounded: emotional intelligence is a nice bonus in a robot co-worker, but competence still wins. People might appreciate a machine that reads the room — they just won’t forgive one that drops the ball.