The Art Behind
Human-Robot Interaction
Aug 01, 2026
While watching Boston Dynamics’ design team discuss human–robot interaction, I found myself returning to a familiar tension: the fundamentals of good design remain remarkably consistent across products, yet applying them to systems with unresolved complexity is remarkably challenging.
Principles such as clarity, feedback, predictability, and recovery still apply. Yet when an interface can move through physical space, carry weight, approach people, and act with partial autonomy, those principles raise a different class of questions. How should a robot communicate what it perceives? How can its movements make its next action understandable? And where should expressive interaction end and certifiable safety begin?
The webinar offered several useful HRI insights, but more importantly, it surfaced questions that extend beyond Atlas—and may become foundational to how we design robots that people can understand, trust, and work alongside.
—————————————
The Abstract Head as Interaction Anchor
Whether to give Atlas a head at all was itself a huge discussion in the team — treated as an opportunity rather than a given. Obviously there are advantage to include head; it can handle multiple functions. On the other hand, it has downside to have a head as a robot; it’s less cost-effective and reliability.
The team chose to include one, but deliberately without a human-like face. Instead, the head carries six jobs:
Sensing space and compute hardware
Clarifying spatial intent with its head orientation
Providing acknowledgement cue as a small nod
Communicates status through a light ring
Most of a robot’s perceived character may come from its bigger cues such as head orientation and intelligently timed movement. When these larger cues are convincing, people can mentally fill in the rest. The team described that a detailed face would be icing on a cake: a secondary layer rather than the foundation of the interaction.
I interpret this to mean that detailed facial expression is optional once the most fundamental interaction cues—such as attention, timing, state and intent—are working convincingly. This is my interpretation, rather than a direct product principle stated by the team.
Questions that arise
How much can an abstract head—through orientation, movement, timing, lighting patterns, and voice—communicate attention, confidence, hesitation, or reassurance?
Could simplifying the face reduce the tendency to over-attribute human emotion or intelligence, while making the robot’s actual attention, intent, and state more legible?
—————————————
Embracing the Mechanical Form
Atlas is positioned as an industrial tool, not a companion — and that framing gives the design permission to depart from human anatomy (an abstract light ring instead of two eyes, for example). Every product that interacts with people has to find its own ratio between mechanical honesty and understandable behavior — Atlas is just a visible version of a design problem that shows up everywhere.
On the uncanny valley specifically: Atlas's movement isn't constrained by human musculature, so it can do things people can't — use extra degrees of freedom with head, arms and legs, and move in ways that read as unfamiliar.
The team's conclusion: it’s ok that people feel the movement unfamiliar, but it needs to become learnable. People adapt to new movement patterns once the purpose behind them is clear. That reframes the real design question: is it worth chasing the uncanny valley at all, or is that effort better spent on safety and efficiency?
Questions that arise:
How to set an acceptable design framework to balance between mechanical honesty and understandable behavior?
Rather than aiming for immediate familiarity, should robots establish consistent and learnable movement languages over time?
Could clear functional principles—such as safety, predictability, and purposefulness—become the foundation of a distinctive and coherent aesthetic language?
—————————————
Communicating Robot’s Intent
Two ways to signal what a robot is about to do:
Conventional UI — explicit lights and sounds
Embodied cues — anticipatory motion built into the robot's own movement vocabulary
From the discussion, I think the team leans toward the latter. A robot's gesture and body function as an interface, previewing its next action. A "thinking" or preparing pose — borrowed from animation and game design — keeps a pause from reading as a failure. This sequencing (cue → anticipation → action) is what builds trust over repeated interactions: once people can predict the robot, they know what to do around it, and that's when they feel at ease.
Questions arise:
Can HRI borrow from established human-machine conventions—such as traffic lights, machine alerts, and flashing and pulsing lights—that people already understand?
How might these conventions be embodied into the robot motion and behavior so the users interpret them without additional training?
—————————————
Designing for Safe, Retaskable
Operation in the Real World
Atlas is designed around flexibility: it should be able to work near people and be retasked across different factory workflows without requiring fixed safety cages. From an HRI and UX perspective, retasking is not limited to issuing a new command. It includes helping users direct the robot to a different location, assign or queue work, and teach it new manipulation skills. The webinar described several interaction channels for these activities, including a handheld controller, conversational commands, and Orbit. Because factories can be extremely noisy, the appropriate channel may depend on the environment and the task rather than relying on voice alone.
Working near people also requires the robot to make its perception and safety response legible. When someone approaches within roughly two meters, Atlas is expected to turn toward the person, blink its light, and likely pause—communicating that the person has been detected. If the person moves within roughly one meter, Atlas may begin rendering itself safe by lowering its body, putting down the object it is carrying, or remaining stationary, depending on its current activity. These embodied cues help people understand what the robot has perceived and what it is likely to do next; however, the underlying safety-critical behavior must be reliable and certifiable rather than merely reassuring in appearance.
The same principle applies during system disruptions. Safety and control of the robot’s physical body are intended to remain onboard, so losing Wi-Fi should not cause Atlas to become physically uncontrolled. It may be unable to proceed with the next manipulation task, but it should still be able to return safely to its dock and report the connectivity problem through a wired connection. From a UX perspective, this means recovery is not only a technical fallback: the robot must also communicate what was interrupted, what remains safe, and what the user should expect next.
Questions that arise
Which interaction channel—physical controller, voice, embodied cues, or Orbit—should take priority at different distances, noise levels, and levels of urgency?
What should or should not the robot communicate before, during, and after an interruption so that users understand what happened, whether it is safe, and what it will do next?
When users teach Atlas a new task, how should the system communicate what it has learned, where that learning applies, and when additional demonstrations are needed?
How should connectivity loss, unexpected human intervention, interrupted tasks, and recovery be designed as first-class UX flows rather than edge cases?