Why robots that can’t communicate naturally won’t be adopted

0
1
Why robots that can’t communicate naturally won’t be adopted


Why robots that can’t communicate naturally won’t be adopted

For humanoid robots to operate in the real world, they’ll need more than a sense of sight. | Source: Adobe Stock

At recent major tech events, humanoid robots were everywhere. Some of them walked, navigated, and manipulated objects with a level of dexterity that would have seemed unrealistic just a few years ago. And yet, despite all of this progress, I find myself consistently underwhelmed when I try to communicate with them at Treble.

Supporting the development of advanced audio and voice technologies has been Treble’s bread and butter for a long time. We’ve already grown used to building powerful audio AI for world-leading products. So, naturally, my perspective is shaped by that. Still, I find it striking how underdeveloped communication and auditory perception remain in robotics.

We’re building astonishing machines that can move like us and, in many cases, see better than us. But these machines can’t interact with us in a natural way yet. In my opinion, this is a central challenge that can determine whether this technology is truly adopted.

In my view, human-machine interaction will become the defining hurdle for widespread acceptance of robots by humans because it heavily impacts trust, safety, efficiency, and ease of use.

The real world is loud, complex and chaotic

Humanoids, and more broadly mobile physical AI, are unlikely to be accepted by people until they can communicate effectively through voice. And not just in controlled, quiet environments, but in the places where real life unfolds.

This includes crowded trade show floors, industrial settings, city streets, and homes filled with noise, movement of sound sources, reverberation, and unpredictability. These are the environments in which humans operate, and they’re precisely the environments where current embodied AI audio and voice systems tend to break down.

At the same time, it’s completely understandable why sound has not been a primary focus. The current competitive landscape in robotics rewards visible, measurable progress.

Locomotion is a clear signal of advancement. Vision-based perception has a mature and powerful ecosystem behind it, with abundant data, well-established models, and scalable training pipelines.

The entire stack, from data collection to simulation, has evolved to support these modalities. Data is abundant, benchmarks are clear, and improvements are easy to demonstrate.

Simulation, in particular, has become a cornerstone of progress. Platforms like NVIDIA Isaac Sim have enabled rapid iteration and large-scale training in ways that were previously impossible. These systems are powerful, well-designed, and aligned with the broader economics of the industry.

They also reveal something important. The environments we use to train intelligent machines are overwhelmingly visual. But, they are, for the most part, silent.

This is not an accident. It reflects a set of rational decisions made under real constraints, including compute limitations, engineering bandwidth, and the need to prioritize what is tractable. It also means that an entire dimension of perception and interaction has been systematically underdeveloped.

Treble examines the evolutionary gap

Looking at this through the lens of human evolution makes the gap even clearer. I originally trained as a biologist. After some unexpected turns, I ended up building audio technology, I still think about these systems from a physiological and evolutionary perspective.

Humans have evolved over millennia to allocate significant energy to processing sensory information. Vision dominates this allocation, accounting for a large portion of the brain’s sensory workload. Hearing, by comparison, consumes less. However, it still represents the second most significant share, roughly in the range of 15% to 20%, depending on context.

It is therefore entirely logical that vision has become the dominant modality in early robotics systems. But, the fact that hearing is the second most energy-intensive sense tells you something important. It reflects how critical sound is to functioning in the real world, and especially within a human environment.

In the brutal calculus of evolution, energy is never wasted. That 15% to 20% allocation isn’t an accident, but a direct result of natural selection optimizing our species to survive, thrive, and prosper on planet Earth. This should be a glaring hint for roboticists. If a biological intelligence needs that much auditory bandwidth just to navigate and survive in the physical world, silicon intelligence will not succeed without it.

Focusing only on how much energy a system consumes misses the more important question: What is the system actually used for? Hearing plays a fundamentally different role than vision. It is central to how we interpret intent, maintain awareness beyond our field of view, and most importantly, communicate.

Through sound, we infer whether something is approaching or moving away, whether a voice is calm or hostile, and whether an environment is safe or unpredictable. It functions as an always-on layer of perception that complements vision in critical ways.

More than that, it underpins human communication. Speech is not simply a sequence of words. It is a complex exchange of timing, rhythm, micro-intonation, and emotional signaling. It is inherently dynamic and remarkably robust.

Humans can communicate effectively in environments that are noisy, reverberant, and chaotic, extracting meaning from sound with a level of resilience that current systems still struggle to match. For robots, this capability is essential.



SITE AD for the 2026 RoboBusiness call for speakers
Register now and save on your pass to RoboBusiness 2026

Is there an even higher standard for audio?

Humans benefit from shared biology and deeply ingrained social patterns. We compensate for imperfections in one another’s communication because we intuitively understand the system we’re part of. Robots don’t have this advantage. As a result, they’re held to a different standard, particularly in the early stages of adoption.

A robot that moves slightly imperfectly can still be perceived as functional. A robot that communicates poorly, for example, one that mishears, responds out of sync, or fails to operate in real-world acoustic conditions, quickly becomes frustrating or even unsettling.

The issue isn’t just technical performance. It’s the breakdown of trust. This is why audio becomes disproportionately important in the context of adoption. We aren’t evaluating robots solely based on their capabilities, but on how it feels to interact with them. And interaction, at its core, is deeply auditory.

The reason this hasn’t been solved isn’t a lack of awareness, but a lack of infrastructure. High-quality audio data is difficult to obtain and even harder to scale. Unlike visual data, it cannot simply be scraped and labeled at scale without losing critical context.

Spatial relationships, environmental acoustics, and device-specific characteristics all play a significant role in how sound is perceived. Capturing and annotating this information in real-world settings is complex and expensive.

Simulation offers a path forward, but it introduces its own challenges. Sound is governed by wave physics, which makes it highly sensitive to geometry, materials, and environmental conditions. Small changes in a scene can lead to large differences in perception. This makes accurate simulation computationally demanding and difficult to approximate.

As a result, many current systems rely on simplified or non-physical models, which limits their ability to generalize. This is a key reason why audio has lagged behind other modalities in the development of synthetic training data.

Treble works to shift the paradigm

However, this situation is beginning to change. Treble is part of a growing shift toward physically accurate simulation of sound as a foundation for training audio systems. Instead of relying on approximations or limited recorded datasets, it’s now becoming possible to generate large-scale, high-fidelity acoustic data that reflects real-world conditions. This includes environmental acoustics, device-specific behavior, and spatial perception.

At Treble, this has been a core focus. We’ve built a platform specifically for generating physically accurate acoustic data with an extremely low sim-to-real gap. It’s already being used by several of the leading technology companies in the world to develop and train audio systems. What was a bottleneck is becoming a scalable component of the AI stack.

For robotics, this represents an inflection point. The industry has, so far, made the right trade-offs. It has focused on vision and locomotion because those were the areas where progress was most achievable. But as these capabilities mature, they will cease to be differentiators.

Robots will increasingly converge on similar levels of visual perception and physical capability. At that point, the limiting factor will shift. It will no longer be about whether a robot can perceive the world, but whether it can interact within it in a way that aligns with human expectations. And that interaction will depend heavily on sound.

I believe robots that ultimately succeed will not only be defined by how well they see or how well they move, but to a large extent on by how natural and intuitive they are to communicate with. They’ll be able to operate in real acoustic environments, understand speech robustly, respond with appropriate timing and tone, and convey intent in ways that humans instinctively understand.

Treble is actively looking to work with teams in robotics and physical AI who recognize this shift and want to push in this direction. There’s an opportunity to build a new class of systems, machines that do not just function in human environments, but genuinely integrate into them.

The difference between a robot that works and a robot that is accepted will not be subtle. It will largely come down to how it communicates. And that, fundamentally, is auditory.

Gunnar Pétur Hauksson is the co-founder and chief commercial officer at Treble TechnologiesAbout the author

Gunnar Pétur Hauksson is the co-founder and chief commercial officer at Treble Technologies. He has over a decade experience in sales and marketing management, entrepreneurship, and business development.

Reykjavík, Iceland-based Treble is building high-fidelity acoustic simulation infrastructure that enables sound to move from simulation to the real world — whether in buildings, devices, vehicles, or intelligent machines.

The post Why robots that can’t communicate naturally won’t be adopted appeared first on The Robot Report.