← Blog

Toward Household Autonomy for Humanoid Robots

We demonstrate our humanoid robot performing everyday household tasks, including watering plants, wiping surfaces, disposing of trash, and interacting with objects in human environments. This demo shows that humanoid robots are not limited to simple dance-like movements or entertainment-oriented behaviors. They have the potential to perform complex, practical household work in real-world settings.

Starting from humanoid robots, our goal is to build a general-purpose model that can perform meaningful real-world tasks in unstructured environments. Our focus is building real-world robot intelligence — systems that can learn, reason, and reliably translate decisions into physical actions.

General Robot Intelligence Begins with Humanoid Mobile Manipulation

Our starting point is a simple but fundamental belief: humanoid robots are the most natural starting point for general robot intelligence.

We are not pursuing humanoids as a narrow platform for motion or manipulation. Instead, we view them as a general-purpose embodiment for physical intelligence — a system capable of expressing a wide range of behaviors in environments designed for humans.

The key reason is that human environments are inherently structured around the human body. Tasks in the real world are not isolated manipulation problems; they require continuous movement, navigation, tool use, and interaction across diverse spatial contexts. This makes mobile manipulation — the coupling of locomotion and dexterous interaction — a core capability for real-world intelligence.

Humanoid robots are uniquely positioned for this setting. They can operate directly within existing human infrastructure without requiring environmental redesign, enabling them to serve as a general interface between intelligence and the physical world.

More importantly, humanoids provide a unified substrate for learning transferable skills. If general mobile manipulation can be achieved on humanoid embodiments, the resulting intelligence is not tied to a single hardware design, but can be extended across different robotic forms with varying levels of mobility and dexterity.

Building Efficient Mobile Manipulation

Reliable mobile manipulation requires both scalable data and a unified intelligence architecture. However, a key bottleneck in existing approaches lies in how data is collected.

Traditional robot learning pipelines rely heavily on teleoperation data. While effective in controlled settings, this approach does not scale well and introduces fundamental limitations. Teleoperation is tightly coupled to specific hardware embodiments, making it difficult to transfer across different robot morphologies. More critically, collecting high-quality humanoid teleoperation data is expensive and unstable. It often suffers from latency, motion inconsistency, and unnatural movement patterns, which become particularly problematic in long-horizon mobile manipulation tasks.

To address these limitations, we shift the data paradigm toward human-centric demonstrations as the primary learning source.

Our approach leverages large-scale human data, including motion capture and egocentric demonstrations using hand-held tools such as UMI grippers. Compared to teleoperation, human-centric data is significantly more scalable and naturally embodiment-independent, capturing high-quality motion priors directly from human behavior.

We then combine this with large-scale simulation-based reinforcement learning and real-world imitation learning, enabling the system to refine human-derived behaviors into stable, precise, and physically consistent policies suitable for real-world deployment.

Toward a Unified Future of Robot Intelligence

The challenge of robotics is not simply making robots move. It is enabling robots to perform useful work reliably in environments built for humans.

Our focus is solving the core problems behind practical deployment: scalable learning, transferable intelligence, and robust physical execution.

The future of robotics will not come from systems that only demonstrate impressive capabilities in controlled settings. It will come from systems that can learn, adapt, and operate reliably in everyday environments. We will extend humanoid mobile manipulation to a broader range of embodiments, including dual-arm systems, mobile manipulators, and future robotic platforms.

Our vision is:

One intelligence. Many embodiments. Real-world autonomy.

By combining scalable human-centered data with a unified robot intelligence architecture, we aim to bridge the gap between demonstration and deployment — toward a single system that generalizes across robot bodies and operates in the real world.