From Capability to Efficiency
Industrial robots have long demonstrated that machines can operate with exceptional speed, precision, and consistency. But that efficiency is largely built on highly structured environments—fixed objects, fixed positions, fixed trajectories, and predefined processes.
Physical AI offers a different path. By combining perception, understanding, and control, robots can adapt to changing environments and handle more flexible and complex tasks that are difficult to solve with traditional automation. As Physical AI continues to advance, robots are becoming capable of performing an increasingly wide range of real-world manipulation tasks.
But as capability improves, a new challenge becomes increasingly important: efficiency. Many embodied AI systems can now complete complex tasks, but often at only a fraction of human speed. For a demo, completing the task may be enough; for real-world deployment, cycle time, throughput, and productivity matter just as much as capability. A robot that can perform a task but takes five or ten times longer than a human is still difficult to integrate into real workflows. This is the gap Mind-1 is designed to close: bringing the flexibility of Physical AI towards human-level execution efficiency.
Why Can’t Robots Simply “Move Faster”?
Making a robot complete a task faster is not as simple as increasing motor speed or replaying the same trajectory at a higher rate. As execution speed increases, many problems that are barely noticeable at low speeds become significantly amplified.
Several factors make high-speed manipulation fundamentally challenging:
- The speed distribution of training data. Much of today’s robot manipulation data comes from teleoperation, which is often significantly slower than the natural speed at which humans perform the same task. Models learn not only how to complete a task, but also the motion tempo and action distribution of their demonstrations. Slow demonstrations naturally lead to slower and more conservative policies.
- Inference latency. Every step from sensing and data processing to model inference and action execution introduces delay. At an end-effector speed of 3 m/s, just 10 ms of latency corresponds to approximately 3 cm of movement—already enough to affect precise grasping, insertion, or contact-rich manipulation.
- Intent consistency. During execution, the model continuously receives new observations and generates new action chunks. If consecutive inference steps produce inconsistent trajectory intents, the robot must repeatedly correct its motion. At low speeds, these may be minor adjustments; at high speeds, they can become abrupt direction changes, repeated corrections, or end-effector oscillation.
- Control and sensor synchronization. High-speed motion places much greater demands on trajectory tracking and state estimation. Small tracking errors can accumulate rapidly, while even minor timing mismatches between sensors can cause the model to act on an inaccurate representation of the current state.
There is also a more fundamental challenge: high-speed manipulation is not simply low-speed manipulation played back faster. When humans perform a task quickly, the strategy, trajectory, transitions between actions, and interactions with objects all change. Simply increasing action speed also changes contact dynamics—even along the same geometric trajectory, objects may move and respond very differently at higher speeds, pushing the robot into states that were never represented in the training data.
Reaching human-level execution efficiency is therefore not a single-point optimization problem. It requires optimizing the entire system—from training data and model behavior to inference infrastructure, sensor synchronization, and low-level control.
Mind-1: A New Framework for Human-Level Efficiency
Achieving human-level efficiency requires more than making individual components faster. Mind-1 is a new Physical AI framework designed around high-speed manipulation from the ground up, spanning the entire execution stack—from how the model learns from human behavior, to how it generates and executes motion, and how the perception-action loop runs in real time.
- Learning from human-speed data. Mind-1 is trained on human-centric manipulation data captured directly from people performing tasks at natural working speeds, rather than relying primarily on slow-speed demonstrations. To support this, we upgraded our human-centric data collection system to capture natural human manipulation at scale with 1 mm data accuracy, preserving not only what actions are performed, but also their natural speed, motion patterns, and transitions. We also developed an automated data quality and filtering pipeline to identify demonstrations that are both efficient and suitable for robot learning.
- A model and control architecture built for high-speed manipulation. Human-speed data alone is not enough. Mind-1 uses a hierarchical architecture designed for high-speed manipulation, where the high-level multimodal model generates temporally consistent trajectory intent and the lower-level system translates that intent into continuous, fast, and smooth robot motion. A key focus is intent consistency: as new observations arrive and new action chunks are generated, the model maintains consistent trajectory intent across consecutive inference steps, reducing unnecessary corrections and oscillation at high speed. At the execution layer, a new whole-body controller is designed for high-speed trajectory tracking, maintaining accuracy and stability while suppressing high-frequency noise and end-effector oscillation.
- A low-latency inference stack. We significantly redesigned Mind-1’s inference engine, using FlashRT as the inference backend and rebuilding key parts of the engine with hand-written and explicitly adapted kernels tailored to our model architecture. This reduced average model inference latency from 82 ms to 32 ms, a 61% reduction. Beyond model inference itself, we optimized the entire observation-to-action pipeline, from the moment the camera begins exposure to the moment the corresponding action is executed on the robot, reducing end-to-end latency from 153 ms to 77.5 ms. We also tightly synchronized observations across different sensors, keeping the average temporal alignment error below 5 ms. Together, faster inference, lower end-to-end latency, and tighter sensor synchronization allow Mind-1 to make decisions based on a more recent and temporally consistent representation of the physical world, which becomes increasingly critical as robot motion approaches human speed.
Together, these components form the foundation of Mind-1. Rather than accelerating an existing policy after training, Mind-1 is designed from the ground up around one goal: enabling Physical AI to perform complex manipulation at human-level efficiency.
Across the logistics and everyday manipulation tasks we have evaluated, Mind-1 achieves task cycle times comparable to human performance, representing a significant improvement over Mind-0. On some tasks, including packaging and parcel sorting, Mind-1 completes the same task even faster than humans.
| Task | Human | Mind-0 | Mind-1 |
|---|---|---|---|
| Packaging | 5s | 10s | 3s |
| Parcel Sorting | 4s | 6s | 3s |
| Folding | 15s | 80s | 25s |
Rather than measuring speed through maximum joint velocity or isolated motion segments, we focus on end-to-end task cycle time—the time required to complete the entire task from start to finish. This is the metric that ultimately matters in real workflows, where productivity depends on how much useful work a robot can complete over time.
High-Speed Capability Across Different Robot Embodiments
Mind-1’s high-speed capability is not an optimization designed for a single robot. In this demo, we deploy Mind-1 across three distinct robot embodiments: a dual-arm robot, a wheeled dual-arm robot, and a humanoid robot. All three platforms are powered by Mind-1.
Each embodiment has its own morphology, workspace, dynamics, and control system. If high-speed execution only works on one hardware platform, it remains closer to a robot-specific system optimization than a generalizable form of intelligence. We want to test something different: Can the ability to perform tasks efficiently transfer together with the model across different robot embodiments?
This means that what we aim to transfer is not a fixed trajectory designed for one robot, but the model’s understanding of how to complete a task efficiently. MindOn’s goal is not to build a separate specialized intelligence stack for every platform, but to build Physical AI that works across different embodiments while preserving the execution efficiency required in the real world.
From Capability to Real Deployment
Over the past few years, Physical AI has demonstrated that robots can perform increasingly complex tasks. But when we begin thinking seriously about deployment, the question changes. It is no longer only “Can the robot do it?”, but “Can it do it reliably, at the speed required by the real world?”
Mind-1 is our answer to that question. Through human-speed training data, a model architecture optimized for high-speed execution, a redesigned whole-body controller, and a low-latency inference stack with strict temporal synchronization, we have validated a technical path toward significantly improving real-world robot execution efficiency.
Achieving human-level task cycle time does not mean that every challenge of real-world deployment has been solved. Long-term reliability, failure recovery, safety, integration with existing workflows, and ultimately deployment economics all need to be validated further in real operating environments. That is exactly where Mind-1 is going next.
We are shifting our focus from “What new demo can the robot perform?” to “Where can the robot create real productivity?” The next destination for Mind-1 is not another demo. It is real-world deployment.