New AI System Cuts Robot Response Delays by Nearly 12 Times Without Extra Computing Power
Researchers have developed a new artificial intelligence (AI) system that could make robots controlled by vision-language-action models respond more quickly. The system, called VLASH, reduces maximum response latency by up to 11.8 times without requiring additional computing power.
Many advanced robots use Vision Language Action (VLA) models to interpret camera images, understand human instructions and convert decisions into physical movements. However, these robots often move in a stop-and-go pattern: they complete one action, pause while the AI calculates the next instruction, and then continue moving.
These delays can limit the usefulness of VLA-controlled robots in situations that require continuous movement, rapid reactions and real-time interaction.
How VLASH Helps Robots React Faster
VLASH is designed to remove the waiting period between robot actions. While the robot carries out its current movement, the system predicts where the robot will finish and prepares the next action in advance.
This asynchronous approach allows the robot to plan continuously instead of waiting for the current instruction sequence to end. In experiments, robots using VLASH completed some tasks between 1.5 and 2 times faster while maintaining most or all of their accuracy.
Depending on the hardware, the system reduced the robot’s maximum response latency by as much as 11.8 times. An earlier version of the research reported speed improvements of more than 30 times, but those tests used longer action sequences and slower computing hardware.
The researchers said the newer 11.8-fold result may provide a more realistic estimate of VLASH’s current performance because the latest experiments used shorter, more general action sequences and a more powerful graphics processing unit (GPU).
AI Gives Robots a More Efficient Planning System
VLA models act as the brains of many modern robots. They combine visual information from cameras with natural-language instructions and data about the robot’s current position. The model then generates a sequence of movements.
VLASH improves this process by using the robot’s scheduled movements to estimate its future state. The AI can then prepare the next action based on that predicted position before the current action has finished.
Unlike a full world model, VLASH does not attempt to predict everything that may happen in the surrounding environment. Modeling the entire world would require additional computation and could slow down the robot’s responses.
Instead, VLASH focuses on predicting the robot itself. The researchers said this narrower approach could eventually work alongside world models, combining fast robot control with broader environmental awareness.
The concept is similar to planning the next step before the current step is complete. Rather than waiting for the robot to finish moving before deciding what comes next, VLASH prepares the next instruction in parallel.
Robots Tested on Sorting, Stacking and Fast-Reaction Tasks
Researchers tested two VLA models on two robotic platforms using a laptop equipped with an Nvidia RTX 5090 GPU. The robots performed pick-and-place, stacking and sorting tasks, with 20 trials conducted for each method.
The team also tested VLASH on fast-reaction tasks, including table tennis and whack-a-mole. In these experiments, targets could move while the robot was calculating its next action.
VLASH does not attempt to predict the movements of those targets. Instead, it processes new visual observations between 15 and 30 times per second, allowing the robot to respond quickly when objects change position.
Because VLASH estimates the robot’s future position from movements that have already been scheduled, it does not require a separate predictive model or additional inference steps during operation.
Faster Robot Training Could Improve Real-World Performance
The researchers also reorganized existing training data to make fine-tuning more efficient. In one benchmark, each training step was 3.26 times faster while reaching comparable accuracy.
However, the researchers cautioned that this does not mean the entire training process is 3.26 times faster or cheaper. Overall performance depends on the AI model, hardware, dataset and total number of training steps.
Faster Robot Reactions Do Not Automatically Mean Safer Robots
Although VLASH can reduce processing delays, faster responses do not necessarily make robots safer around people.
Roshni Lulla, co-founder and chief research officer at the Humane Robotics Institute, was not involved in the study. She said that faster reactions are useful but do not, by themselves, demonstrate improved robot safety.
“Reacting faster is a big advantage, but I don’t see any evidence that it would improve safety,” Lulla said.
The researchers believe VLASH could initially benefit robotic manipulation tasks that involve continuous movement, quick reactions and changing plans. Potential applications include manufacturing, warehouse automation and other environments where robots must handle moving objects.
However, testing robots around people will be essential. “The real question is how does this change the performance of robots in human-centered environments?” Lulla said. “If the state of the environment can be changed by another agent, I’d like to see tests done with a human in the room.”
More Testing Is Needed Before Search-and-Rescue Use
Search-and-rescue applications remain more distant. VLASH has not yet been tested in low light, smoke, unstable terrain or environments where cameras may be damaged and communications unreliable.
Performance in those conditions will depend largely on the underlying VLA model. The researchers said future studies must test VLASH with more robots, AI models, tasks and environments, as well as longer experiments involving unexpected disturbances.
VLASH could eventually be combined with more advanced world models capable of anticipating wider changes in the environment. For now, however, the system’s main achievement is reducing the time robots spend waiting for their next instruction.
Lulla said the findings support a relatively narrow conclusion: VLASH appears to reduce processing bottlenecks while preserving the robot’s existing capabilities. She added that the results do not yet demonstrate meaningful progress in robot safety or intelligence.
“Rather than a single demonstration, we will seek consistent outcomes across many robots, tasks, environments and long-term trials, with appropriate reliability and safety validation,” the study authors said.
VLASH has shown that asynchronous inference can help VLA-controlled robots respond more smoothly and quickly. Whether those gains continue outside controlled laboratory conditions remains an important question for future research.
Source: www.livescience.com


