HiDream-O1-Embodied tops RoboColiseum’s Robustness leaderboard, highlighting the model’s ability to maintain stable performance under complex real-world conditions

BEIJING, CHINA - Media OutReach Newswire - 8 September 2026 - HiDream.ai has officially launched HiDream-O1-Embodied, an embodied world model designed to advance physical interaction for embodied intelligence. Built on HiDream.ai's native omni-modal technology strategy, the model enhances robots' physical perception, dynamic prediction, and execution capabilities, enabling more robust interaction with the physical world.

Picture.png

The launch marks an important step in HiDream.ai's broader effort to connect image, video, 3D, and action modalities within a unified architecture. By extending its world model capabilities from understanding and reasoning to action and execution, HiDream.ai is building a closed-loop technical foundation for native omni-modal intelligence.

Alongside its release, HiDream-O1-Embodied made its debut on RoboColiseum, an embodied intelligence model evaluation platform. The model ranked No. 1 on the platform's Robustness leaderboard, achieving an average score of 0.692.

"We believe a complete world model foundation requires three core capabilities: omni-modal representation, causal reasoning, and physical-world modeling — all centered on the ability to express, understand, and generate within the real world," said Ting Yao, CTO of HiDream.ai. "From the beginning, HiDream.ai's native omni-modal world model architecture was designed to support unified representations across modalities, including action. The release of HiDream-O1-Embodied marks a critical milestone in our technology roadmap, as we move from simulating the world to enabling AI to operate in the real world."

HiDream-O1-Embodied Tops RoboColiseum's Robustness Leaderboard with a Score of 0.692

RoboColiseum is a standardized simulation benchmark for embodied intelligence models, designed to provide a multidimensional and reproducible evaluation framework. Through high-fidelity simulation tasks that closely approximate real-robot performance, the platform helps developers assess model strengths and limitations while continuously tracking progress across the field.

Open to universities, research institutions, model developers, and researchers worldwide, RoboColiseum continuously updates its evaluation results with the goal of establishing a reliable benchmark for embodied models.

Built on high-fidelity simulation environments that closely mirror real-world conditions, RoboColiseum evaluates models across four major dimensions: instruction following, spatial understanding, robustness, and general-purpose manipulation. These dimensions are assessed through four capability leaderboards and 78 high-fidelity simulation tasks.

Since entering internal testing, RoboColiseum has attracted dozens of leading models from China and abroad. Among its evaluation dimensions, Robustness is widely regarded as one of the most challenging. It measures a model's stability and generalization under non-ideal conditions by varying backgrounds, lighting, materials, robot initial states, camera positions, and image quality, while also introducing diverse paraphrases of instructions.

In other words, this is not a test conducted in the "greenhouse" of a lab environment. It is designed to evaluate how well a model performs when faced with the kinds of uncertainty, variation, and interference that robots are likely to encounter in the real world.

HiDream-O1-Embodied ranked first on the Robustness leaderboard with a score of 0.692, supported by HiDream.ai's native omni-modal foundation. A native omni-modal world model provides an inherent basis for cross-modal understanding, generation, and action. At the execution level, HiDream-O1-Embodied introduces advances across three core capabilities, enabling more precise instruction understanding, more reliable perception, and stronger resistance to environmental interference.

Language Understanding: Moving Beyond Keyword Matching

Traditional robots often interpret language instructions at the level of keyword matching. A robot may understand "Bring me the cup," but change the phrasing to "Get me a cup" or "Hand me the cup," and it may fail to respond correctly. HiDream-O1-Embodied covers an equivalent instruction space encompassing diverse verbs, sentence structures, and expressions. Rather than being constrained by specific wording, it focuses on the underlying intent. No matter how an instruction is phrased, the model can move beyond the literal wording and accurately identify what the user actually means.

Visual Perception: Multi-View Collaboration for Greater Reliability

In the physical world, a robot's visual input is rarely ideal. Camera positions may shift, calibration accuracy can change over time, and individual visual feeds may be obstructed or disrupted. HiDream-O1-Embodied integrates information from multiple viewpoints, allowing different visual channels to complement one another rather than relying on a single fixed perspective. When part of the visual information becomes inaccurate or temporarily unavailable, the model can still leverage other viewpoints to understand the scene, assess the task, and continue execution. This transforms the system from one where "a single failure causes the entire system to fail" into one where "local limitations do not prevent the system from operating as a whole."

High Fault Tolerance: Learning to Execute Reliably in an Imperfect World

Most models are trained primarily on "perfect" data — clear images, complete frames, and standardized viewpoints. The real world, however, rarely provides such ideal conditions. Changes in lighting, image degradation, occlusion, signal fluctuations, and scene variation are all common challenges robots face during real-world operation.

HiDream-O1-Embodied proactively introduces a wide range of non-ideal conditions during training. By repeatedly exposing the model to incomplete, noisy, and unstable information, the system learns to make reliable decisions based on limited visual cues.

This approach means the model is not optimized solely for peak performance under ideal conditions. Instead, it is designed to maintain stable task execution in complex, dynamic environments. Its fault tolerance is not limited to any single type of visual anomaly. When faced with changes in lighting, object appearance, scene layout, or visual quality, the model can make more flexible use of available information and reduce the impact of environmental variation on execution.

For HiDream-O1-Embodied, the real measure of capability is not simply whether it performs well when everything is clear, but whether it can continue to complete tasks reliably when conditions are far from ideal.

Model + Data: Building a "Real-World Foundation + Generative Augmentation" Data Production Paradigm

The ability to perform reliably under imperfect conditions does not emerge by chance. It points to a fundamental challenge in embodied intelligence: the cognitive boundaries of a model are largely shaped by the data it can access.

High-quality embodied data remains one of the scarcest and most decisive resources in the field. HiDream.ai's dual-driven "model + data" strategy is a key factor behind the performance of HiDream-O1-Embodied on the Robustness leaderboard.

The core breakthrough lies in making data production an integral part of model iteration. To achieve this, HiDream.ai has developed a "real-world foundation + generative augmentation" data production paradigm. Instead of passively consuming existing data, the model actively participates in creating and refining the data it needs to improve.

A collaboration with Noitom provides a representative example. Using Noitom's high-precision human motion-capture data as the real-world foundation, HiDream.ai leverages its native omni-modal capabilities to achieve 100x-scale data augmentation and refinement.

Starting from a single real-world motion sample, the model can generate physically consistent video variations by changing variables such as background environments, lighting conditions, object forms, and scene configurations. This produces a large and diverse set of training samples while preserving underlying physical constraints.

The key to this mechanism is that the model acts as both the "student" and the "teacher." It generates targeted training samples based on the capabilities it needs to improve, creating a growth flywheel in which data and models continuously reinforce one another. This data-model flywheel helps HiDream-O1-Embodied maintain exceptional robustness when confronted with severe disturbances and real-world variability.

HiDream.ai's Native Omni-Modal World Model Matrix Continues to Take Shape

Less than a month ago, HiDream.ai launched HiDream-O1-World, an interactive world model that took the top spot on the Navi sub-leaderboard of WBench, an interactive video world model benchmark, with an average score of 80.9.

Interactive world models address "understanding and reasoning," enabling AI to develop a comprehensive understanding of space, time, motion, and object relationships in digital environments. Embodied world models, by contrast, address "operation and execution," enabling AI to perform real-world tasks in physical environments. Together, the two form a powerful complement to one another and lay a solid foundation for the development of native omni-modal world models.

As Dr. Tao Mei, Founder and CEO of HiDream.ai, previously noted, the key to the next generation of foundation model competition lies not in improving the capabilities of individual modalities alone, but in moving from single-modal to multimodal intelligence, and ultimately toward natively unified omni-modal intelligence.

From HiDream-O1-Image to HiDream-O1-World and now HiDream-O1-Embodied, HiDream.ai is steadily building a model family spanning vision models, interactive world models, and embodied world models. This not only demonstrates the expanding capabilities of its native omni-modal technology across multiple domains, but also reflects the technology's strong capacity for intrinsic evolution.

At a pivotal moment when advances in world model technology are accelerating, HiDream.ai is continuing to accelerate innovation and push the field forward.


Hashtag: #HiDreamAI

The issuer is solely responsible for the content of this announcement.