Caroline Bishop Aug 04, 2026 16:39
NVIDIA's Cosmos 3 boosts World Action Models (WAMs), enabling robotics to evolve from vision-based to physics-driven decision-making.
NVIDIA is reshaping the robotics landscape with its Cosmos 3 platform, a foundation model designed for World Action Models (WAMs). WAMs represent a cutting-edge shift in robotics, combining physical world modeling and action prediction to enable robots to adapt and operate in dynamic environments. This marks a departure from traditional vision-language-action (VLA) models, which excel at semantic understanding but often struggle with physical generalization.
Unlike VLAs, which map observations and instructions to actions, WAMs are built on video-based world models that predict how physical scenes evolve. This approach equips robots with the ability to anticipate outcomes, such as how objects will move or interact, making them significantly more versatile. NVIDIA's Cosmos 3, launched as an open foundation in 2026, provides the backbone for this transformation.
What Makes Cosmos 3 Stand Out?
Cosmos 3 is based on NVIDIA’s Mixture-of-Transformers (MoT) architecture, enabling multimodal learning across video, text, audio, and action. With an expansive training dataset of 348 million videos and 8 million action samples, Cosmos 3 delivers a physics-driven understanding of real-world dynamics. The model comes in three configurations—Edge (4B parameters), Nano (16B), and Super (64B)—catering to both on-device and workstation deployments.
For robotics developers, Cosmos 3 offers practical benefits. It reduces the amount of task-specific data needed to train robots, enables better generalization to unseen environments, and accelerates adaptation to new robotic platforms, such as different arms or grippers. A recent study revealed that policies trained on Cosmos 3's omni-model architecture outperformed traditional methods, improving task success rates from 28.1% to 36.8% on the RoboLab benchmark.
Why WAMs Are Gaining Traction
World Action Models are rapidly gaining attention in the robotics and AI communities. Unlike VLAs, which rely heavily on predefined demonstrations, WAMs leverage general physical principles. This makes data collection cheaper and more efficient, as any interaction data—such as objects being pushed, grasped, or dropped—becomes valuable training material.
Moreover, WAMs excel in open-world scenarios where semantic understanding alone falls short. For example, a robot equipped with a WAM can predict how a towel will fold or how liquid will pour, even in situations it hasn’t encountered during training. This level of adaptability is critical for real-world applications, from autonomous manufacturing to home assistance.
Deployment and Accessibility
NVIDIA has made Cosmos 3 widely accessible through platforms like Hugging Face and GitHub. Developers can post-train WAMs using their own datasets and deploy them across various hardware setups. For instance, the Edge model can run real-time robotic control tasks on NVIDIA Jetson modules, while the Nano model supports more computationally intensive workloads on RTX GPUs.
By releasing Cosmos 3 under a permissive license for commercial use, NVIDIA is encouraging widespread adoption and innovation. Upcoming events, such as the Cosmos Labs livestream on August 13, aim to further engage the community and demonstrate the platform's capabilities.
Looking Ahead
WAMs represent a paradigm shift in robotics, moving from mimicking human demonstrations to understanding and predicting physical dynamics. NVIDIA’s Cosmos 3 is at the forefront of this evolution, offering a robust foundation for building adaptable, efficient robot policies. For developers and researchers, the opportunity to experiment with Cosmos 3 could unlock new possibilities in embodied AI.
With open tools, state-of-the-art models, and a growing ecosystem, WAMs are poised to redefine what robots can achieve in the real world.
Image source: Shutterstock

By Blockchain News | Created at 2026-08-05 03:46:30 | Updated at 2026-08-05 23:32:27
1 day ago







