Shanghai AI Lab and OpenDataLab open-sourced InternW0-Δ, a unified World Action Model (WAM) pretrained on over 20,000 hours of curated robotic demonstrations, UMI data, and egocentric videos. Built on a Mixture-of-Transformers (MoT) architecture that binds visual predictive dynamics with robot actions, it introduces 'Causal Imprint' to inject future representations directly into the action expert without requiring costly future-video rollouts during inference.

Key Takeaways

  • ✓20,000+ Hours Open-Source Corpus: Unifies robot manipulation, UMI dexterous data, and egocentric human demonstrations into the largest publicly released robotic training corpus.
  • ✓Mixture-of-Transformers (MoT): Coordinates video dynamics experts and action generation experts under semantic guidance from a frozen VLM, with 4D spatial priors distilled during pretraining.
  • ✓Causal Imprint Without Future Rollout: Eliminates the critical latency bottleneck of traditional WAMs that require rendering future video frames at test time, providing predictive latent cues directly for real-time control.
  • ✓Complete Open Source: Training code, model weights, data processing pipelines, and datasets released at internrobotics.github.io/InternW0-Delta/.
🧭

Finished reading? Explore benchmark rankings & pricing

Real-world SWE-bench scores & $20/mo vs API cost break-even calculator

🔬

In-Depth Technical Analysis

核心背景与行业痛点 / Background & Pain Points Embodied robotics has long struggled to bridge large-scale visual dynamics with precise motor action. While World Action Models (WAMs) show great promise by jointly predicting visual evolution and motor commands, existing methods suffer from severely fragmented training data and an intolerable inference latency bottleneck: traditional WAMs require multi-second future video rollouts during inference before emitting control actions, precluding high-frequency real-world deployment. ### 架构亮点与底层机制 / Architectural Highlights Shanghai AI Lab and OpenDataLab released InternW0-Δ, introducing major systems breakthroughs: 1. 20,000+ Hours Open Corpus: Assembles and harmonizes the largest publicly available robotic dataset, unifying arm demonstrations, UMI data, and egocentric videos into a standardized state-action format; 2. Mixture-of-Transformers (MoT) Architecture: Fuses video experts and action generation experts within a unified cross-attention space steered by a frozen vision-language model, with 4D physical priors injected via training-time distillation; 3. Causal Imprint Mechanism: Employs future-conditioned supervision strictly during training, embedding anticipated visual and physical dynamics into a latent representation directly consumed by the action network. At test time, no future video rollout is required, enabling pure real-time policy execution. ### 权威 Benchmark 与实测跑分对比 / Benchmark & Evaluation - Simulation Environments: Outperforms state-of-the-art baselines by +18.6% in average success across CALVIN and RoboSuite long-horizon benchmarks; - Physical Robot Generalization: Achieves >86.4% task completion across uncalibrated multi-object table operations and bimanual handovers; - Inference Latency Breakthrough: Slashes decision latency from 1,800–3,200ms (typical of rollout-based WAMs) down to under 45ms, enabling 20Hz+ closed-loop robot control on commodity hardware. ### 开发者实战落地与开箱指南 / Developer Practical Guide - Project Hub: Documentation, weights, and processing pipelines are available at https://internrobotics.github.io/InternW0-Delta/; - Paper Citation: Detailed theoretical formulations are accessible in arXiv preprint 2609.31394; - Hardware Deployment: Packaged with lightweight ROS2 nodes and Python bindings for direct deployment on RTX 4090/A100 workstations.