English

PLUS ULTRAProduct LaunchesMuJoCo WarpMJWarpGoogle DeepMind

Google DeepMind Releases MuJoCo Warp (MJWarp) for Large-Scale GPU-Accelerated Robotics Simulation

PLUS ULTRA by Amenoyomi

Google DeepMind has released MuJoCo Warp (MJWarp), a new backend for large-scale multibody dynamics built on the NVIDIA Warp framework. MJWarp transitions MuJoCo models from traditional CPU-based simulations to a GPU-accelerated regime, enabling the execution of thousands of parallel environments.

While classic MuJoCo is optimized for fast CPU-based simulation and parallelization across CPU cores, MJWarp utilizes NVIDIA Warp—a Python framework for writing high-performance, GPU-accelerated kernels—to handle large batches of independent simulation states on NVIDIA GPUs. This shift addresses the growing demand for high-throughput data generation required by reinforcement learning and large-scale sampling workloads.

MJWarp works by implementing the MuJoCo physics pipeline within NVIDIA Warp, compiling CUDA kernels to advance simulation states. In testing with an SO-101 follower arm, the technology demonstrated the ability to scale from a single environment to as many as 2,048 parallel environments.

NVIDIA Warp provides developers with the ability to write differentiable, high-performance kernels in Python. By leveraging automatic differentiation, the framework allows simulation kernels to be integrated into optimization and training workflows, such as PyTorch or JAX. This capability is particularly relevant for physical AI, where high-fidelity, physics-compliant data is essential for training foundation models.

PLUS ULTRAby Amenoyomi

The core value of MJWarp lies in shifting the focus from single-world latency—how quickly one simulation runs—to aggregate throughput, defined as the total number of world-steps completed per second. In reinforcement learning and large-scale sampling, the volume of experience data generated is more critical than the speed of a single environment. MJWarp maximizes this by advancing thousands of independent simulation states in large batches on the GPU, ensuring the hardware has sufficient parallel work to optimize data generation.

This throughput is enabled by the NVIDIA Warp framework's use of the Single-Instruction, Multiple-Threads (SIMT) paradigm, which differs fundamentally from tensor-based frameworks. While tensor frameworks express computation as operations on N-dimensional arrays and rely on Boolean masks to handle conditional logic, Warp kernels allow each GPU thread to branch or exit independently. This allows physics engines to implement complex control flows, such as selective updates and early-outs, naturally and without the computational waste associated with masking.

Furthermore, Warp introduces differentiability to the simulation layer through native support for automatic differentiation. By recording kernel launches on a computational tape and replaying them in reverse, the system can compute exact gradients of a loss function with respect to simulation inputs. This allows the physics engine to be integrated directly into the computational graphs of ML frameworks like PyTorch or JAX, enabling end-to-end optimization and more efficient training of physical AI models.

These architectural choices result in significant performance gains over traditional tensor-based alternatives. In computational fluid dynamics, Warp has demonstrated speeds up to 8 times faster than JAX on a single A100 GPU while utilizing significantly less memory. For multibody dynamics, MJWarp achieves speedups between 252x and 475x over JAX on comparable hardware by leveraging sparse matrix operations and speculative execution to optimize compute dispatch.

Sources

  1. How to Use NVIDIA Warp and MjWarp to Accelerate Robotics Simulation and Learning Workflows (Hugging Face Blog, 2026-09-23)
  2. NVIDIA Developer Blog