Robot skills are learned in simulation first
Before a robot arm ever touches a real cube, its skills are usually rehearsed in simulation. A simulated copy of the robot, the table and the objects lets engineers test grasp strategies, control code and eventually learning algorithms without breaking hardware. The catch has always been scale. Classical physics simulation runs on a CPU, and a CPU can only run as many independent simulated scenes as it has cores. If you want thousands of slightly varied practice rooms, you historically needed thousands of cores.
A tutorial published by NVIDIA's team on Hugging Face on September 23, 2026, walks through a workflow that changes that economics for anyone with a modern GPU. It shows how to take an SO-101 follower arm, a popular open robot arm design, from a standard MuJoCo CPU workflow into MuJoCo Warp (MJWarp), a GPU implementation of MuJoCo's physics, and scale it to as many as 2,048 parallel simulated environments on a single GPU.
It is worth being precise about what this tutorial does and does not demonstrate. The authors state the scope directly:
Scope statement
Here, we prepare and scale the simulation environment; we do not train a policy.
Attribution
Johnny Nuñez Cano, Asier Arranz, Rishabh Chadha and Ben Oliveri, NVIDIA authors on Hugging Face, "How to Use NVIDIA Warp and MjWarp to Accelerate Robotics Simulation and Learning Workflows," published September 23, 2026 (introductory section). In other words, this is infrastructure work. No learned robot skill and no simulation-to-reality transfer result is claimed in the article itself.
MuJoCo on a CPU, and why its scale is limited
MuJoCo is a physics engine widely used in robotics research. According to the project's own historical record, MuJoCo was originally created by Emo Todorov at Roboti LLC and was later acquired by Google, where DeepMind maintains it today; the review process for this article identified the original authorship and acquisition, though the MuJoCo documentation itself could not be retrieved in this session. MuJoCo simulates rigid bodies, joints and contacts, and it is fast on a CPU. Its long history means most existing robot models, including open asset collections like the MuJoCo Menagerie, which includes the SO-101 arm, already work with it.
The limitation is architectural. MuJoCo on a CPU advances one world at a time, and parallel sampling across worlds requires parallel CPU cores. As reinforcement learning workloads grow, the tutorial argues, the question shifts from how quickly one world can run to how many worlds can run at once. Collecting millions of practice steps across many varied starting conditions is what learning methods need, and that is fundamentally a throughput problem.
MuJoCo Warp (MJWarp) addresses this by re-implementing MuJoCo's physics pipeline on NVIDIA Warp, a Python framework for writing GPU-accelerated kernels. The model and a batch of independent simulation states are placed on the GPU, and a single call to mjw.step advances every world in the batch at once. The tutorial demonstrates this with an SO-101 pick-and-place scene: the arm must grasp a 44 mm red cube and stack it on a blue cube, with the same robot and scene used for both the CPU baseline and the MJWarp validation.
Throughput, not speed: the key distinction
The central performance concept deserves careful explanation, because it is easy to misunderstand. MJWarp is not claimed to make a single simulation step faster. The tutorial states:
Key quotation
MJWarp's value is not necessarily a faster step for one world. It is the ability to advance hundreds or thousands together, giving the GPU enough parallel work to improve aggregate throughput, the total world-steps completed per second.
Latency versus throughput explained
NVIDIA authors on Hugging Face, "How to Use NVIDIA Warp and MjWarp to Accelerate Robotics Simulation and Learning Workflows," published September 23, 2026 (MJWarp overview section).
Latency and aggregate throughput are different quantities. Latency is the wall-clock time for one simulation step in one world. Aggregate throughput is the total number of world-steps completed per measured second across the whole batch. For a developer interactively inspecting one robot, low latency matters, and a CPU workflow remains a sensible choice. For reinforcement learning or other large-scale sampling, what matters is total experience collected per second, and that is where batching thousands of worlds on one GPU pays off.
This distinction also explains the framing of this article: the specific figure of up to 2,048 parallel worlds comes from the tutorial's own migration walkthrough of the SO-101 scene, not from an independent benchmark. No throughput numbers are asserted here beyond what the tutorial demonstrates and documents.
What Warp itself contributes
NVIDIA Warp is a Python framework for authoring statically typed kernels that compile for CPU or CUDA execution. For robotics, three properties matter. First, parallel work is explicit: each logical thread handles one point, contact, body or world, so the same code scales from a handful of objects to millions. Second, arrays live on the chosen device, with copying to CPU memory handled explicitly. Third, kernel launches are composable, and supported work can be captured into a CUDA graph and replayed, reducing repeated dispatch overhead.
Warp also offers two capabilities the tutorial highlights but does not use in the SO-101 workflow: differentiable kernels, which record forward launches and replay their adjoints for gradients, and deterministic execution, introduced in Warp 1.15. Determinism is opt-in and trades some performance for reproducible ordering, which matters in simulation, validation and regression tests. The tutorial is careful to note that these are Warp capabilities, not guarantees of differentiability or determinism for an entire MJWarp rollout.
For readers trying the tools, the tutorial points to pip install warp-lang, with version 1.15 or later for GPU determinism, and pip install mujoco-warp, plus a Colab tutorial and the mjwarp-viewer for inspecting scenes.
Why this matters beyond the benchmark
For students, small labs and hobbyists, the significance is less about any single benchmark and more about access. GPU-scale parallel simulation was previously associated with large industrial pipelines. A workflow that takes a familiar open robot arm model, moves it onto an off-the-shelf GPU with a small API change, and runs up to 2,048 parallel environments lowers the hardware and expertise barrier for large-scale robot learning experiments.
The tutorial itself frames this as a step in a series. The State of Simulation for Physical AI series from the same NVIDIA authors on Hugging Face began with a landscape overview; the later installments cover Newton and Isaac Lab, which add training loops, sensors and multi-solver integration on top of the physics. The decision table in the article is useful guidance for practitioners: for single-robot control and teleoperation, classical CPU MuJoCo; for maximum throughput on MuJoCo physics, MJWarp or mjlab; for JAX-based training recipes, MuJoCo Playground or MJX; for multi-solver integration with Isaac Lab, Newton.
A reasonable caution follows from all of this. Scaling simulation environments is a prerequisite for scalable robot learning, not a demonstration of it. Whether policies trained across thousands of GPU worlds transfer reliably to real hardware remains the open, hard problem that this infrastructure work aims to make more tractable.
