At the RAI Institute, we are hardware agnostic – our research is conducted on a wide range of robots ranging from humanoids and quadrupeds like Boston Dynamics’ Spot to our own custom-built robots like the Ultra Mobility Vehicle (UMV), Roadrunner and AthenaZero. In modern robotics systems, Reinforcement Learning (RL) has become a popular tool for achieving agile locomotion and manipulation behaviors. However, transferring a policy trained in simulation onto real-world hardware remains a major bottleneck. The sim-to-real gap stems not only from out-of-distribution observations or physics model mismatches, but also from software implementation discrepancies, bugs, and overall code complexity. Training setups typically operate in Python using RL frameworks, whereas real-time physical control loops run in high performance environments. Bridging this divide without introducing silent bugs or spending weeks re-writing infrastructure is a persistent engineering challenge.

Beyond ONNX: Solving the Sim-to-Real Challenges

Common deployment workflows usually export only the neural network’s weights into an open format like ONNX. While this makes the network weights portable, there is much more that needs to be done such as processing robot states to generate network inputs, processing network outputs to generate control signals, and handling an arbitrary range of use cases that involve the use of robot sensors. These requirements  translate into additional work for developers who are typically forced to manually rewrite observation functions, sensor processing, and action post-processing logic.

This status quo introduces critical pain points:

  • Numerical Discrepancies: Minor variations in math calculations or floating-point tolerances between Python and C++ lead to unexpected, dangerous robot behaviors on real hardware.
  • Engineering Friction: Manually recreating environment logic for every new iteration is error-prone, labor-intensive, and hard to audit.
  • Parallel Codebases: Simulation environments and control stacks diverge over time, making it difficult to share policies across teams or replicate research.

Exploy was created to let roboticists focus on rapid iteration and deployment of their work onto physical hardware, rather than how to re-engineer deployment pipelines.

What is Exploy?

Exploy (EXport and dePLOY) is an open-source library that packages environment logic and neural network policies into a single, self-contained computational graph. While developed primarily with reinforcement learning in mind, its framework is generalizable. Instead of exporting only neural network weights, Exploy traces PyTorch operations directly from the source environment to export the entire control pipeline—mapping raw sensor measurements to target outputs ranging from low-level actuator commands to high-level velocity targets.

Exploy consists of two main pillars:

  1. Exporter: Traces the environment’s observation generation, neural network forward pass, and action processing using PyTorch. It registers inputs, outputs, persistent memory (for recurrent models like LSTMs/RNNs), and custom metadata (e.g., control rates, joint stiffness, damping gains) via a unified context manager.
  2. Controller: A lightweight C++ runtime built around ONNX Runtime. It automatically discovers and wires ONNX tensor inputs/outputs to physical hardware state interfaces.

By embedding the complete computational graph and configuration metadata into a single .onnx file, a single C++ deployment stack can load and execute policies across different tasks and robots without manual code changes.

What Frameworks and Architectures Does Exploy Support?

  • Training Frameworks: Built-in adapters for IsaacLab and MjLab, with an extensible interface to support custom PyTorch-based frameworks.
  • Network Architectures: Being built around ONNX, Exploy supports any operation supported by ONNX, including architectures like MLPs, Recurrent Neural Networks (LSTM, GRU), Convolutional Networks (CNNs), Transformers, and Diffusion policies.
  • Hardware Integration: Plain CMake exportable libraries and a native ROS 2 package.

Real-World Examples & Field Deployments

Exploy is actively utilized by research groups at the RAI Institute across a diverse range of hardware platforms and control tasks. Exploy-backed policies can be seen in several videos published this year:

Roadrunner: A custom multimodal robot capable of dynamically transitioning between Segway balancing, bicycle driving, and stepping modes using a unified RL policy. Exploy handles the deployment of its unified control and get-up policies zero-shot from simulation to hardware.

Spot Parkour: Boston Dynamics’ Spot quadruped utilizing depth-based and recurrent policies via Exploy to traverse high obstacles and aggressive parkour terrain.

Exploy is also being used for research across RAI, including with the UMV. Exploy powers autonomous trail racing on RAI’s custom autonomous bike platform. Additionally, Exploy has also been deployed on Unitree G1 Humanoid & Boston Dynamics’ Atlas for autonomous navigation, blind walking, and rough terrain traversal using locomotion policies.

The Future of Sim-to-Real Policy Transfer

Exploy bridges the sim-to-real gap by changing how reinforcement learning policies are transferred to hardware. It reduces deployment effort by eliminating manual C++ re-implementation of observation and action pipelines, reducing control codebase size and iteration time. It offers evaluation tools to verify step-by-step numerical equivalence between original PyTorch policies and exported ONNX artifacts before running code on physical hardware. And its unified control stack enables a single C++ deployment pipeline to run diverse policy architectures across different robots and RL training frameworks.

To start exporting your own environments or integrating the C++ controller, check out the Exploy GitHub Repository.