Interestana
Home/News/Reward AI Releases OM-1 Robot Policy Trained on Human Demonstrations
MarkTechPost3 min read

By Interestana AI Editorial — AI-drafted, human-overseen. How we report

Reward AI Releases OM-1 Robot Policy Trained on Human Demonstrations

Robotics startup Reward AI has released OM-1, short for Omnibody Model 1, a general-purpose manipulation policy designed to learn from human demonstrations. The key innovation of OM-1 is its training methodology, which exclusively utilizes data from humans wearing a sensorized glove and entirely omits teleoperation data or on-robot data. This approach adheres to the principle of 'One Model, One Data Interface, Any Body.' Currently, OM-1 is an in-house policy developed by Reward AI, and its weights, code, dataset, or API have not been released, preventing external developers from running it on their own hardware.

Reward AI's decision to forgo robot data in training stems from the belief that human-level manipulation will not be achieved through more teleoperated or self-collected robot data, which often ties datasets to specific hardware embodiments. Instead, the company emphasizes a unified pipeline for capture, learning, and control, enabling demonstrations recorded today to train future robot bodies that do not yet exist. This philosophy is rooted in the concept that capture, learning, and control are integrated, allowing for adaptability across different robotic platforms. The company cites Anderson's principle of "More Is Different" to support its argument that simply increasing data or compute power is not the sole path to advanced robotics.

The training process begins with the Omnibody Hand, a seven-degree-of-freedom wearable device that builds upon Reward AI's prior work with DexCap for portable motion capture. This wearable is engineered around crucial functional aspects rather than precisely replicating human hand joints. It focuses on enabling users to select contact points, reorient objects within the hand, and transition between precision and power grips. The device captures specific hand movements, including thumb-index pinching, individual thumb and index finger flexion, and the coupled motion of the middle, ring, and little fingers at the MCP joints. Ergonomics is considered a critical factor for data quality, as a device that slips or restricts the wearer's movement can lead to compensated grasps, potentially skewing the learned policy. A distal flexion mechanism is incorporated to accommodate variations in finger length, eliminating the need for per-user adjustments.

The 'One Data Interface' component of the system transforms wearer motion into training data without requiring staged setups or human supervisors. The design objective for this interface is to capture data at human speed, ensuring that the demonstrations are natural and fluid. This seamless integration of capture and learning aims to create a robust and versatile manipulation policy that can be applied to a wide range of robotic hardware. The system's ability to learn from human demonstrations without direct robot interaction or complex calibration represents a significant step towards more adaptable and generalizable robotic control.

Original source — read the full reporting at the publisher:

Read on MarkTechPost

Get the weekly AI digest

AI news + new model releases, weekly. Drafted by our agents, reviewed by humans.

Read next