HandUMI Collection
Open-source, hand-worn UMI variant for collecting bimanual manipulation data without a robot in the loop.
Capture layerWe make data collection cheaper, so you can focus on training robots.
HandUMI is one of the tools that makes this possible.
01 / Problem
Training a manipulation policy takes 300 to 1,200 high-quality demonstrations per task. Most teams still collect them by teleoperating a robot, at over $100 per usable hour.
General-purpose robots require general-purpose human data.
02 / Solution
Human demonstrations are several times more sample-efficient than teleoperation, and a few hundred well-curated demos now beat thousands of mediocre ones. The winning move is better data, collected off-robot.
RoboNet captures demonstrations with HandUMI, collects with trained operators in Peru, and delivers QA-reviewed, policy-ready datasets without tying up robot arms for every run.
Open-source, hand-worn UMI variant for collecting bimanual manipulation data without a robot in the loop.
Capture layerStart with one task and validate signal quality, labels, and training usefulness.
Pilot datasetLower-cost collection nodes with standardized protocols, QA, and versioned releases.
Versioned batches03 / Why HandUMI
HandUMI is a hand-worn, open-source UMI variant for bimanual arms with parallel-jaw grippers. It mounts on the operator's thumb and index or middle fingers, opens and closes with a natural pinch, and swaps printable gripper tips for different robot targets.
04 / How it works
Tell us the manipulation behavior your robot needs to improve.
Task briefWe define objects, setup, success criteria, labels, metadata, and delivery format.
Capture protocolTrained operators collect demonstrations through the HandUMI workflow.
Session captureWe review signals, labels, metadata, and package a versioned dataset.
QA reportYour team evaluates the pilot, then we expand if the data is useful.
Versioned dataset05 / Final vision
The near-term wedge is bimanual manipulation data. The long-term RoboNet vision is a distributed network for full-body human data that can train humanoids, bimanual arms, mobile manipulators, and future robot embodiments.

06 / Use cases
Capture chores such as tidying, laundry handling, dishwasher loading, table clearing, and object retrieval in real home-like layouts.
Build broad human hand/object interaction datasets for embodied AI systems that need manipulation priors across tasks, objects, and environments.
Collect picking, packing, kitting, sorting, box assembly, container unloading, and long-tail SKU interaction data.
Record repetitive packing, portioning, replenishment, bagging, staging, and delicate item handling tasks for deployable robot workers.
Capture vial handling, rack loading, instrument-adjacent movement, sample transfer, disposal workflows, and protocol-driven bench tasks.
Collect demonstrations for high-mix factory tasks, workstation loading, assembly support, inspection, finishing, and repetitive manual workflows.
Record daily-living support tasks such as fetching, feeding-adjacent handling, dressing assistance, medication staging, and mobility support objects.
Capture inspection, maintenance, tool handoff, part handling, and constrained-space manipulation for sites where downtime is expensive.
FAQ
RoboNet collects task-specific manipulation datasets for robot learning teams using HandUMI and QA-reviewed workflows.
The main offer is data collection. HandUMI is the capture layer that enables the workflow.
Human demonstrations are more sample-efficient than teleoperation for the same collection time, and they capture natural dexterity that teleop rigs degrade. HandUMI keeps the robot out of the collection loop so robot hardware can stay focused on validation and embodiment-specific tuning.
The current printable tips target AgileX Piper, ARX X5 2023, Dream Gripper TRLC, Trossen WidowX AI, and UMI Gripper. Comparable parallel-jaw grippers can be supported by designing and printing a matching tip.
RoboNet combines Peru-based operations, lower facility costs, and lower-cost HandUMI hardware.
No. The model is built around cost-efficient quality: SOPs, calibration, task protocols, operator training, metadata, labels, QA review, and dataset acceptance criteria.
Yes. The recommended starting point is one task, one capture protocol, and a focused pilot dataset your team can evaluate before scaling.