Skip to main content

nferent

 UMI is not simply a system for recording videos of people usingheir hands.

The data collection setup is designed around a hand held gripper with a wrist mounted camera. The system records the visual scene while also capturing information about the manipulation action. This combination is important because a robot eventually needs to connect what it sees with what it should do.

Imagine someone demonstrating how to arrange a cup on a saucer.

The useful information is not just the video of the person’s hand. The training example also needs to describe the relationship between the visual observation and the movement being performed.

That is what makes UMI style robot demonstration data different from an ordinary instructional video.

The collected demonstrations can contain information useful for learning:

  • Object and scene appearance
  • Manipulation movements
  • Gripper position and movement
  • Changes in the environment
  • Timing of actions
  • Relationships between observation and action

The exact data pipeline depends on the application and the training system being used.

Why Use Human Demonstrations

Robots are very good at repeating precisely defined movements. The problem starts when the environment changes.

A person can pick up a bottle even when it is not exactly where expected. We naturally adjust our hand position, change our grip and compensate for small differences in the object or surroundings.

Programming all of those possibilities individually would be difficult.

Human demonstrations provide another route.

A person can perform the same task several times while the system records what happens. One demonstration might start with an object on the left. Another might involve a different orientation. A third might be performed at a different speed.

Together, these examples become UMI training data that represents more than one fixed trajectory.

The goal is not simply to reproduce a person’s exact movement. The broader objective is to give a learning system enough examples to discover useful relationships between observations and actions.

How UMI Robot Data Collection WORK        

A typical collection process starts with the task rather than the hardware.

Suppose the objective is to teach a robot to fold a piece of clothing. The first question is not how many recordings should be made. It is what the robot actually needs to learn.

Does it need to recognise different clothing positions? Does it need to handle different sizes? Does the task involve two hands? How much variation should be present in the demonstrations?

Once these questions are clear, the collection setup can be prepared.

The human performs the task

The operator uses the UMI interface to demonstrate the required movement.

Because the interface is portable, demonstrations do not necessarily have to be collected next to the final robot setup. This is one of the ideas behind UMI’s approach to collecting demonstrations in varied environments.

The system records the demonstration

The camera captures the visual information while the hand held gripper provides the action related information.

The result is a demonstration that connects the scene with the manipulation being performed.

The recordings are checked

Not every recording should automatically enter a training dataset.

A demonstration can contain an incomplete action, tracking problem, unwanted interruption, or an unsuccessful attempt. These issues need to be identified before the data is used for model training.

The useful data is organised

After collection, demonstrations can be processed and organised according to the requirements of the robotics project.

This may involve synchronisation, segmentation, metadata, quality checks and dataset formatting.

That part of the workflow is often overlooked. Recording demonstrations is only the beginning. Turning them into dependable robot training data is where much of the practical data engineering work happens.

UMI Demonstration Data for Physical AI

Physical AI needs a different kind of dataset from conventional AI.

An image can tell a model what a cup looks like. It does not, by itself, teach a robot how to pick up that cup.

For a physical task, the model needs information about actions, movement and interaction with the environment.

This is why Physical AI data collection increasingly focuses on real world demonstrations and other forms of action data.

UMI fits into this area because it connects human manipulation with robot learning. The original UMI work demonstrated tasks involving dynamic actions, precise manipulation, bimanual behaviour and longer sequences of actions.

For example, UMI research included demonstrations for tasks such as dynamic object tossing, cup arrangement, cloth folding and dish washing.

These examples also show why collecting only simple pick and place movements is not enough for every robotics application.

A useful dataset may need to represent the messy parts of physical interaction the changes in object position, different movements, contact with surfaces and sequences where one action depends on another.

The Difference Between UMI Data and Regular Robot Data

Robot datasets can be created in several ways.

Some are generated through direct robot teleoperation. Others come from simulation, scripted trajectories, autonomous robot interaction or human demonstrations.

Each method has its own advantages.

Teleoperation gives direct control over the robot, but it can require specialised hardware and an operator who is comfortable controlling the system.

Simulation can generate large amounts of experience, although the simulated environment does not always reproduce every detail of the real world.

Human video provides plenty of visual information, but the connection between a person’s movement and a robot’s action can be difficult to establish.

UMI takes a different approach. The person demonstrates the task using an interface designed around robot manipulation, which helps reduce the gap between the human demonstration and the robot action. The UMI researchers specifically designed the system around this transfer problem.

Why Data Quality Matters

There is a temptation in robotics to focus on the number of demonstrations.

More data can certainly be useful, but volume alone does not make a dataset good.

Imagine collecting 10,000 demonstrations where the camera frequently loses tracking, the task instructions change between operators, or unsuccessful attempts are mixed with successful ones without clear labels.

That dataset can cause problems further down the line in training.

A better collection process looks at:

Consistency: The setup should be the same from session to session.

Variation: Demonstrations should have useful differences, not just do the same movement exactly.

Task definition: Operators need to know what constitutes a successful demonstration.

Quality control: Poor recordings should be identified before they reach the training dataset.

Metadata: Information about the task and recording can make the dataset easier to manage and analyse later.

These details become especially important when a robotics project moves from a small research dataset to large-scale UMI robot data collection.

UMI Data Collection Services

For a robotics company, collecting a few demonstrations internally is one thing. Building a repeatable data pipeline is another.

A larger project may involve dozens of tasks, multiple environments and hundreds or thousands of demonstrations. Operators need training. Recording equipment needs to be maintained. Data needs to be checked, processed and stored properly.

This is where UMI data collection services can support robotics and AI teams.

A data-collection partner can help with areas such as:

  • Designing task-specific collection protocols
  • Setting up demonstration environments
  • Recruiting and managing operators
  • Recording human demonstrations
  • Data validation and quality checks
  • Dataset organisation
  • Annotation and metadata
  • Preparing data for model training

The exact requirements will vary from project to project. A manipulation dataset for a warehouse robot may look very different from one designed for household robotics or dexterous manipulation.

What Kind of Robot Training Data Can Be Built?

UMI demonstrations can be part of a larger dataset rather than existing on their own.

A Physical AI project might combine demonstration data with other sources, including teleoperation, tactile sensing, RGB-D recordings, simulation and autonomous interaction.

For example, a project could begin with human demonstrations to establish how a task should be performed. Additional data could then be collected to cover difficult objects, unusual positions or edge cases.

This approach makes the dataset more closely connected to the actual behaviour the robot needs to learn.

It also avoids treating robot learning data as a single category. Different tasks require different information.

A robot learning to fold clothing has different data requirements from a robot learning to insert an electronic component or manipulate a kitchen tool.

Where UMI Fits Into the Future of Robot Learning

The interest in UMI is part of a wider shift in robotics: moving from robots that follow carefully written instructions toward systems that learn skills from data.

The research community is already exploring extensions and variations of UMI-style collection. For example, later work has looked at faster and more scalable UMI-style systems, including larger real-world demonstration datasets.

More recent research has also explored adding 3D information to UMI-style data collection, addressing situations where depth and geometry are important for manipulation.

This points to a broader direction for Physical AI: robots will need increasingly rich datasets that connect perception, movement and physical interaction.

Building UMI Data for Real Robot Learning

The most useful question is not simply, “How many UMI demonstrations do we need?”

A better question is:

What does the robot need to learn, and what data will show it how to do that?

Once the answer is clear, the collection process becomes much easier to design.

The right tasks can be selected. Demonstrations can be collected with appropriate variation. Quality checks can be defined before recording starts. And the final dataset can be structured around the requirements of the model rather than around the recording process itself.

That is ultimately the value of UMI data collection.

It gives robotics teams a practical way to turn human physical skills into structured examples that can be used for robot learning.

For companies building Physical AI systems, robot foundation models or manipulation policies, this type of UMI demonstration data can become one part of a much larger real-world data strategy.

The robots of the future will need more than images of the world. They will need examples of how to act in it.

 

Leave A Comment