Robotic manipulation · Aston University

UniversalPick

A robot clearing a real tote end to end, picking objects it has never seen, with no CAD models, no teleoperated demonstrations, and depth optional.

The unsolved problem

Most warehouses cannot afford to automate picking.

Not because robot arms are expensive, but because everything around them is. Four requirements sit between a business and a working picking cell. UniversalPick removes all four.

That stack is why roughly 80% of a $26.5B market is locked out of automation, and why millions of dull, hazardous picking tasks are still done by hand.

What it is

One engine. RGB is the only input it requires.

RGB images from multiple views go in. Depth, camera rays and pose are optional, not required. A single multi-modal engine fuses whatever it is given and returns a dense point cloud and ranked grasp vectors.

REQUIRED RGB view 1 RGB view 2 RGB view 3 OPTIONAL depth camera rays pose UniversalPick multi-modal engine SINGLE SET OF WEIGHTS DENSE POINT CLOUD RANKED GRASP VECTORS #1 #2 #3

Sensor-agnostic

The model works from standard 2D cameras, and uses high-end depth hardware when it is present. It never depends on it.

Optical interference is a feature

Transparent glass and reflective metal are treated as distinct visual signals to reason about, rather than the failure case that halts a line.

Retrofits existing arms

Decoupling the intelligence from the embodiment means accessible hardware can be upgraded in place, with no environmental re-engineering.

Training

Trained entirely in simulation.

Nobody teleoperated anything. Every grasp the model knows, it learned from rendered scenes.

0

scenes

0

distinct objects

0

rendered views

0

human demonstrations

See it run

Five real runs. Two robot arms. The same weights.

No retraining and no re-calibration between them. The left-hand panel in every clip is Rerun, logging live off the robot while it runs.

Clips are silent and loop. Click any tile to enlarge.

Performance

Fast, light, and quick to deploy.

< 10 GB

VRAM

Runs on a single consumer GPU

~0.5 s

Median inference

Scene to ranked grasps

~10 min

To a new scene

No retraining, no re-calibration

The full film

Clearing a real tote with a robot that has never seen the objects.

5:44 · Captions available on YouTube · Watch on YouTube

Chapters

Where this is going

The goal is picking that a small business can actually afford.

Retrofit, don't rebuild

Turning accessible arms already on the floor into capable pickers, instead of asking a business to re-engineer the environment around a rigid cell.

Runs at the edge

Under 10 GB of VRAM means inference happens locally. Imagery is processed on site to guide the arm, it does not need to leave the floor.

High-mix, low-volume

The cases nobody builds a CAD library for: mixed totes, waste sorting, medical packs, where every bin looks different from the last.

If you are working on something like this, I would like to hear about it.

Whether that is a pilot, a research collaboration, or just a conversation about what does and does not work in real bins.