We're pioneering the process of web-scale video to finalized CAD schematics for dexterous robotic hands.
Starting with a task demonstration from first-person video, we show a mechanic grasping the impact drill and threading the lug nuts to the wheel. A very typical task a mechanic may encounter in their day-to-day duties.
The video then gets transformed into a parsable trace the remainder of our pipeline can read. We capture where the hand was in space, where the object went over time, and when contact was made with the object of interest. That trace is what captures all relevant information for everything downstream.
The scene trace then composes into simulation-ready assets which define our training environment. We compose fully articulate objects, to mirror behaviors of what is possible in real life. Drag the wrench, or any object, around the tray. This reconstructed environment is where the gripper will train.
Eighteen gripper compositions attempt to successfully reproduce the demonstrated manipulation task in the most efficient manner possible; no wasted movements. Change the task of interest and a different gripper morphology wins. We aim to find the best gripper over a broad scope of tasks in this process, at scale.
The human hand motion maps to the chosen gripper in task space, making sure to preserve all relevant contact points.
Replaying the motion is not a sufficient metric for success. Position alone says nothing about force on the system, hence, a residual policy closes the loop in simulation, rewarded by the object's own trajectory.
Success rate alone would produce a weak hand, which is why every winning gripper takes a static finite element analysis (FEA) pass to determine system load.
Below we have the winning gripper assembled live from consumer packaged parts: four Allegro v5 finger chains, two driven joints each, on a composed radial palm.
Capture one human egocentric video sample in your environment, and we'll handle the rest.
GET IN TOUCH →