Video Foundation Models as
Physically Grounded Hand Trackers for Robot Learning

Seungjun Moon1,2  ·  Subin Jeon3  ·  Sangwoo Kim1  ·  Hanbyul Joo3  ·  Jinwoo Shin1,2

1RLWRLD    2KAIST    3Seoul National University

RLHND delivers accurate hand pose, contact, and force estimates from monocular video, enabling improved robot policy learning.

RLHND on egocentric clips: hand pose (left), contact (middle) and force (right) with the total force over time.

Abstract

Recently, approaches that leverage human video datasets for robot policy training have become increasingly prevalent. However, most existing hand trackers regress pose from cropped frames with limited priors on hand motion and object interaction, resulting in inaccurate and physically inconsistent estimates. Moreover, the lack of physical cues, e.g., contact and force, limits the use of human videos for robot policy training. To this end, we propose RLHND, a video foundation model-based hand tracking model that jointly estimates hand pose and realistic tactile information from monocular egocentric videos. RLHND turns the pre-trained Cosmos 3 video diffusion backbone into a deterministic clip-level feature extractor via clean-latent conditioning, carrying its learned priors on hand motion and hand-object interaction into tracking. For pose estimation, RLHND (i) predicts hand poses with anatomically plausible joint angles and (ii) enables optional conditioning on the shape parameter to maintain consistent hand shape within the same video and even across videos recorded by the same actor. For tactile estimation, a separate tactile expert stream, trained with the pose stream frozen, predicts dense contact and force over the hand surface. We further adopt LBS-based feature spreading to enable vertex-wise feature extraction without costly per-vertex attention. RLHND achieves state-of-the-art performance across various benchmark datasets for pose estimation, while also achieving state-of-the-art performance in contact and force estimation. Moreover, we demonstrate the utility of RLHND for robot learning through retargeting results and real-world robot experiments. Code will be released at github.com/seungjun-moon/rlhnd.

Quantitative results

Key numbers from the paper's main tables, averaged over the benchmarks.

Show the per-benchmark numbers

Motion reconstruction

Extensive qualitative comparison of 2D and 3D hand motion reconstruction.

EgoDex
HOT3D
ARCTIC

EgoDex — arranging dominoes

video overlay
Original
HaMeR
WiLoR
HaWoR
HaPTIC
HandFlow
ACE-Ego-Hand
RLHND (ours)
world-frame trajectory
Ground truth
HaMeR
WiLoR
HaWoR
HaPTIC
HandFlow
ACE-Ego-Hand
RLHND (ours)

Contact and force

Extensive qualitative comparison of per-vertex contact and force estimation.

OpenTouch
PVDB
EgoDex

OpenTouch — reaching into a kitchen sink

contact + force
Input
Ground truth
HACO
HOPE
RLHND (ours)
Force GT
Force (ours)

Real-robot demos

Real-robot rollouts of DPP + RLHND on an RB-Y1 with WUJI hands.

PnP - bottle

Retargeting to robot hands

Predicted hand motion retargeted to a dexterous robot hand with DexPilot vector retargeting.

ARCTIC ego — lifting a box

Inspire RH56
HaWoR
RLHND (ours)