Why first person video may matter for robot learning[D]
Mirrored from r/MachineLearning for archival readability. Support the source by reading on the original site.
I can see why first-person video might help a robot model, but not because the robot can copy a human hand. The joints, reach, timing, and control space are all different. What may transfer is the sequence of visual attention: which object enters view, what changes before contact, and where the actor looks next.
LingBot-VLA 2.0 (arXiv:2607.06403) uses first-person data alongside robot trajectories. A useful ablation separates visual prediction from robot control. First-person pretraining might improve next-state prediction without improving task success, which would still tell us where the information survived. A matched third-person comparison is important too, otherwise viewpoint and video volume are mixed together.
Occlusion is the obvious problem. Hands often cover the object at the exact moment of contact. Would you treat those frames as useful evidence of intent, missing visual data, or both? I have not seen a convincing evaluation that cleanly separates those effects.
[link] [comments]
More from r/MachineLearning
-
TMLR Relevance and Prestige [D]
Aug 13
-
Reproducible canvas-aligned low-level patterns in somerandomllm-generated images and their possible relation to iterative editing artifacts [D]
Aug 13
-
worldproof: diagnosing where world-model predictions break and a measurement of when pixel metrics stop being able to rank models at all [P]
Aug 13
-
UrgenT Help Detecting Performance Regressions Using Machine Learning and Hardware Counters [P]
Aug 13
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.