Hugging Face Daily Papers · · 4 min read

DreamTraj: Generating 6-DoF Object Trajectories by Reading Unrendered Video Diffusion Latents

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

DreamTraj predicts a 6-DoF object trajectory from a single RGB image and one language instruction, by reading motion directly from the internal representations of a frozen image-to-video diffusion model at an early denoising step. The paper also introduces MOVE, 5,038 egocentric 6-DoF trajectories with language annotations.</p>\n","updatedAt":"2026-08-04T08:51:44.977Z","author":{"_id":"697c3bc48632cb493b0e7206","avatarUrl":"/avatars/55d35974706dfa092bddf379a7150a62.svg","fullname":"Tongsheng Ding","name":"dts1347","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.816735565662384},"editors":["dts1347"],"editorAvatarUrls":["/avatars/55d35974706dfa092bddf379a7150a62.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.00486","authors":[{"_id":"6a715895ec5082b9f872ce29","user":{"_id":"697c3bc48632cb493b0e7206","avatarUrl":"/avatars/55d35974706dfa092bddf379a7150a62.svg","isPro":false,"fullname":"Tongsheng Ding","user":"dts1347","type":"user","name":"dts1347"},"name":"Tongsheng Ding","status":"claimed_verified","statusLastChangedAt":"2026-08-04T08:45:04.671Z","hidden":false},{"_id":"6a715895ec5082b9f872ce2a","user":{"_id":"67e22ee3202d6535979d9651","avatarUrl":"/avatars/3615d9a1a62001c6a9a8d852d8b2487e.svg","isPro":false,"fullname":"SII-Zhen","user":"zhen0610","type":"user","name":"zhen0610"},"name":"Zhen Luo","status":"claimed_verified","statusLastChangedAt":"2026-08-04T10:00:46.436Z","hidden":false},{"_id":"6a715895ec5082b9f872ce2b","name":"Yixuan Yang","hidden":false},{"_id":"6a715895ec5082b9f872ce2c","name":"Boyu Wang","hidden":false},{"_id":"6a715895ec5082b9f872ce2d","name":"Luyang Xie","hidden":false},{"_id":"6a715895ec5082b9f872ce2e","name":"Jinyu Yang","hidden":false},{"_id":"6a715895ec5082b9f872ce2f","name":"Feng Zheng","hidden":false}],"publishedAt":"2026-08-01T00:00:00.000Z","submittedOnDailyAt":"2026-08-04T00:00:00.000Z","title":"DreamTraj: Generating 6-DoF Object Trajectories by Reading Unrendered Video Diffusion Latents","submittedOnDailyBy":{"_id":"697c3bc48632cb493b0e7206","avatarUrl":"/avatars/55d35974706dfa092bddf379a7150a62.svg","isPro":false,"fullname":"Tongsheng Ding","user":"dts1347","type":"user","name":"dts1347"},"summary":"Accurate prediction of object trajectories during manipulation is essential for closing the perception-action loop. Progress is limited on two fronts: available datasets lack fine-grained language-to-motion annotations, and existing predictors either rely on privileged inputs such as video, depth, or CAD models, or recover motion from fully generated videos through costly, error-prone perception pipelines. We close the supervision gap with the MOVE dataset, 5,038 object-centric egocentric trajectories, each paired with a fine-grained natural-language instruction rather than a coarse verb-noun label. We further propose DreamTraj, which predicts a 6-DoF object trajectory from a single RGB image and a task instruction, requiring no video, depth, or CAD model at inference: rather than generating a video, it reads motion from the internal representations of a frozen image-to-video diffusion model at an early denoising step. A lightweight flow-matching Reader decodes query-key attention tracks and pooled hidden states into relative 6-DoF poses. To our knowledge, this is the first approach to directly decode object 6-DoF trajectories from intermediate video diffusion representations rather than generated pixels. DreamTraj sets a new state of the art on both translation and rotation against forecasters that consume multi-frame or privileged inputs, and runs 4.6x faster than generate-then-extract pipelines.","upvotes":11,"discussionId":"6a715896ec5082b9f872ce30","projectPage":"https://whathappen0.github.io/DreamTraj/","organization":{"_id":"63072fb21801ecc7d25a4d7a","name":"SUSTech","fullname":"Southern university of science and technology","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/1673355266575-63072c121801ecc7d25a2604.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"697c3bc48632cb493b0e7206","avatarUrl":"/avatars/55d35974706dfa092bddf379a7150a62.svg","isPro":false,"fullname":"Tongsheng Ding","user":"dts1347","type":"user"},{"_id":"6a6a829fa698558ca76157ee","avatarUrl":"/avatars/98e9a42de130397e7f4efce0039cdacd.svg","isPro":false,"fullname":"Richard Williams","user":"richardwilliams","type":"user"},{"_id":"6a6d3d66ef16fe7cceaf0e2a","avatarUrl":"/avatars/9a0c85f1ad367ca2b3909cd536abf6f9.svg","isPro":false,"fullname":"Michael Brown","user":"robert-0248731","type":"user"},{"_id":"6a6c89271569d2cf12aae140","avatarUrl":"/avatars/c1de7a4a0199c61993f9b10a1f66941a.svg","isPro":false,"fullname":"Steven Smith","user":"charles-3412977","type":"user"},{"_id":"6a6da80639f401bb7d2e516d","avatarUrl":"/avatars/60ea984db5052568ea69cd3f0f9a0211.svg","isPro":false,"fullname":"Susan Martin","user":"matthew-0402082","type":"user"},{"_id":"67e22ee3202d6535979d9651","avatarUrl":"/avatars/3615d9a1a62001c6a9a8d852d8b2487e.svg","isPro":false,"fullname":"SII-Zhen","user":"zhen0610","type":"user"},{"_id":"67540bf3bba3a63c32506d1d","avatarUrl":"/avatars/1be0e729a900b3ee2d6634eab648614f.svg","isPro":false,"fullname":"Thomas Young","user":"ThomasYHT-Damn","type":"user"},{"_id":"654c4f6d848e8c8b118e16e3","avatarUrl":"/avatars/b2323e95624e28c2578a3c5d73020c77.svg","isPro":false,"fullname":"Arnold Yang","user":"B3rrYang","type":"user"},{"_id":"65f81b4abbd76dece59b2928","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/05xBAEngudSAlcYXsGsSK.jpeg","isPro":false,"fullname":"Liqiong Wang","user":"Kki11","type":"user"},{"_id":"6a14736cb28ec6a2ad92463d","avatarUrl":"/avatars/4e6b86c838443f10d828b8b18179bbf9.svg","isPro":false,"fullname":"Смирнов Алина","user":"oliviaydwb","type":"user"},{"_id":"64ddbc3f1f2dad27e1a05ac1","avatarUrl":"/avatars/4fb2753b7998c8536bfd4780d3b10a6d.svg","isPro":false,"fullname":"Junrulu","user":"Junrulu","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"63072fb21801ecc7d25a4d7a","name":"SUSTech","fullname":"Southern university of science and technology","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/1673355266575-63072c121801ecc7d25a2604.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2608/2608.00486.md","query":{}}">
Papers
arxiv:2608.00486

DreamTraj: Generating 6-DoF Object Trajectories by Reading Unrendered Video Diffusion Latents

Published on Aug 1
· Submitted by
Tongsheng Ding
on Aug 4
Authors:

Abstract

Accurate prediction of object trajectories during manipulation is essential for closing the perception-action loop. Progress is limited on two fronts: available datasets lack fine-grained language-to-motion annotations, and existing predictors either rely on privileged inputs such as video, depth, or CAD models, or recover motion from fully generated videos through costly, error-prone perception pipelines. We close the supervision gap with the MOVE dataset, 5,038 object-centric egocentric trajectories, each paired with a fine-grained natural-language instruction rather than a coarse verb-noun label. We further propose DreamTraj, which predicts a 6-DoF object trajectory from a single RGB image and a task instruction, requiring no video, depth, or CAD model at inference: rather than generating a video, it reads motion from the internal representations of a frozen image-to-video diffusion model at an early denoising step. A lightweight flow-matching Reader decodes query-key attention tracks and pooled hidden states into relative 6-DoF poses. To our knowledge, this is the first approach to directly decode object 6-DoF trajectories from intermediate video diffusion representations rather than generated pixels. DreamTraj sets a new state of the art on both translation and rotation against forecasters that consume multi-frame or privileged inputs, and runs 4.6x faster than generate-then-extract pipelines.

Community

Paper author Paper submitter about 4 hours ago

DreamTraj predicts a 6-DoF object trajectory from a single RGB image and one language instruction, by reading motion directly from the internal representations of a frozen image-to-video diffusion model at an early denoising step. The paper also introduces MOVE, 5,038 egocentric 6-DoF trajectories with language annotations.

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2608.00486
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

No model linking this paper

Cite arxiv.org/abs/2608.00486 in a model README.md to link it from this page.

Datasets citing this paper

No dataset linking this paper

Cite arxiv.org/abs/2608.00486 in a dataset README.md to link it from this page.

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2608.00486 in a Space README.md to link it from this page.

Collections including this paper

No Collection including this paper

Add this paper to a collection to link it from this page.

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers