We introduce OpenLongTail, an open-source generative data engine that transforms monocular long-tail driving videos into temporally coherent and pose-aligned multi-view training assets. Our approach combines pose-informed extrapolative view synthesis with Plücker ray geometry to generate missing viewpoints while preserving cross-view consistency. Experiments demonstrate that training with the generated data improves closed-loop driving robustness in rare and challenging scenarios.</p>\n","updatedAt":"2026-07-21T03:10:07.466Z","author":{"_id":"671b080780cecd29fed27887","avatarUrl":"/avatars/a00aff671f7490ac32234fc03ab1d768.svg","fullname":"Luuuulinnnn","name":"luuuulinnnn","type":"user","isPro":true,"isHf":false,"isHfAdmin":false,"isMod":false,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8854348063468933},"editors":["luuuulinnnn"],"editorAvatarUrls":["/avatars/a00aff671f7490ac32234fc03ab1d768.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2607.09655","authors":[{"_id":"6a5ee2de4fe5d1d13e84ab62","name":"Lulin Liu","hidden":false},{"_id":"6a5ee2de4fe5d1d13e84ab63","name":"Nuo Chen","hidden":false},{"_id":"6a5ee2de4fe5d1d13e84ab64","name":"Yan Wang","hidden":false},{"_id":"6a5ee2de4fe5d1d13e84ab65","name":"Bangya Liu","hidden":false},{"_id":"6a5ee2de4fe5d1d13e84ab66","name":"Wenyan Cong","hidden":false},{"_id":"6a5ee2de4fe5d1d13e84ab67","name":"Hezhen Hu","hidden":false},{"_id":"6a5ee2de4fe5d1d13e84ab68","name":"Boris Ivanovic","hidden":false},{"_id":"6a5ee2de4fe5d1d13e84ab69","name":"Hao Wang","hidden":false},{"_id":"6a5ee2de4fe5d1d13e84ab6a","name":"Ziyao Zeng","hidden":false},{"_id":"6a5ee2de4fe5d1d13e84ab6b","name":"Xinyu Gong","hidden":false},{"_id":"6a5ee2de4fe5d1d13e84ab6c","name":"Yang Zhou","hidden":false},{"_id":"6a5ee2de4fe5d1d13e84ab6d","name":"Zixiang Xiong","hidden":false},{"_id":"6a5ee2de4fe5d1d13e84ab6e","name":"Dilin Wang","hidden":false},{"_id":"6a5ee2de4fe5d1d13e84ab6f","name":"Zhangyang Wang","hidden":false},{"_id":"6a5ee2de4fe5d1d13e84ab70","name":"Weisong Shi","hidden":false},{"_id":"6a5ee2de4fe5d1d13e84ab71","name":"Ruohan Zhang","hidden":false},{"_id":"6a5ee2de4fe5d1d13e84ab72","name":"Marco Pavone","hidden":false},{"_id":"6a5ee2de4fe5d1d13e84ab73","name":"Zhiwen Fan","hidden":false}],"mediaUrls":["https://cdn-uploads.huggingface.co/production/uploads/671b080780cecd29fed27887/wH5RXgrS7h0zJIYgdZ1XS.mp4"],"publishedAt":"2026-07-10T00:00:00.000Z","submittedOnDailyAt":"2026-07-21T00:00:00.000Z","title":"OpenLongTail: Generative Scaling of Long-Tail Driving Data","submittedOnDailyBy":{"_id":"671b080780cecd29fed27887","avatarUrl":"/avatars/a00aff671f7490ac32234fc03ab1d768.svg","isPro":true,"fullname":"Luuuulinnnn","user":"luuuulinnnn","type":"user","name":"luuuulinnnn"},"summary":"Scaling robust driving policies is fundamentally bottlenecked by the scarcity of edge cases in curated datasets. While the real world continuously captures these critical events, such long-tail events remain underutilized when collected from heterogeneous sources. Specifically, diverse but valuable in-the-wild long-tail videos lack the full view coverage required for training policy models, often missing multi-view poses or originating solely from monocular dash cameras. This modality gap prevents these ubiquitous observations from being converted into scalable training data for long-tail generalization. We introduce OpenLongTail, an open-source generative data engine for scaling autonomous driving policies under long-tail events. To transform heterogeneous data sources into view-aligned and temporally coherent multi-view assets that are useful for policy learning, we develop a pose-informed extrapolative view synthesis pipeline that generates the missing views. We further enhance cross-view consistency and the temporal alignment for the newly generated views by injecting Plücker ray geometry into the scalable generation engine. By synthesizing heterogeneous long-tail data, we observe a significant improvement in closed-loop driving robustness in handling long-tail events. By measuring the extrapolative view synthesis and pose metrics, we validate the effectiveness of OpenLongTail in visual fidelity, cross-view consistency, and ego-trajectory recovery.","upvotes":2,"discussionId":"6a5ee2df4fe5d1d13e84ab74","projectPage":"https://openlongtail.github.io/","githubRepo":"https://github.com/phai-lab/OpenLongTail","githubRepoAddedBy":"user","githubStars":24,"organization":{"_id":"693049768605dfa68334b46d","name":"TexasAMUniversity","fullname":"Texas A&M University","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/68e396f2b5bb631e9b2fac9a/uv9z1cu15X7vyo70DW0tH.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"671b080780cecd29fed27887","avatarUrl":"/avatars/a00aff671f7490ac32234fc03ab1d768.svg","isPro":true,"fullname":"Luuuulinnnn","user":"luuuulinnnn","type":"user"},{"_id":"6a2da6c8ca070ee12c6e396c","avatarUrl":"/avatars/0355287dcabaa67dbc7f0b10b87451f9.svg","isPro":false,"fullname":"Joe Mama","user":"JoeMama123123123","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"693049768605dfa68334b46d","name":"TexasAMUniversity","fullname":"Texas A&M University","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/68e396f2b5bb631e9b2fac9a/uv9z1cu15X7vyo70DW0tH.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2607/2607.09655.md","query":{}}">
OpenLongTail: Generative Scaling of Long-Tail Driving Data
Abstract
Scaling robust driving policies is fundamentally bottlenecked by the scarcity of edge cases in curated datasets. While the real world continuously captures these critical events, such long-tail events remain underutilized when collected from heterogeneous sources. Specifically, diverse but valuable in-the-wild long-tail videos lack the full view coverage required for training policy models, often missing multi-view poses or originating solely from monocular dash cameras. This modality gap prevents these ubiquitous observations from being converted into scalable training data for long-tail generalization. We introduce OpenLongTail, an open-source generative data engine for scaling autonomous driving policies under long-tail events. To transform heterogeneous data sources into view-aligned and temporally coherent multi-view assets that are useful for policy learning, we develop a pose-informed extrapolative view synthesis pipeline that generates the missing views. We further enhance cross-view consistency and the temporal alignment for the newly generated views by injecting Plücker ray geometry into the scalable generation engine. By synthesizing heterogeneous long-tail data, we observe a significant improvement in closed-loop driving robustness in handling long-tail events. By measuring the extrapolative view synthesis and pose metrics, we validate the effectiveness of OpenLongTail in visual fidelity, cross-view consistency, and ego-trajectory recovery.
Community
We introduce OpenLongTail, an open-source generative data engine that transforms monocular long-tail driving videos into temporally coherent and pose-aligned multi-view training assets. Our approach combines pose-informed extrapolative view synthesis with Plücker ray geometry to generate missing viewpoints while preserving cross-view consistency. Experiments demonstrate that training with the generated data improves closed-loop driving robustness in rare and challenging scenarios.
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2607.09655 in a model README.md to link it from this page.
Cite arxiv.org/abs/2607.09655 in a dataset README.md to link it from this page.
Cite arxiv.org/abs/2607.09655 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.