Annotated walking videos with reliable gait ground truth are difficult to collect at scale, partly due to privacy concerns and the need for specialized motion-capture setups. We introduce <strong>SynthGait-19K</strong>, a synthetic dataset of <strong>19,272 RGB walking videos</strong> generated from <strong>6,427 MoCap sequences across 437 subjects</strong>, with ground-truth annotations for six clinically relevant gait parameters.</p>\n<p>Our <strong>Gait2Vid</strong> pipeline unifies MoCap data from multiple datasets and converts it into diverse, controllable RGB walking videos while preserving the underlying gait kinematics. We also introduce <strong>GaitXFormer</strong>, a direct RGB-to-gait model that estimates gait parameters from video without relying on intermediate pose or mesh representations and can run in real time.</p>\n<p>Experiments show that training with SynthGait-19K transfers effectively to real-world videos, providing a scalable way to develop video-based gait analysis systems.</p>\n","updatedAt":"2026-09-09T12:32:12.157Z","author":{"_id":"6375965008eebfdd0a399891","avatarUrl":"/avatars/946768f40a18793ced82f09a1de47952.svg","fullname":"Soroush Mehraban","name":"SoroushMehraban","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":4,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8613103628158569},"editors":["SoroushMehraban"],"editorAvatarUrls":["/avatars/946768f40a18793ced82f09a1de47952.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2609.08108","authors":[{"_id":"6aa151c3ff4bf7311191ab34","name":"Soroush Mehraban","hidden":false},{"_id":"6aa151c3ff4bf7311191ab35","name":"Xin Lei Lin","hidden":false},{"_id":"6aa151c3ff4bf7311191ab36","name":"Vida Adeli","hidden":false},{"_id":"6aa151c3ff4bf7311191ab37","name":"Majid Mirmehdi","hidden":false},{"_id":"6aa151c3ff4bf7311191ab38","name":"Amirhossein Dadashzadeh","hidden":false},{"_id":"6aa151c3ff4bf7311191ab39","name":"Clint Hansen","hidden":false},{"_id":"6aa151c3ff4bf7311191ab3a","name":"Andrea Iaboni","hidden":false},{"_id":"6aa151c3ff4bf7311191ab3b","name":"Babak Taati","hidden":false}],"mediaUrls":["https://cdn-uploads.huggingface.co/production/uploads/6375965008eebfdd0a399891/LmbwnwwXz0znOibP_ZQjM.mp4"],"publishedAt":"2026-09-08T00:00:00.000Z","submittedOnDailyAt":"2026-09-09T00:00:00.000Z","title":"SynthGait-19K: A Physically Grounded Synthetic Video Dataset for Gait Parameter Estimation","submittedOnDailyBy":{"_id":"6375965008eebfdd0a399891","avatarUrl":"/avatars/946768f40a18793ced82f09a1de47952.svg","isPro":false,"fullname":"Soroush Mehraban","user":"SoroushMehraban","type":"user","name":"SoroushMehraban"},"summary":"Accurate estimation of clinically meaningful gait parameters from monocular video is important for scalable mobility assessment, yet progress is limited by the small scale, restricted viewpoints, and limited visual diversity of existing datasets. We introduce SynthGait-19k, a physically grounded synthetic video dataset containing 19,272 walking videos derived from 6,427 MoCap sequences across 437 subjects, with paired SMPL motion and annotations for six gait parameters. To construct the dataset, we develop Gait2Vid, which unifies heterogeneous MoCap recordings through SMPL and synthesizes diverse RGB walking videos under controllable viewpoints and scene appearances. We assess the generated videos for consistency with their conditioning gait kinematics and validate extracted gait events against force-platform measurements. Using SynthGait-19K, we benchmark direct RGB, pose-based, biomechanical, and human-mesh-recovery approaches and analyze viewpoint, training-data scale, and synthetic-to-real domain shift. We also introduce GaitXFormer as a direct RGB reference model for estimating gait parameters. Synthetic supervision transfers effectively to real videos across both GaitXFormer and a pose-based architecture, demonstrating utility across different representations. We further find that spatial gait parameters are more sensitive to visual domain shift and that improved HMR reconstruction alone does not necessarily translate to improved downstream gait estimation.","upvotes":4,"discussionId":"6aa151c3ff4bf7311191ab3c","projectPage":"https://soroushmehraban.github.io/SynthGait-19k/","githubRepo":"https://github.com/TaatiTeam/SynthGait-19k","githubRepoAddedBy":"user","ai_summary":"SynthGait-19k is a large synthetic video dataset for gait analysis that enables benchmarking of video-based gait estimation and shows synthetic supervision transfers to real data.","ai_keywords":["SynthGait-19k","Gait2Vid","SMPL","MoCap","gait parameters","GaitXFormer","human-mesh-recovery","synthetic-to-real domain shift"],"ai_summary_model":"thinkingmachines/Inkling-Small","githubStars":0,"organization":{"_id":"62c5000b4d3cf26ce7c62822","name":"uoft","fullname":"University of Toronto","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/1657077766523-62c4ff85cb7033fd49b7a559.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"6375965008eebfdd0a399891","avatarUrl":"/avatars/946768f40a18793ced82f09a1de47952.svg","isPro":false,"fullname":"Soroush Mehraban","user":"SoroushMehraban","type":"user"},{"_id":"64dd788122f93c32882c0c9c","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/kVVoiT21S3CrvvbPq7uzF.jpeg","isPro":false,"fullname":"Babak Taati","user":"BabakTaati","type":"user"},{"_id":"6a0383b4acd78d3c3bed489d","avatarUrl":"/avatars/fb49d80bcc8c8b08423221d0ff6073b1.svg","isPro":false,"fullname":"Andrew Peng","user":"andrew-canada","type":"user"},{"_id":"673256564c2f18a60e44f77c","avatarUrl":"/avatars/e06edd4f288f70acfb86b12d2a82a82b.svg","isPro":false,"fullname":"Arman Hatami","user":"armanhatami","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"62c5000b4d3cf26ce7c62822","name":"uoft","fullname":"University of Toronto","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/1657077766523-62c4ff85cb7033fd49b7a559.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2609/2609.08108.md","query":{}}">
SynthGait-19K: A Physically Grounded Synthetic Video Dataset for Gait Parameter Estimation
Abstract
SynthGait-19k is a large synthetic video dataset for gait analysis that enables benchmarking of video-based gait estimation and shows synthetic supervision transfers to real data.
Accurate estimation of clinically meaningful gait parameters from monocular video is important for scalable mobility assessment, yet progress is limited by the small scale, restricted viewpoints, and limited visual diversity of existing datasets. We introduce SynthGait-19k, a physically grounded synthetic video dataset containing 19,272 walking videos derived from 6,427 MoCap sequences across 437 subjects, with paired SMPL motion and annotations for six gait parameters. To construct the dataset, we develop Gait2Vid, which unifies heterogeneous MoCap recordings through SMPL and synthesizes diverse RGB walking videos under controllable viewpoints and scene appearances. We assess the generated videos for consistency with their conditioning gait kinematics and validate extracted gait events against force-platform measurements. Using SynthGait-19K, we benchmark direct RGB, pose-based, biomechanical, and human-mesh-recovery approaches and analyze viewpoint, training-data scale, and synthetic-to-real domain shift. We also introduce GaitXFormer as a direct RGB reference model for estimating gait parameters. Synthetic supervision transfers effectively to real videos across both GaitXFormer and a pose-based architecture, demonstrating utility across different representations. We further find that spatial gait parameters are more sensitive to visual domain shift and that improved HMR reconstruction alone does not necessarily translate to improved downstream gait estimation.
Community
Annotated walking videos with reliable gait ground truth are difficult to collect at scale, partly due to privacy concerns and the need for specialized motion-capture setups. We introduce SynthGait-19K, a synthetic dataset of 19,272 RGB walking videos generated from 6,427 MoCap sequences across 437 subjects, with ground-truth annotations for six clinically relevant gait parameters.
Our Gait2Vid pipeline unifies MoCap data from multiple datasets and converts it into diverse, controllable RGB walking videos while preserving the underlying gait kinematics. We also introduce GaitXFormer, a direct RGB-to-gait model that estimates gait parameters from video without relying on intermediate pose or mesh representations and can run in real time.
Experiments show that training with SynthGait-19K transfers effectively to real-world videos, providing a scalable way to develop video-based gait analysis systems.
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2609.08108 in a model README.md to link it from this page.
Cite arxiv.org/abs/2609.08108 in a dataset README.md to link it from this page.
Cite arxiv.org/abs/2609.08108 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.