We introduce <strong>V2N (Video to Notes)</strong>, the first complete visual piano transcription (VPT) system. From top-view video alone, with no audio, it predicts <em>onset</em>, <em>offset</em>, <em>key hold</em>, and <em>velocity</em> for every note.</p>\n<p><strong>Why video?</strong> When the sustain pedal is down, a note keeps sounding long after the key lifts, so audio systems predict pedal-extended offsets rather than the physical key release. The camera sees the key itself, including how hard it is struck. Prior visual systems, though, focus on onset from short windows, so offset accuracy lags onset by a wide margin and note-level velocity has not been reported. V2N closes both gaps with a shared temporal backbone and task-specific heads for onset, offset, key hold, and velocity, trained with per-frame supervision instead of only at the window center.</p>\n<p>Ablations show multi-task supervision enables offset and velocity while improving onset, and longer temporal context helps further. V2N sets new state of the art on <a href=\"https://huggingface.co/datasets/PianoVAM/PianoVAM_v1\">PianoVAM</a> and <a href=\"https://zenodo.org/records/15921293\" rel=\"nofollow\">R3</a>. To appear at <a href=\"https://ismir2026.ismir.net/\" rel=\"nofollow\">ISMIR 2026</a>.</p>\n<p>🎬 Demo: <a href=\"https://huggingface.co/spaces/PianoVAM/V2N\">https://huggingface.co/spaces/PianoVAM/V2N</a><br>📄 Paper: <a href=\"https://arxiv.org/abs/2608.03419\" rel=\"nofollow\">https://arxiv.org/abs/2608.03419</a><br>💻 Code: <a href=\"https://github.com/yonghyunk1m/V2N\" rel=\"nofollow\">https://github.com/yonghyunk1m/V2N</a></p>\n","updatedAt":"2026-08-05T11:53:08.240Z","author":{"_id":"67dc6c6e8fc6577e1851b36e","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/67dc6c6e8fc6577e1851b36e/Sb0-erpwIC4Qutty5KEln.jpeg","fullname":"Yonghyun Kim","name":"yonghyunk1m","type":"user","isPro":true,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":3,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8133805394172668},"editors":["yonghyunk1m"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/67dc6c6e8fc6577e1851b36e/Sb0-erpwIC4Qutty5KEln.jpeg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.03419","authors":[{"_id":"6a728dd01a375f948521c306","user":{"_id":"67dc6c6e8fc6577e1851b36e","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/67dc6c6e8fc6577e1851b36e/Sb0-erpwIC4Qutty5KEln.jpeg","isPro":true,"fullname":"Yonghyun Kim","user":"yonghyunk1m","type":"user","name":"yonghyunk1m"},"name":"Yonghyun Kim","status":"claimed_verified","statusLastChangedAt":"2026-08-05T16:45:04.532Z","hidden":false},{"_id":"6a728dd01a375f948521c307","name":"Hoyeol Sohn","hidden":false},{"_id":"6a728dd01a375f948521c308","name":"Juhan Nam","hidden":false},{"_id":"6a728dd01a375f948521c309","name":"Alexander Lerch","hidden":false}],"mediaUrls":["https://cdn-uploads.huggingface.co/production/uploads/67dc6c6e8fc6577e1851b36e/ZSWF9IlOdw7vPBQbmcYYI.mp4","https://cdn-uploads.huggingface.co/production/uploads/67dc6c6e8fc6577e1851b36e/WkD2o1lb3zqD7S9Cde2UG.png"],"publishedAt":"2026-08-04T00:00:00.000Z","submittedOnDailyAt":"2026-08-05T00:00:00.000Z","title":"Multi-Task Multi-Frame Visual Piano Transcription","submittedOnDailyBy":{"_id":"67dc6c6e8fc6577e1851b36e","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/67dc6c6e8fc6577e1851b36e/Sb0-erpwIC4Qutty5KEln.jpeg","isPro":true,"fullname":"Yonghyun Kim","user":"yonghyunk1m","type":"user","name":"yonghyunk1m"},"summary":"Audio-based piano transcription performs well on onset, pitch, and velocity, but the sustain pedal lets sound persist long after key release, so audio systems predict pedal-extended offsets rather than physical key release. Yet existing Visual Piano Transcription (VPT) systems focus on onset detection from short video windows, offset accuracy lags onset by a wide margin, and note-level velocity has not been reported. To address these gaps, we present V2N (Video to Notes), the first complete VPT system: a shared temporal backbone feeds task-specific heads for onset, offset, key hold, and velocity, jointly trained with per-frame supervision rather than only at the window center. Ablations show that multi-task supervision enables offset and velocity prediction while improving onset accuracy; longer temporal context yields further improvements. V2N sets new state-of-the-art results on PianoVAM and R3.","upvotes":3,"discussionId":"6a728dd01a375f948521c30a","projectPage":"https://huggingface.co/spaces/PianoVAM/V2N","githubRepo":"https://github.com/yonghyunk1m/V2N","githubRepoAddedBy":"user","githubStars":3,"organization":{"_id":"68be25560a3fcebdcad4574a","name":"PianoVAM","fullname":"PianoVAM","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/67dc6c6e8fc6577e1851b36e/Ywr6km4_XAMzIFJylVWGp.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"6a2fc978ca5ded0a5ca01e98","avatarUrl":"/avatars/510951b9752601a35f01712a0d9c9ee1.svg","isPro":false,"fullname":"Changjae Yi","user":"paulyi0250","type":"user"},{"_id":"6974eca019d789b7f50400a4","avatarUrl":"/avatars/c98f7f2e3077f9b46b31613754126eaf.svg","isPro":false,"fullname":"Gatsby Abessolo","user":"gatsby13579","type":"user"},{"_id":"6a14c51fd222ecc8fcf518bc","avatarUrl":"/avatars/13d2a2bf9d044050d43f878c50a69fe8.svg","isPro":false,"fullname":"Bastián Núñez-Boettiger","user":"nunezboettiger","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"68be25560a3fcebdcad4574a","name":"PianoVAM","fullname":"PianoVAM","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/67dc6c6e8fc6577e1851b36e/Ywr6km4_XAMzIFJylVWGp.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2608/2608.03419.md","query":{}}">
Multi-Task Multi-Frame Visual Piano Transcription
Abstract
Audio-based piano transcription performs well on onset, pitch, and velocity, but the sustain pedal lets sound persist long after key release, so audio systems predict pedal-extended offsets rather than physical key release. Yet existing Visual Piano Transcription (VPT) systems focus on onset detection from short video windows, offset accuracy lags onset by a wide margin, and note-level velocity has not been reported. To address these gaps, we present V2N (Video to Notes), the first complete VPT system: a shared temporal backbone feeds task-specific heads for onset, offset, key hold, and velocity, jointly trained with per-frame supervision rather than only at the window center. Ablations show that multi-task supervision enables offset and velocity prediction while improving onset accuracy; longer temporal context yields further improvements. V2N sets new state-of-the-art results on PianoVAM and R3.
Community
We introduce V2N (Video to Notes), the first complete visual piano transcription (VPT) system. From top-view video alone, with no audio, it predicts onset, offset, key hold, and velocity for every note.
Why video? When the sustain pedal is down, a note keeps sounding long after the key lifts, so audio systems predict pedal-extended offsets rather than the physical key release. The camera sees the key itself, including how hard it is struck. Prior visual systems, though, focus on onset from short windows, so offset accuracy lags onset by a wide margin and note-level velocity has not been reported. V2N closes both gaps with a shared temporal backbone and task-specific heads for onset, offset, key hold, and velocity, trained with per-frame supervision instead of only at the window center.
Ablations show multi-task supervision enables offset and velocity while improving onset, and longer temporal context helps further. V2N sets new state of the art on PianoVAM and R3. To appear at ISMIR 2026.
🎬 Demo: https://huggingface.co/spaces/PianoVAM/V2N
📄 Paper: https://arxiv.org/abs/2608.03419
💻 Code: https://github.com/yonghyunk1m/V2N
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2608.03419 in a model README.md to link it from this page.
Cite arxiv.org/abs/2608.03419 in a dataset README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.