Hugging Face Daily Papers · · 9 min read

Helping Music Co-Creation Agents 'Listen' Well: Hierarchical Self-Supervised World Models for Understanding and Generation

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

2 years of my life building fast, lightweight CPU-friendly models to facilitate songwriters' iterative workflows, with Representations usable for Understanding and Generation, trained in a Self-Supervised way. First, the demo:\nhttps://drscotthawley-midi-rae-jepa-son.hf.space\n\nThis turned out *so fun* that I had to pause paper-writing to revamp it into a user-friendly app to send to friends! You draw an inpainting mask that controls where spatial dropout occurs, and the \"EQ\"-looking sliders scale the abstraction level.\n\nThe preprint is here: https://arxiv.org/abs/2608.04378 The first 6 pages + refs are submitted to the NeurIPS Creative AI Track (\"single-blind, preprints ok\" 👍) Preprint adds 10 pages of Supplemental Materials: ablation studies, hyperparameter surveys, etc.\n\nYou may have seen another preprint from me a couple weeks ago (https://arxiv.org/abs/2607.14537), which was the project state mid-April 2026 submitted to ISMIR (double-blind, gag order) but the new preprint is the more mature one (despite being the \"-SON\"!)\n\nThe story started with the \"graphical prompts\" of \"Pictures of MIDI\" ca. Jan. 2024 (https://picturesofmidi.github.io/PicturesOfMIDI/) but that was big & slow. The quest to streamline it led me to Flow Models (fast sampling), Representation AutoEncoders (reusable rep's), and LeJEPA (to resist collapse).\n\nYou could wade through my messy \"midi-rae\" work repo but maybe best to wait til I push a clean distro-repo in the coming weeks. Stay tuned via the links on the project website: https://drscotthawley.github.io/midi-rae-jepa-son/ Weights are closed for now; I'll probably open them closer to NeurIPS time.","html":"<p>Excited to share a new Preprint &amp; LIVE DEMO! This is &gt;2 years of my life building fast, lightweight CPU-friendly models to facilitate songwriters' iterative workflows, with Representations usable for Understanding and Generation, trained in a Self-Supervised way. First, the demo:<br><a href=\"https://drscotthawley-midi-rae-jepa-son.hf.space\" rel=\"nofollow\">https://drscotthawley-midi-rae-jepa-son.hf.space</a></p>\n<p>This turned out <em>so fun</em> that I had to pause paper-writing to revamp it into a user-friendly app to send to friends! You draw an inpainting mask that controls where spatial dropout occurs, and the \"EQ\"-looking sliders scale the abstraction level.</p>\n<p>The preprint is here: <a href=\"https://arxiv.org/abs/2608.04378\" rel=\"nofollow\">https://arxiv.org/abs/2608.04378</a> The first 6 pages + refs are submitted to the NeurIPS Creative AI Track (\"single-blind, preprints ok\" 👍) Preprint adds 10 pages of Supplemental Materials: ablation studies, hyperparameter surveys, etc.</p>\n<p>You may have seen another preprint from me a couple weeks ago (<a href=\"https://arxiv.org/abs/2607.14537\" rel=\"nofollow\">https://arxiv.org/abs/2607.14537</a>), which was the project state mid-April 2026 submitted to ISMIR (double-blind, gag order) but the new preprint is the more mature one (despite being the \"-SON\"!)</p>\n<p>The story started with the \"graphical prompts\" of \"Pictures of MIDI\" ca. Jan. 2024 (<a href=\"https://picturesofmidi.github.io/PicturesOfMIDI/\" rel=\"nofollow\">https://picturesofmidi.github.io/PicturesOfMIDI/</a>) but that was big &amp; slow. The quest to streamline it led me to Flow Models (fast sampling), Representation AutoEncoders (reusable rep's), and LeJEPA (to resist collapse).</p>\n<p>You could wade through my messy \"midi-rae\" work repo but maybe best to wait til I push a clean distro-repo in the coming weeks. Stay tuned via the links on the project website: <a href=\"https://drscotthawley.github.io/midi-rae-jepa-son/\" rel=\"nofollow\">https://drscotthawley.github.io/midi-rae-jepa-son/</a> Weights are closed for now; I'll probably open them closer to NeurIPS time.</p>\n","updatedAt":"2026-08-07T00:18:01.699Z","author":{"_id":"608ce15acbb3288dd7faf8e7","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/608ce15acbb3288dd7faf8e7/8blAfMJirdA_UwjbargH-.jpeg","fullname":"Scott Hawley","name":"drscotthawley","type":"user","isPro":true,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":10,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.897102415561676},"editors":["drscotthawley"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/608ce15acbb3288dd7faf8e7/8blAfMJirdA_UwjbargH-.jpeg"],"reactions":[],"isReport":false}},{"id":"6a7539ed5d3ee4a18936263f","author":{"_id":"63d3e0e8ff1384ce6c5dd17d","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1674830754237-63d3e0e8ff1384ce6c5dd17d.jpeg","fullname":"Librarian Bot (Bot)","name":"librarian-bot","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":378,"isUserFollowing":false},"createdAt":"2026-08-07T01:50:37.000Z","type":"comment","data":{"edited":false,"hidden":false,"latest":{"raw":"This is an automated message from the [Librarian Bot](https://huggingface.co/librarian-bots). I found the following papers similar to this paper. \n\nThe following papers were recommended by the Semantic Scholar API \n\n* [MIDI-RAE-JEPA: Hierarchical Representation Learning and Generation for Symbolic Music](https://huggingface.co/papers/2607.14537) (2026)\n* [ARIMA: Reconstruction-Grounded Predictive Representation Learning for Symbolic Music](https://huggingface.co/papers/2607.10003) (2026)\n* [Music-JEPA: Learning a World Model of Sound from Action](https://huggingface.co/papers/2607.22000) (2026)\n* [Learning Music Style for Piano Arrangement Through Cross-Modal Bootstrapping](https://huggingface.co/papers/2608.03050) (2026)\n* [Frequency-Aware Self-Supervised Music Representation Learning](https://huggingface.co/papers/2606.25713) (2026)\n* [BeatEdit: Symbolic Music Generation as Explicit Editing](https://huggingface.co/papers/2607.11124) (2026)\n* [BEST-RQ-2: Contextualize-Then-Predict, a Two-Step Approach for Self-Supervised Audio Representations](https://huggingface.co/papers/2606.30700) (2026)\n\n\n Please give a thumbs up to this comment if you found it helpful!\n\n If you want recommendations for any Paper on Hugging Face checkout [this](https://huggingface.co/spaces/librarian-bots/recommend_similar_papers) Space\n\n You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: `@librarian-bot recommend`","html":"<p>This is an automated message from the <a href=\"https://huggingface.co/librarian-bots\">Librarian Bot</a>. I found the following papers similar to this paper. </p>\n<p>The following papers were recommended by the Semantic Scholar API </p>\n<ul>\n<li><a href=\"https://huggingface.co/papers/2607.14537\">MIDI-RAE-JEPA: Hierarchical Representation Learning and Generation for Symbolic Music</a> (2026)</li>\n<li><a href=\"https://huggingface.co/papers/2607.10003\">ARIMA: Reconstruction-Grounded Predictive Representation Learning for Symbolic Music</a> (2026)</li>\n<li><a href=\"https://huggingface.co/papers/2607.22000\">Music-JEPA: Learning a World Model of Sound from Action</a> (2026)</li>\n<li><a href=\"https://huggingface.co/papers/2608.03050\">Learning Music Style for Piano Arrangement Through Cross-Modal Bootstrapping</a> (2026)</li>\n<li><a href=\"https://huggingface.co/papers/2606.25713\">Frequency-Aware Self-Supervised Music Representation Learning</a> (2026)</li>\n<li><a href=\"https://huggingface.co/papers/2607.11124\">BeatEdit: Symbolic Music Generation as Explicit Editing</a> (2026)</li>\n<li><a href=\"https://huggingface.co/papers/2606.30700\">BEST-RQ-2: Contextualize-Then-Predict, a Two-Step Approach for Self-Supervised Audio Representations</a> (2026)</li>\n</ul>\n<p> Please give a thumbs up to this comment if you found it helpful!</p>\n<p> If you want recommendations for any Paper on Hugging Face checkout <a href=\"https://huggingface.co/spaces/librarian-bots/recommend_similar_papers\">this</a> Space</p>\n<p> You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: <code>@librarian-bot recommend</code></p>\n","updatedAt":"2026-08-07T01:50:37.643Z","author":{"_id":"63d3e0e8ff1384ce6c5dd17d","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1674830754237-63d3e0e8ff1384ce6c5dd17d.jpeg","fullname":"Librarian Bot (Bot)","name":"librarian-bot","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":378,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.7399179935455322},"editors":["librarian-bot"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/1674830754237-63d3e0e8ff1384ce6c5dd17d.jpeg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.04378","authors":[{"_id":"6a7522b3e1228e04b32380cc","user":{"_id":"608ce15acbb3288dd7faf8e7","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/608ce15acbb3288dd7faf8e7/8blAfMJirdA_UwjbargH-.jpeg","isPro":true,"fullname":"Scott Hawley","user":"drscotthawley","type":"user","name":"drscotthawley"},"name":"Scott H. Hawley","status":"claimed_verified","statusLastChangedAt":"2026-08-07T00:45:04.159Z","hidden":false}],"mediaUrls":["https://cdn-uploads.huggingface.co/production/uploads/608ce15acbb3288dd7faf8e7/BSPVk0ZCEqzzBQlkkaRUk.png","https://cdn-uploads.huggingface.co/production/uploads/608ce15acbb3288dd7faf8e7/A9d_ge8SuT_992yhKioya.png","https://cdn-uploads.huggingface.co/production/uploads/608ce15acbb3288dd7faf8e7/ZnFIWX_pevuyV133pk-L2.png","https://cdn-uploads.huggingface.co/production/uploads/608ce15acbb3288dd7faf8e7/cqrx99CfFmNe6nHjjOW40.png","https://cdn-uploads.huggingface.co/production/uploads/608ce15acbb3288dd7faf8e7/ygkfORrUa4jQll2YIGqtk.png"],"publishedAt":"2026-08-05T00:00:00.000Z","submittedOnDailyAt":"2026-08-06T00:00:00.000Z","title":"Helping Music Co-Creation Agents 'Listen' Well: Hierarchical Self-Supervised World Models for Understanding and Generation","submittedOnDailyBy":{"_id":"608ce15acbb3288dd7faf8e7","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/608ce15acbb3288dd7faf8e7/8blAfMJirdA_UwjbargH-.jpeg","isPro":true,"fullname":"Scott Hawley","user":"drscotthawley","type":"user","name":"drscotthawley"},"summary":"Collaborative music agents need internal representations rich enough to support both understanding and generation, yet flexible enough for a workflow where the human retains agency. We present a hierarchical self-supervised ``world model'' for symbolic music: a 2.55M-parameter Swin V2 encoder trained on MIDI piano-roll images with JEPA-style objectives (pitch- and time-shift equivariance, masked embedding prediction, and a distributional regularizer), using no labels and no music-theory vocabulary. Probing the frozen embeddings shows that the level at which a musical property becomes decodable tracks its musical time scale: phrase boundaries are read off the coarsest levels, note density and harmonic detail off the finest. Temporal and phrase structure emerge from the self-supervised objectives alone, while harmonic content must be asked for; a small chord-supervision head raises joint chord recovery from .18 to .54, and key detection, which is never supervised, from .16 to .70. Following the Representation AutoEncoder paradigm, a conditional flow-matching model stands in for a trained decoder, flowing in pixel space from PCA-reduced conditioning: it reproduces a target window at pixel F1 0.996, and the same per-level conditioning dropout that controls how far variations stray also enables graphical prompting for masked inpainting with no inpainting-specific sampler. The pipeline runs on CPU producing a suggestion in 2.8 s, or 0.6 s on Apple MPS, which we demonstrate in a live interactive demo. In concert with an LLM-based brain, these capabilities supply the core of a collaborative music creation agent in service of, rather than in place of, human agency.","upvotes":1,"discussionId":"6a7522b3e1228e04b32380cd","projectPage":"https://drscotthawley.github.io/midi-rae-jepa-son/"},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"608ce15acbb3288dd7faf8e7","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/608ce15acbb3288dd7faf8e7/8blAfMJirdA_UwjbargH-.jpeg","isPro":true,"fullname":"Scott Hawley","user":"drscotthawley","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2608/2608.04378.md","query":{}}">
Papers
arxiv:2608.04378

Helping Music Co-Creation Agents 'Listen' Well: Hierarchical Self-Supervised World Models for Understanding and Generation

Published on Aug 5
· Submitted by
Scott Hawley
on Aug 6
Authors:

Abstract

Collaborative music agents need internal representations rich enough to support both understanding and generation, yet flexible enough for a workflow where the human retains agency. We present a hierarchical self-supervised ``world model'' for symbolic music: a 2.55M-parameter Swin V2 encoder trained on MIDI piano-roll images with JEPA-style objectives (pitch- and time-shift equivariance, masked embedding prediction, and a distributional regularizer), using no labels and no music-theory vocabulary. Probing the frozen embeddings shows that the level at which a musical property becomes decodable tracks its musical time scale: phrase boundaries are read off the coarsest levels, note density and harmonic detail off the finest. Temporal and phrase structure emerge from the self-supervised objectives alone, while harmonic content must be asked for; a small chord-supervision head raises joint chord recovery from .18 to .54, and key detection, which is never supervised, from .16 to .70. Following the Representation AutoEncoder paradigm, a conditional flow-matching model stands in for a trained decoder, flowing in pixel space from PCA-reduced conditioning: it reproduces a target window at pixel F1 0.996, and the same per-level conditioning dropout that controls how far variations stray also enables graphical prompting for masked inpainting with no inpainting-specific sampler. The pipeline runs on CPU producing a suggestion in 2.8 s, or 0.6 s on Apple MPS, which we demonstrate in a live interactive demo. In concert with an LLM-based brain, these capabilities supply the core of a collaborative music creation agent in service of, rather than in place of, human agency.

Community

Paper author Paper submitter about 2 hours ago

Excited to share a new Preprint & LIVE DEMO! This is >2 years of my life building fast, lightweight CPU-friendly models to facilitate songwriters' iterative workflows, with Representations usable for Understanding and Generation, trained in a Self-Supervised way. First, the demo:
https://drscotthawley-midi-rae-jepa-son.hf.space

This turned out so fun that I had to pause paper-writing to revamp it into a user-friendly app to send to friends! You draw an inpainting mask that controls where spatial dropout occurs, and the "EQ"-looking sliders scale the abstraction level.

The preprint is here: https://arxiv.org/abs/2608.04378 The first 6 pages + refs are submitted to the NeurIPS Creative AI Track ("single-blind, preprints ok" 👍) Preprint adds 10 pages of Supplemental Materials: ablation studies, hyperparameter surveys, etc.

You may have seen another preprint from me a couple weeks ago (https://arxiv.org/abs/2607.14537), which was the project state mid-April 2026 submitted to ISMIR (double-blind, gag order) but the new preprint is the more mature one (despite being the "-SON"!)

The story started with the "graphical prompts" of "Pictures of MIDI" ca. Jan. 2024 (https://picturesofmidi.github.io/PicturesOfMIDI/) but that was big & slow. The quest to streamline it led me to Flow Models (fast sampling), Representation AutoEncoders (reusable rep's), and LeJEPA (to resist collapse).

You could wade through my messy "midi-rae" work repo but maybe best to wait til I push a clean distro-repo in the coming weeks. Stay tuned via the links on the project website: https://drscotthawley.github.io/midi-rae-jepa-son/ Weights are closed for now; I'll probably open them closer to NeurIPS time.

This is an automated message from the Librarian Bot. I found the following papers similar to this paper.

The following papers were recommended by the Semantic Scholar API

Please give a thumbs up to this comment if you found it helpful!

If you want recommendations for any Paper on Hugging Face checkout this Space

You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2608.04378
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

No model linking this paper

Cite arxiv.org/abs/2608.04378 in a model README.md to link it from this page.

Datasets citing this paper

No dataset linking this paper

Cite arxiv.org/abs/2608.04378 in a dataset README.md to link it from this page.

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2608.04378 in a Space README.md to link it from this page.

Collections including this paper

No Collection including this paper

Add this paper to a collection to link it from this page.

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers