","html":"<p>This work shows a very cool application of origami representation reconstruction from a natural video. The domain is fun, practical and underexplored. Its use of an agentic framework is practically relevant to the field.<br><img src=\"https://maya-moriya.github.io/origami-page/static/figures/Input%20vs%20Output.svg\" width=\"600\"></p>\n","updatedAt":"2026-09-03T10:14:41.002Z","author":{"_id":"60fc61d3afd4bcbfc22c005e","avatarUrl":"/avatars/814ad58676168476879afc5ed54200a9.svg","fullname":"Kate","name":"kate-feingold","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.7471358180046082},"editors":["kate-feingold"],"editorAvatarUrls":["/avatars/814ad58676168476879afc5ed54200a9.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2609.00377","authors":[{"_id":"6a985f04fea818274321fd25","user":{"_id":"68b81c06a3db45c2ab8b87f2","avatarUrl":"/avatars/81c858c2c30a91176cf4b8830152db23.svg","isPro":false,"fullname":"Maya Moriya","user":"mayaweiz","type":"user","name":"mayaweiz"},"name":"Maya Moriya","status":"claimed_verified","statusLastChangedAt":"2026-09-03T00:45:04.474Z","hidden":false},{"_id":"6a985f04fea818274321fd26","name":"Sigal Raab","hidden":false},{"_id":"6a985f04fea818274321fd27","name":"Yael Vinker","hidden":false},{"_id":"6a985f04fea818274321fd28","name":"Tali Dekel","hidden":false}],"publishedAt":"2026-08-31T00:00:00.000Z","submittedOnDailyAt":"2026-09-03T00:00:00.000Z","title":"FoldingAgent: Inferring Parametric Origami Procedures from Demonstration Videos","submittedOnDailyBy":{"_id":"60fc61d3afd4bcbfc22c005e","avatarUrl":"/avatars/814ad58676168476879afc5ed54200a9.svg","isPro":false,"fullname":"Kate","user":"kate-feingold","type":"user","name":"kate-feingold"},"summary":"We present FoldingAgent, an agentic framework for inferring explicit parametric folding programs directly from origami demonstration videos. Our framework leverages the reasoning power of a pre-trained Vision-Language Model (VLM) equipped with a suite of specialized tools that enable the agent to simulate geometric transitions, verify physical plausibility, retrieve and compare visual content, and evaluate its own predictions. To translate visual content into folding programs, we define a parametric space that consists of the paper's geometry and a set of parametric folding actions. Unlike models that predict static crease patterns, our agent operates sequentially and possesses the ability to re-plan its actions, effectively mitigating the compounding errors inherent in multi-step folding. Our approach takes a step toward closing the gap between human origami knowledge, which is primarily shared through unstructured visual demonstrations, and computational methods, which typically rely on structured, parametric representations such as a crease pattern or an executable parametric plan. We evaluate our approach on PurelandFold, a newly curated benchmark of diverse Pureland origami videos with ground-truth geometry and action labels. Our results demonstrate that by combining VLM reasoning with a set of specialized tools and physical simulation, we can successfully transform unstructured visual demonstrations into executable, physically plausible folding procedures.","upvotes":3,"discussionId":"6a985f04fea818274321fd29","projectPage":"https://maya-moriya.github.io/origami-page/","ai_summary":"FoldingAgent uses a vision-language model with specialized tools to convert origami videos into executable parametric folding programs via sequential reasoning and physical verification.","ai_keywords":["Vision-Language Model","parametric folding programs","origami demonstration videos","geometric transitions","physical plausibility","cross-modal retrieval","parametric folding actions","crease patterns","multi-step folding","PurelandFold"],"ai_summary_model":"thinkingmachines/Inkling-Small"},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"60fc61d3afd4bcbfc22c005e","avatarUrl":"/avatars/814ad58676168476879afc5ed54200a9.svg","isPro":false,"fullname":"Kate","user":"kate-feingold","type":"user"},{"_id":"62a315a03621790e22aafd68","avatarUrl":"/avatars/4a8a01bbe807f494a03c7a49de3f6f97.svg","isPro":false,"fullname":"Raab","user":"Sigal","type":"user"},{"_id":"684d57f26e04c265777ead3f","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/cuOj-bQqukSZreXgUJlfm.png","isPro":false,"fullname":"Joakim Lee","user":"Reinforcement4All","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2609/2609.00377.md","query":{}}">
FoldingAgent: Inferring Parametric Origami Procedures from Demonstration Videos
Published on Aug 31
· Submitted by Kate on Sep 3 Abstract
FoldingAgent uses a vision-language model with specialized tools to convert origami videos into executable parametric folding programs via sequential reasoning and physical verification.
We present FoldingAgent, an agentic framework for inferring explicit parametric folding programs directly from origami demonstration videos. Our framework leverages the reasoning power of a pre-trained Vision-Language Model (VLM) equipped with a suite of specialized tools that enable the agent to simulate geometric transitions, verify physical plausibility, retrieve and compare visual content, and evaluate its own predictions. To translate visual content into folding programs, we define a parametric space that consists of the paper's geometry and a set of parametric folding actions. Unlike models that predict static crease patterns, our agent operates sequentially and possesses the ability to re-plan its actions, effectively mitigating the compounding errors inherent in multi-step folding. Our approach takes a step toward closing the gap between human origami knowledge, which is primarily shared through unstructured visual demonstrations, and computational methods, which typically rely on structured, parametric representations such as a crease pattern or an executable parametric plan. We evaluate our approach on PurelandFold, a newly curated benchmark of diverse Pureland origami videos with ground-truth geometry and action labels. Our results demonstrate that by combining VLM reasoning with a set of specialized tools and physical simulation, we can successfully transform unstructured visual demonstrations into executable, physically plausible folding procedures.
Community
This work shows a very cool application of origami representation reconstruction from a natural video. The domain is fun, practical and underexplored. Its use of an agentic framework is practically relevant to the field.

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2609.00377 in a model README.md to link it from this page.
Cite arxiv.org/abs/2609.00377 in a dataset README.md to link it from this page.
Cite arxiv.org/abs/2609.00377 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.