News / #robotics Tag Robotics 362 articles archived under #robotics · RSS Sign in to follow Hugging Face Daily Papers research 1mo ago PoLAR: Factorizing Extent and Mode in Latent Actions for Robot Policy Learning Abstract PoLAR introduces a geometrically structured latent action representation in hyperbolic space that separates transition extent from transition mode, improving robotic policy learning performance. Generated by Qwen/Qwen2.5-Coder-32B-Instruct Latent action pretraining… 12 MIT News — AI research 1mo ago New chip could help tiny robots traverse complex environments Researchers combined an efficient algorithm with dedicated hardware to rapidly generate 3D maps for navigation using minimal memory and power. 10 Ars Technica — AI news-outlet 1mo ago GM installs robots at flagship EV factory after laying off 1,300 workers US autoworkers union warns of robot automation as dark factory future looms. 23 NVIDIA Developer Blog official-blog 1mo ago Inside NVIDIA Halos for Robotics: A Full-Stack Functional Safety System for Physical AI Physical AI—robots working autonomously alongside people in factories, warehouses, hospitals, and homes—is arriving faster than most expected. Traditional... 12 Hugging Face Daily Papers research 1mo ago GeneralVLA-2: Geometry-Aware Reconstruction and Governed Memory for Robot Planning Abstract GeneralVLA-2 addresses limitations in vision-language-action systems by introducing GeoFuse-MV3D for improved 3D reconstruction and an enhanced KnowledgeBank for better memory management in robotic manipulation tasks. Generated by Qwen/Qwen2.5-Coder-32B-Instruct… 32 Hacker News — AI on Front Page community 1mo ago Hyundai buys Boston Dynamics Article URL: https://startupfortune.com/hyundai-takes-full-control-of-boston-dynamics-as-softbank-exits-for-325-million/ Comments URL: https://news.ycombinator.com/item?id=48600312 Points: 227 # Comments: 118 12 Hugging Face Daily Papers research 1mo ago ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? Abstract ImageWAM demonstrates that pretrained image editing models can effectively replace video generation in world action models for robot control, achieving better performance with reduced computational costs. Generated by Qwen/Qwen2.5-Coder-32B-Instruct World Action Models… 25 Hugging Face Daily Papers research 1mo ago ENPIRE: Agentic Robot Policy Self-Improvement in the Real World Abstract ENPIRE framework enables autonomous robotics research through a closed-loop system that automates policy improvement via environment feedback, policy refinement, and evolutionary code optimization. Generated by Qwen/Qwen2.5-Coder-32B-Instruct Achieving dexterous robotic… 27 Hugging Face Daily Papers research 1mo ago DragMesh-2: Physically Plausible Dexterous Hand-Object Interaction with Articulated Objects Abstract DragMesh-2 enables dexterous hand-object interaction through contact-driven manipulation, with PICA enhancing robustness under varying contact loads without tactile feedback. Generated by Qwen/Qwen2.5-Coder-32B-Instruct Dexterous interaction with articulated objects is… 19 Hugging Face Daily Papers research 1mo ago HumanScale: Egocentric Human Video Can Outperform Real-Robot Data for Embodied Pretraining Abstract Egocentric human video can effectively replace teleoperated robot trajectories for embodied model pretraining, achieving better performance with reduced data collection costs. Generated by Qwen/Qwen2.5-Coder-32B-Instruct Embodied foundation models are expected to… 22 Hugging Face Daily Papers research 1mo ago Playful Agentic Robot Learning Abstract Embodied robots learn reusable skills through self-directed play and exploration, then apply these skills to improve performance on downstream tasks without additional training. Generated by Qwen/Qwen2.5-Coder-32B-Instruct Current agentic robot systems can write… 4 Hugging Face Daily Papers research 1mo ago Reinforcement Learning-Guided Retrieval with Soft Fusion for Robust Multimodal Imitation Learning under Missing Modalities Abstract RL4IL enables robust robotic manipulation under sensor dropout by using reinforcement learning to retrieve relevant demonstrations and cross-attention fusion to impute missing modalities without retraining. Generated by Qwen/Qwen2.5-Coder-32B-Instruct Robotic systems… 23 r/LocalLLaMA community 1mo ago My suitcase robot gets high now off a real gas sensor wired straight into the LLM sampler. Smoke raises temperature/top_p/top_k live, so his speech genuinely gets loopier and never repeats. Follow-up on Sparky, my offline suitcase robot I keep overdeveloping. He gets high now, and there's no scripted "stoned mode" anywhere in it. A real MQ-2 gas sensor sits in the case. Every 0.5s I read it against an adaptive clean-air baseline and turn a smoke hit into a 0 to 10… 30 Hugging Face Daily Papers research 1mo ago MolmoMotion: Forecasting Point Trajectories in 3D with Language Instruction Abstract 3D point motion forecasting model predicts object trajectories from visual history and language goals, demonstrating superior performance on benchmarks and transferring effectively to robot manipulation and video generation tasks. Generated by… 4 Hugging Face Daily Papers research 1mo ago PAIWorld: A 3D-Consistent World Foundation Model for Robotic Manipulation Abstract PAIWorld enhances diffusion-transformer world models with geometric awareness and cross-view attention to improve multi-view 3D consistency for robotic manipulation tasks. Generated by Qwen/Qwen2.5-Coder-32B-Instruct World foundation models (WFMs) are powerful… 18 arXiv — Machine Learning research 1mo ago Stealthy World Model Manipulation via Data Poisoning arXiv:2606.18697v1 Announce Type: new Abstract: Model-based learning agents use learned world models to predict future states, plan actions, and adapt to new environments. However, the process of updating world models from collected experience creates a training-time attack… 18 arXiv — Machine Learning research 1mo ago Strategic Feature Selection arXiv:2606.18867v1 Announce Type: new Abstract: When algorithmic predictors inform resource allocation in high-stakes domains such as healthcare, these predictors must account for strategic manipulation of input features. The typical solution is to redesign the predictor itself… 35 Hugging Face Daily Papers research 1mo ago Kairos: A Native World Model Stack for Physical AI Abstract Kairos is a native world model framework that learns from diverse experiences, maintains persistent states through hybrid temporal attention, and supports efficient deployment for physical AI applications. Generated by Qwen/Qwen2.5-Coder-32B-Instruct World models are… 33 Hugging Face Daily Papers research 1mo ago Guava: An Effective and Universal Harness for Embodied Manipulation Abstract A harness framework for embodied tool use combines high-level reasoning with external modules, enabling compact models to perform complex manipulation tasks with minimal training data. Generated by Qwen/Qwen2.5-Coder-32B-Instruct Language models trained on large-scale… 15 Hacker News — AI on Front Page community 1mo ago A robot is sprinting towards you. Do you want it running on Claude or Grok? Article URL: https://openrouter.ai/blog/insights/royale-last-agent-standing/ Comments URL: https://news.ycombinator.com/item?id=48576824 Points: 244 # Comments: 189 25 Ars Technica — AI news-outlet 1mo ago AI coding agents taught robots how to install GPUs and cut zip-ties NVIDIA’s self-improvement program for robots enlists teams of AI coding agents. 13 TechCrunch — AI news-outlet 1mo ago Collecting robot training data is dirty, unglamorous work. Some AI labs are already paying XDOF to do it If physical AI is going to match the accomplishments of LLMs, there's a data problem that needs to be solved. 33 Hugging Face official-blog 1mo ago From the Hugging Face Hub to robot hardware with Strands Agents and LeRobot Back to Articles a]:hidden"> From the Hugging Face Hub to robot hardware with Strands Agents and LeRobot Enterprise Article Published June 17, 2026 Upvote 4 Sundar Raghavan rsundaraws amazon Cagatay Cali cagataydev amazon A walkthrough of the LeRobot integration in Strands… 28 Hugging Face Daily Papers research 1mo ago Text-Vision Co-Instructed Image Editing Abstract A unified text-visual image editing framework is presented that combines semantic intent from textual instructions with spatial guidance from visual prompts to achieve more precise and faithful image manipulation. Generated by Qwen/Qwen2.5-Coder-32B-Instruct Existing… 16 arXiv — NLP / Computation & Language research 1mo ago Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems arXiv:2606.17443v1 Announce Type: cross Abstract: Large language models (LLMs) are becoming a major way for consumers to find products, but we do not yet understand how brands compete in this new channel. We study brand dynamics in LLM recommendations using skincare products --… 11 MIT News — AI research 1mo ago Could AI tell you where you left your keys? A new spatial memory system for robots efficiently captures details about the objects they see while exploring their environment. 17 Hugging Face Daily Papers research 1mo ago MotionVLA: Vision-Language-Action Model for Humanoid Motion Abstract A dual-stream frequency tokenizer and autoregressive model are proposed to improve humanoid motion generation by separately encoding pose and physical dynamics, achieving better diversity and consistency compared to single-codebook approaches. Generated by… 11 Hugging Face Daily Papers research 1mo ago ACE-Ego-0: Unifying Egocentric Human and Robotic Data for VLA Pretraining Abstract A unified Vision-Language-Action pretraining framework leverages heterogeneous data sources including human egocentric videos and robot trajectories through a reliability-aware training approach that improves performance on embodied AI tasks. Generated by… 6 r/MachineLearning community 1mo ago I built a leakage-clean verifier for robot manipulation, is this useful? Am I solving a non-problem? [D] Spent the last few weeks on a benchmark/harness that tries to answer one question honestly: did a robot arm actually do the demonstrated task, or did the success metric just get fooled? The setup: compile a human demo into an object-centric graph (what changed in the world:… 7 Hugging Face Daily Papers research 1mo ago Human Universal Grasping Abstract A flow-matching model generates diverse human grasps from RGB-D images, enabling zero-shot robotic grasping with improved performance over existing methods. Generated by Qwen/Qwen2.5-Coder-32B-Instruct Humans can grasp objects effortlessly, whereas multi-fingered robots… 25 Hugging Face Daily Papers research 1mo ago LaWAM: Latent World Action Models for Efficient Dynamics-Aware Robot Policies Abstract LaWAM enables efficient robot control by predicting compact latent visual subgoals instead of expensive video generation, achieving high performance with reduced computational latency. Generated by Qwen/Qwen2.5-Coder-32B-Instruct Vision-Language-Action models (VLAs)… 33 r/LocalLLaMA community 1mo ago Qwen Robot Suite Looks pretty cool... https://qwen.ai/blog?id=qwen-robotsuite   submitted by   /u/Snoo_27681 [link]   [comments] 8 Hugging Face Daily Papers research 1mo ago Qwen-RobotWorld Technical Report: Unifying Embodied World Modeling through Language-Conditioned Video Generation Abstract Qwen-RobotWorld is a language-conditioned video world model that predicts future visual trajectories across multiple robotic domains using a double-stream diffusion transformer and embodied world knowledge corpus. Generated by Qwen/Qwen2.5-Coder-32B-Instruct We… 5 Hugging Face Daily Papers research 1mo ago Hierarchical Advantage Weighting for Online RL Fine-Tuning of VLAs from Sparse Episode Outcomes Abstract Hierarchical Advantage-Weighted Behavior Cloning (HABC) addresses sparse reward challenges in robot learning by separately optimizing viability and efficiency objectives through adaptive critic heads and intervention-aware credit assignment, significantly improving… 9 Hugging Face Daily Papers research 1mo ago Geometric Action Model for Robot Policy Learning Abstract A geometric action model leverages pretrained geometric foundation models to enable language-conditioned manipulation policies with improved accuracy, robustness, and efficiency in 3D physical environments. Generated by Qwen/Qwen2.5-Coder-32B-Instruct Generalist robot… 21 arXiv — Machine Learning research 1mo ago Phase-Localized Curation Does Not Help: A Negative Result on Per-Phase Metric Selection for Demonstration Filtering arXiv:2606.15064v1 Announce Type: new Abstract: Manipulation demonstrations have temporal phase structure, and a natural hypothesis is that demonstration-curation metrics should be applied within phases rather than globally. The idea is to segment each trajectory into phases,… 10 arXiv — NLP / Computation & Language research 1mo ago Beyond English: Uncovering the Multilingual Gap in Vision-Language-Action Models arXiv:2606.15714v1 Announce Type: new Abstract: Vision-Language-Action models have recently demonstrated promising capabilities in learning generalist robot policies from large-scale multimodal data. However, most existing VLA systems are trained and evaluated primarily with… 12 NVIDIA Developer Blog official-blog 2mo ago Pretrained to Imagine, Fine-Tuned to Act: The Rise of World-Action Models Quick glossary for readers new to VLA/WAM terminology VLA Vision-Language-Action model: a robot policy that starts from a pretrained VLM backbone and adapts it... 22 arXiv — Machine Learning research 2mo ago SpikF-GO: Spiking Fourier Graph Operators for Multivariate Time Series Forecasting arXiv:2606.13901v1 Announce Type: new Abstract: Spiking Neural Networks (SNNs) have emerged as an energy-efficient alternative to conventional neural networks, demonstrating strong performance in computer vision and robotics. More recently, SNNs have been applied to time series… 30 arXiv — Machine Learning research 2mo ago More with LESS -- Local Scene Representations for Tactile Imaging arXiv:2606.14344v1 Announce Type: new Abstract: Tactile imaging seeks to reconstruct the internal structure of soft objects through touch sensing, with applications in medical diagnosis and robotic manipulation. Recent self-supervised learning approaches have shown promising… 33 arXiv — NLP / Computation & Language research 2mo ago Persuasion Index: A Theory-Guided Framework for Persuasion Analysis arXiv:2606.14580v1 Announce Type: new Abstract: Identifying persuasive rhetorical cues is critical across domains, from detecting information manipulation and improving AI safety to advancing public health communication. We propose Persuasion Index (PI), a taxonomy of 15… 36 Ars Technica — AI news-outlet 2mo ago Here's what Jeff Bezos' new startup Prometheus will do It isn't the only startup tackling physical AI, but it's one of the best-funded. 5 Ars Technica — AI news-outlet 2mo ago Ukraine's one-time test used fully autonomous drones to kill Russian soldiers Full autonomy is rare, but Ukraine is installing AI modules on drones and robots. 32 Hugging Face Daily Papers research 2mo ago WEAVER, Better, Faster, Longer: An Effective World Model for Robotic Manipulation Abstract WEAVER is a multi-view world model architecture that achieves high fidelity, consistency, and efficiency in robotic manipulation tasks through flow-matching loss and demonstrates superior performance in policy evaluation, improvement, and test-time planning. Generated… 27 Hugging Face Daily Papers research 2mo ago Revisiting Articulated Parts Perception in Robot Manipulation Abstract A new geometric representation called Geometric Primary Structure (GPS) is introduced for articulated parts perception, enabling efficient data collection through VR annotation and achieving high manipulation success rates without fine-tuning. Generated by… 27 Hugging Face Daily Papers research 2mo ago LabVLA: Grounding Vision-Language-Action Models in Scientific Laboratories Abstract LabVLA, a vision-language-action model trained with a two-stage approach combining action token pretraining and flow matching, demonstrates superior performance on laboratory automation tasks through simulated data generation and robot-specific learning. Generated by… 18 arXiv — NLP / Computation & Language research 2mo ago Does AI Reviewer See the Full Picture? Attacking and Defending Multimodal Peer Review arXiv:2606.12716v1 Announce Type: new Abstract: The integration of Large Language Models (LLMs) and Multimodal LLMs (MLLMs) into scientific peer-review workflows introduces novel and significant risks for adversarial manipulation, especially given the multimodal nature of… 8 arXiv — NLP / Computation & Language research 2mo ago ComAct: Reframing Professional Software Manipulation via COM-as-Action Paradigm arXiv:2606.13239v1 Announce Type: cross Abstract: Existing computer-use agents remain fundamentally limited in professional software manipulation: GUI-based agents suffer from fragile visual grounding and long-horizon error accumulation, while API-basedapproaches struggle with… 34 Hugging Face Daily Papers research 2mo ago MuJoCo-Drones-Gym: A GPU-Accelerated Multi-Drone Simulator for Control and Reinforcement Learning Abstract A Gymnasium-compatible multi-drone simulation environment built on MuJoCo physics engine that supports flexible physics models, action interfaces, and observation spaces for reinforcement learning applications. Generated by Qwen/Qwen2.5-Coder-32B-Instruct Robotic… 35 TechCrunch — AI news-outlet 2mo ago Theker just raised $85M to build the factory robot that doesn’t specialize in anything Unlike humanoid robots designed around a fixed form — think Boston Dynamics — Theker's machines are built to be reconfigured. 18 Page 5 of 8 · 362 articles ← Newer Older →