News / #robotics Tag Robotics 360 articles archived under #robotics · RSS Sign in to follow arXiv — Machine Learning research 3h ago Towards Socially Compliant Navigation in Deep Reinforcement Learning via Proxemics-Based Reward Modeling arXiv:2608.12917v1 Announce Type: new Abstract: Developing effective robot navigation methods in crowded environments is essential for real-world applications. Although recent deep reinforcement learning (DRL) methods have improved navigation performance in crowded environments,… 31 Hugging Face official-blog 14h ago Record, train, and deploy from one place with Strands Agents, LeRobot, and Hugging Face Storage Buckets Back to Articles a]:hidden"> Record, train, and deploy from one place with Strands Agents, LeRobot, and Hugging Face Storage Buckets Enterprise Article Published August 13, 2026 Upvote 4 Sundar Raghavan rsundaraws amazon Steven Palma imstevenpmwork amazon Cagatay Cali cagataydev… 34 Hugging Face Daily Papers research 1d ago AtlasVLA: Persistent World-Ego State Modeling for Vision-Language-Action Models Abstract AtlasVLA improves embodied AI by replacing reactive control with proactive reasoning via persistent world-ego memory, enabling robust long-horizon manipulation from a single wrist camera. Generated by thinkingmachines/Inkling-Small While Vision-Language-Action (VLA)… 36 arXiv — Machine Learning research 1d ago MOON: Multi-Objective OrthoNormalized Updates for Multitask Learning arXiv:2608.11749v1 Announce Type: new Abstract: Multi-objective optimization (MOO) has demonstrated significant success in multi-task learning by mitigating task conflicts through gradient manipulation. However, most existing methods flatten model parameters into vectors and… 8 NVIDIA Developer Blog official-blog 2d ago NVIDIA JetPack 7.2.1 Adds Agentic Video Skills and T3000 Emulation Video is a core data path across NVIDIA Jetson applications, from robotics and intelligent video analytics to industrial automation, healthcare, media... 5 Hugging Face Daily Papers research 3d ago RynnValue: Scaling Robotic Value Foundation Models with Temporal Distance Abstract RynnValue is a scalable open-source value foundation model for robot manipulation that uses temporal distance instead of preferences or progress to learn generalizable value predictions and improve real-world policy success. Generated by thinkingmachines/Inkling-Small… 28 arXiv — Machine Learning research 3d ago LUCID: Latent-Skill Unified Control via Imagined Dynamics for Long-Horizon Humanoid Loco-Manipulation arXiv:2608.07746v1 Announce Type: new Abstract: Long-horizon humanoid loco-manipulation requires composing versatile whole-body skills and reliable high-level decision making. Existing methods often coordinate pretrained skills with scripted planners, finite-state machines or… 30 arXiv — Machine Learning research 3d ago V-Simba: Unleashing the Architectural Potential of RL in Visual Continuous Control arXiv:2608.07870v1 Announce Type: new Abstract: Improving sample efficiency remains a core challenge in reinforcement learning (RL), especially in real-world settings like robotics, where data collection is costly. This challenge is pronounced in visual RL, where… 14 Hacker News — AI on Front Page community 3d ago Show HN: Needle2: 14MB agentic LLM for phones, wearables, smart home and robots Hey HN, Henry from Cactus here! We previously released Cactus Needle, a 14MB agentic LLM for tool call, device use, and structured extraction for phones, wearables, smart homes, small robots and microcontrollers. We got really great feedback here, and have now incorporated the… 17 r/LocalLLaMA community 3d ago Needle 2: 14MB agentic LLM for phones, wearables, smart home and robots. Hey LocalLlaMa, Henry from Cactus here! We previously released Cactus Needle, a 14MB agentic LLM for tool call, device use, and structured extraction for phones, wearables, smart homes, small robots and microcontrollers. We got really great feedback here, and have now… 15 Hugging Face Daily Papers research 3d ago Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence Abstract Vision-language models are increasingly serving as the reasoning core of embodied agents. Robot execution is inherently iterative: each action reshapes the scene and physical state, continually renewing what must be perceived, reasoned about, and verified. Meeting these… 25 r/LocalLLaMA community 4d ago omlab/VLX-Seek-1.5-10B · Hugging Face VLX-Seek-1.5-10B VLX-Seek-1.5-10B is the open-source 10B model in the VLX-Seek 1.5 family, designed for fine-grained perception and visual grounding in embodied scenarios. It targets practical settings such as drones, robots, robotic dogs, surveillance cameras, inspection… 27 arXiv — Machine Learning research 4d ago Fairis: Fairness-Aware Aggregation with Provable Influence Containment against Fairness Poisoning Attacks in Collaborative Machine Learning arXiv:2608.06469v1 Announce Type: cross Abstract: Collaborative machine learning among financial institutions must be both group-fair and robust against deliberate adversarial manipulation. Existing fairness-aware aggregation methods remain formally vulnerable to fairness… 21 arXiv — NLP / Computation & Language research 4d ago How Should I Pick a Foundation Model for My Robot? In Favor of a Community Evaluation Framework for Social Robots arXiv:2608.06898v1 Announce Type: cross Abstract: Researchers who seek to build social robot applications on foundation models are faced with a difficult question: how should we pick a model? Public leaderboards offer little guidance: the demands of real-time, embodied social… 14 r/MachineLearning community 4d ago Non-Physical Intelligence Has A Ceiling [D] Reasoning alone cannot predict the chaotic physical world. Without a sensory and motor interface to reality, non-physical AI will not deliver the scientific and technological breakthroughs we expect.   submitted by   /u/dontkry4me [link]   [comments] 6 Hugging Face Daily Papers research 6d ago Weights or Skills? A Survey of Robot-Learning Techniques: from Action-Predicting Weights to Robots that Write their Own Skills Abstract Robot learning is splitting into two bets: policies that bake competence into frozen weights (vision-language-action, or VLA, models), and agents that write and refine their own executable skills as code. This survey organises the field around that axis of weights… 38 Hugging Face Daily Papers research 6d ago DyPES-VLA: Learning Shared Dynamics Priors and Embodiment-Specific Control for Cross-Embodiment Manipulation Abstract Vision-Language-Action (VLA) models have become a powerful paradigm for robot manipulation, but training a single generalist policy for heterogeneous robot embodiments remains an open problem. Existing methods have two main limitations. First, they underuse dynamics… 38 arXiv — Machine Learning research 7d ago Failing Gracefully: Mitigating Impact of Inevitable Robot Failures arXiv:2608.05313v1 Announce Type: cross Abstract: Service robots operate in household environments shared with humans, pets, and everyday objects, where they are highly susceptible to failures such as software crashes, hardware degradation, or unpredictable interactions. While… 25 arXiv — Machine Learning research 7d ago Velocity- and Regime-Aware Detection of Intraday Options Market Manipulation, with Explainable Attribution arXiv:2608.05373v1 Announce Type: cross Abstract: Intraday market manipulation is hard to detect because its footprint is brief, buried in millions of quotes, and statistically similar to ordinary volatility. Detectors reach high recall only by flagging so many other days that… 6 Hugging Face Daily Papers research 7d ago World-to-Wrist: Task-Conditioned Future Wrist Modeling for Fine-Grained Robot Manipulation Abstract Vision-language-action (VLA) models often treat main-view and wrist-view observations as parallel visual inputs, overlooking their distinct roles in robot manipulation. Fine-grained manipulation, however, benefits from anticipating how wrist-local interactions may… 38 Hugging Face Daily Papers research 7d ago DRIFT: Derailing Denoising Trajectories of Flow-Matching VLAs with Adversarial Patch Attack Abstract Flow-matching vision-language-action (VLA) models such as pi0 generate robot actions by integrating a learned denoising velocity field, and have been reported to resist adversarial perturbations that readily fool autoregressive VLAs. We show that this robustness is… 25 arXiv — Machine Learning research 8d ago Manipulation-Proof Oblivious Audits against Deceptive Model Providers arXiv:2608.04365v1 Announce Type: new Abstract: Audits have emerged as a critical instrument for algorithmic governance, providing a mechanism for external scrutiny and governance of machine learning models. However, ensuring the integrity of such assessments remains a… 19 Hugging Face Daily Papers research 8d ago Ego2Robot: Scalable Robot Data Synthesis from Egocentric Human Data Abstract Learning generalizable robot manipulation policies requires large-scale and diverse demonstration data. Egocentric human manipulation videos offer rich scene and task diversity, and prior work has shown that retargeting and rendering such videos into robot-format data… 12 Hugging Face Daily Papers research 8d ago BridgeVLA++: A Data-Efficient, Generalizable, and Memory-Augmented Vision-Language-Action Framework for 3D Manipulation Abstract Leveraging pre-trained vision-language models (VLMs) to construct vision-language-action (VLA) models has emerged as a promising paradigm for 3D robot manipulation. However, existing 3D VLA methods remain data-hungry, exhibit limited generalization under distribution… 4 Marcus on AI community 8d ago Elon Musk’s preposterous and possibly harmful prediction about robotic surgery Some predictions are off. This one is way off. 13 r/LocalLLaMA community 8d ago Xiaomi-Robotics-1: New robotics model released Xiaomi-Robotics-1 is a robot foundation model trained on over 100K hours of real-world manipulation trajectories. It is a Vision-Language-Action (VLA) model engineered for out-of-the-box mobile manipulation in unseen environments and efficient adaptation to new tasks. XR-1… 14 TechCrunch — AI news-outlet 8d ago TechCrunch Disrupt 2026’s Real World AI Stage features robots, automated factories, and extinct animals On our new Real World AI stage, we’ll be focusing on the intersection between the digital and physical, and all the ways we’ll continue to see a blending of the two. 24 Hugging Face Daily Papers research 9d ago Push-Wiper: Toward General-Purpose Robotic Cleaning across Varied Stains and Surfaces with Segmented Pushing Trajectories Abstract Viscous stains, characterized by high viscosity and complex rheological properties, remain a major challenge for robotic surface cleaning. Conventional wiping often spreads the stain, while scrubbing provides stronger friction but risks damaging the surface. In this… 12 Hugging Face Daily Papers research 9d ago ST-WAM: Semantic-Temporal World Action Model for Robust Manipulation under Visual Distribution Shifts Abstract World Action Models (WAMs) have emerged as a promising paradigm by jointly modeling robot actions and future visual dynamics. However, their reliance on pixel-generative future supervision can entangle action-relevant state transitions with task-irrelevant visual… 5 NVIDIA Developer Blog official-blog 9d ago Beyond VLAs: How World Action Models Reshape Robot Manipulation A central challenge in robotics is building policies that generalize beyond the demonstrations they’re trained on. A policy that succeeds in a training scene... 16 TechCrunch — AI news-outlet 9d ago Elon Musk spends half his time talking robots and AI on Tesla earnings calls An analysis of the last seven years of Tesla earnings calls shows just little attention Musk pays to Tesla's car business. 30 Hugging Face Daily Papers research 9d ago DreamTraj: Generating 6-DoF Object Trajectories by Reading Unrendered Video Diffusion Latents Abstract Accurate prediction of object trajectories during manipulation is essential for closing the perception-action loop. Progress is limited on two fronts: available datasets lack fine-grained language-to-motion annotations, and existing predictors either rely on privileged… 19 arXiv — Machine Learning research 10d ago Fairness Auditing: Lower Bounds on Company Manipulation arXiv:2608.00568v1 Announce Type: new Abstract: Fairness audits are increasingly mandated in high-stakes applications such as hiring, lending, and automated decision-making. Recent work has established fundamental impossibility results for black-box fairness auditing, showing… 32 arXiv — NLP / Computation & Language research 10d ago DeBERTa-Sentinel: Toward Transparent and Trustworthy Detection of AI-Generated Text arXiv:2608.01046v1 Announce Type: new Abstract: The rapid spread of large language models (LLMs) across the web raises concerns about misinformation, academic integrity, automated content manipulation, and risks to vulnerable online communities. Existing transformer-based… 7 Hugging Face Daily Papers research 10d ago WCM: A World Critic Model for Vision-Language-Action Reinforcement Learning Abstract Reinforcement learning (RL) post-training of Vision-Language-Action (VLA) models has shown strong promise for robotic manipulation. Among RL methods, critic-based approaches rely on a value estimator that predominantly operates on single-frame observations or… 24 MIT Technology Review — AI news-outlet 10d ago Trump’s AI protectionism has come for robotics This story originally appeared in The Algorithm, our weekly newsletter on AI. To get stories like this in your inbox first, sign up here. Humanoid robots usually elicit more cringe than awe: They stumble, kick children, and despite advances are still worse at using their hands… 8 Hugging Face Daily Papers research 10d ago One Future, Every Robot: Label-Efficient Collective-State Prediction with Decentralized JEPA Abstract Can every robot in a swarm predict the same future collective state from only local observations and bandwidth-limited messages? We formulate this as decentralized shared-state prediction and introduce Collective-State JEPA (CS-JEPA), a recurrent joint-embedding… 29 Hugging Face Daily Papers research 11d ago N_0-TWAM: Scaling Tactile-Native World-Action Model for Contact-Rich Manipulation Abstract We present N_0-TWAM, a tactile-native world-action model for contact-rich manipulation that predicts both future vision and future contact. To our knowledge, it is the first tactile world-action model trained at large scale, and it shows strong capability on… 31 arXiv — Machine Learning research 11d ago When Does On-Policy Interaction Help? Representational Tradeoffs in Value-Based Imitation Learning arXiv:2607.29617v1 Announce Type: new Abstract: Imitation learning (IL)---training an agent to replicate expert behavior from demonstrations---underpins applications from robotics to language model training. Standard approaches such as Behavior Cloning (BC) are known to suffer… 24 arXiv — NLP / Computation & Language research 11d ago WCM: A World Critic Model for Vision-Language-Action Reinforcement Learning arXiv:2607.29613v1 Announce Type: cross Abstract: Reinforcement learning (RL) post-training of Vision-Language-Action (VLA) models has shown strong promise for robotic manipulation. Among RL methods, critic-based approaches rely on a value estimator that predominantly operates… 34 Hugging Face Daily Papers research 11d ago N_0-VTLA: Scaling Vision-Tactile-Language-Action Model with Latent Tactile Tokens Abstract We present N_0-VTLA, a vision-tactile-language-action (VTLA) foundation model capable of (1) fine-grained contact-rich manipulation with tactile perception and tactile-feedback control, and (2) offline policy improvement from stored deployment data. Building on current… 24 arXiv — Machine Learning research 14d ago MUGEN: A Unified Framework for Efficient Motion Understanding and Generation arXiv:2607.27581v1 Announce Type: new Abstract: Grounding human motion in language, and language in motion, is a central step toward physical AI systems that can understand, generate, and communicate human behavior. Unified motion--language systems first coupled the two… 4 Hugging Face Daily Papers research 14d ago ACE-Data-0: Human-Centric Ambient Capture as Embodied Data Engine Abstract Embodied intelligence faces a fundamental data bottleneck. Models must capture how first-person perception, whole-body motion, dexterous manipulation, object state, sound, and touch evolve together as humans pursue goals over time. Existing datasets fragment this… 24 MIT News — AI research 14d ago Daniela Rus receives Bavarian Minister-President's High-Tech Prize Director of CSAIL and MIT professor honored for her contributions to robotics, artificial intelligence, and autonomous systems. 21 Ars Technica — AI news-outlet 14d ago Google reveals Gemini Robotics 2.0, promising improved dexterity and safety Gemini Robotics 2 includes three models, but only one is publicly available right now. 22 r/LocalLLaMA community 14d ago How close are we to local llama robotics for consumer price point? I'm guessing 3 years, what do you think? In other words: many of us will be able to afford a general purpose robot in 3 years to experiment with in the home. Cost roughly $5k? Probably small size, but hopefully still able to do the dishes and operate a vacuum.   submitted by… 13 Hugging Face Daily Papers research 14d ago πR^2: Reactive Real-time Flow Policies Abstract Generalist manipulation policies increasingly take the form of action-chunking flow policies built on large pretrained backbones. Such chunks run open-loop, so the policy cannot react to sensory input arriving mid-execution, sacrificing reactivity. Replanning more often… 4 Google DeepMind official-blog 14d ago Gemini Robotics ER 2: powering robotics with video understanding, task orchestration, and multi-robot collaboration Gemini Robotics ER 2 helps robots reason, collaborate, and solve real-world tasks. It represents a step change in video understanding, tool orchestration, and multi-robot collaboration for robotic applications. 5 arXiv — Machine Learning research 15d ago Conformal Changepoint Localization and Root Cause Analysis with Corrupted Observations arXiv:2607.26481v1 Announce Type: new Abstract: Detecting when the statistical behavior of an engineered system changes, and identifying which component is responsible, are core problems in the monitoring of telecommunication networks, robotic platforms, security infrastructure,… 10 arXiv — NLP / Computation & Language research 15d ago The Reliability of LLMs for Medical Diagnosis: An Examination of Consistency, Manipulation, and Contextual Awareness arXiv:2503.10647v2 Announce Type: replace Abstract: This study evaluated the diagnostic reliability of two Large Language Models (LLMs), Google Gemini 2.0 Flash and OpenAI ChatGPT-4o, across three dimensions: consistency under rephrased inputs, susceptibility to irrelevant… 10 Page 1 of 8 · 360 articles Older →