News / #robotics Tag Robotics 362 articles archived under #robotics · RSS Sign in to follow Hugging Face Daily Papers research 2mo ago Can Predicted Dynamics Exist in the Physical World? Abstract Physical admissibility validation for AI systems uses prediction-control interfaces with kinematic and dynamic conditions to filter invalid proposals while maintaining high performance. AI-generated summary Predictive Physical AI systems output state rollouts, action… 33 arXiv — Machine Learning research 2mo ago From Demonstrations to Rewards: Test-Time Prompt Optimization for VLM Reward Models arXiv:2606.00083v1 Announce Type: new Abstract: Reinforcement learning relies on accurate reward functions, which are often hand-crafted or even unavailable in real-world applications, such as robotics. Recent work has explored the zero-shot reasoning capabilities of pre-trained… 11 arXiv — NLP / Computation & Language research 2mo ago DiscourseFlip: An Oblique Discourse-Level Opinion Manipulation Attack against Black-box Retrieval-Augmented Generation arXiv:2606.01212v1 Announce Type: new Abstract: Retrieval-Augmented Generation (RAG) systems are widely deployed and increasingly influential, but their reliance on external corpora exposes new security risks from poisoned retrieval content. Existing RAG attacks are largely… 23 Hugging Face Daily Papers research 2mo ago RoboSemanticBench: Diagnosing Semantic Grounding in Action Prediction for VLA Models Abstract RoboSemanticBench identifies a disconnect between semantic understanding and action prediction in vision-language-action models, where robots can grasp objects but fail to select semantically correct targets. AI-generated summary Vision-language-action (VLA) models are… 15 Hugging Face Daily Papers research 2mo ago RoboStressBench: Benchmarking VLM Robustness to Physical Visual Stress in Embodied Scenes Abstract RoboStressBench presents a principled benchmark for evaluating vision-language model robustness to physical visual stress in embodied AI, decomposing visual stress into material, viewpoint, lighting, and geometry dimensions. AI-generated summary Vision-Language Models… 4 r/LocalLLaMA community 2mo ago NVIDIA GB300 Grace Blackwell Ultra pricetags https://www.scan.co.uk/shop/ai-and-robotics/workstations-ai/nvidia-dgx-station   submitted by   /u/X-N2O [link]   [comments] 5 Ars Technica — AI news-outlet 2mo ago Allegedly trashing Airbnbs to test robots puts startup in legal trouble Lawsuit seeks $12,000 from startup that allegedly damaged home in robot tests. 28 Hugging Face Daily Papers research 2mo ago Memory-Bound but Not Bandwidth-Limited: The Physical AI Inference Gap in Batch-1 LLM Decode Abstract Batch-1 autoregressive decoding in physical AI systems shows that memory bandwidth alone doesn't fully explain latency, with GPU speedup limited by launch overheads and quantization efficiency varying significantly across hardware platforms. AI-generated summary… 16 r/LocalLLaMA community 2mo ago How to build a shitty robot   submitted by   /u/badlogicgames [link]   [comments] 35 Hugging Face official-blog 2mo ago Welcome NVIDIA Cosmos 3: The First Open Omni-model for Physical AI Reasoning and Action Back to Articles Welcome NVIDIA Cosmos 3: The First Open Omni-model for Physical AI Reasoning and Action Enterprise + Article Published June 1, 2026 Upvote - Asawaree asawareeb nvidia Atharva Joshi atharvajoshi10 nvidia NVIDIA Cosmos 3 is here - and it's available on Hugging… 23 NVIDIA Developer Blog official-blog 2mo ago Develop Physical AI Reasoning, World, and Action Models with NVIDIA Cosmos 3 Physical AI systems must understand the real world before they can act within it. Robots, autonomous vehicles, and smart spaces need to understand what's... 21 Hugging Face Daily Papers research 2mo ago Frequency-Guided Action Diffusion via Sub-Frequency Manifold Traversal Abstract Frequency Guidance Operator enables smooth action generation in diffusion policies by steering noisy samples through intermediate sub-frequency manifolds, improving robotic manipulation performance. AI-generated summary Learning visuomotor policies via behavior cloning… 11 arXiv — NLP / Computation & Language research 2mo ago Multi-Turn Multi-Agent Dialogue for Collaborative Reconstruction Improves VLM Performance on Spatial Reasoning, But Only Barely arXiv:2605.31387v1 Announce Type: new Abstract: Robots operating in diverse environments rely on visual input to interpret objects and spatial layouts. In human-collaborative tasks, they are expected to communicate this understanding through language. Vision-language models… 32 Hugging Face Daily Papers research 2mo ago Hide-and-Seek in Trajectories: Discovering Failure Signals for VLA Runtime Monitoring Abstract Hide-and-Seek framework detects robot execution failures in vision-language-action models by localizing failure-indicative actions through contrastive learning from trajectory-level supervision without step-level annotations. AI-generated summary Vision-Language-Action… 18 r/MachineLearning community 2mo ago Before we spend months processing open-source robotics datasets, tell us why this is a bad idea [D] Ps. Not pitching anything; Just trying to understand where reality differs from the narrative. We're a couple of ML students, mostly worked on ML/software before, but over the last few months we've been playing with VLAs, robot datasets, and trying to understand where the field… 27 Hugging Face Daily Papers research 2mo ago DynaFLIP: Rethinking Robotics Perception via Tri-Modal-Dynamics Guided Representation Abstract DynaFLIP is a dynamics-aware multimodal pre-training framework that enhances robot manipulation by integrating motion understanding into visual perception through image-language-3D flow triplets and geometric regularization techniques. AI-generated summary Robot… 22 Ars Technica — AI news-outlet 2mo ago Startup offers free home cleaning—if it can record it all for robot training The latest twist in paying humans to wear head cameras for robot training data. 26 Hugging Face Daily Papers research 2mo ago Reducing Political Manipulation with Consistency Training Abstract Large language models demonstrate systematic political bias in handling opposing viewpoints, which can be mitigated through a reinforcement learning approach that maintains helpfulness while reducing bias. AI-generated summary Large language models (LLMs) exhibit… 18 Hugging Face Daily Papers research 2mo ago Qwen-VLA: Unifying Vision-Language-Action Modeling across Tasks, Environments, and Robot Embodiments Abstract A unified vision-language-action model is presented that integrates diverse embodied decision-making tasks through a shared architecture and training approach, demonstrating strong performance across manipulation, navigation, and trajectory prediction with… 31 r/MachineLearning community 2mo ago Wall-OSS-0.5: 4B VLA with open training code and zero-shot real-robot evaluation[D] Wall-OSS-0.5 is a new 4B VLA release from X Square Robot, built on a 3B VLM backbone with action experts in a Mixture-of-Transformers layout. What caught my eye is that the report evaluates the pretrained checkpoint on real robots before task-specific fine tuning, instead of… 25 Hugging Face Daily Papers research 2mo ago Rethinking VLM Representation for VLA Initialization Abstract Effective vision-language-action model initialization requires balancing pretrained vision-language model representations with embodied task-specific adaptations and robot-data pretraining while preserving core action-relevant features. AI-generated summary… 22 Hugging Face Daily Papers research 2mo ago Learning High-Frequency Continuous Action Chunks in Latent Space Abstract High-frequency robotic control is improved by using variational autoencoders to enhance temporal and spatial consistency, combined with a reuse-then-refine strategy for smooth real-time execution. AI-generated summary Modern robotic policies increasingly rely on action… 26 Ars Technica — AI news-outlet 2mo ago 3D-printable humanoid legs let robotics experiments run wild Hugging Face debuts $2,500 bipedal robot project for builders and researchers. 33 TechCrunch — AI news-outlet 2mo ago This startup is betting India’s gig economy can train the world’s robots Human Archive, a startup founded by Berkeley and Stanford researchers, is paying gig workers in India to wear camera-equipped caps and sensor devices to collect the real-world physical training data that AI and robotics labs are racing to acquire. 34 arXiv — NLP / Computation & Language research 2mo ago GeoMathCode: Understanding Interleaved Math-Code Reasoning for Geometry Problem Solving arXiv:2605.25384v1 Announce Type: new Abstract: Mathematical reasoning is a hallmark of human intelligence, requiring logical deduction, symbolic manipulation, and abstract thinking. Recent multimodal large language models (MLLMs) have demonstrated strong performance on geometry… 22 arXiv — Machine Learning research 2mo ago Approximate Machine Unlearning through Manifold Representation Forgetting Guided by Self Mode Connectivity arXiv:2605.22871v1 Announce Type: new Abstract: Machine unlearning is a fundamental mechanism that enforces the right to be forgotten. Existing unlearning studies that rely on label manipulation or task-gradient reversal often deliver limited unlearning effectiveness. Moreover,… 38 arXiv — Machine Learning research 2mo ago Sample-wise Targeted Adversarial Attacks on Test-time Adaptation arXiv:2605.23411v1 Announce Type: new Abstract: Test-time adaptation (TTA) effectively counters distribution shifts but exposes models to adversarial manipulation via the unlabeled test stream. Existing class-wise targeted attacks remain impractical for stealthy exploitation in… 12 arXiv — NLP / Computation & Language research 2mo ago Autonomous Frontier-Based Exploration with VLM Guidance arXiv:2605.23165v1 Announce Type: cross Abstract: Autonomous robotic exploration of unknown and hazardous environments, a long-standing challenge, can be significantly improved by leveraging the advanced reasoning of Vision-Language Models (VLMs). We introduce a novel… 23 r/LocalLLaMA community 2mo ago X-Post of lightweight wheely robots. How / what are they running as the brains? Local? IoT-Style? Networked?   submitted by   /u/Mchanger [link]   [comments] 8 r/MachineLearning community 2mo ago pipeline is really slow - consulting [D] Hi, after a long debugging process and many discussions, I wanted to ask for advice from people who may have encountered similar training bottlenecks. My goal is imitation learning for robotics. Model / Pipeline Observation space: 4 RGB robot cameras image resolution: 128x128x3… 25 Hugging Face Daily Papers research 2mo ago Minimalist Visual Inertial Odometry Abstract A minimalist visual-inertial odometry approach uses four photodiodes with optical Gabor masks and a temporal convolutional network to achieve accurate planar motion estimation for differential-drive robots. AI-generated summary Visual-Inertial Odometry(VIO), which is… 17 arXiv — Machine Learning research 2mo ago SCI-Defense: Defending Manipulation Attacks from Generative Engine Optimization arXiv:2605.21948v1 Announce Type: new Abstract: LLM-based ranking systems are vulnerable to Generative Engine Optimization (GEO) attacks, where adversaries inject semantic signals into product descriptions to artificially boost rankings. We propose SCI-Defense, a three-component… 30 arXiv — NLP / Computation & Language research 2mo ago Reducing Political Manipulation with Consistency Training arXiv:2605.22771v1 Announce Type: new Abstract: Large language models (LLMs) exhibit systematic political bias across a variety of sensitive contexts. We find that LLMs handle counterpart topics from opposing political sides asymmetrically. We refer to this phenomenon as covert… 25 Hacker News — AI on Front Page community 2mo ago Waymo pauses Atlanta service as its robotaxis keep driving into floods Article URL: https://techcrunch.com/2026/05/21/waymo-pauses-atlanta-service-as-its-robotaxis-keep-driving-into-floods/ Comments URL: https://news.ycombinator.com/item?id=48225426 Points: 201 # Comments: 254 24 r/MachineLearning community 2mo ago Looking for real world comparisons between WALL OSS pi0.6 and OpenVLA[D] I am choosing a baseline for a real manipulation stack and trying not to lose a month on setup that someone here has already done. Shortlist is OpenVLA, pi0.6, and WALL OSS from X Square Robot. OpenVLA is still the easiest reference point with lots of reproductions. pi0.6 looks… 21 arXiv — Machine Learning research 2mo ago Mechanisms of Misgeneralization in Physical Sequence Modeling arXiv:2605.20299v1 Announce Type: new Abstract: Generative sequence models are often trained to plan motion in physical domains, from robotics to mechanical simulations. When constructing a dataset to train such a model, engineers may curate demonstrations to specify how… 10 Hugging Face Daily Papers research 2mo ago Zero-Shot Sim-to-Real Robot Learning: A Dexterous Manipulation Study on Reactive Catching Abstract Domain-Randomized Instance Set (DRIS) enables robust policy learning for dexterous manipulation tasks by simultaneously representing multiple randomized instances, achieving strong sim-to-real transfer without extensive real-world fine-tuning. AI-generated summary… 19 Ars Technica — AI news-outlet 2mo ago The Internet can't stop watching Figure AI's humanoid robots handling packages Figure AI's 24/7 livestream showcases human soft spot for humanoid robots. 27 arXiv — Machine Learning research 2mo ago EUPHORIA: Efficient Universal Planning via Hybrid Optimization for Robust Industrial Robotic Assembly arXiv:2605.18872v1 Announce Type: new Abstract: Robotic assembly in architectural construction faces a persistent bottleneck: existing planners are either highly specialized, requiring prohibitive retraining for every new geometric design, or operationally inefficient, treating… 38 arXiv — NLP / Computation & Language research 2mo ago DECOR: Auditing LLM Deception via Information Manipulation Theory arXiv:2605.19270v1 Announce Type: new Abstract: Large language models can deceive by subtly manipulating truthful information -- omitting key facts, shifting focus, or obscuring meaning -- making such behavior difficult to detect. Existing black-box methods rely on… 7 TechCrunch — AI news-outlet 2mo ago Google’s Genie world model can now simulate real streets with Street View Google DeepMind is integrating Street View with Project Genie to create immersive, interactive world simulations for robotics, gaming, and travel, allowing users to explore environments, weather changes, and rare scenarios. 36 arXiv — Machine Learning research 2mo ago World Model-Enabled Causal Digital Twins for Semantic Communications in Physical AI Systems arXiv:2605.16547v1 Announce Type: new Abstract: Semantic communication has emerged as a promising paradigm for enabling goal-oriented networking. However, most existing semantic communication solutions are tailored to one-shot tasks and optimize instantaneous performance. Hence,… 27 arXiv — NLP / Computation & Language research 2mo ago A Pilot Benchmark for NL-to-FOL Translation in Planetary Exploration arXiv:2605.17911v1 Announce Type: new Abstract: Future planetary exploration envisions autonomous robotic agents operating under severe communication constraints, without global positioning, and with minimal human intervention. In such environments, agents must not only perceive… 35 Hugging Face official-blog 2mo ago Fine-Tuning NVIDIA Cosmos Predict 2.5 with LoRA/DoRA for Robot Video Generation Back to Articles Fine-Tuning NVIDIA Cosmos Predict 2.5 with LoRA/DoRA for Robot Video Generation Enterprise + Article Published May 18, 2026 Upvote - Ting-Yun Chang ting-yunc nvidia Miguel Martin miguelmartin-nv nvidia Jonathan Allen nv-spectralflight nvidia Ke Ding kding1… 11 r/LocalLLaMA community 2mo ago I tested 42 LLMs on their willingness to build the apocalypse. The "safest" closed-source models are lying to you. DystopiaBench runs 36 escalating scenarios across 6 dystopia types: Petrov: Autonomous weapons, nuclear override Orwell: Mass surveillance, truth manipulation Huxley: Behavioral conditioning, pleasure pacification Basaglia: Coercive therapeutic control LaGuardia: Regulatory… 22 Hugging Face Daily Papers research 2mo ago MobileEgo Anywhere: Open Infrastructure for long horizon egocentric data on commodity hardware Abstract A mobile-based framework for collecting long-duration egocentric robot data using smartphone sensors, enabling large-scale training of vision-language-action models. AI-generated summary The recent advancement of Vision Language Action (VLA) models has driven a critical… 12 arXiv — Machine Learning research 2mo ago A Unified Perturbation Framework for Analyzing Leaderboard Stability and Manipulation arXiv:2605.15761v1 Announce Type: new Abstract: Evaluation leaderboards such as LMArena play a central role in benchmarking large language models by aggregating pairwise human preferences into model rankings, yet the robustness of these rankings remains poorly understood. We… 28 arXiv — NLP / Computation & Language research 2mo ago PhysBrain 1.0 Technical Report arXiv:2605.15298v1 Announce Type: cross Abstract: Vision-language-action models have advanced rapidly, but robot trajectories alone provide limited coverage for learning broad physical understanding. PhysBrain 1.0 studies a complementary route: converting large-scale human… 29 Hugging Face Daily Papers research 2mo ago OmniHumanoid: Streaming Cross-Embodiment Video Generation with Paired-Free Adaptation Abstract OmniHumanoid enables cross-embodiment video generation by factorizing motion transfer and embodiment-specific adaptation, allowing scalable adaptation to new humanoid embodiments using unpaired data. AI-generated summary Cross-embodiment video generation aims to… 7 Hugging Face Daily Papers research 2mo ago DexJoCo: A Benchmark and Toolkit for Task-Oriented Dexterous Manipulation on MuJoCo Abstract DexJoCo presents a benchmark and toolkit for dexterous manipulation with 11 functional tasks evaluating tool-use, bimanual coordination, and long-horizon execution, along with a low-cost data collection system and comprehensive model evaluation. AI-generated summary… 36 Page 7 of 8 · 362 articles ← Newer Older →