News / #robotics Tag Robotics 500 articles archived under #robotics · RSS Sign in to follow Hacker News — AI on Front Page community 1mo ago Show HN: Needle2: 14MB agentic LLM for phones, wearables, smart home and robots Hey HN, Henry from Cactus here! We previously released Cactus Needle, a 14MB agentic LLM for tool call, device use, and structured extraction for phones, wearables, smart homes, small robots and microcontrollers. We got really great feedback here, and have now incorporated the… 17 r/LocalLLaMA community 1mo ago Needle 2: 14MB agentic LLM for phones, wearables, smart home and robots. Hey LocalLlaMa, Henry from Cactus here! We previously released Cactus Needle, a 14MB agentic LLM for tool call, device use, and structured extraction for phones, wearables, smart homes, small robots and microcontrollers. We got really great feedback here, and have now… 15 Hugging Face Daily Papers research 1mo ago Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence Abstract Vision-language models are increasingly serving as the reasoning core of embodied agents. Robot execution is inherently iterative: each action reshapes the scene and physical state, continually renewing what must be perceived, reasoned about, and verified. Meeting these… 25 r/LocalLLaMA community 1mo ago omlab/VLX-Seek-1.5-10B · Hugging Face VLX-Seek-1.5-10B VLX-Seek-1.5-10B is the open-source 10B model in the VLX-Seek 1.5 family, designed for fine-grained perception and visual grounding in embodied scenarios. It targets practical settings such as drones, robots, robotic dogs, surveillance cameras, inspection… 27 arXiv — Machine Learning research 1mo ago Fairis: Fairness-Aware Aggregation with Provable Influence Containment against Fairness Poisoning Attacks in Collaborative Machine Learning arXiv:2608.06469v1 Announce Type: cross Abstract: Collaborative machine learning among financial institutions must be both group-fair and robust against deliberate adversarial manipulation. Existing fairness-aware aggregation methods remain formally vulnerable to fairness… 21 arXiv — NLP / Computation & Language research 1mo ago How Should I Pick a Foundation Model for My Robot? In Favor of a Community Evaluation Framework for Social Robots arXiv:2608.06898v1 Announce Type: cross Abstract: Researchers who seek to build social robot applications on foundation models are faced with a difficult question: how should we pick a model? Public leaderboards offer little guidance: the demands of real-time, embodied social… 14 r/MachineLearning community 1mo ago Non-Physical Intelligence Has A Ceiling [D] Reasoning alone cannot predict the chaotic physical world. Without a sensory and motor interface to reality, non-physical AI will not deliver the scientific and technological breakthroughs we expect.   submitted by   /u/dontkry4me [link]   [comments] 6 Hugging Face Daily Papers research 1mo ago Weights or Skills? A Survey of Robot-Learning Techniques: from Action-Predicting Weights to Robots that Write their Own Skills Abstract Robot learning is splitting into two bets: policies that bake competence into frozen weights (vision-language-action, or VLA, models), and agents that write and refine their own executable skills as code. This survey organises the field around that axis of weights… 38 Hugging Face Daily Papers research 1mo ago DyPES-VLA: Learning Shared Dynamics Priors and Embodiment-Specific Control for Cross-Embodiment Manipulation Abstract Vision-Language-Action (VLA) models have become a powerful paradigm for robot manipulation, but training a single generalist policy for heterogeneous robot embodiments remains an open problem. Existing methods have two main limitations. First, they underuse dynamics… 38 arXiv — Machine Learning research 1mo ago Failing Gracefully: Mitigating Impact of Inevitable Robot Failures arXiv:2608.05313v1 Announce Type: cross Abstract: Service robots operate in household environments shared with humans, pets, and everyday objects, where they are highly susceptible to failures such as software crashes, hardware degradation, or unpredictable interactions. While… 25 arXiv — Machine Learning research 1mo ago Velocity- and Regime-Aware Detection of Intraday Options Market Manipulation, with Explainable Attribution arXiv:2608.05373v1 Announce Type: cross Abstract: Intraday market manipulation is hard to detect because its footprint is brief, buried in millions of quotes, and statistically similar to ordinary volatility. Detectors reach high recall only by flagging so many other days that… 6 Hugging Face Daily Papers research 1mo ago World-to-Wrist: Task-Conditioned Future Wrist Modeling for Fine-Grained Robot Manipulation Abstract Vision-language-action (VLA) models often treat main-view and wrist-view observations as parallel visual inputs, overlooking their distinct roles in robot manipulation. Fine-grained manipulation, however, benefits from anticipating how wrist-local interactions may… 38 Hugging Face Daily Papers research 1mo ago DRIFT: Derailing Denoising Trajectories of Flow-Matching VLAs with Adversarial Patch Attack Abstract Flow-matching vision-language-action (VLA) models such as pi0 generate robot actions by integrating a learned denoising velocity field, and have been reported to resist adversarial perturbations that readily fool autoregressive VLAs. We show that this robustness is… 25 arXiv — Machine Learning research 1mo ago Manipulation-Proof Oblivious Audits against Deceptive Model Providers arXiv:2608.04365v1 Announce Type: new Abstract: Audits have emerged as a critical instrument for algorithmic governance, providing a mechanism for external scrutiny and governance of machine learning models. However, ensuring the integrity of such assessments remains a… 19 Hugging Face Daily Papers research 1mo ago Ego2Robot: Scalable Robot Data Synthesis from Egocentric Human Data Abstract Learning generalizable robot manipulation policies requires large-scale and diverse demonstration data. Egocentric human manipulation videos offer rich scene and task diversity, and prior work has shown that retargeting and rendering such videos into robot-format data… 12 Hugging Face Daily Papers research 1mo ago BridgeVLA++: A Data-Efficient, Generalizable, and Memory-Augmented Vision-Language-Action Framework for 3D Manipulation Abstract Leveraging pre-trained vision-language models (VLMs) to construct vision-language-action (VLA) models has emerged as a promising paradigm for 3D robot manipulation. However, existing 3D VLA methods remain data-hungry, exhibit limited generalization under distribution… 4 Marcus on AI community 1mo ago Elon Musk’s preposterous and possibly harmful prediction about robotic surgery Some predictions are off. This one is way off. 13 r/LocalLLaMA community 1mo ago Xiaomi-Robotics-1: New robotics model released Xiaomi-Robotics-1 is a robot foundation model trained on over 100K hours of real-world manipulation trajectories. It is a Vision-Language-Action (VLA) model engineered for out-of-the-box mobile manipulation in unseen environments and efficient adaptation to new tasks. XR-1… 14 TechCrunch — AI news-outlet 1mo ago TechCrunch Disrupt 2026’s Real World AI Stage features robots, automated factories, and extinct animals On our new Real World AI stage, we’ll be focusing on the intersection between the digital and physical, and all the ways we’ll continue to see a blending of the two. 24 Hugging Face Daily Papers research 1mo ago Push-Wiper: Toward General-Purpose Robotic Cleaning across Varied Stains and Surfaces with Segmented Pushing Trajectories Abstract Viscous stains, characterized by high viscosity and complex rheological properties, remain a major challenge for robotic surface cleaning. Conventional wiping often spreads the stain, while scrubbing provides stronger friction but risks damaging the surface. In this… 12 Hugging Face Daily Papers research 1mo ago ST-WAM: Semantic-Temporal World Action Model for Robust Manipulation under Visual Distribution Shifts Abstract World Action Models (WAMs) have emerged as a promising paradigm by jointly modeling robot actions and future visual dynamics. However, their reliance on pixel-generative future supervision can entangle action-relevant state transitions with task-irrelevant visual… 5 NVIDIA Developer Blog official-blog 1mo ago Beyond VLAs: How World Action Models Reshape Robot Manipulation A central challenge in robotics is building policies that generalize beyond the demonstrations they’re trained on. A policy that succeeds in a training scene... 16 TechCrunch — AI news-outlet 1mo ago Elon Musk spends half his time talking robots and AI on Tesla earnings calls An analysis of the last seven years of Tesla earnings calls shows just little attention Musk pays to Tesla's car business. 30 Hugging Face Daily Papers research 1mo ago DreamTraj: Generating 6-DoF Object Trajectories by Reading Unrendered Video Diffusion Latents Abstract Accurate prediction of object trajectories during manipulation is essential for closing the perception-action loop. Progress is limited on two fronts: available datasets lack fine-grained language-to-motion annotations, and existing predictors either rely on privileged… 19 arXiv — Machine Learning research 1mo ago Fairness Auditing: Lower Bounds on Company Manipulation arXiv:2608.00568v1 Announce Type: new Abstract: Fairness audits are increasingly mandated in high-stakes applications such as hiring, lending, and automated decision-making. Recent work has established fundamental impossibility results for black-box fairness auditing, showing… 32 arXiv — NLP / Computation & Language research 1mo ago DeBERTa-Sentinel: Toward Transparent and Trustworthy Detection of AI-Generated Text arXiv:2608.01046v1 Announce Type: new Abstract: The rapid spread of large language models (LLMs) across the web raises concerns about misinformation, academic integrity, automated content manipulation, and risks to vulnerable online communities. Existing transformer-based… 7 Hugging Face Daily Papers research 1mo ago WCM: A World Critic Model for Vision-Language-Action Reinforcement Learning Abstract Reinforcement learning (RL) post-training of Vision-Language-Action (VLA) models has shown strong promise for robotic manipulation. Among RL methods, critic-based approaches rely on a value estimator that predominantly operates on single-frame observations or… 24 MIT Technology Review — AI news-outlet 1mo ago Trump’s AI protectionism has come for robotics This story originally appeared in The Algorithm, our weekly newsletter on AI. To get stories like this in your inbox first, sign up here. Humanoid robots usually elicit more cringe than awe: They stumble, kick children, and despite advances are still worse at using their hands… 8 Hugging Face Daily Papers research 1mo ago One Future, Every Robot: Label-Efficient Collective-State Prediction with Decentralized JEPA Abstract Can every robot in a swarm predict the same future collective state from only local observations and bandwidth-limited messages? We formulate this as decentralized shared-state prediction and introduce Collective-State JEPA (CS-JEPA), a recurrent joint-embedding… 29 Hugging Face Daily Papers research 1mo ago N_0-TWAM: Scaling Tactile-Native World-Action Model for Contact-Rich Manipulation Abstract We present N_0-TWAM, a tactile-native world-action model for contact-rich manipulation that predicts both future vision and future contact. To our knowledge, it is the first tactile world-action model trained at large scale, and it shows strong capability on… 31 arXiv — Machine Learning research 1mo ago When Does On-Policy Interaction Help? Representational Tradeoffs in Value-Based Imitation Learning arXiv:2607.29617v1 Announce Type: new Abstract: Imitation learning (IL)---training an agent to replicate expert behavior from demonstrations---underpins applications from robotics to language model training. Standard approaches such as Behavior Cloning (BC) are known to suffer… 24 arXiv — NLP / Computation & Language research 1mo ago WCM: A World Critic Model for Vision-Language-Action Reinforcement Learning arXiv:2607.29613v1 Announce Type: cross Abstract: Reinforcement learning (RL) post-training of Vision-Language-Action (VLA) models has shown strong promise for robotic manipulation. Among RL methods, critic-based approaches rely on a value estimator that predominantly operates… 34 Hugging Face Daily Papers research 1mo ago N_0-VTLA: Scaling Vision-Tactile-Language-Action Model with Latent Tactile Tokens Abstract We present N_0-VTLA, a vision-tactile-language-action (VTLA) foundation model capable of (1) fine-grained contact-rich manipulation with tactile perception and tactile-feedback control, and (2) offline policy improvement from stored deployment data. Building on current… 24 arXiv — Machine Learning research 1mo ago MUGEN: A Unified Framework for Efficient Motion Understanding and Generation arXiv:2607.27581v1 Announce Type: new Abstract: Grounding human motion in language, and language in motion, is a central step toward physical AI systems that can understand, generate, and communicate human behavior. Unified motion--language systems first coupled the two… 4 Hugging Face Daily Papers research 1mo ago ACE-Data-0: Human-Centric Ambient Capture as Embodied Data Engine Abstract Embodied intelligence faces a fundamental data bottleneck. Models must capture how first-person perception, whole-body motion, dexterous manipulation, object state, sound, and touch evolve together as humans pursue goals over time. Existing datasets fragment this… 24 MIT News — AI research 1mo ago Daniela Rus receives Bavarian Minister-President's High-Tech Prize Director of CSAIL and MIT professor honored for her contributions to robotics, artificial intelligence, and autonomous systems. 21 Ars Technica — AI news-outlet 1mo ago Google reveals Gemini Robotics 2.0, promising improved dexterity and safety Gemini Robotics 2 includes three models, but only one is publicly available right now. 22 r/LocalLLaMA community 1mo ago How close are we to local llama robotics for consumer price point? I'm guessing 3 years, what do you think? In other words: many of us will be able to afford a general purpose robot in 3 years to experiment with in the home. Cost roughly $5k? Probably small size, but hopefully still able to do the dishes and operate a vacuum.   submitted by… 13 Hugging Face Daily Papers research 1mo ago πR^2: Reactive Real-time Flow Policies Abstract Generalist manipulation policies increasingly take the form of action-chunking flow policies built on large pretrained backbones. Such chunks run open-loop, so the policy cannot react to sensory input arriving mid-execution, sacrificing reactivity. Replanning more often… 4 Google DeepMind official-blog 1mo ago Gemini Robotics ER 2: powering robotics with video understanding, task orchestration, and multi-robot collaboration Gemini Robotics ER 2 helps robots reason, collaborate, and solve real-world tasks. It represents a step change in video understanding, tool orchestration, and multi-robot collaboration for robotic applications. 5 arXiv — Machine Learning research 2mo ago Conformal Changepoint Localization and Root Cause Analysis with Corrupted Observations arXiv:2607.26481v1 Announce Type: new Abstract: Detecting when the statistical behavior of an engineered system changes, and identifying which component is responsible, are core problems in the monitoring of telecommunication networks, robotic platforms, security infrastructure,… 10 arXiv — NLP / Computation & Language research 2mo ago The Reliability of LLMs for Medical Diagnosis: An Examination of Consistency, Manipulation, and Contextual Awareness arXiv:2503.10647v2 Announce Type: replace Abstract: This study evaluated the diagnostic reliability of two Large Language Models (LLMs), Google Gemini 2.0 Flash and OpenAI ChatGPT-4o, across three dimensions: consistency under rephrased inputs, susceptibility to irrelevant… 10 Hugging Face Daily Papers research 2mo ago TurboVLA: Real-Time Vision-Language-Action Model at 32 Hz on an RTX 4090 with <1 GB VRAM Abstract Vision-language-action (VLA) models commonly adopt an LLM-centric V to L to A pathway, where visual observations are projected into the representation space of a large language model before being decoded into robot actions. Although effective, this design incurs… 36 Hugging Face Daily Papers research 2mo ago Explicit Layer Modeling for Video Object Insertion and Layer Decomposition Abstract Most video editing systems still lack explicit layered video representations, limiting their ability to perform realistic compositing, object reuse, and consistent manipulation. This limitation is especially pronounced in video object insertion and video layer… 26 Ars Technica — AI news-outlet 2mo ago Who wins and who loses after US bans foreign robots? Government ban on foreign-made robots may hinder instead of help US robotics. 18 Hugging Face Daily Papers research 2mo ago HiFi-UMI: Learning Deployable Manipulation Policies from High-Fidelity UMI Data Alone Abstract Learning deployable manipulation policies is bottlenecked by the scarcity of data that is both high-fidelity and scalable. Real-robot teleoperation is accurate but costly to scale; robot-free UMI capture scales readily, and current practice uses the resulting data… 13 r/LocalLLaMA community 2mo ago I got Kimi-k3 running..... Results: prompt eval: 40 tokens / 97.5s → 0.41 tok/s eval: 400 tokens / 1769.9s → 0.23 tok/s total: 440 tokens / 1867s (31 min) Prompt: "Write a C++ function that reverses a linked list in place. Explain the pointer manipulation." How I ran it: Using PR#26185 from llama.cpp… 32 NVIDIA Developer Blog official-blog 2mo ago Developing Healthcare Robotics with GPU-Native Medical Physics Simulation Unlike autonomous driving or industrial robotics, healthcare robotics can’t rely on internet-scale data collection or unlimited real-world experimentation.... 28 r/MachineLearning community 2mo ago NeurIPS-side prompt injection triggering ethics reviewers? [D] Does anyone experience a similar story that some reviewers reporting ethical issue due to NeurIPS-side prompt injection for catching LLM-reviewers? Even ethics reviewers were not informed about this conference-side manipulation…   submitted by   /u/dontknowwhattoplay… 19 Hugging Face Daily Papers research 2mo ago WorldDiT: A Unified Diffusion Architecture for World and Action Modeling Abstract Many recent robot policies pursue stronger control by using large pretrained vision-language models (VLMs) as the action backbone. We introduce WorldDiT, a unified diffusion transformer architecture that couples action generation with visual world modeling and achieves… 35 Page 4 of 10 · 500 articles ← Newer Older →