News / #robotics Tag Robotics 362 articles archived under #robotics · RSS Sign in to follow TechCrunch — AI news-outlet 2mo ago Jeff Bezos’s Prometheus raises $12B to build an ‘artificial general engineer’ for the physical world The new round values the physical AI startup that aims to automate heavy engineering and drug design at $41 billion. 31 r/LocalLLaMA community 2mo ago Refiner: Robotics library from the ex-Hugging Face pre-training team ex-Huggingface pre-training team just announce a new library create for robotics data refinment! It supports ingestion of all robotics formats (Parquet, HDF5, MCAP, Zarr, RLDS, and LeRobot), as well as the common processing flows like visual hand-tracking, subtask annotations… 26 arXiv — Machine Learning research 2mo ago Implicit Neural Representations of Individual Behavior arXiv:2606.12200v1 Announce Type: new Abstract: We study policy representation learning from unlabeled multi-policy behavioral data. Each episode is generated by a fixed policy, but policy labels are unavailable. This setting appears in robotics play, demonstrations, games,… 28 arXiv — Machine Learning research 2mo ago Fourier Features Let Agents Learn High Precision Policies with Imitation Learning arXiv:2606.12334v1 Announce Type: new Abstract: High-precision robotic manipulation requires fine-grained spatial reasoning that is often difficult to achieve with RGB-only policies due to depth ambiguity and perspective scale issues. Policies that leverage 3D information… 14 arXiv — NLP / Computation & Language research 2mo ago Detecting AI-Generated Content on Social Media with Multi-modal Language Models arXiv:2606.11200v1 Announce Type: new Abstract: Generative AI has enabled the creation of photorealistic images and videos that are increasingly disseminated on social media, often used for spam, misinformation, manipulation, and fraud. Existing AI-generated content (AIGC)… 36 arXiv — NLP / Computation & Language research 2mo ago When Does Language Matter? Multilingual Instructions Reveal Step-wise Language Sensitivity in Vision-Language-Action Models arXiv:2606.11906v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models have shown strong performance in language-conditioned robotic manipulation, yet their robustness to linguistic variation remains poorly understood. In this work, we present the first systematic… 17 Hugging Face Daily Papers research 2mo ago World Pilot: Steering Vision-Language-Action Models with World-Action Priors Abstract World Pilot enhances Vision-Language-Action models by incorporating dynamic scene evolution and trajectory priors from a World-Action Model, achieving superior performance in zero-shot out-of-distribution manipulation tasks. Generated by Qwen/Qwen2.5-Coder-32B-Instruct… 10 Hugging Face Daily Papers research 2mo ago BrainSurgery: Reproducible and Reliable Declarative Weight Manipulations for Model Editing and Upcycling Abstract BrainSurgery is a tool for robust and reproducible tensor manipulation of neural network checkpoints through declarative YAML plans with built-in validation. Generated by Qwen/Qwen2.5-Coder-32B-Instruct As deep learning models scale, managing, inspecting, and modifying… 12 arXiv — Machine Learning research 2mo ago Co-GLANCE: Uncertainty-Aware Active Perception for Heterogeneous Robot Teaming arXiv:2606.09919v1 Announce Type: new Abstract: Perceptual uncertainty is a central challenge for heterogeneous robot teams operating in unstructured outdoor environments, where no single viewpoint affords reliable scene understanding. Perceptual uncertainty, arising from… 36 arXiv — NLP / Computation & Language research 2mo ago TabClaw: An Interactive and Self-Evolving Agent for Spreadsheet Manipulation and Table Reasoning arXiv:2606.10316v1 Announce Type: new Abstract: Spreadsheets and tables are widely used representations for structured data analysis, but effective analysis still requires substantial manual effort and domain expertise. Recent large language model (LLM) agents can automate parts… 31 arXiv — NLP / Computation & Language research 2mo ago Beyond APIs: Probing the Limits of MLLMs in Physical Tool Use arXiv:2606.10803v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) excel at utilizing digital APIs and increasingly serve as the "brain" of embodied AI, instructing robots to interact with the physical world. In such embodied settings, a central capability… 38 Hugging Face Daily Papers research 2mo ago ABot-Earth 0.5: Generative 3D Earth Model Abstract ABot-Earth 0.5 generates realistic 3D environments from satellite imagery using 3D Gaussian Splatting representation, enabling fast synthesis and real-time visualization for Embodied AI applications. Generated by Qwen/Qwen2.5-Coder-32B-Instruct We present ABot-Earth… 22 Hugging Face Daily Papers research 2mo ago VoLo: A Physical Orchestrator for Open-Vocabulary Long-Horizon Manipulation Abstract VoLoAgent enables physical orchestration by integrating vision-language models with robot capabilities for open-vocabulary long-horizon manipulation tasks. Generated by Qwen/Qwen2.5-Coder-32B-Instruct Open-vocabulary long-horizon manipulation requires robots to reason… 24 The Information — AI news-outlet 2mo ago Kalshi Asks Some Customers For Employer Information Prediction markets platform Kalshi is asking customers in some wagers to provide the name of their employer, industry and job function before making bets, to help the company crack down on potential insider trading. “For markets with heightened insider or manipulation risk, we… 23 TechCrunch — AI news-outlet 2mo ago Hey Siri, here’s what I actually want from AI I'm desperate for a personal AI assistant, but do I really want to become the kind of person who can't function without the friendly robot voice in my phone? 4 Hugging Face Daily Papers research 2mo ago Robotic Policy Adaptation via Weight-Space Meta-Learning Abstract WIZARD is a weight-space meta-learning framework that generates task-specific LoRA parameters for frozen VLA policies using language instructions and demonstration videos, enabling efficient task adaptation without fine-tuning. Generated by… 31 Hugging Face Daily Papers research 2mo ago Light-WAM: Efficient World Action Models with State-Fusion Action Decoding Abstract Light-WAM is a lightweight world action model for robot manipulation that uses a compact video backbone and downsampled latent space for efficient future-video supervision, combined with a StateFusionActionExpert for direct action prediction. Generated by… 25 Google DeepMind official-blog 2mo ago Powering the future of robotics in Europe Powering the future of robotics in Europe Jun 09, 2026 · Share x.com Facebook LinkedIn Mail Google DeepMind Accelerator selects 15 robotics companies from across Europe to join the program. Providing 3 months of intensive mentorship and technical support, enabling the… 22 Hugging Face Daily Papers research 2mo ago WorldCraft: From Camera Navigation to Object Manipulation in Interactive Video World Models Abstract WorldCraft extends interactive video world models to enable object-level trajectory control while maintaining camera navigation capabilities through specialized control pipelines. Generated by Qwen/Qwen2.5-Coder-32B-Instruct Recent video-based world models have made… 27 r/LocalLLaMA community 2mo ago Jetson Orin NX Build for Hermes Agent + Benchmarking I had a huge LLM server , and now I have a tiny one! I had a Jetson Orin NX gathering dust from a long dead robotics project, from back in the Llama-7B days. I figured now with MoE and smaller models doing well, it was time to mess with it again. Goal: As silent as possible… 34 The Information — AI news-outlet 2mo ago U.S. Accuses Alibaba, Baidu, Others of Aiding Chinese Military in Blacklist Move The U.S. Department of Defense on Monday added more than a dozen Chinese tech companies including Alibaba and Baidu to a blacklist, a move that could further escalate tensions between the world’s two largest economies. Electric vehicle makers Byd and Nio, humanoid maker Unitree,… 25 Hugging Face Daily Papers research 2mo ago AHA-WAM:Asynchronous Horizon-Adaptive World-Action Modeling with Observation-Guided Context Routing Abstract AHA-WAM is an asynchronous world-action model that uses dual Diffusion Transformers to enable efficient long-horizon planning and real-time action execution in robotic manipulation tasks. Generated by Qwen/Qwen2.5-Coder-32B-Instruct World-action models have emerged as a… 4 Hugging Face Daily Papers research 2mo ago OASIS: From Simulation Data Collection to Real-World Humanoid Loco-Manipulation Abstract A simulation-data-driven framework for humanoid loco-manipulation that uses 3D generative models to create realistic assets and hierarchical visuomotor policies trained on simulated data achieves better zero-shot performance than real-robot training. Generated by… 24 Hugging Face Daily Papers research 2mo ago Robots Need More than VLA and World Models Abstract Robot intelligence advancement requires integrating unstructured behavioral data through specialized interfaces for labeling, embodiment mapping, world modeling, and reward inference rather than relying solely on policy scaling. Generated by… 27 arXiv — NLP / Computation & Language research 2mo ago The Sim-to-Real Gap of Foundation Model Agents: A Unified MDP Perspective arXiv:2606.07017v1 Announce Type: cross Abstract: Foundation model agents are increasingly deployed for real-world decision-making, but suffer from the sim-to-real gap. While robotics and classical control have mature frameworks to address this gap, the foundation model… 5 Hugging Face Daily Papers research 2mo ago LIMMT: Less is More for Motion Tracking Abstract Training with high-quality motion data improves tracking policy optimization trajectories, with minimal data subsets outperforming full datasets in physics-based humanoid motion tracking. Generated by Qwen/Qwen2.5-Coder-32B-Instruct We argue that high-quality motion… 24 Hugging Face Daily Papers research 2mo ago The Road Ahead in Autonomous Driving: The KITScenes Multimodal Dataset Abstract KITScenes Multimodal dataset provides high-fidelity European driving data with comprehensive 3D maps and diverse urban environments for embodied AI research. Generated by Qwen/Qwen2.5-Coder-32B-Instruct Existing autonomous driving datasets have enabled major progress,… 36 Hugging Face Daily Papers research 2mo ago AffordanceVLA: A Vision-Language-Action Model Empowering Action Generation through Affordance-Aware Understanding Abstract AffordanceVLA introduces a unified framework that uses structured affordance forecasting as an intermediate representation to improve the precision of perception-action mapping in robotic manipulation by leveraging vision-language models. Generated by… 4 r/MachineLearning community 2mo ago I'm looking to join/form a team working on physical AI robotics challenge [P] Hey all, I'm a robotics engineer by training turned ML/AI engineer because of passion right after school. I want to start combining these skills together and I think a competition is the best way of doing it. Here's an example of a challenge I'm talking about to set expectations… 18 Hugging Face Daily Papers research 2mo ago World-Language-Action Model for Unified World Modeling, Language Reasoning, and Action Synthesis Abstract World-language-action models combine textual instruction processing with robot state prediction through an autoregressive transformer backbone, enabling efficient long-horizon task execution and cross-embodiment learning. Generated by Qwen/Qwen2.5-Coder-32B-Instruct We… 7 Hugging Face Daily Papers research 2mo ago Dream.exe: Can Video Generation Models Dream Executable Robot Manipulation? Abstract Video generation models were evaluated through robotic manipulation tasks to assess their ability to reflect physical reality, revealing that visual quality does not predict executable motion accuracy. Generated by Qwen/Qwen2.5-Coder-32B-Instruct Video generation models… 20 Hugging Face Daily Papers research 2mo ago SEAOTTER: Sensor Embedded Autoencoding with One-Time Transcode for Efficient Reconstruction Abstract A compression framework for cloud robotics combines learned latent representations with standard JPEG compatibility to achieve faster encoding and decoding while maintaining high perceptual quality. Generated by Qwen/Qwen2.5-Coder-32B-Instruct In robotics systems, vast… 31 r/MachineLearning community 2mo ago Would you say capture-time semantic annotation for robot trajectories is a solved problem? [R] It seems raw teleoperation data (RGB + joint states) structurally lacks affordance, contact intent, and embodiment-specific kinematic context. (information that can't be reliably recovered post-hoc once the demonstration is recorded) Most current approaches either filter/clean… 11 Hugging Face Daily Papers research 2mo ago RobotValues: Evaluating Household Robots When Human Values Conflict Abstract RobotValues benchmark evaluates household robot planners in value-conflict scenarios, revealing that vision-language models exhibit default value preferences and struggle to override them when instructed to prioritize conflicting values. Generated by… 8 arXiv — Machine Learning research 2mo ago Flash-WAM: Modality-Aware Distillation for World Action Models arXiv:2606.05254v1 Announce Type: new Abstract: World-action models (WAMs) jointly generate future video and robot actions through iterative diffusion, achieving strong performance on manipulation benchmarks but requiring tens of denoising steps, a cost that precludes real-time… 13 arXiv — Machine Learning research 2mo ago What Objects Enable, Not What They Are: Functional Latent Spaces for Affordance Reasoning arXiv:2606.05533v1 Announce Type: new Abstract: Existing robot planning systems rely on appearance-based reasoning, where visual observations are encoded into latent spaces organized around object appearances (e.g., recognizing a "cart" based on how it looks). However, planning… 13 Ars Technica — AI news-outlet 2mo ago The skeptic’s guide to humanoid robots going viral on the Internet Robot demonstrations can distort public perceptions of robotic capabilities. 9 Dwarkesh Podcast news-outlet 2mo ago Alex Imas and Phil Trammell – What remains scarce after AGI? “One robot now turns into many robots next year, but the number of ballerinas is the same.” 37 TechCrunch — AI news-outlet 2mo ago Is Silicon Valley ready to put robots in people’s homes? Hello Robot is. The California startup released the fourth-generation of its home assistance robot, Stretch. 30 Hugging Face Daily Papers research 2mo ago PaintBench: Deterministic Evaluation of Precise Visual Editing Abstract PaintBench presents a scalable benchmark for precise visual editing tasks, revealing low performance across models and identifying key challenges in geometric transformations and structural manipulations. Generated by Qwen/Qwen2.5-Coder-32B-Instruct While current… 12 Hugging Face Daily Papers research 2mo ago Cosmos 3: Omnimodal World Models for Physical AI Abstract Cosmos 3 is an omnimodal world model that processes and generates multiple data types through a unified mixture-of-transformers architecture, achieving state-of-the-art performance in various understanding and generation tasks. Generated by… 38 Hugging Face Daily Papers research 2mo ago OVO-S-Bench: A Hierarchical Benchmark for Streaming Spatial Intelligence in Multimodal LLMs Abstract OVO-S-Bench presents a comprehensive benchmark for evaluating streaming spatial intelligence in multimodal language models through human-annotated questions spanning multiple abstraction levels. Generated by Qwen/Qwen2.5-Coder-32B-Instruct Multimodal agents in robotics,… 23 arXiv — NLP / Computation & Language research 2mo ago Hybrid Adversarial Defence for Natural Language Understanding Tasks arXiv:2606.04612v1 Announce Type: new Abstract: Large Language Models (LLMs) are vulnerable both to hallucination and adversarial manipulation. Although these problems are closely related, existing defences typically address them separately. We investigate a hybrid defence… 21 arXiv — NLP / Computation & Language research 2mo ago Dive into the Scene: Breaking the Perceptual Bottleneck in Vision-Language Decision Making via Focus Plan Generation arXiv:2606.04046v1 Announce Type: cross Abstract: In embodied vision-language decision making tasks such as robotic manipulation and navigation, Vision-Language and Vision-Language-Action Models (VLMs & VLAs) are powerful tools with different benefits: VLMs are better at… 31 Hugging Face Daily Papers research 2mo ago GRAIL: Generating Humanoid Loco-Manipulation from 3D Assets and Video Priors Abstract GRAIL generates diverse humanoid manipulation and locomotion data through 3D asset composition and video foundation models, enabling effective sim-to-real transfer for robot control. Generated by Qwen/Qwen2.5-Coder-32B-Instruct Scaling humanoid loco-manipulation… 9 Hugging Face Daily Papers research 2mo ago AURA: Action-Gated Memory for Robot Policies at Constant VRAM Abstract AURA-Mem is a recurrent memory system that adapts to embodied AI constraints by writing only when observations affect actions, significantly reducing memory writes compared to traditional KV-cache approaches. Generated by Qwen/Qwen2.5-Coder-32B-Instruct The KV-cache is… 7 Hugging Face Daily Papers research 2mo ago Humanoid-GPT: Scaling Data and Structure for Zero-Shot Motion Tracking Abstract Humanoid-GPT is a GPT-style Transformer with causal attention trained on a billion-scale motion corpus that achieves zero-shot generalization to unseen motions and control tasks through scalable pre-training on diverse motion data. Generated by… 29 Hugging Face Daily Papers research 2mo ago AFUN: Towards an Affordance Foundation Model for Functionality Understanding Abstract Affordance understanding model predicts functional masks and 3D motion curves from RGB-D observations and language descriptions, enabling generalizable robot manipulation across diverse environments. Generated by Qwen/Qwen2.5-Coder-32B-Instruct Affordance understanding… 36 Hugging Face Daily Papers research 2mo ago τ_0-WM: A Unified Video-Action World Model for Robotic Manipulation Abstract A unified video-action world model integrates policy learning, video prediction, and action evaluation using a shared video diffusion backbone for robotic manipulation tasks. Generated by Qwen/Qwen2.5-Coder-32B-Instruct Robotic manipulation requires models that generate… 22 Hugging Face Daily Papers research 2mo ago Silent Failures in Physical AI: A Literature Review of Runtime Action Authorization for Autonomous Systems Abstract Physical AI systems face safety challenges where black-box models can execute harmful actions without detection, necessitating comprehensive runtime guardrail mechanisms for safe operation. AI-generated summary Physical AI systems increasingly map multimodal… 12 Page 6 of 8 · 362 articles ← Newer Older →