PromptAI

Robotics — AI news & research

Embodied AI — robots, drones, self-driving systems and real-world manipulation. Updated continuously from 20+ curated sources.

AgentsVisionRoboticsPolicyResearchToolsLLMs
arxiv · Robotics

Skill-Space Shooting for Autonomous Robot Policy Improvement

Robots deployed in the physical world must be able to improve beyond their initial training as they encounter new situations and failures. For this improvement to scale across tasks, it must make effective use of experi…

arxiv · Robotics

Counterfactual Video Generation Enables Scalable Humanoid Loco-Manipulation

Teaching humanoids loco-manipulation skills, such as carrying diverse objects, via visual imitation is a promising path toward generalist robots. However, collecting diverse, high-quality interaction videos, such as cli…

arxiv · Robotics

Estimate, Don't Imitate: Reusing Differentiable State-Based Policies for Visuomotor Control

Simulation-trained manipulation policies can exploit privileged state information to learn effective contact-rich behaviours, but deployment requires acting from partial observations such as noisy camera images. A commo…

arxiv · Robotics

ZeroBot: Learning from Scratch in Minutes with Generative Real2Sim

We present ZeroBot, a real2sim framework for learning a robot manipulation task from scratch in minutes under challenging conditions: zero human demonstrations, zero policy pre-training, and zero known object models. Gi…

arxiv · Robotics

Towards VLA-Dreamer: Refining VLA Behavior Using World Models

Vision-Language-Action models (VLAs), while showing strong potential for robot control, require massive amounts of high-quality imitation learning data. Moreover, the absence of an explicit world model casts further dou…

arxiv · Robotics

RAPID: Robot Agentic Programming from Demonstrations

Coding agents have demonstrated enormous success in solving complex programming problems. To leverage their potential for robot systems, this work introduces Robot Agentic Programming from Demonstrations (RAPID), which…

arxiv · Robotics

Rolling-WAM: World Action Models with Rolling Imagination

World Action Models (WAMs) couple action generation with future visual prediction for robotic manipulation. However, completing the joint video-action denoising process at each replanning cycle incurs substantial latenc…

arxiv · Robotics

Coding Agents for Generalized Task and Motion Planning Problems

Task and motion planning (TAMP) problems remain difficult even with full observability and object-centric states because discrete decisions are tightly coupled to geometric, kinematic, and dynamic constraints. Generaliz…

arxiv · Robotics

Underwater C3-JEPA: An Object-Centric Cross-View World Model for ROV Salvage

We present Underwater C$^{3}$-JEPA (cross-view, control-conditioned, context-extended), an object-centric multi-view predictive world model for near-field heavy-load underwater ROV salvage. Without contact sensors, it p…

arxiv · Robotics

Where Should I Join? Robot Group Joining via Language-Guided Goal Prediction

Social navigation typically assumes a specified goal and focuses on reaching it while respecting social conventions, whereas robot group joining requires predicting where to join based on the group's real-time activity…

arxiv · Robotics

ForgetMimic: Motion Unlearning for Reinforcement Learning Humanoid Control

Humanoid control, leveraging human demonstrations, has achieved diverse, agile, and natural locomotion behaviors through reinforcement learning (RL). While this paradigm has yielded remarkable performance in physical hu…

arxiv · Robotics

LEAP-CBF: A Safety Filter for Uncertain Systems with Least-Effort Adversarial Potentials

Control barrier functions (CBF) are a popular safety filter to ensure safety for nonlinear dynamical systems. However, when the system is subject to uncertainties and disturbances, this requires the use of robust varian…

arxiv · Robotics

Less Language, More Latents: Annotation-Efficient VLAs for Driving

Vision-language-action models (VLA) promise human-steerable autonomous driving, but their training is bottlenecked by the scarcity of frames paired with natural-language instructions: while camera streams and expert tra…

arxiv · Robotics

InfiNoVA: Infinite Novel View Augmentation for Viewpoint Invariant Robot Policies

Vision-Language-Action (VLA) policies often rely strongly on the camera viewpoints seen during training, causing substantial performance degradation when deployed from unseen perspectives. Collecting demonstrations from…

arxiv · Robotics

DAVIO: Dense Monocular-Inertial SLAM with Feed-Forward Initialization and Pose-Conditioned Mapping

A camera and an IMU are the minimal sensor setup for metric localization and dense mapping, yet classical visual--inertial filters must wait for parallax before they start and then retain only sparse landmarks. Feed-for…

hf · Robotics

How to Use NVIDIA Warp and MjWarp to Accelerate Robotics Simulation and Learning Workflows

arxiv · Robotics

φ-RIE: From Photorealistic Reconstruction to Interactive Environments

3D Gaussian Splatting (3DGS) can reconstruct a captured scene photorealistically, but the resulting representation does not by itself support physical interaction. Robot simulation instead requires object-level change,…

arxiv · Robotics

DreamStream: Towards Policy-Oriented Generative Simulation for End-to-End Driving

Faithfully evaluating end-to-end driving policies in simulation requires observations that are not merely photo-realistic, but preserve the scene features a policy relies on to make decisions. Existing platforms, howeve…

arxiv · Robotics

MAVP: Map-Aware Visuomotor Policies for Mobile Manipulation

Successful mobile manipulation requires coordinated base and arm motion while maintaining accurate spatial positioning. However, demonstration-trained policies can struggle to realise the intended base motion reliably,…

arxiv · Robotics

TriWorldBench: A Tri-View Consistency Perspective on Embodied World Models

Embodied world models predict the outcomes of robot actions to support learning and planning. For robots equipped with head and wrist cameras, this requires complementary views: the head view captures the overall task,…

tc · Robotics

TechCrunch Disrupt 2026: Aaron Edsinger brings Hello Robot’s Stretch 4 to life onstage

Hello Robot CEO and co-founder Aaron Edsinger will bring Stretch 4 for a live demo on the Real World AI Stage at TechCrunch Disrupt 2026. Register before September 25 to save up to $200, plus get a second pass at 50% of…

arxiv · Robotics

DexTacWAM: A Visuo-Tactile World-Action Model for Dexterous Manipulation

Dexterous manipulation depends on contact dynamics that are often only partially observable from vision. Recent World-Action Models (WAMs) couple predictive video world modeling with action generation, but remain largel…

arxiv · Robotics

Visuomotor Robotic Pruning in Planar Orchards Using Hybrid Reinforcement Learning

Dormant tree pruning is labor-intensive yet essential for maintaining modern high-productivity fruit orchards. In this work, we focus on pruning of modern planar tree training systems - V-Trellis apples and UFO cherries…

hn · Robotics

Go-based Robotics Framework built around NATS.io

arxiv · Robotics

Coding Agents with an Obstacle-Aware Harness for Safe Robot Manipulation

Coding agents have emerged as a promising paradigm for robot manipulation: a language model writes the robot controller as a program, and agents built in this way now operate robots without robot-specific training.Wheth…

arxiv · Robotics

Workspace Models: Lightweight Robotic Memory via Saliency-Driven Supervision

Complex robotic manipulation tasks frequently require a long-term memory of past events and actions. As conditioning on full histories renders policies prone to spurious correlations and degrades performance, many appro…

arxiv · Robotics

GeoAAC: Geometry-Based Adaptive Action Chunking from Denoising Trajectories in VLA Policies

Action chunking is widely used for action generation and execution in Vision-Language-Action (VLA) policies, yet existing approaches commonly use a fixed action horizon. During a rollout, different task stages may requi…

arxiv · Robotics

Agile-WAM: An Agile Tactile World Action Model for Contact-Rich Robot Control

World Action Models (WAMs) advance beyond conventional visuomotor policies by jointly predicting future world states and robot actions, enabling the policy to learn physical dynamics that support effective control. Howe…

arxiv · Robotics

OPTED: On-Policy Fine-Tuning for End-to-End Driving using a Render-Free Teacher

As scaling pre-training data alone yields diminishing returns, post-training is becoming increasingly important across physical AI domains such as autonomous driving. End-to-end driving policies are pre-trained in open…

arxiv · Robotics

MILER: Semantic Mid-Level Representation for Sim-to-Real Reinforcement Learning in Unstructured Autonomous Driving

Reinforcement learning constitutes a promising approach owing to its potential for superhuman performance and self-learned policies. However, its application to real-world autonomous driving remains scarce, particularly…

tc · Robotics

Iceland-based Treble raises $18 million for its voice simulation platform

Treble's voice simulation platform is used by voice AI model developers, AI wearable, and robotics companies

arxiv · Robotics

Dreaming the Sound of Contact: Leveraging Video and Audio Generation for Zero-Shot Force-Aware Manipulation and Data Generation

Recent advances in video generation allow robots to learn manipulation trajectories from generated videos. However, these approaches produce purely kinematic trajectories that lack force information, causing failures in…

arxiv · Robotics

rMuscle: Robotic Muscle Memory for Efficient Vision-Language-Action Model Inference

Factory work is a promising early scenario for embodied AI: assigning repetitive manual jobs to robots has clear economic payoff, and a structured station keeps the jobs tractable for current policies. Vision-Language-A…

tc · Robotics

Robots are waiting for a ChatGPT moment: Nvidia’s Les Karpas explains why at TechCrunch Disrupt 2026

The robotics industry is still waiting for their breakthrough into day-to-day life. Nvidia's Les Karpas has an answer as to why at TechCrunch Disrupt 2026. Register before September 25 to save up to $200 on your pass.

arxiv · Robotics

SlipSense: Multimodal Tactile Learning for Low-Latency and Generalized Slip Detection

Slip detection is fundamental to dexterous manipulation, yet existing systems often lack precise characterization of detection latency and cross-platform generalization. We present SlipSense, a multimodal tactile slip-d…

tc · Robotics

Discover how to take your startup from prototype to production at TechCrunch Disrupt 2026

Learn how to scale your startup breakthrough from prototype to production at TechCrunch Disrupt 2026 with scaling leaders, Adrian Macneil (Foxglove), John Mackey (MBRYONICS), and Boris Sofman (Bedrock Robotics. Register…

arxiv · Robotics

ASTRIL-MPC: Autonomous Traversal Framework of Articulated Tracked Robots with Language-Guided Neural-Kinematic MPC

In urban search and rescue, articulated tracked robots (ATRs) must traverse structured but contact-rich environments such as stairwells and cluttered building interiors. Reliable autonomy remains challenging because rob…

arxiv · Robotics

Dynin-Robotics: Omnimodal Unified Diffusion Vision-Language-Action Model

Visual goal and dynamics prediction can provide language-conditioned robot policies with both a target outcome and a representation of action-dependent scene changes. We bring these predictions into action generation an…

marktechpost · Robotics

NVIDIA Open-Sources OSMO: One YAML Orchestrates Physical AI Training, Simulation, and Robot Testing

NVIDIA has open-sourced OSMO, the Kubernetes-native workflow orchestrator it uses internally for Project GR00T, Isaac Lab, and Isaac Sim. OSMO lets robotics teams define training, simulation, and hardware-in-the-loop ta…

tc · Robotics

Mecka AI nears $500M valuation in Sequoia-led deal amid rush for robot training data

The round for the two-year-old startup is coming together months after Mecka announced its Series A.

tc · Robotics

Maven Robotics wants to steal your robot deployment deal

Maven Robotics emerged from stealth today with a $100 million Series A and active deployments.

arxiv · Robotics

Show-Harness: Just a VLM Agent Can Play Robots

Foundation vision-language models (VLMs) exhibit broad intelligence about the world, yet translating this intelligence into robot control remains challenging. We present Show-Harness, an Embodied Harness that enables VL…

arxiv · Robotics

DUET-DINO: Simultaneous Cross-View World Modeling for Latent Planning in Robot Manipulation

Action-conditioned latent world models predict future visual representations, enabling zero-shot goal-conditioned robot planning and control. However, their predictions for fine-grained spatial and rotational actions ar…

arxiv · Robotics

Multi-Agent Reinforcement Learning for Autonomous UAV Exploration in Wildfire Response

This study develops a deep reinforcement learning framework for training Unmanned Aerial Vehicle (UAV) agents to navigate and monitor simulated wildfire environments. Results show that agents learn increasingly stable a…

arxiv · Robotics

TANGO: Humanoid Navigation in Cluttered Environments with a Whole-Body Vision-Language-Action Model

We study the problem of navigating cluttered indoor environments with a humanoid robot. Unlike conventional methods that model navigation as a 2D path planning problem, humanoid traversal in cluttered environments requi…

arxiv · Robotics

DeCAL: Towards Physically-Grounded Dexterous Vision-Language-Action Models via Contact-Aware Latent Co-Imagination

Dexterous manipulation involves contact-rich and fine-grained interactions with the physical world, posing significant challenges for existing vision-language-action (VLA) models due to severe visual occlusions and comp…

arxiv · Robotics

Rethinking Learned Occupancy in Autonomous Active Mapping with Observation-Gated Filtering

Autonomous 3D active mapping requires a space robot to choose where to sense while building the geometry needed for navigation. Learned occupancy completion extends spatial context beyond the current field of view, but…

arxiv · Robotics

Same Trajectory, Contradictory Rewards (ROBORMBENCH): Paraphrase Fragility in Vision Language Reward Models

Vision-language models are increasingly used as reward functions for robotic learning, but this role requires paraphrase invariance: the same trajectory should receive the same reward under semantically equivalent goal…

arxiv · Robotics

What Matters, When? Diagnosing and Improving Conditional Visual Grounding in Visuomotor Imitation Policies

Visuomotor imitation policies can achieve high performance under in-distribution visual conditions yet fail when visually similar objects or receptacles are introduced. We study this behavior as a problem of conditional…

arxiv · Robotics

Towards Neuro-Symbolic Procedural Reasoning for Long-Horizon Vision-Language-Action Manipulation

Vision-language-action (VLA) models can execute short manipulation skills, but remain brittle in long-horizon procedures requiring persistent task state, dependency-aware reasoning, conditional decisions, and reliable g…

hn · Robotics

GPT-6 Astra on robot arms

arxiv · Robotics

FWBC-VLA: Force-Aware Whole-Body Compensation for Contact-Rich Loco-Manipulation

Contact-rich loco-manipulation requires a bridge between semantic action generation and physical interaction control. Existing Vision-language-action (VLA) models generate task-level actions from visual and linguistic o…

arxiv · Robotics

Towards Trustworthy Autonomous Robots: An Explainable AI-Based Decision Framework

Autonomous robots powered by deep learning face a fundamental auditability challenge: when incidents occur, investigators cannot reconstruct why the system made specific decisions. This paper presents TRACE (Transparent…

hn · Robotics

Reasons Robotics Is Hard

arxiv · Robotics

Facet-0: A Robotic Foundation Model for Contact-Rich Precise Manipulation

Real-world robotic assembly at sub-millimeter tolerances demands spatial precision, compliant interaction, and robustness to contact failures. We present Facet-0, a robotic foundation model that predicts and values the…

arxiv · Robotics

SUN: Persistent Programs For Language-Grounded Control-to-Learning-to-Real Policies

Bridging model-based control and learned policies in long-horizon manipulation has harbored a silent disagreement: control executes specified objectives, learning amortizes that behavior into a reactive policy, yet exis…

arxiv · Robotics

$\mathcal{N}_0$-Foundation: Towards the Age of Tactile Intelligence

We present $\mathcal{N}_0$-Foundation, a paradigm for tactile-enabled embodied manipulation, which integrates tactile sensing hardware, large-scale multimodal data, tactile representation learning, and standardized eval…

arxiv · Robotics

Aero Hand Open: A Simulation-Ready Tendon-Driven Hand for Dexterous Manipulation Learning

Tendon-driven hands are anthropomorphic, and moving the actuators off the joints is what makes a hand of this capability affordable to build. Two effects produce that saving. Routing force through a cable removes the re…

arxiv · Robotics

CLAP: Cross-Embodiment Video World Models are Zero-Shot Physical Simulators

State-of-the-art action-conditioned video models are typically restricted to a single robot embodiment, preventing them from leveraging the vast corpus of heterogeneous video data that contains rich signals for learning…

hn · Robotics

Pollen Robotics (Hugging Face) Microduck

Browse briefing issues →

Get the briefing

The one story that matters, 5 headlines and the paper everyone's citing — every Tuesday, free.

Subscribe free