<?xml version="1.0" encoding="UTF-8"?>
<?xml-stylesheet type="text/xsl" href="/feed.xsl"?>
<feed xmlns="http://www.w3.org/2005/Atom">
  <title>Khenda Resources Radar</title>
  <subtitle>Recent humanoid-robotics research from arXiv: manipulation, VLA models, imitation learning, locomotion, and more.</subtitle>
  <id>https://www.khendarobotics.com/resources</id>
  <link href="https://www.khendarobotics.com/resources" />
  <link rel="self" href="https://www.khendarobotics.com/research.xml" />
  <updated>2026-07-24T08:07:33.984Z</updated>
  <entry>
    <title>AXIS: A Growable Community-Driven Data Engine for Scalable Robot Manipulation</title>
    <id>arxiv:2607.21588</id>
    <link href="https://arxiv.org/abs/2607.21588" />
    <published>2026-07-23T17:58:08.000Z</published>
    <updated>2026-07-23T17:58:08.000Z</updated>
    <author><name>Mengfei Zhao</name></author>
    <author><name>Dihong Huang</name></author>
    <author><name>Yikai Tang</name></author>
    <author><name>Peihao Li</name></author>
    <author><name>Mingxuan Yan</name></author>
    <author><name>Ruiqi Zhuang</name></author>
    <category term="cs.RO" />
    <summary>Learning effective robot manipulation policies requires diverse, high-quality demonstrations, yet existing data pipelines are often difficult to scale because they rely on specialized hardware, centralized operators, or fixed task suites. We present AXIS, a growable community-driven data engine and benchmark for scalable robot learning, which enables browser-based teleoperation for large-scale demonstration collection, automatically generates and validates new manipulation tasks, and transforms community-collected demonstrations into training-ready data through automated success checking, qua…</summary>
  </entry>
  <entry>
    <title>GS-Agent: Creating 4D Physical Worlds With Generative Simulation</title>
    <id>arxiv:2607.21522</id>
    <link href="https://arxiv.org/abs/2607.21522" />
    <published>2026-07-23T17:04:36.000Z</published>
    <updated>2026-07-23T17:04:36.000Z</updated>
    <author><name>Hongxin Zhang</name></author>
    <author><name>Chunru Lin</name></author>
    <author><name>Junyan Li</name></author>
    <author><name>Zhou Xian</name></author>
    <author><name>Tsun-Hsuan Wang</name></author>
    <author><name>Chuang Gan</name></author>
    <category term="cs.RO" />
    <summary>Creating dynamic and physically realistic 4D worlds from natural language descriptions is both fascinating and challenging. Traditional computer graphics methods rely on manual creation, requiring extensive human effort to fine-tune materials, motions, and visual fidelity. Recent advances in generative foundation models have sparked interest in learning to generate such 4D worlds from large-scale data; however, existing methods still struggle to ensure physical plausibility and controllability. In this work, we take a different path by leveraging foundation models to construct an agentic syst…</summary>
  </entry>
  <entry>
    <title>TransBiolab: A Real-World Multi-View Dataset of Cluttered Transparent Biomedical Objects</title>
    <id>arxiv:2607.21071</id>
    <link href="https://arxiv.org/abs/2607.21071" />
    <published>2026-07-23T09:04:11.000Z</published>
    <updated>2026-07-23T09:04:11.000Z</updated>
    <author><name>Ke Ma</name></author>
    <author><name>Yifei Wang</name></author>
    <author><name>Meng Wang</name></author>
    <author><name>Tian Xia</name></author>
    <category term="cs.CV" />
    <summary>Autonomous biomedical laboratories increasingly rely on visual perception to recognize, localize, and manipulate transparent plasticware, yet high-quality real-world datasets for this setting remain limited. The scarcity of domain-relevant data is particularly restrictive in cluttered multi-object scenes, where mutual occlusion and view-dependent appearance changes remain challenging even for contemporary visual foundation models. Existing transparent-object datasets have advanced segmentation, depth, and pose estimation, but they usually do not evaluate the combined setting of multi-object c…</summary>
  </entry>
  <entry>
    <title>GuidedAttention: Interpretable and Correctable Visual Attention for OOD-Robust Robot Manipulation via Imitation Learning</title>
    <id>arxiv:2607.21049</id>
    <link href="https://arxiv.org/abs/2607.21049" />
    <published>2026-07-23T08:33:40.000Z</published>
    <updated>2026-07-23T08:33:40.000Z</updated>
    <author><name>Masaki Murooka</name></author>
    <author><name>Ryoichi Nakajo</name></author>
    <author><name>Keisuke Shirai</name></author>
    <author><name>Tomohiro Motoda</name></author>
    <author><name>Hanbit Oh</name></author>
    <author><name>Ryo Hanai</name></author>
    <category term="cs.RO" />
    <summary>End-to-end visuomotor policies provide little opportunity for humans to understand or correct the policy&apos;s visual attention. We propose GuidedAttention, a visuomotor imitation learning framework that introduces interpretable and correctable visual attention as an explicit intermediate representation. Task-relevant attention keypoints are predicted from camera images and condition a diffusion-based action policy. Users can inspect and optionally correct selected keypoints once at rollout initialization, after which the corrected attention is automatically propagated throughout execution by a t…</summary>
  </entry>
  <entry>
    <title>TableVerse: A Large-scale Tabletop Dataset with Real-world Grounded Layouts for Generalizable Manipulation</title>
    <id>arxiv:2607.21017</id>
    <link href="https://arxiv.org/abs/2607.21017" />
    <published>2026-07-23T08:02:53.000Z</published>
    <updated>2026-07-23T08:02:53.000Z</updated>
    <author><name>Boyuan Wang</name></author>
    <author><name>Yue Zhang</name></author>
    <author><name>Xutao Xue</name></author>
    <author><name>Xueyu Song</name></author>
    <author><name>Yu Sun</name></author>
    <category term="cs.RO" />
    <summary>The development of generalizable robotic manipulation policies is inherently bounded by the availability of large-scale, high-fidelity scene data. While recent automated synthesis methods attempt to bridge this gap via text-to-layout hallucination or simplified procedural generation, they frequently suffer from physical implausibility and fail to capture the complex, dense clutter of actual human environments. In this paper, we introduce TableVerse, a fully automated Real2Sim pipeline that shifts the paradigm from imaginative layout generation to deterministic reconstruction from unstructured…</summary>
  </entry>
  <entry>
    <title>HyWorldVLA: A Vision-Language-Action Model with Hybrid World Modeling for Autonomous Driving</title>
    <id>arxiv:2607.20988</id>
    <link href="https://arxiv.org/abs/2607.20988" />
    <published>2026-07-23T07:11:41.000Z</published>
    <updated>2026-07-23T07:11:41.000Z</updated>
    <author><name>Quanfu Yu</name></author>
    <author><name>Xian Wu</name></author>
    <author><name>Hao Xu</name></author>
    <author><name>Liulong Ma</name></author>
    <category term="cs.CV" />
    <summary>Vision-Language-Action (VLA) models augmented with world modeling represent a promising paradigm for end-to-end autonomous driving. While pixel-level future prediction enables fine-grained spatiotemporal reasoning, it compromises robustness in noisy driving scenarios. Conversely, latent-based world models alleviate this sensitivity but often incur limited interpretability and representational degradation due to absent pixel-level grounding. To reconcile this trade-off, we propose HyWorldVLA, a hybrid world-VLA framework that unifies pixel-level supervision and latent representation learning.…</summary>
  </entry>
  <entry>
    <title>URF: A Unified Robot Control-Policy Framework for Stable Contact Aware Manipulation</title>
    <id>arxiv:2607.20912</id>
    <link href="https://arxiv.org/abs/2607.20912" />
    <published>2026-07-23T04:46:19.000Z</published>
    <updated>2026-07-23T04:46:19.000Z</updated>
    <author><name>Jiyou Shin</name></author>
    <author><name>Youngjin Seo</name></author>
    <author><name>Jaeseog Won</name></author>
    <author><name>Sungwon Seo</name></author>
    <author><name>Hyunjun Kim</name></author>
    <author><name>Seokmin Yoon</name></author>
    <category term="cs.RO" />
    <summary>Learning-based manipulation policies usually predict robot actions from sensory observations and leave their execution to a separate low-level controller. In rigid contact, this separation can be problematic: the same motion to a virtual target or compliant motion command can lead to unstable contact, tracking error, excessive loading, or tool damage, depending on the low-level controller. In this paper, we propose a \textit{Unified Robot Control-Policy Framework} (URF), which connects compliant action prediction with unified impedance-admittance control. Given multimodal observations, URF pr…</summary>
  </entry>
  <entry>
    <title>Offline RL with Hierarchical Action Chunking</title>
    <id>arxiv:2607.20834</id>
    <link href="https://arxiv.org/abs/2607.20834" />
    <published>2026-07-23T01:48:46.000Z</published>
    <updated>2026-07-23T01:48:46.000Z</updated>
    <author><name>Ahad Jawaid</name></author>
    <category term="cs.LG" />
    <summary>Offline goal-conditioned reinforcement learning (RL) holds the promise of learning general-purpose policies from static datasets. However, scaling these methods to long-horizon tasks remains a challenge due to the curse of horizon, where value estimation errors can compound through long chains of bootstrapped Bellman backups. Existing hierarchical approaches mitigate this by decomposing tasks into subgoals, yet they often rely on low-level controllers that suffer from myopic execution and biased value estimates. In this work, we propose Hierarchical Implicit Q-Chunking (HiQC), an offline goal…</summary>
  </entry>
  <entry>
    <title>Emergent Compositional Skills in Mixture-of-Experts VLAs</title>
    <id>arxiv:2607.20771</id>
    <link href="https://arxiv.org/abs/2607.20771" />
    <published>2026-07-22T22:36:52.000Z</published>
    <updated>2026-07-22T22:36:52.000Z</updated>
    <author><name>Shlok Shah</name></author>
    <author><name>Rhiaan Jhaveri</name></author>
    <author><name>Tharun Kumar Tiruppali Kalidoss</name></author>
    <author><name>Chirayu Nimonkar</name></author>
    <author><name>Ishaan Javali</name></author>
    <category term="cs.RO" />
    <summary>We consider the problem of learning compositional robot policies end-to-end from expert demonstrations, without any pre-specified notion of task decomposition or hierarchy. We ask whether a VLA trained with a simplified Mixture-of-Experts (MoE) action head can emergently learn to decompose tasks into reusable, interpretable primitives. We find that learned experts are heavily reused across tasks and consistently correspond to qualitatively distinct low-level behaviors, suggesting that the router implicitly learns to perform high-level sequencing while experts serve as compositional primitives…</summary>
  </entry>
  <entry>
    <title>A real-time RGB-D perception pipeline for autonomous impact hammers in mining: self-filtering, rock segmentation and rock-breaking poses generation</title>
    <id>arxiv:2607.20748</id>
    <link href="https://arxiv.org/abs/2607.20748" />
    <published>2026-07-22T22:00:29.000Z</published>
    <updated>2026-07-22T22:00:29.000Z</updated>
    <author><name>Martín Gallegos</name></author>
    <author><name>Francisco Leiva</name></author>
    <author><name>Patricio Loncomilla</name></author>
    <author><name>Michelle Cortés</name></author>
    <author><name>Javier Ruiz-del-Solar</name></author>
    <category term="cs.RO" />
    <summary>Impact hammers, also known as rock-breakers, are essential machines in mining operations, where they perform secondary reduction. In underground mining, these machines are typically teleoperated, limiting operational efficiency. This paper presents a real-time RGB-D perception pipeline as a step towards automating the operation of hydraulic impact hammers used in mining. The proposed system simultaneously generates operationally feasible rock-breaking poses and a robot-free 3D representation of the workspace. The proposed approach combines image-based instance segmentation with geometric poin…</summary>
  </entry>
  <entry>
    <title>FELT: Generating Tactile Signals from Vision for Visuo-Tactile Manipulation</title>
    <id>arxiv:2607.20683</id>
    <link href="https://arxiv.org/abs/2607.20683" />
    <published>2026-07-22T19:32:42.000Z</published>
    <updated>2026-07-22T19:32:42.000Z</updated>
    <author><name>Zinan Li</name></author>
    <author><name>Yiyang Ling</name></author>
    <author><name>Yuming Gu</name></author>
    <author><name>Binghao Huang</name></author>
    <author><name>Chenhao Liang</name></author>
    <author><name>Sharfin Islam</name></author>
    <category term="cs.RO" />
    <summary>The sense of touch is central to manipulation, especially when vision is occluded or ambiguous. Although combining vision and touch improves manipulation, learning robust visuo-tactile policies requires substantial tactile data. Such data remains scarcer than visual data, because tactile sensors are fragile, specialized, and hard to standardize. To address this, we present Feature-Extracted Latent Tactile (FELT), a learning-based framework that synthesizes per-finger pressure tactile images from RGB observations, reducing the need for tactile-equipped data collection. FELT uses a large frozen…</summary>
  </entry>
  <entry>
    <title>PhysCoRe: Physics-Corrected Residual World Models for Material-Aware Deformable Dynamics</title>
    <id>arxiv:2607.20653</id>
    <link href="https://arxiv.org/abs/2607.20653" />
    <published>2026-07-22T18:25:57.000Z</published>
    <updated>2026-07-22T18:25:57.000Z</updated>
    <author><name>Haocheng Yin</name></author>
    <author><name>Shuohan Tao</name></author>
    <author><name>Yongsheng Chen</name></author>
    <author><name>Lu Gan</name></author>
    <category term="cs.RO" />
    <summary>Predicting how deformable objects evolve under robotic manipulation is a longstanding challenge. Existing approaches typically rely on per-object optimization to fit material parameters, which can be slow and cannot generalize, while end-to-end learned alternatives extrapolate poorly and often violate basic physical structure. We present PhysCoRe, a physics-corrected residual world model that couples a differentiable Material Point Method (MPM) simulator with two feed-forward neural networks. A material refinement module, Material from Motion (MfM), infers per-particle elasticity from visual…</summary>
  </entry>
  <entry>
    <title>Towards Miniature Humanoid Tele-Loco-Manipulation Using Virtual Reality and Reinforcement Learning</title>
    <id>arxiv:2607.20399</id>
    <link href="https://arxiv.org/abs/2607.20399" />
    <published>2026-07-22T17:35:43.000Z</published>
    <updated>2026-07-22T17:35:43.000Z</updated>
    <author><name>Nicolas Kosanovic</name></author>
    <author><name>Jordan Dowdy</name></author>
    <author><name>Jean Chagas Vaz</name></author>
    <category term="cs.RO" />
    <summary>Full-sized humanoid robot capabilities have grown exponentially in recent years, aiming towards general-purpose deployment in human environments. A popular control method used by manufacturers utilizes Virtual Reality for upper-body teleoperation and Reinforcement Learning for lower-body balance and locomotion control. As a result, a single remote operator can see, manipulate, and navigate about a real, distant physical environment. This powerful control stack is often relegated to expensive full-sized robots, many of which are inaccessible to the research community. Miniature humanoids are m…</summary>
  </entry>
  <entry>
    <title>Closing the Lab-to-Store Gap: A Data-Efficient Post-Training and Experience-Driven Learning VLA Framework for Retail Humanoids</title>
    <id>arxiv:2607.20345</id>
    <link href="https://arxiv.org/abs/2607.20345" />
    <published>2026-07-22T16:30:51.000Z</published>
    <updated>2026-07-22T16:30:51.000Z</updated>
    <author><name>Roger Sala Sisó</name></author>
    <author><name>Tiago Silvério</name></author>
    <author><name>Jakob Sand</name></author>
    <author><name>Tran Nguyen Le</name></author>
    <category term="cs.RO" />
    <summary>Closing the gap between benchmark performance and reliable real-world operation remains a central challenge for Vision-Language-Action (VLA) humanoid robots, which must handle execution errors, distribution shifts, and environmental variability. This paper presents DEED (Data-Efficient Post-Training and Experience-Driven Learning), a systems-level approach evaluated on a supermarket chip-restocking task using a Unitree G1-Edu humanoid robot and the GR00T N1.6 foundation model. DEED comprises three key components: (1) a data-efficient post-training pipeline with control-frequency alignment, da…</summary>
  </entry>
  <entry>
    <title>Extreme-RGMT: Continual Learning of Highly Dynamic Skills for Robust Generalist Humanoid Control</title>
    <id>arxiv:2607.20110</id>
    <link href="https://arxiv.org/abs/2607.20110" />
    <published>2026-07-22T13:10:22.000Z</published>
    <updated>2026-07-22T13:10:22.000Z</updated>
    <author><name>Yubiao Ma</name></author>
    <author><name>Han Yu</name></author>
    <author><name>Kai Guo</name></author>
    <author><name>Changtai Lv</name></author>
    <author><name>Zhengquan Mao</name></author>
    <author><name>Boyang Xing</name></author>
    <category term="cs.RO" />
    <summary>Humans can progressively acquire highly dynamic motor skills while preserving reliable everyday motor abilities. In contrast, existing humanoid controllers face a trade-off between generalist and specialist capabilities: generalist motion tracking policies struggle to reliably execute rare highly dynamic motions, whereas specialist training can degrade previously acquired behaviors. We introduce Extreme-RGMT, a two-stage continual learning framework for robust generalist humanoid control. The method first learns a generalist motion-tracking base policy from diverse multi-source motion data, t…</summary>
  </entry>
  <entry>
    <title>ReferTrack: Referring Then Tracking for Embodied Visual Tracking</title>
    <id>arxiv:2607.20061</id>
    <link href="https://arxiv.org/abs/2607.20061" />
    <published>2026-07-22T12:05:13.000Z</published>
    <updated>2026-07-22T12:05:13.000Z</updated>
    <author><name>Hanjing Ye</name></author>
    <author><name>Tianle Zeng</name></author>
    <author><name>Jiazhao Zhang</name></author>
    <author><name>Shaoan Wang</name></author>
    <author><name>Zibo Zhang</name></author>
    <author><name>Weisi Situ</name></author>
    <category term="cs.RO" />
    <summary>Embodied visual tracking (EVT) requires a mobile agent to continuously follow a specific target described in natural language using only onboard vision. While recent vision-language-action (VLA) policies unify target identification and trajectory planning, their chain-of-thought (CoT) reasoning often operates in abstract spatial latents that are difficult to supervise and weakly aligned with explicit image-space detections. To address this, we introduce ReferTrack, a referring-then-tracking paradigm that grounds EVT using a single forward-facing camera. Our model first selects the target from…</summary>
  </entry>
  <entry>
    <title>Diffusion ReRoll: Revisable Denoising for Robotic Sequential Prediction</title>
    <id>arxiv:2607.19919</id>
    <link href="https://arxiv.org/abs/2607.19919" />
    <published>2026-07-22T08:50:40.000Z</published>
    <updated>2026-07-22T08:50:40.000Z</updated>
    <author><name>Seonsoo Kim</name></author>
    <author><name>Seongil Hong</name></author>
    <author><name>Jun-Gill Kang</name></author>
    <category term="cs.RO" />
    <summary>We propose Diffusion ReRoll, a diffusion-based framework for robotic sequential prediction that enables revisable denoising over horizons. Existing diffusion-based sequence predictors typically perform a single monotonic denoising process. In contrast, Diffusion ReRoll selectively re-noises regions that have become locally stable while the remaining regions continue denoising, so the re-noised regions can be refined again using context from the rest of the horizon. This structured re-noising enables iterative cross-horizon revision, allowing earlier and later segments to revise one another, w…</summary>
  </entry>
  <entry>
    <title>What Matters in Humanoid General Motion Tracking? An Empirical Study</title>
    <id>arxiv:2607.19903</id>
    <link href="https://arxiv.org/abs/2607.19903" />
    <published>2026-07-22T08:37:19.000Z</published>
    <updated>2026-07-22T08:37:19.000Z</updated>
    <author><name>Fabio Amadio</name></author>
    <author><name>Enrico Mingo Hoffman</name></author>
    <category term="cs.RO" />
    <summary>Humanoid general motion tracking requires policies that can follow diverse whole-body references while maintaining balance. Building such policies involves many practical design choices, and their individual effects are often hard to assess. We address this issue with an empirical study of common modeling and training factors used in recent humanoid motion-imitation pipelines. To make the study controlled and reproducible, we developed YAHMP, an open-source modular framework for training, evaluating, and deploying whole-body motion tracking policies on the Unitree G1. Within YAHMP, we define…</summary>
  </entry>
  <entry>
    <title>EA-Nav: Learning Safe Visual Navigation Policies with Embodiment Awareness</title>
    <id>arxiv:2607.19880</id>
    <link href="https://arxiv.org/abs/2607.19880" />
    <published>2026-07-22T08:10:49.000Z</published>
    <updated>2026-07-22T08:10:49.000Z</updated>
    <author><name>Jialu Zhang</name></author>
    <author><name>Yong Du</name></author>
    <author><name>Xianda Guo</name></author>
    <author><name>Shunwang Sun</name></author>
    <author><name>Xinqi Liu</name></author>
    <author><name>Yue Sun</name></author>
    <category term="cs.RO" />
    <summary>Cross-embodiment navigation is a key challenge in embodied intelligence. Due to differences in embodiment, the same visual observation may imply different actions for different agents, making prediction ambiguous when relying solely on vision. Existing studies mainly rely on reinforcement learning, which requires large-scale interaction and careful reward design, making it difficult to support scalable pretraining and real-world adaptation. In contrast, imitation-learning-based approaches remain limited. To address these challenges, we propose an imitation-learning-based embodiment-aware navi…</summary>
  </entry>
  <entry>
    <title>KineBench: Benchmarking Embodied World Models via IDM-Free Kinematic Grounding</title>
    <id>arxiv:2607.19876</id>
    <link href="https://arxiv.org/abs/2607.19876" />
    <published>2026-07-22T08:04:17.000Z</published>
    <updated>2026-07-22T08:04:17.000Z</updated>
    <author><name>Zeyu Liu</name></author>
    <author><name>Zhangzhe Zhu</name></author>
    <author><name>Yang Zhang</name></author>
    <author><name>Chenyou Fan</name></author>
    <author><name>Chenjia Bai</name></author>
    <author><name>Xuelong Li</name></author>
    <category term="cs.RO" />
    <summary>Evaluating the physical consistency of embodied world models(EWMs) is a critical open challenge. While closed-loop evaluation via simulator rollouts offers a more faithful assessment of physical plausibility than open-loop alternatives, existing frameworks almost exclusively rely on Inverse Dynamics Models(IDMs) for action extraction. Due to the intricate mapping from 2D pixel space to 3D kinematic space, the learned IDMs can be brittle to data outside their training distribution, resulting in unreliable action extraction from the generated videos with novel objects and scenarios. This create…</summary>
  </entry>
  <entry>
    <title>V2F: Vision-Informed Grasp Force Prediction for Damage-Aware Robotic Handling of Date Fruits</title>
    <id>arxiv:2607.19804</id>
    <link href="https://arxiv.org/abs/2607.19804" />
    <published>2026-07-22T06:36:55.000Z</published>
    <updated>2026-07-22T06:36:55.000Z</updated>
    <author><name>Shahd Shami</name></author>
    <author><name>Obadah Wali</name></author>
    <author><name>Eric Feron</name></author>
    <author><name>Shinkyu Park</name></author>
    <category term="cs.RO" />
    <summary>This paper presents a vision-informed grasp force prediction framework for robotic handling of date fruits. Addressing the dual challenge of high detachment forces and low bruise thresholds, we first conduct mechanical characterization on date samples to define a safe grasping envelope and quantify the relationship between fruit geometry and bioyield stress. In this work, we develop a Vision-to-Force (V2F) pipeline that combines computer vision-based segmentation, active-contour refinement, and geometric feature extraction with a physics-informed residual neural network that augments a Hertz…</summary>
  </entry>
  <entry>
    <title>EgoRecovery: Acquiring Failure Recovery Ability Through Human Recovery Demonstration</title>
    <id>arxiv:2607.19745</id>
    <link href="https://arxiv.org/abs/2607.19745" />
    <published>2026-07-22T04:44:26.000Z</published>
    <updated>2026-07-23T16:20:02.000Z</updated>
    <author><name>Zuhao Ge</name></author>
    <author><name>Yuchen Zhou</name></author>
    <author><name>Weitao Zhou</name></author>
    <author><name>Minglei Li</name></author>
    <author><name>Xinyu Li</name></author>
    <author><name>Chao Wu</name></author>
    <category term="cs.RO" />
    <summary>Robust embodied robots should be able to recover from failures and retry tasks in order to operate reliably in unstructured and noisy real-world environments. Achieving this capability requires training policies on data that captures recovery behaviors. However, collecting such data through robot teleoperation is difficult to scale, as it is time-consuming to induce diverse failure states, perform corrective actions, and reset the environment. This challenge is further exacerbated by the high diversity of failure modes, which demands substantially more recovery data than success demonstration…</summary>
  </entry>
  <entry>
    <title>Morphing MILR: Design and control of a cable-driven limbless robot with rolling joints for maneuvering in complex environments</title>
    <id>arxiv:2607.19714</id>
    <link href="https://arxiv.org/abs/2607.19714" />
    <published>2026-07-22T03:29:39.000Z</published>
    <updated>2026-07-22T03:29:39.000Z</updated>
    <author><name>Donoven Dortilus</name></author>
    <author><name>Tianyu Wang</name></author>
    <author><name>Galen Tunnicliffe</name></author>
    <author><name>Matthew Fernandez</name></author>
    <author><name>Daniel I. Goldman</name></author>
    <category term="cs.RO" />
    <summary>Limbless robots offer exceptional mobility in confined and cluttered environments due to their slender bodies and their ability to exploit body-terrain interactions. Recent designs incorporating compliance demonstrate robust locomotion without complex sensing or control; however, these systems typically rely on fixed body configurations, with each morphology specialized for a single locomotion mode or environment. This raises a key challenge: how can a single limbless robot achieve versatile locomotion while preserving the robustness of compliance-mediated locomotion? To address this challeng…</summary>
  </entry>
  <entry>
    <title>NavVerse: Benchmarking Indoor-to-Outdoor Embodied Navigation in Continuous Robot Simulation</title>
    <id>arxiv:2607.19695</id>
    <link href="https://arxiv.org/abs/2607.19695" />
    <published>2026-07-22T02:53:46.000Z</published>
    <updated>2026-07-22T02:53:46.000Z</updated>
    <author><name>Junzhe Wu</name></author>
    <author><name>Yue Hu</name></author>
    <author><name>Zeyu Han</name></author>
    <author><name>Po-Hsun Chang</name></author>
    <author><name>Yinan Dong</name></author>
    <author><name>Behrad Rabiei</name></author>
    <category term="cs.RO" />
    <summary>Robots deployed in delivery, campus, and emergency-response settings often need to navigate from buildings to streets within a single continuous episode. Existing benchmarks usually evaluate indoor and outdoor navigation separately, and many abstract away robot execution, leaving exit finding, boundary traversal, adaptation, and kinodynamic failures underexplored. We introduce NavVerse, a physics-enabled benchmark for indoor-to-outdoor embodied navigation. NavVerse contains 100 indoor scenes, 50 urban outdoor scenes, and 50 indoor-to-outdoor scenes, and 10,000 episodes spanning Object Navigat…</summary>
  </entry>
  <entry>
    <title>LENS: LLM-guided Environment Simplification for Planning and Control in Clutter</title>
    <id>arxiv:2607.19633</id>
    <link href="https://arxiv.org/abs/2607.19633" />
    <published>2026-07-22T00:05:00.000Z</published>
    <updated>2026-07-22T00:05:00.000Z</updated>
    <author><name>Aileen Liao</name></author>
    <author><name>Rachel Holladay</name></author>
    <author><name>Dinesh Jayaraman</name></author>
    <author><name>Michael Posa</name></author>
    <category term="cs.RO" />
    <summary>Despite recent advances in general-purpose robotic manipulation, real-world multi-object clutter remains challenging to handle for today&apos;s prevalent approaches. The problem scales in complexity due to more objects and collisions, more unpredictable contact physics, distractors, and task ambiguity. Bridging this gap to real-world deployment requires effective scene abstractions; yet today, producing such abstractions requires extensive task-specific manual engineering, which does not scale. These abstractions are costly to generate and difficult to adjust or fine-tune. We instead propose a plu…</summary>
  </entry>
  <entry>
    <title>Learning Personalized Safety Interventions for Haptic Human-Robot Shared Control</title>
    <id>arxiv:2607.19534</id>
    <link href="https://arxiv.org/abs/2607.19534" />
    <published>2026-07-21T19:28:41.000Z</published>
    <updated>2026-07-21T19:28:41.000Z</updated>
    <author><name>Dawei Zhang</name></author>
    <author><name>Roberto Tron</name></author>
    <category term="cs.RO" />
    <summary>Haptic feedback provides an implicit channel for communicating safety intentions during human-robot shared control. Existing haptic guidance systems typically employ predefined intervention strategies that cannot accommodate the diverse safety preferences of individual users or application scenarios. To address this limitation, we propose a Learning from Haptics (LfH) framework that learns user-preferred safety interventions from sparse demonstrations, eliminating the need for manual trial-and-error design. Our framework is built on a differentiable Control Barrier Function (CBF)-based optimi…</summary>
  </entry>
  <entry>
    <title>ModPack: An Extensible Teleoperation Interface for Bimanual Mobile Manipulation</title>
    <id>arxiv:2607.19479</id>
    <link href="https://arxiv.org/abs/2607.19479" />
    <published>2026-07-21T18:02:20.000Z</published>
    <updated>2026-07-21T18:02:20.000Z</updated>
    <author><name>Joshua Citron</name></author>
    <author><name>Renee Zbizika</name></author>
    <author><name>Zeyi Liu</name></author>
    <author><name>Shuran Song</name></author>
    <category term="cs.RO" />
    <summary>Existing teleoperation systems are often tailored to specific robot hardware and task domains, limiting their scalability and adaptability. We present ModPack, a modular and extensible teleoperation system designed to support diverse robot embodiments and task requirements within a unified framework. At the core of ModPack is a self-contained wearable &quot;backpack&quot; that integrates onboard computation, power, communication, and data storage. Built on top of this shared interface, the system supports plug-and-play capability modules including joint-level teleoperation with haptic feedback, mobile…</summary>
  </entry>
  <entry>
    <title>Masked Visual Actions for Unified World Modeling</title>
    <id>arxiv:2607.19343</id>
    <link href="https://arxiv.org/abs/2607.19343" />
    <published>2026-07-21T17:59:11.000Z</published>
    <updated>2026-07-21T17:59:11.000Z</updated>
    <author><name>Hadi Alzayer</name></author>
    <author><name>Wenlong Huang</name></author>
    <author><name>Haonan Chen</name></author>
    <author><name>Christopher Luey</name></author>
    <author><name>Lvmin Zhang</name></author>
    <author><name>Maneesh Agrawala</name></author>
    <category term="cs.CV" />
    <summary>Video models absorb rich priors over how the visual world moves, interacts, and responds to contact, making them promising substrates for robotic world modeling. The central challenge is how to communicate action to such models in a form aligned with the visual space in which they learned these interaction priors, yet still grounded in physical manipulation. We introduce Masked Visual Actions, a pixel-space control interface that expresses action as a partially revealed trajectory of an arbitrary entity in a video. Revealing robot motion makes the model act as a forward dynamics model that pr…</summary>
  </entry>
  <entry>
    <title>Agentic Real2Sim: Physics-based World Modeling with Vision-Language Agents</title>
    <id>arxiv:2607.19190</id>
    <link href="https://arxiv.org/abs/2607.19190" />
    <published>2026-07-21T15:23:38.000Z</published>
    <updated>2026-07-22T03:43:37.000Z</updated>
    <author><name>Guanxiong Chen</name></author>
    <author><name>Qianjun Xia</name></author>
    <author><name>Jiawei Peng</name></author>
    <author><name>Heng Zhang</name></author>
    <author><name>Bole Ma</name></author>
    <author><name>Justin Qian</name></author>
    <category term="cs.RO" />
    <summary>Real-to-sim conversion for robotic interaction with objects remains labor-intensive because it requires more than visual reconstruction: a streamlined real2sim process must recover scene geometries and object states, infer physical parameters, and assemble actors, objects, cameras, poses, and trajectories into a runnable physical simulation. Today this process still depends on manual tuning of visual foundation models, mesh cleanup, coordinate-frame alignment, and brittle workflow glue across visual perception tools and simulators. We introduce \textit{Agentic Real2Sim}, a framework for gener…</summary>
  </entry>
  <entry>
    <title>Design and stability analysis of an underactuated hand with passively rotating fingers</title>
    <id>arxiv:2607.18950</id>
    <link href="https://arxiv.org/abs/2607.18950" />
    <published>2026-07-21T10:37:18.000Z</published>
    <updated>2026-07-21T10:37:18.000Z</updated>
    <author><name>Léonie Plancoulaine</name></author>
    <author><name>Sylvain Guégan</name></author>
    <author><name>Franck Plestan</name></author>
    <author><name>Damien Chablat</name></author>
    <category term="cs.RO" />
    <summary>This paper presents an innovative design and stability analysis of an underactuated robotic finger with spatial mobility, designed to enhance gripping dexterity in robotic hands. The finger architecture incorporates a revolute joint at its base, enabling passive spatial rotation that facilitates both cylindrical and spherical grasping. With only two phalanges per finger, the design simplifies kinematic complexity while supporting precision and enveloping grasps. Stability criteria, based on the moment at the finger base joint induced by contact forces, are introduced to ensure reliable object…</summary>
  </entry>
  <entry>
    <title>The Twist Decomposition of Serial Robots Under Lower-Mobility Tasks</title>
    <id>arxiv:2607.18940</id>
    <link href="https://arxiv.org/abs/2607.18940" />
    <published>2026-07-21T10:23:23.000Z</published>
    <updated>2026-07-21T10:23:23.000Z</updated>
    <author><name>Luc Baron</name></author>
    <author><name>Damien Chablat</name></author>
    <category term="cs.RO" />
    <summary>This paper introduces a twist decomposition framework for serial manipulators performing lower mobility tasks. Rather than relying on Jacobian null-space projections, the method separates the end-effector twist into task and redundant components using geometrically defined twist projectors. This formulation provides a direct and intuitive distinction between task-relevant and task-irrelevant motions in operational space, enabling a compact inverse kinematics scheme that naturally handles both manipulator and task redundancy.</summary>
  </entry>
  <entry>
    <title>Pose-Parameterized Motion Planning and CBF-QP Self-Collision Filtering for a Long-Reach Drilling Boom</title>
    <id>arxiv:2607.18855</id>
    <link href="https://arxiv.org/abs/2607.18855" />
    <published>2026-07-21T08:41:49.000Z</published>
    <updated>2026-07-21T08:41:49.000Z</updated>
    <author><name>Mehdi Heydari Shahna</name></author>
    <author><name>Tuomo Kivelä</name></author>
    <author><name>Jouni Mattila</name></author>
    <category term="cs.RO" />
    <summary>Long-reach drilling booms must reach successive poses without self-collision. Moving from operator-supervised control toward autonomy requires collision-aware motion planning and execution. For the Sandvik SB60, this study adapts established methods by integrating pose-parameterized planning with a capsule-based control barrier function quadratic program (CBF-QP) in measured-state inverse kinematics (IK). A fixed task-specific parameter set within each task generates waypoints, detours, timed references, and chained motion without target-specific retuning. The offline detour planner screens c…</summary>
  </entry>
  <entry>
    <title>WorldScape Policy 2.0: Empowering Steerable World Action Modeling with Reasoning-Augmented Memory</title>
    <id>arxiv:2607.18840</id>
    <link href="https://arxiv.org/abs/2607.18840" />
    <published>2026-07-21T08:25:37.000Z</published>
    <updated>2026-07-21T08:25:37.000Z</updated>
    <author><name>Haisheng Su</name></author>
    <author><name>Zongdai Liu</name></author>
    <author><name>Xin Jin</name></author>
    <author><name>Haoxuan Dou</name></author>
    <author><name>Chengming Hu</name></author>
    <author><name>Baorun Li</name></author>
    <category term="cs.RO" />
    <summary>World Action Models (WAMs) offer a promising paradigm for robotic manipulation by jointly modeling visual state transitions and robot actions. However, existing WAMs are constrained by limited temporal context, coarse episode-level language supervision, and predominantly text-only conditioning, which hinder task-progress tracking and fine-grained language-video-action grounding while limiting visual-context reasoning and cross-embodiment transfer. In this paper, we introduce WorldScape Policy 2.0, a controllable WAM with reasoning-augmented long short-term memory. Its causal short-term visual…</summary>
  </entry>
  <entry>
    <title>Koopman DCM: Unstable Eigenfunctions as Data-driven Representations for Legged Balancing</title>
    <id>arxiv:2607.18760</id>
    <link href="https://arxiv.org/abs/2607.18760" />
    <published>2026-07-21T06:34:35.000Z</published>
    <updated>2026-07-21T06:34:35.000Z</updated>
    <author><name>Stéphane Caron</name></author>
    <category term="cs.RO" />
    <summary>In legged locomotion, divergent components of motion (DCMs) have emerged as characteristic states for balance control. They isolate the unstable mode of the dynamics but, in existing formulations, apply only to reduced models such as the linear inverted pendulum. In this study, we show how DCMs can be more generally formulated as Koopman eigenfunctions. Whereas Koopman analysis typically targets eigenvalues near zero, which capture conserved or slowly varying quantities, our investigation leads us to deliberately search for unstable eigenpairs with large eigenvalues. The resulting Koopman DCM…</summary>
  </entry>
  <entry>
    <title>Motion Primitive Discovery in a Humanoid Robot via Self-Organising Maps for Phase Recognition</title>
    <id>arxiv:2607.18737</id>
    <link href="https://arxiv.org/abs/2607.18737" />
    <published>2026-07-21T05:53:29.000Z</published>
    <updated>2026-07-21T05:53:29.000Z</updated>
    <author><name>Radovan Gregor</name></author>
    <author><name>Igor Farkaš</name></author>
    <category term="cs.RO" />
    <summary>Understanding the computational basis of action recognition is a central challenge in social cognition as well as in human-robot interaction. Inspired by the Mirror Neuron System (MNS), we propose a two-level architecture for motor primitive discovery and online phase recognition applied to the NICO humanoid robot. At the first level, two Self-Organising Maps (SOMs) learn topographic representations of arm kinematics (A-SOM) and hand kinematics (H-SOM) from simulated trials covering seven motor actions. The maps are trained on non-redundant features identified through hierarchical correlation…</summary>
  </entry>
  <entry>
    <title>RoboInter1.5: A Holistic Intermediate Representation Suite for Embodied World Modeling and Robotic Manipulation</title>
    <id>arxiv:2607.18709</id>
    <link href="https://arxiv.org/abs/2607.18709" />
    <published>2026-07-21T05:05:01.000Z</published>
    <updated>2026-07-22T05:33:17.000Z</updated>
    <author><name>Ziqin Wang</name></author>
    <author><name>Hao Li</name></author>
    <author><name>Weijun Wang</name></author>
    <author><name>Junhao Cai</name></author>
    <author><name>Jia Zeng</name></author>
    <author><name>Yilun Chen</name></author>
    <category term="cs.RO" />
    <summary>Existing robot datasets remain expensive to curate, embodiment-specific, and insufficiently annotated with the fine-grained structure required for generalizable reasoning, execution, or long-horizon environment dynamics simulation. Building on our prior work, RoboInter1.0, we present RoboInter1.5, an extended and holistic suite of intermediate representations for both robotic manipulation and embodied world modeling. RoboInter1.5 provides a unified resource of data, benchmarks, and models centered on dense manipulation-oriented intermediate representations. Specifically, RoboInter-Data contai…</summary>
  </entry>
  <entry>
    <title>End-to-end Conditional Diffusion for Realistic and Controllable Visual Traffic Scenario Generation</title>
    <id>arxiv:2607.18637</id>
    <link href="https://arxiv.org/abs/2607.18637" />
    <published>2026-07-21T02:13:25.000Z</published>
    <updated>2026-07-21T02:13:25.000Z</updated>
    <author><name>Jingzheng Li</name></author>
    <author><name>Yufei Ge</name></author>
    <author><name>Zhijun Chen</name></author>
    <author><name>Qianren Mao</name></author>
    <author><name>Zizhe Wang</name></author>
    <author><name>Binhang Qi</name></author>
    <category term="cs.RO" />
    <summary>Generating closed-loop traffic scenarios that are both realistic and controllable is crucial for evaluating autonomous driving systems, especially under rare safety-critical interactions. Existing learning-based methods often struggle to balance controllability and realism, offering either limited fine-grained control over traffic behavior or controllable scenarios at the expense of behavioral plausibility. This paper presents E2E-CDiff, an end-to-end conditional diffusion framework for controllable and realistic scenario generation. Conditioned on front-view visual observations, E2E-CDiff jo…</summary>
  </entry>
  <entry>
    <title>STeP: Signal Temporal Logic for Precise Specifications for Action Generation with Vision Language Models</title>
    <id>arxiv:2607.18580</id>
    <link href="https://arxiv.org/abs/2607.18580" />
    <published>2026-07-20T23:26:32.000Z</published>
    <updated>2026-07-20T23:26:32.000Z</updated>
    <author><name>Kasra Torshizi</name></author>
    <author><name>Anukriti Singh</name></author>
    <author><name>Sidharth Mathur</name></author>
    <author><name>Khuzema Habib</name></author>
    <author><name>Leo Du</name></author>
    <author><name>Pratap Tokekar</name></author>
    <category term="cs.RO" />
    <summary>Vision-language-action (VLA) models have shown impressive generalization, but often lack interpretability and can struggle to follow precise natural language instructions that encode spatial, temporal, and logical requirements. We propose a hierarchical framework that uses Signal Temporal Logic (STL) as a shared representation connecting high-level language understanding with low-level robot execution. A high-level policy leverages a VLM to decompose language instructions into high-level subtasks, generate STL specifications for each subtask, and choose a low-level policy for executing each s…</summary>
  </entry>
  <entry>
    <title>DASH Robot: Minimalistic Design and Optimal Aerial-Terrestrial Locomotion via Contact-Implicit Control</title>
    <id>arxiv:2607.18527</id>
    <link href="https://arxiv.org/abs/2607.18527" />
    <published>2026-07-20T21:39:25.000Z</published>
    <updated>2026-07-20T21:39:25.000Z</updated>
    <author><name>Ryan Gomes Paiva</name></author>
    <author><name>Conrad Ho</name></author>
    <author><name>Jiarong Kang</name></author>
    <author><name>Kunzhao Ren</name></author>
    <author><name>Xiangru Xu</name></author>
    <author><name>Xiaobin Xiong</name></author>
    <category term="cs.RO" />
    <summary>We present a novel and minimalistic design of an aerial-terrestrial robot DASH: Ducted Aerial Spring Hopper. The goal is to enable both aerial and ground locomotion capabilities on a unified mobile robot that is mechanically-minimalistic, locomotion-versatile, and energy-efficient. We propose an organic integration of ducted fan co-axial body with a springy leg at the bottom for realization. The ducted fan module provides thrust-vectoring as the main actuation for agile flying; when it is combined with the light-weight spring leg, the robot realizes highly efficient ground hopping with energy…</summary>
  </entry>
  <entry>
    <title>Patch Policy: Efficient Embodied Control via Dense Visual Representations</title>
    <id>arxiv:2607.18236</id>
    <link href="https://arxiv.org/abs/2607.18236" />
    <published>2026-07-20T17:59:41.000Z</published>
    <updated>2026-07-20T17:59:41.000Z</updated>
    <author><name>Gaoyue Zhou</name></author>
    <author><name>Zichen Jeff Cui</name></author>
    <author><name>Ada Langford</name></author>
    <author><name>Bowen Tan</name></author>
    <author><name>Yann LeCun</name></author>
    <author><name>Lerrel Pinto</name></author>
    <category term="cs.RO" />
    <summary>Pretrained dense visual features from Vision Transformers (ViTs) are powerful yet have been underutilized in robot learning. Modern robot policies either compress each observation into a single global token, or rely on visual backbones trained from scratch, sacrificing both fine-grained spatial detail and the benefits of large-scale visual pre-training. While there exist policies that do operate on dense patch features like large vision-language-action models (VLAs), they tend to be heavy and slow, inheriting the full cost of a billion-parameter vision-language model (VLM) backbone. We close…</summary>
  </entry>
  <entry>
    <title>FM-VLA: Force-based Memory for Vision-Language-Action Models in Contact-Rich Manipulation</title>
    <id>arxiv:2607.18231</id>
    <link href="https://arxiv.org/abs/2607.18231" />
    <published>2026-07-20T17:58:31.000Z</published>
    <updated>2026-07-20T17:58:31.000Z</updated>
    <author><name>Ruicheng Li</name></author>
    <author><name>Qixiu Li</name></author>
    <author><name>Ruichun Ma</name></author>
    <author><name>Yu Deng</name></author>
    <author><name>Lin Luo</name></author>
    <author><name>Zhiying Du</name></author>
    <category term="cs.RO" />
    <summary>Vision-language-action (VLA) models have achieved impressive generalization in robotic manipulation, and recent memory-augmented VLAs have relaxed the Markovian assumption by conditioning on past images or language summaries. Vision-based memory approaches address this by conditioning on sampled past image frames, but they are computationally expensive and fundamentally limited when temporal events are visually ambiguous, e.g., pushing a button multiple times with small movements. We propose FM-VLA, a VLA model with force-based memory, enabling temporal context reasoning for non-Markovian, co…</summary>
  </entry>
  <entry>
    <title>Optimization of sim-to-real transfer in the humanoid robot NICO</title>
    <id>arxiv:2607.18210</id>
    <link href="https://arxiv.org/abs/2607.18210" />
    <published>2026-07-20T17:45:04.000Z</published>
    <updated>2026-07-20T17:45:04.000Z</updated>
    <author><name>Juraj Gavura</name></author>
    <author><name>Igor Farkaš</name></author>
    <category term="cs.RO" />
    <summary>Robotic grasping requires accurate coordination between visual perception, object localization, inverse kinematics, and hand control. However, when movements planned in simulation are executed on a physical robot, the sim-to-real gap can cause small positioning errors that prevent successful grasping. In our previous work, we introduced a low-cost haptic calibration method that improved 2D reaching accuracy of the humanoid robot NICO. In this paper, we extend this approach from reaching to tabletop object grasping by adding YOLO-based object and hand detection, stereo vision-based localizatio…</summary>
  </entry>
  <entry>
    <title>Learning Adaptive Safety Margins for Visual Navigation</title>
    <id>arxiv:2607.18200</id>
    <link href="https://arxiv.org/abs/2607.18200" />
    <published>2026-07-20T17:40:25.000Z</published>
    <updated>2026-07-20T17:40:25.000Z</updated>
    <author><name>Junyi Hu</name></author>
    <author><name>Shuaihang Yuan</name></author>
    <author><name>Geeta Chandra Raju Bethala</name></author>
    <author><name>Anthony Tzes</name></author>
    <author><name>Yi Fang</name></author>
    <category term="cs.RO" />
    <summary>Robots in cluttered indoor spaces often fail not because they cannot generate collision-free paths, but because a fixed safety margin is mis-calibrated: conservative margins cause detours and timeouts, while permissive margins lead to near-boundary shortcuts under perception bias. Diffusion-based planners propose diverse trajectory candidates from egocentric RGB-D, yet reliable selection remains the bottleneck. We propose a context-conditioned safety critic that learns an adaptive clearance preference for ranking diffusion proposals, decomposed into three complementary terms: (i) a safety ter…</summary>
  </entry>
  <entry>
    <title>Imitation of Arm Gestures by the Semi-Humanoid Robot NICO</title>
    <id>arxiv:2607.18197</id>
    <link href="https://arxiv.org/abs/2607.18197" />
    <published>2026-07-20T17:37:57.000Z</published>
    <updated>2026-07-20T17:37:57.000Z</updated>
    <author><name>Anastasiya Ihnatovich</name></author>
    <author><name>Igor Farkaš</name></author>
    <category term="cs.RO" />
    <summary>Seamless human-robot interaction (HRI) requires a number of perceptual and motor abilities from the robot, one of them being the imitation of human gestures. Humanoid robots have an advantage in HRI thanks to their anthropomorphic features. In this work, we develop a system for imitation of human arm gestures by the semi-humanoid robot NICO based on analytical geometry and a pretrained MediaPipe pose-estimation model. For each input RGB frame, 3D coordinates of relevant human body landmarks, including arm joints and hand keypoints, are obtained using the MediaPipe framework. Joint angles are…</summary>
  </entry>
  <entry>
    <title>World Translation: Minimizing Sim-to-Real Gap with Backward Dynamics Extraction and Unpaired Domain Translation</title>
    <id>arxiv:2607.18154</id>
    <link href="https://arxiv.org/abs/2607.18154" />
    <published>2026-07-20T16:52:58.000Z</published>
    <updated>2026-07-20T16:52:58.000Z</updated>
    <author><name>Xinchen Yao</name></author>
    <author><name>Leixin Chang</name></author>
    <author><name>Hua Chen</name></author>
    <category term="cs.RO" />
    <summary>The gap between simulation and reality remains a fundamental challenge in deploying simulation-trained robotic policies in the real world. Real-to-sim methods narrow this gap from the real side, learning transition dynamics from real data to build a more realistic digital world. Learned dynamics models are their dominant instance. Such methods, however, face a partial observability problem: the same observation may branch to different transitions due to unobservable factors. Existing methods assume these factors can be recovered from observation history. However, this may fail whenever observ…</summary>
  </entry>
  <entry>
    <title>Towards Torque-Driven Reinforcement Learning for Quadruped Locomotion</title>
    <id>arxiv:2607.18365</id>
    <link href="https://arxiv.org/abs/2607.18365" />
    <published>2026-07-20T16:35:53.000Z</published>
    <updated>2026-07-20T16:35:53.000Z</updated>
    <author><name>Jordan Dowdy</name></author>
    <author><name>Jean Chagas Vaz</name></author>
    <category term="cs.RO" />
    <summary>Reinforcement learning (RL) for legged robots is advancing locomotion, demonstrating its ability to adapt to new and challenging terrain. Traditionally, these RL locomotion frameworks are position-based, making the policy less adaptable to terrain types and requiring state estimation techniques in the observation space, i.e., linear velocity. Moreover, these RL frameworks often use small, lightweight quadrupeds that are limited in their viability for high-complexity tasks due to hardware constraints. This work explores an RL torque control framework for heavyweight high-torque quadrupeds. The…</summary>
  </entry>
  <entry>
    <title>Isaac Sim-to-Real: Reinforcement Learning based Locomotion for Quadrupeds</title>
    <id>arxiv:2607.18135</id>
    <link href="https://arxiv.org/abs/2607.18135" />
    <published>2026-07-20T16:27:56.000Z</published>
    <updated>2026-07-20T16:27:56.000Z</updated>
    <author><name>Jordan Dowdy</name></author>
    <author><name>Jean Chagas Vaz</name></author>
    <category term="cs.RO" />
    <summary>Learning-based approaches to locomotion have risen in popularity in recent years, showing the capability for complex legged locomotion and whole-body control. Reinforcement learning (RL), the primary learning-based approach for locomotion, often utilizes a high-performance simulation tool, providing a controlled and efficient training and development environment. However, policies that perform well in simulation frequently encounter unexpected challenges when deployed on a physical system, known as the sim-to-real gap. This work presents a robust RL locomotion framework capable of whole-body…</summary>
  </entry>
  <entry>
    <title>Technical Design Review of Duke Robotics Club&apos;s Oogway &amp; Crush: AUVs for RoboSub 2026</title>
    <id>arxiv:2607.18075</id>
    <link href="https://arxiv.org/abs/2607.18075" />
    <published>2026-07-20T15:44:15.000Z</published>
    <updated>2026-07-20T15:44:15.000Z</updated>
    <author><name>Patrick Zheng</name></author>
    <author><name>Saagar Arya</name></author>
    <author><name>Hung Le</name></author>
    <author><name>Mathew Chu</name></author>
    <author><name>Nathanael Ren</name></author>
    <author><name>Niko Weaver</name></author>
    <category term="cs.RO" />
    <summary>The Duke Robotics Club presents Oogway and Crush, our AUVs for RoboSub 2026. This year&apos;s strategy expands on our previously narrowed scope, targeting all four of RoboSub&apos;s design goals for the first time: movement, vision, manipulation, and acoustic tracking. This expansion is based on sustained reliability investment across all three subsystems. Mechanically, Crush gained two additional thrusters and a CFD-optimized case, providing pitch stability. Electrically, we addressed accumulated failure points by repairing unreliable connections and upgraded thruster control hardware. We also redesig…</summary>
  </entry>
  <entry>
    <title>RoboHarness: Memory-Driven Orchestration of Heterogeneous Robot Policies for Long-Horizon Planning</title>
    <id>arxiv:2607.18060</id>
    <link href="https://arxiv.org/abs/2607.18060" />
    <published>2026-07-20T15:27:13.000Z</published>
    <updated>2026-07-20T15:27:13.000Z</updated>
    <author><name>Jinbang Huang</name></author>
    <author><name>Yuanzhao Hu</name></author>
    <author><name>Zhiyuan Li</name></author>
    <author><name>Ran Qi</name></author>
    <author><name>Yixin Xiao</name></author>
    <author><name>Zhanguang Zhang</name></author>
    <category term="cs.RO" />
    <summary>Long-horizon robotic tasks require diverse capabilities that no single policy can reliably provide. Heterogeneous policies offer complementary strengths, but orchestrating them requires reasoning over uncertain capability boundaries and cross-policy distribution mismatch, which are largely overlooked by existing planning methods built on homogeneous, predefined skills with fixed applicability. We propose RoboHarness, a unified framework that encapsulates independently developed robot control systems as reusable agentic skills. Although instantiated in this work with VLAs, RL policies, and tas…</summary>
  </entry>
  <entry>
    <title>Closing the Loop in Humanoid VLA: Persistent 3D Object Tokens for Verifiable Loco-Manipulation</title>
    <id>arxiv:2607.18016</id>
    <link href="https://arxiv.org/abs/2607.18016" />
    <published>2026-07-20T14:52:46.000Z</published>
    <updated>2026-07-20T14:52:46.000Z</updated>
    <author><name>Peng Ren</name></author>
    <author><name>Haoyang Ge</name></author>
    <author><name>Jiang Zhao</name></author>
    <author><name>Cong Huang</name></author>
    <author><name>Yukun Shi</name></author>
    <author><name>Pei Chi</name></author>
    <category term="cs.RO" />
    <summary>Vision-language-action policies are a promising foundation for general robot control, but long-horizon humanoid loco-manipulation requires the robot to treat task objects as persistent physical entities across movement, contact, occlusion, and recovery. We study this problem as object-state divergence: the object state used to condition a whole-body action can differ from the state used to decide whether the action achieved the intended physical relation. We propose \emph{Persistent Object Tokenization} (POT), which maintains role-indexed 3D object records from RGB-D observations and converts…</summary>
  </entry>
</feed>
