Abstract Paper Portal of IEEE Transactions on Robotics (TRO) 2025
PaperID: 1,
Authors: Yichen Xiang, Lifeng Zhu, Aiguo Song, Yongjie Jessica Zhang
Affiliations: State Key Laboratory of Digital Medical Engineering, Jiangsu Key Lab of Robot Sensing and Control, School of Instrument Science and Engineering, Southeast University, Nanjing, China; Computational Bio-modeling Laboratory, Department of Mechanical Engineering, Carnegie Mellon University, Pittsburgh, PA, USA
Abstract: Elasticity is one of the representative parameters that reflect the mechanical properties of soft materials. Detecting the underneath elasticity distribution called elastography is a key step for understanding and interacting with objects. Existing solutions for capturing the interior elasticity distribution typically rely on expensive apparatus. In this work, the dense tactile signal captured by the high-resolution vision-based tactile sensor is introduced as a new modality for reconstructing 3-D elasticity distribution. We propose a model-based method, which exploits the tactile maps from active pressing trials for the elastography task. The interior elasticity distribution for nonrigid objects is reconstructed from an inverse physics model. We analyze the credibility of the estimated elasticity distribution obtained from our method. Varying design factors are also discussed. We experiment our method on a set of synthesized 3-D models and physical models in robot-assisted scenes. Various experimental results have been gathered, demonstrating the efficacy of our approach in perceiving elasticity distribution.
PaperID: 2,
Authors: Shenli Yuan, Shaoxiong Wang, Radhen Patel, Megha Tippur, Connor L. Yako, Mark R. Cutkosky, Edward H. Adelson, J. Kenneth Salisbury
Affiliations: Stanford University, Stanford, CA, USA; Massachusetts Institute of Technology, Cambridge, MA, USA
Abstract: Manipulation of objects within a robot's hand is one of the most important challenges in achieving robot dexterity. To address this challenge, Roller Graspers use steerable rolling fingertips. The fingertips impart motions and exert forces to achieve six degree of freedom mobility and closed-loop grasp force control. The design reported here uses image processing from cameras placed inside steerable compliant rollers to track contact conditions and locations. Integration of this data into a controller enables a variety of robust in-hand manipulation capabilities. We demonstrate that the same information can be used to reconstruct object shape. In addition, we show that by converting in-hand manipulation from a discontinuous process, with fingers frequently attaching and detaching from the object surface, to a continuous process, we can implement a convergent control loop that minimizes errors that otherwise accumulate during large object motions. The difference is apparent when comparing the results of an object rotation using a discontinuous finger-gaiting approach, as would be required without rolling fingertips, to the results obtained with continuous rolling. The results suggest that hybrid rolling fingertip and finger-gaiting approaches to manipulation may be a promising future research direction.
PaperID: 3,
Authors: Alejandro M. Castro, Xuchen Han, Joseph Masterjohn
Affiliations: Toyota Research Institute, Cambridge, MA, USA
Abstract: In this article, we present irrotational contact fields, a framework for generating convex approximations of complex contact models, incorporating experimentally validated models like Hunt and Crossley coupled with Coulomb’s law of friction alongside the principle of maximum dissipation. Our approach is robust across a wide range of stiffness values, making it suitable for both compliant surfaces and rigid approximations. We evaluate these approximations across a wide variety of test cases, detailing properties and limitations. We implement a fully differentiable solution in the open-source robotics toolkit, Drake. Our novel hybrid approach enables efficient computation of gradients for complex geometric models by reusing factorizations from contact resolution. We demonstrate robust simulation of robotic tasks at interactive rates, with accurately resolved stiction and contact transitions, supporting effective sim-to-real transfer.
PaperID: 4,
Authors: Oliver Limoyo, Filip Maric, Matthew Giamou, Petra Alexson, Ivan Petrovic, Jonathan Kelly
Affiliations: Space and Terrestrial Autonomous Robotic Systems Laboratory, Institute for Aerospace Studies, University of Toronto, Toronto, ON, Canada; Autonomous Robotics and Convex Optimization Laboratory, Department of Computing and Software, McMaster University, Hamilton, ON, Canada; Laboratory for Autonomous Systems and Mobile Robotics, Faculty of Electrical Engineering and Computing, University of Zagreb, Zagreb, Croatia
Abstract: Quickly and reliably finding accurate inverse kinematics (IK) solutions remains a challenging problem for many robot manipulators. Existing numerical solvers are broadly applicable but typically only produce a single solution and rely on local search techniques to minimize nonconvex objective functions. Recent learning-based approaches that approximate the entire feasible set of solutions have shown promise in generating multiple fast and accurate IK results in parallel. However, existing learning-based techniques have a significant drawback: each robot of interest requires a specialized model that must be trained from scratch. To address this key shortcoming, we propose a novel distance-geometric robot representation coupled with a graph structure that allows us to leverage the generalizability of graph neural networks (GNNs). Our approach, which we call generative graphical IK (GGIK), is the first learned IK solver that is able to efficiently yield a large number of diverse solutions in parallel while also displaying the ability to generalize—a single learned model can be used to produce IK solutions for a variety of different robots. When compared to several other learned IK methods, GGIK provides more accurate solutions with the same amount of training data. GGIK can also generalize reasonably well to robot manipulators unseen during training. In addition, GGIK is able to learn a constrained distribution that encodes joint limits and scales well with the number of robot joints and sampled solutions. Finally, GGIK can be used to complement local IK solvers by providing a reliable initialization for the local optimization process.
PaperID: 5,
Authors: Shan Luo, Nathan F. Lepora, Wenzhen Yuan, Kaspar Althoefer, Gordon Cheng, Ravinder Dahiya
Affiliations: Department of Engineering, King’s College London, London, U.K.; School of Engineering Mathematics and Bristol Robotics Laboratory, University of Bristol, Bristol, U.K.; Department of Computer Science, University of Illinois Urbana-Champaign, Champaign, IL, USA; School of Engineering and Materials Science, Queen Mary University of London, London, U.K.; Institute for Cognitive Systems (ICS), Technische Universität München, München, Germany; Department of Electrical and Computer Engineering, Northeastern University, Boston, MA, USA
Abstract: Robotics research has long sought to give robots the ability to perceive the physical world through touch in an analogous manner to many biological systems. Developing such tactile capabilities is important for numerous emerging applications that require robots to co-exist and interact closely with humans. Consequently, there has been growing interest in tactile sensing, leading to the development of various technologies, including piezoresistive and piezoelectric sensors, capacitive sensors, magnetic sensors, and optical tactile sensors. These diverse approaches utilize different transduction methods and materials to equip robots with distributed sensing capabilities, enabling more effective physical interactions. These advances have been supported in recent years by simulation tools that generate large-scale tactile datasets to support sensor designs and algorithms to interpret and improve the utility of tactile data. The integration of tactile sensing with other modalities, such as vision, as well as with action strategies for active tactile perception highlights the growing scope of this field. To further the transformative progress in tactile robotics, a holistic approach is essential. In this outlook article, we examine several challenges associated with the current state of the art in tactile robotics and explore potential solutions to inspire innovations across multiple domains, including manufacturing, healthcare, recycling, and agriculture.
PaperID: 6,
Authors: Chunran Zheng, Wei Xu, Zuhao Zou, Tong Hua, Chongjian Yuan, Dongjiao He, Bingyang Zhou, Zheng Liu, Jiarong Lin, Fangcheng Zhu, Yunfan Ren, Rong Wang, Fanle Meng, Fu Zhang
Affiliations: Mechatronics and Robotic Systems (MaRS) Laboratory, Department of Mechanical Engineering, University of Hong Kong, Hong Kong, SAR, China; Information Science Academy, China Electronics Technology Group Corporation, Beijing, China
Abstract: This paper presents FAST-LIVO2, a fast and direct LiDAR-inertial-visual odometry framework designed for accurate and robust state estimation in SLAM tasks, enabling real-time robotic applications. FAST-LIVO2 integrates IMU, LiDAR, and image data through an efficient error-state iterated Kalman filter (ESIKF). To address the dimensional mismatch between LiDAR and image measurements, we adopt a sequential update strategy. Efficiency is further enhanced using direct methods for LiDAR and visual data fusion: the LiDAR module registers raw points without extracting features, while the visual module minimizes photometric errors without relying on feature extraction. Both LiDAR and visual measurements are fused into a unified voxel map. The LiDAR module constructs the geometric structure, while the visual module links image patches to LiDAR points, enabling precise image alignment. Plane priors from LiDAR points improve alignment accuracy and are refined dynamically during the process. Additionally, an on-demand raycast operation and real-time image exposure estimation enhance robustness. Extensive experiments on benchmark and custom datasets demonstrate that FAST-LIVO2 outperforms state-of-the-art systems in accuracy, robustness, and efficiency. Key modules are validated, and we showcase three applications: UAV navigation highlighting real-time capabilities, airborne mapping demonstrating high accuracy, and 3D model rendering (mesh-based and NeRF-based) showcasing suitability for dense mapping. Code and datasets are open-sourced on GitHub to benefit the robotics community.
PaperID: 7,
Authors: Ajay Suresha Sathya, Justin Carpentier
Affiliations: Inria - Département d'Informatique de l'École normale supérieure, PSL Research University, Paris, France
Abstract: Rigid-body dynamics algorithms have played an essential role in robotics development. By finely exploiting the underlying robot structure, they allow the computation of the robot kinematics, dynamics, and related physical quantities with low complexity, enabling their integration into chipsets with limited resources or their evaluation at very high frequency for demanding applications (e.g., model predictive control, large-scale simulation, reinforcement learning, etc.). While most of these algorithms operate on constraint-free settings, only a few have been proposed so far to adequately account for constrained dynamical systems while depicting low algorithmic complexity. In this article, we introduce a series of new algorithms with reduced (and lowest) complexity for the forward simulation of constrained dynamical systems. Notably, we revisit the so-called articulated body algorithm (ABA) and the Popov–Vereshchagin algorithm (PV) in the light of proximal-point optimization and introduce two new algorithms, called constrained ABA and proxPV. These two new algorithms depict linear complexities while being robust to singular cases (e.g., redundant constraints, singular constraints, etc.). We establish the connection with existing literature formulations, especially the relaxed formulation at the heart of the MuJoCo and Drake simulators. We also propose an efficient and new algorithm to compute the damped Delassus inverse matrix with the lowest known computational complexity. All these algorithms have been implemented inside the open-source framework Pinocchio and depict, on a wide range of robotic systems ranging from robot manipulators to complex humanoid robots, state-of-the-art performances compared to alternative solutions of the literature.
PaperID: 8,
Authors: Luqi Wang, Yan Ning, Hongming Chen, Peize Liu, Yang Xu, Hao Xu, Ximin Lyu, Shaojie Shen
Affiliations: Department of Electronic and Computer Engineering, Hong Kong University of Science and Technology, Hong Kong, China; School of Intelligent Systems Engineering, SunYat-sen University, Guangzhou, China; Institute of Unmanned Systems, Beihang University, Beijing, China
Abstract: Multirotors are usually desired to enter confined narrow tunnels that are barely accessible to humans in various applications including inspection, search and rescue, and so on. This task is extremely challenging since the lack of geometric features and illuminations, together with the limited field of view, cause problems in perception; the restricted space and significant ego airflow disturbances induce control issues. This article introduces an autonomous aerial system designed for navigation through tunnels as narrow as 0.5 m in diameter. The real-time and online system includes a virtual omni-directional perception module tailored for the mission and a novel motion planner that incorporates perception and ego airflow disturbance factors modeled using camera projections and computational fluid dynamics analyses, respectively. Extensive flight experiments on a custom-designed quadrotor are conducted in multiple realistic narrow tunnels to validate the superior performance of the system, even over human pilots, proving its potential for real applications. In addition, a deployment pipeline on other multirotor platforms is outlined and open-source packages are provided for future developments.
PaperID: 9,
Authors: Shmuel David Alpert, Kiril Solovey, Itzik Klein, Oren Salzman
Affiliations: Technion–Israel Institute of Technology, Haifa, Israel; The Hatter Department of Marine Technologies, Charney School of Marine Sciences, University of Haifa, Haifa, Israel
Abstract: Autonomous inspection tasks require path-planning algorithms to efficiently gather observations from points of interest (POIs). However, localization errors in urban environments introduce execution uncertainty, posing challenges to successfully completing such tasks. The existing inspection-planning algorithms do not explicitly address this uncertainty, which can hinder their performance. To overcome this, in this article, we introduce incremental random inspection-roadmap search (IRIS)-under uncertainty (IRIS-U^2), an inspection-planning algorithm that provides statistical assurances regarding coverage, path length, and collision probability. Our approach builds upon IRIS—our framework for deterministic, highly efficient, and provably asymptotically optimal framework. This extension adapts IRIS to uncertain settings using a refined search procedure that estimates POI coverage probabilities through Monte Carlo (MC) sampling. We demonstrate IRIS-U^2 through a case study on bridge inspections, achieving improved expected coverage, reduced collision probability, and increasingly precise statistical guarantees as MC samples grow. In addition, we explore bounded suboptimal solutions to reduce computation time while preserving statistical assurances.
PaperID: 10,
Authors: Wilson Jallet, Antoine Bambade, Etienne Arlaud, Sarah El Kazdadi, Nicolas Mansard, Justin Carpentier
Affiliations: Inria—Département d'Informatique de l'École normale supérieure, PSL Research University, Paris, France; LAAS-CNRS, Toulouse, France
Abstract: Trajectory optimization has been a popular choice for motion generation and control in robotics for at least a decade. Several numerical approaches have exhibited the required speed to enable online computation of trajectories for real-time of various systems, including complex robots. Many of these said are based on the differential dynamic programming (DDP) algorithm—initially designed for unconstrained trajectory optimization problems—and its variants, which are relatively easy to implement and provide good runtime performance. However, several problems in robot control call for using constrained formulations (e.g., torque limits, obstacle avoidance), from which several difficulties arise when trying to adapt DDP-type methods: numerical stability, computational efficiency, and constraint satisfaction. In this article, we leverage proximal methods for constrained optimization and introduce a DDP-type method for fast, constrained trajectory optimization suited for model-predictive control (MPC) applications with easy warm-starting. Compared to earlier solvers, our approach effectively manages hard constraints without warm-start limitations and exhibits good convergence behavior. We provide a complete implementation as part of an open-source and flexible C++ trajectory optimization library called aligator. These algorithmic contributions are validated through several trajectory planning scenarios from the robotics literature and the real-time whole-body MPC of a quadruped robot.
PaperID: 11,
Authors: Md. Abir Hossen, Sonam Kharade, Jason M. O'Kane, Bradley R. Schmerl, David Garlan, Pooyan Jamshidi
Affiliations: University of South Carolina, Columbia, SC, USA; Texas A&M University, College Station, TX, USA; Carnegie Mellon University, Pittsburgh, PA, USA
Abstract: Robotic systems are typically composed of various subsystems, such as localization and navigation, each encompassing numerous configurable components (e.g., selecting different planning algorithms). Once an algorithm has been selected for a component, its associated configuration options must be set to the appropriate values. Configuration options across the system stack interact nontrivially. Finding optimal configurations for highly configurable robots to achieve desired performance poses a significant challenge due to the interactions between configuration options across software and hardware that result in an exponentially large and complex configuration space. These challenges are further compounded by the need for transferability between different environments and robotic platforms. Data efficient optimization algorithms (e.g., Bayesian optimization) have been increasingly employed to automate the tuning of configurable parameters in cyber-physical systems. However, such optimization algorithms converge at later stages, often after exhausting the allocated budget (e.g., optimization steps, allotted time) and lacking transferability. This article proposes causal understanding and remediation for enhancing robot performance (CURE)—a method that identifies causally relevant configuration options, enabling the optimization process to operate in a reduced search space, thereby enabling faster optimization of robot performance. CURE abstracts the causal relationships between various configuration options and the robot performance objectives by learning a causal model in the source (a low-cost environment such as the Gazebo simulator) and applying the learned knowledge to perform optimization in the target (e.g., Turtlebot 3 physical robot). We demonstrate the effectiveness and transferability of CURE by conducting experiments that involve varying degrees of deployment changes in both physical robots and simulation.
PaperID: 12,
Authors: James P. Wilson, Shalabh Gupta, Thomas A. Wettergren
Affiliations: Department of Electrical and Computer Engineering, University of Connecticut, Storrs, CT, USA; Naval Undersea Warfare Center, Newport, RI, USA
Abstract: The article develops a novel motion model, called generalized multispeed Dubins motion model (GMDM), which extends the Dubins model by considering multiple speeds. While the Dubins model produces time-optimal paths under a constant-speed constraint, these paths could be suboptimal if this constraint is relaxed to include multiple speeds. This is because a constant speed results in a large minimum turning radius, thus producing paths with longer maneuvers and larger travel times. In contrast, multispeed relaxation allows for slower speed sharp turns, thus producing more direct paths with shorter maneuvers and smaller travel times. Furthermore, the inability of the Dubins model to reduce speed could result in fast maneuvers near obstacles, thus producing paths with high collision risks. In this regard, GMDM provides the motion planners the ability to jointly optimize time and risk by allowing the change of speed along the path. GMDM is built upon the six Dubins path types considering the change of speed on path segments. It is theoretically established that GMDM provides full reachability of the configuration space for any speed selections. Furthermore, it is shown that the Dubins model is a specific case of GMDM for constant speeds. The solutions of GMDM are analytical and suitable for real-time applications. The performance of GMDM in terms of solution quality (i.e., time/time-risk cost) and computation time is comparatively evaluated against the existing motion models in obstacle-free as well as obstacle-rich environments via extensive Monte Carlo simulations. The results show that in obstacle-free environments, GMDM produces near time-optimal paths with significantly lower travel times than the Dubins model while having similar computation times. In obstacle-rich environments, GMDM produces time-risk optimized paths with substantially lower collision risks.
PaperID: 13,
Authors: Sheng Zhong, Nima Fazeli, Dmitry Berenson
Affiliations: Department of Robotics, University of Michigan, Ann Arbor, MI, USA
Abstract: In this article, we present rummaging using mutual information (RUMI), a method for online generation of robot action sequences to gather information about the pose of a known movable object in visually occluded environments. Focusing on contact-rich rummaging, our approach leverages mutual information between the object pose distribution and robot trajectory for action planning. From an observed partial point cloud, RUMI deduces the compatible object pose distribution and approximates the mutual information of it with workspace occupancy in real time. Based on this, we develop an information gain cost function and a reachability cost function to keep the object within the robot’s reach. These are integrated into a model predictive control (MPC) framework with a stochastic dynamics model, updating the pose distribution in a closed loop. Key contributions include a new belief framework for object pose estimation, an efficient information gain computation strategy, and a robust MPC-based control scheme. RUMI demonstrates superior performance in both simulated and real tasks compared to baseline methods.
PaperID: 14,
Authors: Shamsa Al Harthy, S. M. Hadi Sadati, Cédric Girerd, Sukjun Kim, Alessio Mondini, Zicong Wu, Brandon Saldarriaga, Carlo A. Seneci, Barbara Mazzolai, Tania K. Morimoto, Christos Bergeles
Affiliations: School of Biomedical Engineering and Imaging Sciences, King’s College London, London, U.K.; LIRMM, University of Montpellier, CNRS, Montpellier, France; Department of Mechanical and Aerospace Engineering, University of California, San Diego, La Jolla, CA, USA; Bioinspired Soft Robotics Laboratory, Istituto Italiano di Tecnologia, Genova, Italy
Abstract: Growing robots apically extend through material eversion or deposition at their tip. This endows them with unique capabilities, such as follow the leader navigation, long-reach, inherent compliance, and large force delivery bandwidth. Tip-growing robots can therefore conform to sensitive, intricate, and difficult-to-access environments. This review article categorizes, compares, and critically evaluates state-of-the-art growing robots with emphasis on their designs, fabrication processes, actuation and steering mechanisms, mechanics models, controllers, and applications. Finally, this article discusses the main challenges that the research area still faces and proposes future directions.
PaperID: 15,
Authors: Matous Vrba, Viktor Walter, Václav Pritzl, Michal Pliska, Tomás Báca, Vojtech Spurný, Daniel Hert, Martin Saska
Affiliations: Faculty of Electrical Engineering, Czech Technical University in Prague, Prague, Czech Republic
Abstract: A new robust and accurate approach for the detection and localization of flying objects with the purpose of highly dynamic aerial interception and agile multirobot interaction is presented in this article. The approach is proposed for use on board of autonomous aerial vehicles equipped with a 3-D LiDAR sensor. It relies on a novel 3-D occupancy voxel mapping method for the target detection that provides high localization accuracy and robustness with respect to varying environments and appearance changes of the target. In combination with a proposed cluster-based multitarget tracker, sporadic false positives are suppressed, state estimation of the target is provided, and the detection latency is negligible. This makes the system suitable for tasks of agile multirobot interaction, such as autonomous aerial interception or formation control where fast, precise, and robust relative localization of other robots is crucial. We evaluate the viability and performance of the system in simulated and real-world experiments which demonstrate that at a range of \text20 \,\textm, our system is capable of reliably detecting a microscale UAV with an almost \text100 % recall, \text0.2 \,\textm accuracy, and \text20 \,\textms delay.
PaperID: 16,
Authors: Guozheng Lu, Yunfan Ren, Fangcheng Zhu, Haotian Li, Ruize Xue, Yixi Cai, Ximin Lyu, Fu Zhang
Affiliations: Department of Mechanical Engineering, The University of Hong Kong, Hong Kong; School of Intelligent System Engineering, Sun Yat-sen University, Shenzhen, China
Abstract: Trajectory generation for fully autonomous flights of tail-sitter unmanned aerial vehicles (UAVs) presents substantial challenges due to their highly nonlinear aerodynamics. In this article, we introduce, to the best of the authors' knowledge, the world's first fully autonomous tail-sitter UAV capable of high-speed navigation in unknown, cluttered environments. The UAV autonomy is enabled by cutting-edge technologies including LiDAR-based sensing, differential-flatness-based trajectory planning and control with purely onboard computation. In particular, we propose an optimization-based tail-sitter trajectory planning framework that generates high-speed, collision-free, and dynamically-feasible trajectories. To efficiently and reliably solve this nonlinear, constrained problem, we develop an efficient feasibility-assured solver, Efficient Feasibility-assured OPTimization solver (EFOPT), tailored for the online planning of tail-sitter UAVs. We conduct extensive simulation studies to benchmark EFOPT's superiority in planning tasks against conventional nonlinear programming solvers. We also demonstrate exhaustive experiments of aggressive autonomous flights with speeds up to 15 m/s in various real-world environments, including indoor laboratories, underground parking lots, and outdoor parks.
PaperID: 17,
Authors: Cem Bilaloglu, Tobias Löw, Sylvain Calinon
Affiliations: Idiap Research Institute, Martigny, Switzerland
Abstract: In this article, we present a feedback control method for tactile coverage tasks such as cleaning or surface inspection. Although these tasks are challenging to plan due to the complexity of continuous physical interactions, the coverage target and progress can be effectively measured using a camera and encoded in a point cloud. We propose an ergodic coverage method that operates directly on point clouds, guiding the robot to spend more time on regions requiring more coverage. For robot control and contact behavior, we use geometric algebra to formulate a task-space impedance controller that tracks a line while simultaneously exerting a desired force along that line. We evaluate the performance of our method in kinematic simulations and demonstrate its applicability in real-world experiments on kitchenware.
PaperID: 18,
Authors: Muchen Sun, Ayush Gaggar, Pete Trautman, Todd D. Murphey
Affiliations: Department of Mechanical Engineering, Northwestern University, Evanston, IL, USA; Honda Research Institute, San Jose, CA, USA
Abstract: Ergodic search enables optimal exploration of an information distribution with guaranteed asymptotic coverage of the search space. However, current methods typically have exponential computational complexity and are limited to Euclidean space. We introduce a computationally efficient ergodic search method. Our contributions are two-fold as follows: First, we develop a kernel-based ergodic metric, generalizing it from Euclidean space to Lie groups. We prove this metric is consistent with the exact ergodic metric and ensures linear complexity. Second, we derive an iterative optimal control algorithm for trajectory optimization with the kernel metric. Numerical benchmarks show our method is two orders of magnitude faster than the state-of-the-art method. Finally, we demonstrate the proposed algorithm with a peg-in-hole insertion task. We formulate the problem as a coverage task in the space of SE(3) and use a 30-s-long human demonstration as the prior distribution for ergodic coverage. Ergodicity guarantees the asymptotic solution of the peg-in-hole problem so long as the solution resides within the prior information distribution, which is seen in the 100% success rate.
PaperID: 19,
Authors: Javier Laserna Moratalla, Pablo San Segundo, David Álvarez
Affiliations: Centre for Automation and Robotics (CAR) ETSIDI, UPMCSIC, Universidad Politécnica de Madrid, Madrid, Spain
Abstract: We propose a branch-and-bound algorithm for robust rigid registration of two point clouds in the presence of a large number of outlier correspondences. For this purpose, we consider a maximum consensus formulation of the registration problem and reformulate it as a (large) maximal clique search in a correspondence graph, where a clique represents a complete rigid transformation. Specifically, we use a maximum clique algorithm to enumerate large maximal cliques and a fitness procedure that evaluates each clique by solving a least-squares optimization problem. The main advantages of our approach are 1) it is possible to exploit the cutting-edge optimization techniques employed by current exact maximum clique algorithms, such as partial maximum satisfiability-based bounds, branching by partitioning or the use of bitstrings, etc.; 2) the correspondence graphs are expected to be sparse in real problems (confirmed empirically in our tests), and, consequently, the maximum clique problem is expected to be easy; 3) it is possible to have a good control of suboptimality with a k-nearest neighbor analysis that determines the size of the correspondence graph as a function of k. The new algorithm is called CliReg and has been implemented in C++. To evaluate CliReg, we have carried out extensive tests both on synthetic and real public datasets. The results show that CliReg clearly dominates the state of the art (e.g., RANSAC, FGR, and TEASER++) in terms of robustness, with a running time comparable to TEASER++ and RANSAC. In addition, we have implemented a fast variant called CliRegMutual that performs similarly to the fastest heuristic FGR.
PaperID: 20,
Authors: Grazia Zambella, Danilo Caporale, Giorgio Grioli, Lucia Pallottino, Antonio Bicchi
Affiliations: Centro di Ricerca “Enrico Piaggio” Università di Pisa, Pisa, Italy; Technology Innovation Institute, Abu Dhabi, UAE; Centro di Ricerca “Enrico Piaggio,” Università di Pisa, Pisa, Italy
Abstract: Due to their fast and efficient locomotion, two-wheeled humanoids are fascinating systems with the potential to be involved in many application domains, including healthcare, manufacturing, and many others. However, these robots constitute a challenging case of study for control purposes due to the two-wheeled inverted pendulum dynamics that characterizes their mobility and support, as it is underactuated and unstable. In this article, we propose a novel whole-body control approach to stabilize two-wheeled humanoids. To tackle the control problem of their forward motion and pitch equilibrium, leveraging on the observation that such systems are usually characterized by a faster and a slower dynamics (being the pitch angle faster and the forward displacement slower), we design a composite whole-body control that combines two computed-torque control loops to stabilize both dynamics to the desired trajectories. The control approach is introduced and its derivation is described for the simpler case of a two-wheeled inverted pendulum first, and for a whole two-wheeled humanoid after. To prove its validity, the control approach is tested experimentally on the two-wheeled humanoid robot Alter-Ego. The robot proves to be able to perform complicated interaction tasks, including opening a door, grasping a heavy object, and resisting to external dynamic disturbances.
PaperID: 21,
Authors: Dennis Benders, Johannes Köhler, Thijs Niesten, Robert Babuska, Javier Alonso-Mora, Laura Ferranti
Affiliations: Department of Cognitive Robotics, Delft University of Technology, Delft, CD, The Netherlands; Institute for Dynamic Systems and Control, Zürich, Switzerland
Abstract: To efficiently deploy robotic systems in society, mobile robots must move autonomously and safely through complex environments. Nonlinear model predictive control (MPC) methods provide a natural way to find a dynamically feasible trajectory through the environment without colliding with nearby obstacles. However, the limited computation power available on typical embedded robotic systems, such as quadrotors, poses a challenge to running MPC in real time, including its most expensive tasks: constraints generation and optimization. To address this problem, we propose a novel hierarchical MPC scheme that consists of a planning and a tracking layer. The planner constructs a trajectory with a long prediction horizon at a slow rate, while the tracker ensures trajectory tracking at a relatively fast rate. We prove that the proposed framework avoids collisions and is recursively feasible. Furthermore, we demonstrate its effectiveness in simulations and lab experiments with a quadrotor that needs to reach a goal position in a complex static environment. The code is efficiently implemented on the quadrotor's embedded computer to ensure real-time feasibility. Compared to a state-of-the-art single-layer MPC formulation, this allows us to increase the planning horizon by a factor of 5, which results in significantly better performance.
PaperID: 22,
Authors: Kaicheng Zhang, Shida Xu, Yining Ding, Xianwen Kong, Sen Wang
Affiliations: Department of Electrical and Electronic Engineering, Imperial College London, London, U.K.; School of Engineering and Physical Sciences, Heriot-Watt University, Edinburgh, U.K.
Abstract: This article studies 3-D light detection and ranging (LiDAR) mapping with a focus on developing an updatable and localizable map representation that enables continuity, compactness, and consistency in 3-D maps. Traditional LiDAR simultaneous localization and mapping (SLAM) systems often rely on 3-D point cloud maps, which typically require extensive storage to preserve structural details in large-scale environments. In this article, we propose a novel paradigm for LiDAR SLAM by leveraging the continuous and ultracompact representation of LiDAR (CURL). Our proposed LiDAR mapping approach, CURL-SLAM, produces compact 3-D maps capable of continuous reconstruction at variable densities using CURL’s spherical harmonics implicit encoding, and achieves global map consistency after loop closure. Unlike popular iterative-closest-point-based LiDAR odometry techniques, CURL-SLAM formulates LiDAR pose estimation as a unique optimization problem tailored for CURL and extends it to local bundle adjustment, enabling simultaneous pose refinement and map correction. Experimental results demonstrate that CURL-SLAM achieves state of the art 3-D mapping quality and competitive LiDAR trajectory accuracy, delivering sensor-rate real-time performance (10 Hz) on a CPU. We will release the CURL-SLAM implementation to the community.
PaperID: 23,
Authors: Taerim Yoon, Dongho Kang, Seungmin Kim, Jin Cheng, Min Sung Ahn, Stelian Coros, Sungjoon Choi
Affiliations: Department of Artificial Intelligence, Korea University, Seoul, Seongbuk-gu, South Korea; Department of Computer Science, ETH Zurich, Zurich, Switzerland; Department of Mechanical and Aerospace Engineering, UCLA, Los Angeles, CA, USA
Abstract: This work presents a motion retargeting approach for legged robots, aimed at transferring the dynamic and agile movements to robots from source motions. In particular, we guide the imitation learning procedures by transferring motions from source to target, effectively bridging the morphological disparities while ensuring the physical feasibility of the target system. In the first stage, we focus on motion retargeting at the kinematic level by generating kinematically feasible whole-body motions from keypoint trajectories. Following this, we refine the motion at the dynamic level by adjusting it in the temporal domain while adhering to physical constraints. This process facilitates policy training via reinforcement learning, enabling precise and robust motion tracking. We demonstrate that our approach successfully transforms noisy motion sources, such as hand-held camera videos, into robot-specific motions that align with the morphology and physical properties of the target robots. Moreover, we demonstrate terrain-aware motion retargeting to perform BackFlip on top of a box. We successfully deployed these skills to four robots with different dimensions and physical properties in the real world through hardware experiments.
PaperID: 24,
Authors: Reza Vafaee, Kian Behzad, Milad Siami, Luca Carlone, Ali Jadbabaie
Affiliations: Department of Electrical & Computer Engineering, Northeastern University, Boston, MA, USA; Laboratory for Information & Decision Systems, Massachusetts Institute of Technology, Cambridge, MA, USA
Abstract: This article presents a task-oriented computational framework to enhance visual-inertial navigation (VIN) in robots, addressing challenges such as limited time and energy resources. The framework strategically selects visual features using a mean squared error (MSE)-based, nonsubmodular objective function and a simplified dynamic anticipation model. To address the NP-hardness of this problem, we introduce four polynomial-time approximation algorithms: a classic greedy method with constant-factor guarantees; a low-rank greedy variant that significantly reduces computational complexity; a randomized greedy sampler that balances efficiency and solution quality; and a linearization-based selector based on a first-order Taylor expansion for near-constant-time execution. We establish rigorous performance bounds by leveraging submodularity ratios, curvature, and elementwise curvature analyses. Extensive experiments on both standardized benchmarks and a custom control-aware platform validate our theoretical results, demonstrating that these methods achieve strong approximation guarantees while enabling real-time deployment.
PaperID: 25,
Authors: Oscar de Groot, Laura Ferranti, Dariu M. Gavrila, Javier Alonso-Mora
Affiliations: Department of Cognitive Robotics, TU Delft, Delft, The Netherlands
Abstract: Ground robots navigating in complex, dynamic environments must compute collision-free trajectories to avoid obstacles safely and efficiently. Nonconvex optimization is a popular method to compute a trajectory in real time. However, these methods often converge to locally optimal solutions and frequently switch between different local minima, leading to inefficient and unsafe robot motion. In this work, we propose a novel topology-driven trajectory optimization strategy for dynamic environments that plans multiple distinct evasive trajectories to enhance the robot's behavior and efficiency. A global planner iteratively generates trajectories in distinct homotopy classes. These trajectories are then optimized by local planners working in parallel. While each planner shares the same navigation objectives, they are locally constrained to a specific homotopy class, meaning each local planner attempts a different evasive maneuver. The robot then executes the feasible trajectory with the lowest cost in a receding horizon manner. We demonstrate on a mobile robot navigating among pedestrians that our approach leads to faster trajectories than existing planners.
PaperID: 26,
Authors: Zihang Zhao, Yuyang Li, Wanlin Li, Zhenghao Qi, Lecheng Ruan, Yixin Zhu, Kaspar Althoefer
Affiliations: Institute for Artificial Intelligence, Peking University, Beijing, China; Beijing Institute for General Artificial Intelligence, Beijing, China; College of Engineering, Peking University, Beijing, China; Centre for Advanced Robotics @ Queen Mary, School of Engineering and Materials Science, Queen Mary University of London, London, U.K.
Abstract: Integrating robots into human-centric environments such as homes, necessitates advanced manipulation skills as robotic devices will need to engage with articulated objects such as doors and drawers. Key challenges in robotic manipulation of articulated objects are the unpredictability and diversity of these objects' internal structures, which render models based on object kinematics priors, both explicit and implicit, and inadequate. Their reliability is significantly diminished by pre-interaction ambiguities, imperfect structural parameters, encounters with unknown objects, and unforeseen disturbances. Here, we present a prior-free strategy, Tac-Man, focusing on maintaining stable robot-object contact during manipulation. Without relying on object priors, Tac-Man leverages tactile feedback to enable robots to proficiently handle a variety of articulated objects, including those with complex joints, even when influenced by unexpected disturbances. Demonstrated in both real-world experiments and extensive simulations, it consistently achieves near-perfect success in dynamic and varied settings, outperforming existing methods. Our results indicate that tactile sensing alone suffices for managing diverse articulated objects, offering greater robustness and generalization than prior-based approaches. This underscores the importance of detailed contact modeling in complex manipulation tasks, especially with articulated objects. Advancements in tactile-informed approaches significantly expand the scope of robotic applications in human-centric environments, particularly where accurate models are difficult to obtain.
PaperID: 27,
Authors: Haonan Duan, Peng Wang, Yifan Yang, Daheng Li, Wei Wei, Yongkang Luo, Guoqiang Deng
Affiliations: State Key Laboratory of Multimodal Artificial Intelligence Systems, Institute of Automation, Chinese Academy of Sciences, Beijing, China
Abstract: Human-robot object handovers are essential for robots to effectively serve human needs in various domains of human–robot interaction and collaboration, yet remain a significant challenge. Remarkable progress has been made by parallel-jaw gripper robots in grasp generation and motion planning for handovers, while few studies address this issue using anthropomorphic hands, necessitating the ability to handle higher collision probabilities and lower approaching space under occluded situations. In this article, we present a reactive human-to-robot dexterous handover framework for anthropomorphic hands. The closed-loop framework employs an effective collision detection and grasp selection approach to ensure safe and smooth motion in unstructured environments. We implement a handover system using a UR5 robot arm and a Schunk SVH Hand based on the presented framework, which can react to human motion during the handover process, generalize to diverse objects with different 6-DoF poses, and execute suitable grasp configurations. The generalizability, reliability, and efficiency of our method are demonstrated through the handover of 30 novel objects, a system ablation study for submodule evaluation, and a user study assessment involving eight participants.
PaperID: 28,
Authors: Federico Pratissoli, Mattia Mantovani, Amanda Prorok, Lorenzo Sabattini
Affiliations: Department of Sciences and Methods for Engineering, University of Modena and Reggio Emilia, Modena, Italy; Department of Computer Science, University of Cambridge, Cambridge, U.K.
Abstract: Multirobot systems are essential for environmental monitoring, particularly for tracking spatial phenomena like pollution, soil minerals, and water salinity, and more. This study addresses the challenge of deploying a multirobot team for optimal coverage in environments where the density distribution, describing areas of interest, is unknown and changes over time. We propose a fully distributed control strategy that uses Gaussian processes (GPs) to model the spatial field and balance the tradeoff between learning the field and optimally covering it. Unlike existing approaches, we address a more realistic scenario by handling time-varying spatial fields, where the exploration-exploitation tradeoff is dynamically adjusted over time. Each robot operates locally, using only its own collected data and the information shared by the neighboring robots. To address the computational limits of GPs, the algorithm efficiently manages the volume of data by selecting only the most relevant samples for the process estimation. The performance of the proposed algorithm is evaluated through several simulations and experiments, incorporating real-world data phenomena to validate its effectiveness.
PaperID: 29,
Authors: Shiming He, Alexander von Rohr, Dominik Baumann, Ji Xiang, Sebastian Trimpe
Affiliations: School of Information and Electrical Engineering, Hangzhou City University, Hangzhou, China; Institute for Data Science in Mechanical Engineering, RWTH Aachen University, Aachen, Germany; Department of Electrical Engineering and Automation, Aalto University, Espoo, Finland; College of Electrical Engineering, Zhejiang University, Hangzhou, China
Abstract: How can robots learn and adapt to new tasks and situations with little data? Systematic exploration and simulation are crucial tools for efficient robot learning. We present a novel black-box policy search algorithm focused on data-efficient policy improvements. The algorithm learns directly on the robot and treats simulation as an additional information source to speed up the learning process. At the core of the algorithm, a probabilistic model learns the dependence between the policy parameters and the robot learning objective not only by performing experiments on the robot, but also by leveraging data from a simulator. This substantially reduces interaction time with the robot. Using the model, we can guarantee improvements with high probability for each policy update, thereby facilitating fast, goal-oriented learning. We evaluate our algorithm on simulated fine-tuning tasks and demonstrate the data-efficiency of the proposed dual-information source optimization algorithm. In a real robot learning experiment, we show fast and successful task learning on a robot manipulator with the aid of an imperfect simulator.
PaperID: 30,
Authors: Liangming Chen, Chenyang Liang, Shenghai Yuan, Muqing Cao, Lihua Xie
Affiliations: School of System Design and Intelligent Manufacturing and Guangdong Provincial Key Laboratory of Fully Actuated System Control Theory and Technology, Southern University of Science and Technology, Shenzhen, China; School of Mechanical Engineering and Automation, Harbin Institute of Technology, Shenzhen, China; School of Electrical and Electronic Engineering, Nanyang Technological University, Singapore; Robotics Institute, Carnegie Mellon University, Pittsburgh, PA, USA
Abstract: Inter-robot relative positions are crucial for executing various multirobot missions, such as formation maneuvering and collaborative inspection. However, the current sensing technology usually provides part of relative position information, such as inter-robot distances, bearings and angles. This prompts the study of determining inter-robot relative positions, i.e., relative localization, from these partial measurements. Based on the existing results of static networks' localizability and mobile robots' relative localization, we propose a novel concept, relative localizability to describe whether a multirobot system is relatively localizable. Given each robot's self-displacement measurements and inter-robot partial measurements in d (d\leq 4) sampling instants, we show that a multirobot system's relative localization can be achieved in a purely algebraic and distributed manner, in which the multirobot system is said to be d-step relatively localizable. To make the results more general, we consider that the multirobot system consists of landmarks, leaders, and followers, and that the inter-robot measurements can be distances, bearings or angles. When robots' coordinate frames have different orientations, we show that the given local measurements can be used to determine robots' relative positions and their coordinate frames' relative orientations simultaneously. Simulations and experiments of relative localization for ground robots are conducted to validate the obtained results.
PaperID: 31,
Authors: Peng Yin, Jianhao Jiao, Shiqi Zhao, Lingyun Xu, Guoquan Huang, Howie Choset, Sebastian A. Scherer, Jianda Han
Affiliations: Department of Mechanical Engineering, City University of Hong Kong, Hong Kong; Department of Computer Science, University College London, London, U.K.; Robotics Institute, Carnegie Mellon University, Pittsburgh, PA, USA; Robot Perception and Navigation Group, University of Delaware, Newark, DE, USA; Nankai University, Tianjin, China
Abstract: In the realm of robotics, the quest for achieving real-world autonomy, capable of executing large-scale and long-term operations, has positioned place recognition (PR) as a cornerstone technology. Despite the PR community's remarkable strides over the past two decades, garnering attention from fields like computer vision and robotics, the development of PR methods that sufficiently support real-world robotic systems remains a challenge. This article aims to bridge this gap by highlighting the crucial role of PR within the framework of simultaneous localization and mapping 2.0. This new phase in robotic navigation calls for scalable, adaptable, and efficient PR solutions by integrating advanced artificial intelligence technologies. For this goal, we provide a comprehensive review of the current state-of-the-art advancements in PR, alongside the remaining challenges, and underscore its broad applications in robotics. This article begins with an exploration of PR's formulation and key research challenges. We extensively review literature, focusing on related methods on place representation and solutions to various PR challenges. Applications showcasing PR's potential in robotics, key PR datasets, and open-source libraries are discussed.
PaperID: 32,
Authors: Chuhao Liu, Zhijian Qiao, Jieqi Shi, Ke Wang, Peize Liu, Shaojie Shen
Affiliations: Department of Electronic and Computer Engineering, The Hong Kong University of Science and Technology, Hong Kong; School of Intelligence Science and Technology, Nanjing University, Nanjing, China; School of Information Engineering, Chang’an University, Xi’an, China
Abstract: This article addresses the challenges of registering two rigid semantic scene graphs, an essential capability when an autonomous agent needs to register its map against a remote agent, or against a prior map. The handcrafted descriptors in classical semantic-aided registration, or the ground-truth annotation reliance in learning-based scene graph registration, impede their application in practical real-world environments. To address the challenges, we design a scene graph network to encode multiple modalities of semantic nodes: open-set semantic feature, local topology with spatial awareness, and shape feature. These modalities are fused to create compact semantic node features. The matching layers then search for correspondences in a coarse-to-fine manner. In the back end, we employ a robust pose estimator to decide transformation according to the correspondences. We manage to maintain a sparse and hierarchical scene representation. Our approach demands fewer GPU resources and fewer communication bandwidth in multiagent tasks. Moreover, we design a new data generation approach using vision foundation models and a semantic mapping module to reconstruct semantic scene graphs. It differs significantly from previous works, which rely on ground-truth semantic annotations to generate data. We validate our method in a two-agent simultaneous localization and mapping benchmark. It significantly outperforms the handcrafted baseline in terms of registration success rate. Compared to visual loop closure networks, our method achieves a slightly higher registration recall while requiring only 52 kB of communication bandwidth for each query frame.
PaperID: 33,
Authors: Dingqi Zhang, Antonio Loquercio, Jerry Tang, Ting-Hao Wang, Jitendra Malik, Mark W. Mueller
Affiliations: High Performance Robotics Lab, Department of Mechanical Engineering, UC Berkeley, Berkeley, CA, USA; University of Pennsylvania, Philadelphia, CA, USA; Department of Electrical Engineering and Computer Science, University of California at Berkeley, Berkeley, CA, USA
Abstract: This article introduces a learning-based low-level controller for quadcopters, which adaptively controls quadcopters with significant variations in mass, size, and actuator capabilities. Our approach leverages a combination of imitation learning and reinforcement learning, creating a fast-adapting and general control framework for quadcopters that eliminates the need for precise model estimation or manual tuning. The controller estimates a latent representation of the vehicle’s system parameters from sensor-action history, enabling it to adapt swiftly to diverse dynamics. Extensive evaluations in simulation demonstrate the controller’s ability to generalize to unseen quadcopter parameters, with an adaptation range up to 16 times broader than the training set. In real-world tests, the controller is successfully deployed on quadcopters with mass differences of 3.7 times and propeller constants varying by more than 100 times, while also showing rapid adaptation to disturbances such as off-center payloads and motor failures. These results highlight the potential of our controller to simplify the design process and enhance the reliability of autonomous drone operations in unpredictable environments.
PaperID: 34,
Authors: Fabian Schmidt, Julian Daubermann, Marcel Mitschke, Constantin Blessing, Stephan Meyer, Markus Enzweiler, Abhinav Valada
Affiliations: Institute for Intelligent Systems, Esslingen University of Applied Sciences, Esslingen, Germany; ANDREAS STIHL AG and Company KG, Waiblingen, Germany; Department of Computer Science, University of Freiburg, Freiburg, Germany
Abstract: Robust simultaneous localization and mapping (SLAM) is a crucial enabler for autonomous navigation in natural, semistructured environments such as parks and gardens. However, these environments present unique challenges for SLAM due to frequent seasonal changes, varying light conditions, and dense vegetation. These factors often degrade the performance of visual SLAM algorithms originally developed for structured urban environments. To address this gap, we present robot outdoor visual SLAM dataset for environmental robustness (ROVER), a comprehensive benchmark dataset tailored for evaluating visual SLAM algorithms under diverse environmental conditions and spatial configurations. We captured the dataset with a robotic platform equipped with monocular, stereo, and RGBD cameras, as well as inertial sensors. It covers 39 recordings across five outdoor locations, collected through all seasons and various lighting scenarios, i.e., day, dusk, and night with and without external lighting. With this novel dataset, we evaluate several traditional and deep learning-based SLAM methods and study their performance in diverse challenging conditions. The results demonstrate that while stereo-inertial and RGBD configurations generally perform better under favorable lighting and moderate vegetation, most SLAM systems perform poorly in low-light and high-vegetation scenarios, particularly during summer and autumn. Our analysis highlights the need for improved adaptability in visual SLAM algorithms for outdoor applications, as current systems struggle with dynamic environmental factors affecting scale, feature extraction, and trajectory consistency. This dataset provides a solid foundation for advancing visual SLAM research in real-world, semistructured environments, fostering the development of more resilient SLAM systems for long-term outdoor localization and mapping.
PaperID: 35,
Authors: Vignesh Prasad, Lea Heitlinger, Dorothea Koert, Ruth Stock-Homburg, Jan Peters, Georgia Chalvatzaki
Affiliations: Interactive Robot Perception and Learning Group (PEARL), Department of Computer Science, TU Darmstadt, Darmstadt, Germany; Chair for Marketing and Human Resource Management, Department of Law and Economics, TU Darmstadt, Darmstadt, Germany; Interactive AI Algorithms & Cognitive Models for Human-AI Interaction (IKIDA), Department of Computer Science, TU Darmstadt, Darmstadt, Germany; Centre for Cognitive Science, TU Darmstadt, Darmstadt, Germany
Abstract: This article presents a method for learning well-coordinated human–robot interaction (HRI) from human–human interactions (HHI). We devise a hybrid approach using hidden Markov models (HMMs) as the latent space priors for a variational autoencoder to model a joint distribution over the interacting agents. We leverage the interaction dynamics learned from HHI to learn HRI and incorporate the conditional generation of robot motions from human observations into the training, thereby predicting more accurate robot trajectories. The generated robot motions are further adapted with inverse kinematics to ensure the desired physical proximity with a human, combining the ease of joint space learning and accurate task space reachability. For contact-rich interactions, we modulate the robot’s stiffness using HMM segmentation for a compliant interaction. We verify the effectiveness of our approach deployed on a humanoid robot via a user study. Our method generalizes well to various humans despite being trained on data from just two humans. We find that users perceive our method as more human-like, timely, and accurate and rank our method with a higher degree of preference over other baselines. We additionally show the ability of our approach to generate successful interactions in a more complex scenario of bimanual robot-to-human handovers.
PaperID: 36,
Authors: Xiuyuan Lu, Yi Zhou, Jiayao Mai, Kuan Dai, Yang Xu, Shaojie Shen
Affiliations: Department of Electronic and Computer Engineering, The Hong Kong University of Science and Technology, Hong Kong; School of Robotics, Hunan University, Changsha, China; Division of Emerging Interdisciplinary Areas, The Hong Kong University of Science and Technology, Hong Kong
Abstract: Neuromorphic event-based cameras are bioinspired visual sensors with asynchronous pixels and extremely high temporal resolution. Such favorable properties make them an excellent choice for solving state estimation tasks under high-speed maneuvers. However, failures of camera pose tracking are frequently witnessed in state-of-the-art event-based visual odometry systems when the local map cannot be updated timely or feature matching is unreliable. One of the biggest roadblocks in this field is the absence of efficient and robust methods for data association without imposing any assumptions on the environment. This problem seems, however, unlikely to be addressed as in standard vision because of the motion-dependent nature of event data. To address this, we propose a map-free design for event-based visual-inertial state estimation in this article. Instead of estimating camera position, we find that recovering the instantaneous linear velocity aligns better with event cameras’ differential working principle. The proposed system uses raw data from a stereo event camera and an inertial measurement unit (IMU) as input, and adopts a dual-end architecture. The front-end preprocesses raw events and executes the computation of normal flow and depth information. To handle the temporally nonequispaced event data and establish association with temporally nonaligned IMU’s measurements, the back-end employs a continuous-time formulation and a sliding-window scheme that can progressively estimate the linear velocity and IMU’s bias. Experiments on synthetic and real data show our method achieves low-latency, metric-scale velocity estimation. To the best of the authors’ knowledge, this is the first real-time, purely event-based visual-inertial state estimator for high-speed maneuvers, requiring only sufficient textures and imposing no additional constraints on either the environment or motion pattern.
PaperID: 37,
Authors: Julio L. Paneque, J. Ramiro Martinez de Dios, Aníbal Ollero
Affiliations: GRVC Robotics Lab Sevilla, Universidad de Sevilla, Seville, Spain
Abstract: Although a good variety of successful LiDAR-based mapping schemes have been developed, these methods present shortcomings when mapping geometrically poor environments. In these scenarios, the chosen map structure and the consideration of map uncertainty are particularly relevant for providing a robust robot motion estimation, critically affecting the quality of the resulting map. This article introduces the use of probabilistic 3-D triangle meshes in LiDAR-based mapping. Our approach combines: 1) meshes, which consistently represent planar surfaces and enable the use of decimation techniques to reduce the influence of the measurement noise in the map and improve map fidelity, while strongly reducing the map size; with 2) a probabilistic on-manifold formulation of planar objects, which naturally reflects the measurement uncertainty in the mesh map avoiding inconsistencies in state estimation. The proposed methods are experimentally evaluated both individually and jointly integrated in a generic mapping scheme in different scenarios, showing the improvement in robustness and accuracy in geometrically poor environments and providing strong reductions in map size over existing schemes. We release the used datasets and C++ implementations of the proposed methods.
PaperID: 38,
Authors: Jiaxi Wu, Mingxin Wu, Chen Wang, Guangming Xie
Affiliations: State Key Laboratory for Turbulence and Complex Systems, Intelligent Biomimetic Design Lab, College of Engineering, Peking University, Beijing, China; National Engineering Research Center of Software Engineering, Peking University, Beijing, China
Abstract: Inspired by natural organisms, soft robots have showcased remarkable performance across various functions. However, creating multifunctional soft robotic systems typically increases manufacturing complexity, resulting in a cumbersome fabrication workflow with low programmability. Here, we present a monolithic fabric-based approach for the programmable fabrication of multifunctional soft robots. Our method involves programming bonding paths and sequentially attaching fabric layers to directly manufacture monolithic robots. The fabric is precisely shaped using a laser cutter, while a 3-D printer follows predesigned bonding paths to ensure repeatable manufacturing. By programming the contours of each fabric layer and their corresponding bonding paths, we create versatile soft robots that integrate expected functionalities, including large-range manipulation, multimodal locomotion, and their harmonious combination. Our approach offers a promising avenue to efficiently create multifunctional soft robots via monolithic and customized fabrication, which will accelerate the proliferation of soft robots and open the doors to a wide range of applications.
PaperID: 39,
Authors: Armand Jordana, Sébastien Kleff, Avadesh Meduri, Justin Carpentier, Nicolas Mansard, Ludovic Righetti
Affiliations: Machines in Motion Laboratory, New York University, New York, NY, USA; Inria - Département d’Informatique de l’École normale supérieure, PSL Research University, Paris, France; LAAS-CNRS, Université de Toulouse, CNRS, Toulouse, France
Abstract: The promise of model-predictive control (MPC) in robotics has led to extensive development of efficient numerical optimal control solvers in line with differential dynamic programming because it exploits the sparsity induced by time. In this work, we argue that this effervescence has hidden the fact that sparsity can be equally exploited by standard nonlinear optimization. In particular, we show how a tailored implementation of sequential quadratic programming (QP) achieves state-of-the-art MPC. Then, we clarify the connections between popular algorithms from the robotics community and well-established optimization techniques. Further, the sequential quadratic program formulation naturally encompasses the constrained case, a notoriously difficult problem in the robotics community. Specifically, we show that it only requires a sparsity-exploiting implementation of a state-of-the-art QP solver. We illustrate the validity of this approach in a comparative study and experiments on a torque-controlled manipulator. To the best of our knowledge, this is the first demonstration of closed loop nonlinear MPC with constraints on a real robot.
PaperID: 40,
Authors: Hao Liu, Changchun Wu, Senyuan Lin, Yunquan Li, Yonghua Chen, James Lam, Ning Xi
Affiliations: Department of Mechanical Engineering, The University of Hong Kong, Hong Kong; Shien-Ming Wu School of Intelligent Engineering, South China University of Technology, Guangzhou, China; Department of Data and Systems Engineering, The University of Hong Kong, Hong Kong
Abstract: Soft pneumatic actuators (SPAs) are widely used in robotic applications due to their inherent compliance and outstanding mechanical performance. However, a tradeoff between load capacity and deformation capacity is required when selecting material hardness for SPAs. Too hard material usually results in limited deformation, such as small extension, bending, and twisting, while too soft a material decreases robustness. This study introduces coil-reinforced flat tube actuators (CFTAs) that exhibit excellent flexibility and high load capacity by braiding flat tubes in coil springs. By adjusting the braiding pattern of the flat tube, the CFTAs can realize extending, in-plane bending, and out-of-plane helical bending motions. In addition, analytical models are proposed to predict the deformation behavior of the CFTAs and verified by experiments. The bending type CFTA deforms with an excellent curvature (0.516 mm−1) under 160 kPa input pressure, five times larger than the reported SPAs on the same scale. The CFTAs show high design flexibility by programming the flat tube pattern and using multiple coiled springs for wide application scenarios. Based on the CFTAs, this study shows wearable upper limb robots, a soft entanglement gripper, and a rob-climbing robot. CFTAs provide design insight for applications requiring dexterous and versatile deformations.
PaperID: 41,
Authors: Pujie Xin, Zhanteng Xie, Philip M. Dames
Affiliations: School of Communication Engineering, Hangzhou Dianzi University, Hangzhou, China; Department of Mechanical Engineering, Temple University, Philadelphia, PA, USA
Abstract: The increased deployment of multirobot systems (MRS) in various fields has led to the need to analyze system-level performance. However, creating consistent metrics for MRS is challenging due to the wide range of team and task parameters, such as the number of robots and the size of the environment. This article presents a new analytical framework for MRS based on dimensionless variable analysis that effectively condenses the complex relationships between the team and task parameters that influence MRS performance into a manageable set of dimensionless variables. Then, we use these dimensionless variables to fit a predictive parameteric model of team performance. We apply our methodology to two MRS applications: multirobot multitarget tracking and multiagent path finding. The application of dimensionless variable analysis to MRS offers a promising method for MRS analysis that effectively reduces complexity, improves understanding of system behavior, and can inform the design and management of future MRS deployments.
PaperID: 42,
Authors: Ryo Kikuuwe
Affiliations: Machinery Dynamics Laboratory, Hiroshima University, Hiroshima, Japan
Abstract: This article presents a task-space admittance controller applicable to redundant manipulators equipped with torque sensors. It extends Kikuuwe’s (2019) torque-bounded admittance controller, which allows for imposing explicit limits on the joint actuator torques without causing unsafe behaviors, such as oscillation and overshoots. The proposed controller enforces that the end-effector follows predefined task-space dynamics as long as the joint torques are unsaturated and the configuration is away from singularities. The behavior in the nullspace, which arises from the redundant degrees of freedom and singular configurations, is governed by predefined joint-space dynamics. The task-space and joint-space dynamics are combined through a newly proposed continualized pseudoinverse, which employs the singular value decomposition. Results of experiments using a seven-degree-of-freedom Kinova Gen3 robot illustrate the validity of the proposed admittance controller in various scenarios, including the case where the robot is fully stretched.
PaperID: 43,
Authors: Weiliang Deng, Hongming Chen, Biyu Ye, Haoran Chen, Ziliang Li, Ximin Lyu
Affiliations: School of Intelligent Systems Engineering, Sun Yat-sen University, Guangzhou, China
Abstract: Expressive motion planning for aerial manipulators (AMs) is essential for tackling complex manipulation tasks, yet achieving coupled trajectory planning adaptive to various tasks remains challenging, especially for those requiring aggressive maneuvers. In this work, we propose a novel whole-body integrated motion planning framework for quadrotor-based AMs that leverages flexible waypoint constraints to achieve versatile manipulation capabilities. These waypoint constraints enable the specification of individual position requirements for either the quadrotor or end-effector, while also accommodating higher order velocity and orientation constraints for complex manipulation tasks. To implement our framework, we exploit spatio-temporal trajectory characteristics and formulate an optimization problem to generate feasible trajectories for both the quadrotor and manipulator while ensuring collision avoidance considering varying robot configurations, dynamic feasibility, and kinematic feasibility. Furthermore, to enhance the maneuverability for specific tasks, we employ imitation learning to facilitate the optimization process to avoid poor local optima. The effectiveness of our framework is validated through comprehensive simulations and real-world experiments, where we successfully demonstrate nine fundamental manipulation skills across various environments.
PaperID: 44,
Authors: Hyunsoo Sun, Sungwoo Park, Donghyun Hwang
Affiliations: Center for Robotics Research, KIST, Seoul, South Korea
Abstract: We have developed a two-degree-of-freedom robotic wrist with variable stiffness capability, designed for situations where collisions between the end-effector and the environment are inevitable. To enhance environmental adaptability and prevent physical damage, the wrist can operate in a low-stiffness mode. However, the flexibility of this mode might negatively impact stable and precise manipulation. To address this, we proposed a robotic wrist that switches between a passive low-stiffness mode for environmental adaptation and an active high-stiffness mode for precise manipulation. Initially, we developed a functional prototype that could manually switch between these modes, demonstrating the wrist's passive low-stiffness and active high-stiffness states. This prototype was designed as a lightweight, flat-type modular device, incorporating a sheet-type flexure as the motion guide and embedding all essential components, including actuators, sensors, and a control unit, into the wrist module. Based on the functional prototype, we developed an improved version to enhance durability and functionality. The resulting wrist module incorporates a three-axis force/torque sensor and an impedance control system to control the stiffness. It measures 55 mm in height, weighs 200 g, and offers a 232.4-fold active stiffness variation.
PaperID: 45,
Authors: Daniel J. Lynch, Jason L. Pusey, Sean W. Gart, Paul B. Umbanhowar, Kevin M. Lynch
Affiliations: Center for Robotics and Biosystems and the Department of Mechanical Engineering, Northwestern University, Evanston, IL, USA; DEVCOM Army Research Lab, Adelphi, MD, USA
Abstract: Legged robot locomotion is hindered by a mismatch between applications where legs can outperform wheels or treads, most of which feature deformable substrates, and existing tools for planning and control, most of which assume flat, rigid substrates. In this study, we focus on the ramifications of plastic terrain deformation on the hop-to-hop energy dynamics of a spring-legged monopedal hopping robot animated by a switched-compliance energy injection controller. From this deliberately simple robot-terrain template, we derive a hop-to-hop energy return map, and we use physical experiments and simulations to validate the hop-to-hop energy map for a real robot hopping on a real deformable substrate. The dynamical properties (fixed points, eigenvalues, basins of attraction) of this map provide insights into efficient, responsive, and robust locomotion on deformable terrain. Specifically, we identify constant-fixed-point surfaces in a controller parameter space that suggest it is possible to tune control parameters for efficiency or responsiveness while targeting a desired gait energy level. We also identify conditions under which fixed points of the energy map are globally stable, and we further characterize the basins of attraction of fixed points when these conditions are not satisfied. We conclude by discussing the implications of this hop-to-hop energy map for planning, control, and estimation for efficient, agile, and robust legged locomotion on deformable terrain.
PaperID: 46,
Authors: Sven Lilge, Timothy D. Barfoot, Jessica Burgner-Kahrs
Affiliations: Robotics Institute, University of Toronto, Toronto, ON, Canada
Abstract: In contrast to conventional robots, accurately modeling the kinematics and statics of continuum robots is challenging due to partially unknown material properties, parasitic effects, or unknown forces acting on the continuous body. Consequentially, state estimation approaches that utilize additional sensor information to predict the shape of continuum robots have garnered significant interest. This article presents a novel approach to state estimation for systems with multiple coupled continuum robots, which allows estimating the shape and strain variables of multiple continuum robots in an arbitrary coupled topology. Simulations and experiments demonstrate the capabilities and versatility of the proposed method, while achieving accurate and continuous estimates for the state of such systems, resulting in average end-effector errors of 3.3 mm and 5.02^\circ depending on the sensor setup. It is further shown, that the approach offers fast computation times of below 10 ms, enabling its utilization in quasi-static real-time scenarios with average update rates of 100–200 Hz. An open-source C++ implementation of the proposed state estimation method is made publicly available to the community.
PaperID: 47,
Authors: Arash Asgharivaskasi, Fritz Girke, Nikolay Atanasov
Affiliations: Department of Electrical and Computer Engineering, University of California San Diego, La Jolla, CA, USA; School of Computation, Information and Technology, Technical University of Munich, Munich, Germany
Abstract: Autonomous exploration of unknown environments using a team of mobile robots demands distributed perception and planning strategies to enable efficient and scalable performance. Ideally, each robot should update its map and plan its motion not only relying on its own observations, but also considering the observations of its peers. Centralized solutions to multirobot coordination are susceptible to central node failure and require a sophisticated communication infrastructure for reliable operation. Current decentralized active mapping methods consider simplistic robot models with linear-Gaussian observations and Euclidean robot states. In this work, we present a distributed multirobot mapping and planning method, called Riemannian optimization for active mapping (ROAM). We formulate an optimization problem over a graph with node variables belonging to a Riemannian manifold and a consensus constraint requiring feasible solutions to agree on the node variables. We develop a distributed Riemannian optimization algorithm that relies only on one-hop communication to solve the problem with consensus and optimality guarantees. We show that multirobot active mapping can be achieved via two applications of our distributed Riemannian optimization over different manifolds: distributed estimation of a 3-D semantic map and distributed planning of \textSE(3) trajectories that minimize map uncertainty. We demonstrate the performance of ROAM in simulation and real-world experiments using a team of robots with RGB-D cameras.
PaperID: 48,
Authors: Jing Xu, Weihang Chen, Hongyu Qian, Dan Wu, Rui Chen
Affiliations: Department of Mechanical Engineering, Tsinghua University, Beijing, China
Abstract: Vision-based tactile sensors have drawn increasing interest in the robotics community. However, traditional lens-based designs impose minimum thickness constraints on these sensors, limiting their applicability in space-restricted settings. In this article, we propose ThinTact, a novel lensless vision-based tactile sensor with a sensing field of over 200 mm^2 and a thickness of less than 10 mm. ThinTact utilizes the mask-based lensless imaging technique to map the contact information to CMOS signals. To ensure real-time tactile sensing, we propose a real-time lensless reconstruction algorithm that leverages a frequency-spatial-domain joint filter based on discrete cosine transform. This algorithm achieves computation significantly faster than existing optimization-based methods. In addition, to improve the sensing quality, we develop a mask optimization method based on the generic algorithm and the corresponding system matrix calibration algorithm. We evaluate the performance of our proposed lensless reconstruction and tactile sensing through qualitative and quantitative experiments. Furthermore, we demonstrate ThinTact's practical applicability in diverse applications, including texture recognition and contact-rich object manipulation.
PaperID: 49,
Authors: Gang Chen, Zhaoying Wang, Wei Dong, Javier Alonso-Mora
Affiliations: Autonomous Multi-Robots Lab, Department of Cognitive Robotics, School of Mechanical Engineering, Delft University of Technology, Delft, The Netherlands; State Key Laboratory of Mechanical System and Vibration, School of Mechanical Engineering, Shanghai Jiaotong University, Shanghai, China
Abstract: Representing the 3-D environment with instance-aware semantic and geometric information is crucial for interaction-aware robots in dynamic environments. Nevertheless, creating such a representation poses challenges due to sensor noise, instance segmentation and tracking errors, and the objects' dynamic motion. This article introduces a novel particle-based instance-aware semantic occupancy map to tackle these challenges. Particles with an augmented instance state are used to estimate the probability hypothesis density (PHD) of the objects and implicitly model the environment. Utilizing a state-augmented sequential Monte Carlo PHD filter, these particles are updated to jointly estimate occupancy status, semantic, and instance IDs, mitigating noise. In addition, a memory module is adopted to enhance the map's responsiveness to previously observed objects. Experimental results on the Virtual KITTI 2 dataset demonstrate that the proposed approach surpasses state-of-the-art methods across multiple metrics under different noise conditions. Subsequent tests using real-world data further validate the effectiveness of the proposed approach.
PaperID: 50,
Authors: Jia Liu, Guoyao Ma, Shixiong Fu, Chenyang Huang, Xin-Yu Wu, Tiantian Xu
Affiliations: Guangdong Provincial Key Laboratory of Robotics and Intelligent System, Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences, Shenzhen, China
Abstract: Magnetic microrobots have garnered significant attention and hold great potential for biomedical research applications. However, achieving precise manipulation in vivo poses significant challenges, particularly in medical image-based real-time feedback control, because it is difficult for a visual camera to track the motion of magnetic microrobots inside the body in biomedical applications. To realize the precise control of magnetic microrobots, it is also necessary to design and implement a simple and powerful control method. This approach allows for avoiding resource-intensive and complex control strategies. In this article, we present a learning-based real-time control method utilizing ultrasound images. Inspired by the ADboost concept, we use a reinforcement learning approach to integrate two simple control methods: a proportional-integral-derivative controller and a guiding vector field controller. We develop a novel Q-learning method called average Q-learning that incorporates average operation and n-step bootstraps. Its primary objective is to dynamically adjust the outputs of the different simple controllers. While each controller individually offers a straightforward solution, their integration contributes to a powerful control approach. To demonstrate its scalability, a nonsmooth path is utilized to investigate the integration performance of three simple controllers. In addition, we enhance a classic segmentation module, U-net, by incorporating an atrous spatial pyramid pooling module. To validate the effectiveness of the proposed control method, we conduct simulations and experiments using various planar paths. The quantitative analysis of the results demonstrates the efficacy of our approach in achieving precise manipulation, leveraging real-time control based on medical images for magnetic microrobots. Overall, this study provides a preliminary investigation into the field of medical image-based precise manipulation of magnetic microrobots in vivo applications.
PaperID: 51,
Authors: Jun Chen, Mohammed Abugurain, Philip M. Dames, Shinkyu Park
Affiliations: School of Electrical and Automation Engineering, Nanjing Normal University, Nanjing, China; Computer, Electrical, and Mathematical Science and Engineering Division, King Abdullah University of Science and Technology, Thuwal, Saudi Arabia; Department of Mechanical Engineering, Temple University, Philadelphia, PA, USA
Abstract: Utilizing heterogeneous mobile sensors to actively gather information improves adaptability and reliability in extended environments. This article presents a cooperative multirobot multitarget search and tracking framework aimed at enhancing the efficiency of the heterogeneous sensor network, and consequently, improving the overall target tracking accuracy. The concept of normalized unused sensing capacity is introduced to quantify the information a sensor is currently gathering relative to its theoretical maximum. This measurement can be computed using entirely local information and is applicable to various sensor models, distinguishing it from previous literature on the subject. It is then utilized to develop a heuristics distributed coverage control strategy for a heterogeneous sensor network, adaptively balancing the workload based on each sensor's current unused capacity. The algorithm is validated through a series of robot operating system (ROS) and MATLAB simulations, demonstrating superior results compared to standard approaches that do not account for heterogeneity or current usage rates.
PaperID: 52,
Authors: Jingwen Zhao, Leone Costi, Luca Scimeca, Fumiya Iida
Affiliations: Department of Engineering, University of Cambridge, Cambridge, U.K.; MILA AI Institute, Montréal, QC, Canada
Abstract: Teleoperated medical robots have the potential to revolutionize healthcare. However, when developing systems for tasks like remote palpation, state-of-the-art literature still uses test phantoms of oversimplified geometries, due to the complexity of the required mechanical robot–patient interaction. In reality, human bodies have complex 3-D shapes and require fine-tuning of all six manipulator's degrees of freedom, controlled by the user. In this article, we argue that the implementation of depth-vision-driven autonomous dimensionality-reduction (DVD ADR) shared control can greatly improve the users' performance. The proposed control method keeps the user in control of the end-effector’s position, while automatically adjusting its orientation in order to maintain the tactile sensor normal to the phantom's surface. A depth camera and a computer vision algorithm are used to infer the phantom's shape and achieve DVD ADR shared control. Experimental results showcase how this leads to statistically significant performance improvement. Not only were the participants able to achieve more precise palpations, with up to 29.5% and 22.4% more accuracy in position and orientation, respectively, but the DVD ADR shared control allowed them to achieve a 8.8% better detection accuracy while needing 13.8% less time. The abovementioned results are all tested for statistical significance and achieved a p-value lower than 0.05.
PaperID: 53,
Authors: Junkai Niu, Sheng Zhong, Xiuyuan Lu, Shaojie Shen, Guillermo Gallego, Yi Zhou
Affiliations: Neuromorphic Automation and Intelligence Lab (NAIL) at School of Robotics, Hunan University, Changsha, China; Department of Electronic and Computer Engineering, Hong Kong University of Science and Technology, Hong Kong, China; TU Berlin, the Science of Intelligence Excellence Cluster, the Robotics Institute Germany and the Einstein Center Digital Future, Berlin, Germany
Abstract: Event-based visual odometry is a specific branch of visual simultaneous localization and mapping (SLAM) techniques, which aims at solving tracking and mapping subproblems (typically in parallel), by exploiting the special working principles of neuromorphic (i.e., event-based) cameras. Due to the motion-dependent nature of event data, explicit data association (i.e., feature matching) under large-baseline viewpoint changes is difficult to establish, making direct methods a more rational choice. However, state-of-the-art direct methods are limited by the high computational complexity of the mapping subproblem and the degeneracy of camera pose tracking in certain degrees of freedom (DoF) in rotation. In this article, we tackle these issues by building an event-based stereo visual-inertial odometry system, which is built upon a direct pipeline known as event-based stereo visual odometry (ESVO). Specifically, to speed up the mapping operation, we propose an efficient strategy for sampling contour points according to the local dynamics of events. The mapping performance is also improved in terms of structure completeness and local smoothness by merging the temporal stereo and static stereo results. To circumvent the degeneracy of camera pose tracking in recovering the pitch and yaw components of general 6-DoF motion, we introduce IMU measurements as motion priors via preintegration. To this end, a compact back-end is proposed for continuously updating the IMU bias and predicting the linear velocity, enabling an accurate motion prediction for camera pose tracking. The resulting system scales well with modern high-resolution event cameras and leads to better global positioning accuracy in large-scale outdoor environments. Extensive evaluations on five publicly available datasets featuring different resolutions and scenarios justify the superior performance of the proposed system against five state-of-the-art methods. Compared to ESVO, our new pipeline significantly reduces the camera pose tracking error by 40%–80% and 20%–80% in terms of absolute trajectory error and relative pose error, respectively; at the same time, the mapping efficiency is improved by a factor of five. We release our pipeline as an open-source software for future research in this field.
PaperID: 54,
Authors: Jeonghan Yu, Seok Won Kang, Yoon Young Kim
Affiliations: Department of Mechanical Engineering, Seoul National University, Seoul, South Korea; Department of Mechanical Systems, Sookmyung Women's University, Seoul, South Korea
Abstract: Self-aligning mechanisms are essential components in facilitating adaptability in wearable robots, but their synthesis from scratch is very challenging. To overcome this hurdle, we propose a so-far-unprecedented autonomous method to synthesize self-aligning knee joint mechanisms, requiring neither a baseline design nor human intervention during synthesis. Our method transforms the synthesis problem into an optimization problem amenable to an efficient gradient-based algorithm using a discretized ground mechanism model. The main challenge in the conversion lies in how to define the objective and constraint functions in order to ensure the fundamental self-aligning capability and also to impose a desired force transmittance profile. Several design cases were considered to show the effectiveness of the newly proposed functions for the optimization-based synthesis formulation, notably in addressing degree-of-freedom requirements. Although this study focuses primarily on knee joint mechanisms assisting gait motion and aligning with the flexion axis, the developed method can be applied to other self-aligning robot mechanisms.
PaperID: 55,
Authors: Janine Hoelscher, Inbar Fried, Spiros Tsalikis, Jason A. Akulian, Robert J. Webster III, Ron Alterovitz
Affiliations: Department of Bioengineering, Clemson University, Clemson, SC, USA; Department of Computer Science, University of North Carolina at Chapel Hill, Chapel Hill, NC, USA; Division of Pulmonary Diseases and Critical Care Medicine, University of North Carolina at Chapel Hill, NC, USA; Department of Mechanical Engineering, Vanderbilt University, Nashville, TN, USA
Abstract: Steerable needles are minimally invasive devices that can enable novel medical procedures by following curved paths to avoid critical anatomical obstacles. We introduce a new start pose robustness metric for steerable needle motion plans. A steerable needle deployment typically consists of a physician manually placing a steerable needle at a precomputed start pose on the surface of tissue and handing off control to a robot, which then autonomously steers the needle through the tissue to the target. The handoff between humans and robots is critical for procedure success, as even small deviations from a planned start pose change the steerable needle's reachable workspace. Our metric is based on a novel geometric method to efficiently compute how far the physician can deviate from the planned start pose in both position and orientation such that the steerable needle can still reach the target. We evaluate our metric through simulation in liver and lung scenarios. Our evaluation shows that our metric can be applied to plans computed by different steerable needle motion planners and that it can be used to efficiently select plans with large safe start regions.
PaperID: 56,
Authors: Diego Martinez-Baselga, Eduardo Sebastián, Eduardo Montijano, Luis Riazuelo, Carlos Sagüés, Luis Montano
Affiliations: Instituto de Investigación en Ingeniería de Aragón (IA), Universidad de Zaragoza, Zaragoza, Spain
Abstract: We present AdaptiVe Optimal Collision Avoidance Driven by Opinion (AVOCADO), a novel navigation approach to address holonomic robot collision avoidance when the robot does not know how cooperative the other agents in the environment are. AVOCADO departs from a velocity obstacle's (VO) formulation akin to the optimal reciprocal collision avoidance method. However, instead of assuming reciprocity, it poses an adaptive control problem to adapt to the cooperation level of other robots and agents in real time. This is achieved through a novel nonlinear opinion dynamics design that relies solely on sensor observations. As a by-product, we leverage tools from the opinion dynamics formulation to naturally avoid the deadlocks in geometrically symmetric scenarios that typically suffer VO-based planners. Extensive numerical simulations show that AVOCADO surpasses existing motion planners in mixed cooperative/noncooperative navigation environments in terms of success rate, time to goal and computational time. In addition, we conduct multiple real experiments that verify that AVOCADO is able to avoid collisions in environments crowded with other robots and humans.
PaperID: 57,
Authors: Fouad Sukkar, Jennifer Wakulicz, Ki Myung Brian Lee, Weiming Zhi, Robert Fitch
Affiliations: School for Mechanical and Mechatronic Engineering, University of Technology Sydney, Ultimo, NSW, Australia; Robotics Institute, Carnegie Mellon University, Pittsburgh, PA, USA
Abstract: Robotic manipulator applications often require efficient online motion planning. When completing multiple tasks, sequence order and choice of goal configuration can have a drastic impact on planning performance. This is well known as the robot task sequencing problem (RTSP). Existing general-purpose RTSP algorithms are susceptible to producing poor-quality solutions or failing entirely when available computation time is restricted. We propose a new multiquery task sequencing method designed to operate in semistructured environments with a combination of static and nonstatic obstacles. Our method intentionally trades off workspace generality for planning efficiency. Given a user-defined task space with static obstacles, we compute a subspace decomposition. The key idea is to establish approximate isometries known as \epsilon-Gromov-Hausdorff approximations that identify points that are close to one another in both task and configuration space. Importantly, we prove bounded suboptimality guarantees on the lengths of paths within these subspaces. These bounding relations further imply that paths within the same subspace can be smoothly concatenated, which we show is useful for determining efficient task sequences. We evaluate our method with several kinematic configurations in a complex simulated environment, achieving up to 3× faster motion planning and 5× lower maximum trajectory jerk compared to baselines.
PaperID: 58,
Authors: Zean Yuan, Jiaxing Li, Lifu Liu, Xinyu Zhu, Wenbiao Wang, Michael D. Dickey, Guo Zhan Lum, Pakpong Chirarattananon, Jun Luo, Rui Chen
Affiliations: State Key Laboratory of Mechanical Transmission for Advanced Equipment, Chongqing University, Chongqing, China; Department of Biomedical Engineering, City University of Hong Kong, Hong Kong, China; College of Mechanical Engineering, Zhejiang University of Technology, Hangzhou, China; Department of Chemical and Biomolecular Engineering, North Carolina State University, Raleigh, NC, USA; School of Mechanical and Aerospace Engineering, Nanyang Technological University, Singapore
Abstract: Rigid robots can achieve precise motions but expose shortcomings in system complexity, fabrication cost, and human–robot interaction, which motivates researchers to develop various soft robots to fill these gaps. Electro-hydraulic actuators (EHAs) have received widespread attention and been used in many soft robots due to impressive high-strain, fast-speed, and rapid-response characteristics. However, existing EHAs face challenges in achieving large-deformation, high-robustness, and low-weight simultaneously. This limits the application of EHAs in robotic systems that are weight-sensitive or require fail-safe and fault-tolerant behavior. Here, we present a lightweight (0.98 g) electro-pneumatic actuator (EPA) filled with air and only 0.1-mL liquid dielectric, which achieves high-speed bending from 11° to 93.5° in 60 ms, large-angle bending from 11° to 104° in 2 s (the largest in current EHAs), and high-frequency swing at 20 Hz. The EPA is ultrarobust and can operate properly after being punctured by four needles or crushed twice by a 1500-kg vehicle. Furthermore, to validate the above features of EPAs, three applications are demonstrated at a voltage of 6 kV, including four-finger grippers, fast-crawling robots, and water-walking robots. This work pushes the boundaries of robustness and lightweight for EHAs, providing a foundation for the application of electro-pneumatic actuation in soft robotics.
PaperID: 59,
Authors: Xinghua Liu, Ming Cao
Affiliations: Engineering and Technology Institute, University of Groningen, Groningen, The Netherlands
Abstract: In this work, we propose a high-order regularization method to solve the ill-conditioned problems in robot localization. Numerical solutions to robot localization problems are often unstable when the problems are ill-conditioned. A typical way to solve ill-conditioned problems is regularization, and a classical regularization method is the Tikhonov regularization. It is shown that the Tikhonov regularization is a low-order case of our method. We find that the proposed method is superior to the Tikhonov regularization in approximating some ill-conditioned inverse problems, such as some basic robot localization problems. The proposed method overcomes the oversmoothing problem in the Tikhonov regularization as it uses more than one term in the approximation of the matrix inverse, and an explanation for the oversmoothing of the Tikhonov regularization is given. Moreover, one a priori criterion, which improves the numerical stability of the ill-conditioned problem, is proposed to obtain an optimal regularization matrix. As most of the regularization solutions are biased, we also provide two bias-correction techniques for the proposed high-order regularization. The simulation and experimental results using an ultra-wideband sensor network in a 3-D environment are discussed, demonstrating the performance of the proposed method.
PaperID: 60,
Authors: Philippe Nadeau, Jonathan Kelly
Affiliations: STARS Laboratory, Institute for Aerospace Studies, University of Toronto, Toronto, ON, Canada
Abstract: We introduce a planner designed to guide robot manipulators in stably placing objects within complex scenes. Our proposed method reverses the traditional approach to object placement: our planner selects contact points first and then determines a placement pose that solicits the selected points. This is instead of sampling poses, identifying contact points, and evaluating pose quality. Our algorithm facilitates stability-aware object placement planning, imposing no restrictions on object shape, convexity, or mass density homogeneity, while avoiding combinatorial computational complexity. Our proposed stability heuristic enables our planner to find a solution about 20 times faster when compared to the same algorithm not making use of the heuristic and eight times faster than a state-of-the-art method that uses the traditional sample-and-evaluate approach. The proposed planner is also more successful in finding stable placements than the five other benchmarked algorithms. Derived from first principles and validated in ten real robot experiments, our approach provides a general and scalable solution to the problem of rigid object placement planning.
PaperID: 61,
Authors: Wenhui Wei, Yangfan Zhou, Yimin Hu, Zhi Li, Sen Wang, Xin Liu, Jiadong Li
Affiliations: School of Nano-Tech and Nano-Bionics, University of Science and Technology of China, Hefei, China; Suzhou Institute of Nano-Tech and Nano-Bionics, Chinese Academy of Sciences, Suzhou, China
Abstract: Visual–inertial odometry (VIO) provides a robust localization solution for simultaneous localization and mapping systems. Self-supervised VIO, a leading approach, has the advantage of not requiring extensive ground-truth labels. Regrettably, this method still poses challenges for robotic applications, particularly uncrewed aerial vehicles, due to its computational complexity arising from inadequate model designs. To address this bottleneck, we introduce BotVIO (where “Bot” refers to “robotics”), a transformer-based self-supervised VIO model, offering an excellent solution to alleviate computational burdens for robotics. Our lightweight backbone combines shallow CNNs with spatial–temporal-enhanced transformers to replace conventional architectures, while the minimalist cross-fusion module uses single-layer cross-attention to enhance multimodal interaction. Extensive experiments show that, during pose estimation, BotVIO achieves a remarkable 70.37% reduction in trainable parameters and a 74.85% decrease in inference speed, reaching up to 57.80 fps on an NVIDIA Jetson NX (10W&2CORE), while improving pose accuracy and robustness. For the benefit of the community, we make public the source code.1
PaperID: 62,
Authors: Jeongmin Lee, Minji Lee, Sunkyung Park, Jinhee Yun, Dongjun Lee
Affiliations: Department of Mechanical Engineering, IAMD and IOER, Seoul National University, Seoul, Republic of Korea
Abstract: The multicontact nonlinear complementarity problem (NCP) is a naturally arising challenge in robotic simulations. Achieving high performance in terms of both accuracy and efficiency remains a significant challenge, particularly in scenarios involving intensive contacts and stiff interactions. In this article, we introduce a new class of multicontact NCP solvers based on the theory of the augmented Lagrangian (AL). We detail how the standard derivation of AL in convex optimization can be adapted to handle multicontact NCP through the iteration of surrogate problem solutions and the subsequent update of primal-dual variables. Specifically, we present two tailored variations of AL for robotic simulations: the cascaded Newton-based augmented Lagrangian (CANAL) and the subsystem-based alternating direction method of multipliers (SubADMM). We demonstrate how CANAL can manage multicontact NCP in an accurate and robust manner, while SubADMM offers superior computational speed, scalability, and parallelizability for high degrees-of-freedom multibody systems with numerous contacts. Our results showcase the effectiveness of the proposed solver framework, illustrating its advantages in various robotic manipulation scenarios.
PaperID: 63,
Authors: Cristian-Ioan Vasile, Jana Tumova, Sertac Karaman, Calin Belta, Daniela Rus
Affiliations: Lehigh University, Bethlehem, PA, USA; KTH Royal Institute of Technology,Stockholm, Stockholm, Sweden; Department of Aeronautics and Astronautics, Massachusetts Institute of Technology, Cambridge, MA, USA; University of Maryland, College Park, MD, USA; Computer Science and Artificial Intelligence Laboratory, Massachusetts Institute of Technology, Cambridge, MA, USA
Abstract: This article considers the route planning problem for a vehicle with limited capacity operating in a road network. The vehicle is assigned a set of transportation requests that are more complex than traveling between two locations, may involve dependencies between their subtasks, and include deadlines and priorities. The requests arrive gradually over the deployment time-horizon, and thus replanning is needed for new requests. We address cases when not all requests can be serviced by their deadlines despite car sharing. We introduce multiple quality measures for plans that account for requests’ delays with respect to deadlines and priorities. We formalize the problem as planning in a weighted transition system under syntactically cosafe LTL formulas. We develop an online planning and replanning algorithm based on the automata-based approach to least-violating plan synthesis and on translation to a mixed integer linear program (MILP). Furthermore, we show that the MILP reduces to graph search for a subclass of quality measures that satisfy a monotonicity property. We show the approach in simulations, including a case study on the mid-Manhattan road network over the span of 24 h.
PaperID: 64,
Authors: Eduardo Sebastián, Thai Duong, Nikolay Atanasov, Eduardo Montijano, Carlos Sagüés
Affiliations: Department of Computer Science and Systems Engineering (DIIS) and the Engineering Research Institute of Aragon (IA), Universidad de Zaragoza, Zaragoza, Spain; Department of Electrical and Computer Engineering, University of California San Diego, La Jolla, CA, USA
Abstract: The networked nature of multirobot systems presents challenges in the context of multiagent reinforcement learning. Centralized control policies do not scale with increasing numbers of robots, whereas independent control policies do not exploit the information provided by other robots, exhibiting poor performance in cooperative-competitive tasks. In this work, we propose a physics-informed reinforcement learning approach able to learn distributed multirobot control policies that are both scalable and make use of all the available information to each robot. Our approach has three key characteristics. First, it imposes a port-Hamiltonian structure on the policy representation, respecting energy conservation properties of physical robot systems and the networked nature of robot team interactions. Second, it uses self-attention to ensure a sparse policy representation able to handle time-varying information at each robot from the interaction graph. Third, we present a soft actor–critic reinforcement learning algorithm parameterized by our self-attention port-Hamiltonian control policy, which accounts for the correlation among robots during training while overcoming the need of value function factorization. Extensive simulations in different multirobot scenarios demonstrate the success of the proposed approach, surpassing previous multirobot reinforcement learning solutions in scalability, while achieving similar or superior performance (with averaged cumulative reward up to × \text2 greater than the state-of-the-art with robot teams × \text6 larger than the number of robots at training time). We also validate our approach on multiple real robots in the Georgia Tech Robotarium under imperfect communication, demonstrating zero-shot sim-to-real transfer and scalability across number of robots.
PaperID: 65,
Authors: Antonin Dallard, Mehdi Benallegue, Nicola Scianca, Fumio Kanehiro, Abderrahmane Kheddar
Affiliations: Wandercraft, Paris, France; CNRS-AIST Joint Robotics Laboratory, IRL, Tsukuba, Japan; Dipartimento di Ingegneria Informatica, Automatica e Gestionale, Sapienza University of Rome, Rome, Italy
Abstract: In this article, we propose a novel walking control scheme based on the dynamics of the linear inverted pendulum (LIP) model. Pattern generation incorporates a model of contact forces, enabling closed-loop control of the humanoid robot’s state, including the center-of-mass position, velocity, and zero moment point. No additional control policies are required to maintain static and dynamic balance. Our approach also includes dynamic replanning of step locations and timings, thus preserving the LIP’s boundedness condition. We validated this controller on five different humanoid robots, testing its robustness through various disturbances, including sudden pushes during walking and static phases. In addition, our controller demonstrated effective locomotion over uneven and compliant terrain. Both simulation and experimental results confirm the effectiveness and robustness of this controller.
PaperID: 66,
Authors: William Talbot, Julian Nubert, Turcan Tuna, Cesar Cadena, Frederike Dümbgen, Jesus Tordesillas, Timothy D. Barfoot, Marco Hutter
Affiliations: Robotic Systems Lab (RSL), ETH Zürich, Zürich, Switzerland; Computer Science Department of ENS, Willow, Inria, PSL Research University, Paris, France; Institute for Research in Technology, ICAI School of Engineering, Comillas Pontifical University, Madrid, Spain; Autonomous Space Robotics Laboratory (ASRL), University of Toronto, Toronto, ON, Canada
Abstract: Accurate, efficient, and robust state estimation is more important than ever in robotics as the variety of platforms and complexity of tasks continue to grow. Historically, discrete-time filters and smoothers have been the dominant approach, in which the estimated variables are states at discrete sample times. The paradigm of continuous-time state estimation proposes an alternative strategy by estimating variables that express the state as a continuous function of time, which can be evaluated at any query time. Not only can this benefit downstream tasks such as planning and control, but it also significantly increases estimator performance and flexibility, as well as reduces sensor preprocessing and interfacing complexity. Despite this, continuous-time methods remain underutilized, potentially because they are less well-known within robotics. To remedy this, this work presents a unifying formulation of these methods and the most exhaustive literature review to date, systematically categorizing prior work by methodology, application, state variables, historical context, and theoretical contribution to the field. By surveying splines and Gaussian process together and contextualizing works from other research domains, this work identifies and analyzes open problems in continuous-time state estimation and suggests new research directions.
PaperID: 67,
Authors: Song Li, Songnan Bai, Ruihan Jia, Yixi Cai, Runze Ding, Yu Shi, Fu Zhang, Pakpong Chirarattananon
Affiliations: Department of Biomedical Engineering, City University of Hong Kong, Hong Kong; Mechatronics and Robotic Systems Laboratory, Department of Mechanical Engineering, University of Hong Kong, Hong Kong
Abstract: Mobile robots have revolutionized various fields, offering solutions for manipulation, environmental monitoring, and exploration. However, payload capacity remains a limitation. This article presents a novel thrust-based robotic hopper capable of carrying payloads up to nine times its own weight while maintaining agile mobility over less structured terrain. The 220 g robot carries upto 2 kg while hopping—–a capability that bridges the gap between high-payload ground robots and agile aerial platforms. Key advancements that enable this high-payload capacity include the integration of bidirectional thrusters, allowing for both upward and downward thrust generation to enhance energy management while hopping. In addition, we present a refined model of dynamics that accounts for heavy payload conditions, particularly for large jumps. To address the increased computational demands, we employ a neural network compression technique, ensuring real-time onboard control. The robot’s capabilities are demonstrated through a series of experiments, including leaping over a high obstacle, executing sharp turns with large steps, as well as performing simple autonomous navigation while carrying a 730 g LiDAR payload. This showcases the robot’s potential for applications, such as mobile sensing and mapping, in challenging environments.
PaperID: 68,
Authors: Haoran Ding, Noémie Jaquier, Jan Peters, Leonel Rozo
Affiliations: Bosch Center for Artificial Intelligence, Renningen, Germany; Division of Robotics, Perception, and Learning, KTH Royal Institute of Technology, Stockholm, Sweden; Computer Science Department, Technische Universität Darmstadt, Darmstadt, Germany
Abstract: Diffusion-based visuomotor policies excel at learning complex robotic tasks by effectively combining visual data with high-dimensional, multimodal action distributions. However, diffusion models often suffer from slow inference due to costly denoising processes or require complex sequential training arising from recent distilling approaches. This article introduces Riemannian flow matching policy (RFMP), a model that inherits the easy training and fast inference capabilities of flow matching. Moreover, RFMP inherently incorporates geometric constraints commonly found in realistic robotic applications, as the robot state resides on a Riemannian manifold. To enhance the robustness of RFMP, we propose stable RFMP (SRFMP), which leverages LaSalle’s invariance principle to equip the dynamics of FM with stability to the support of a target Riemannian distribution. Rigorous evaluation on ten simulated and real-world tasks show that RFMP successfully learns and synthesizes complex sensorimotor policies on Euclidean and Riemannian spaces with efficient training and inference phases, outperforming diffusion policies and consistency policies.
PaperID: 69,
Authors: Shuli Lv, Yan Gao, Quan Quan
Affiliations: School of Automation Science and Electrical Engineering, Beihang University, Beijing, China; Tianmushan Laboratory, Hangzhou, China
Abstract: This article presents a novel model-free spatial iterative learning (IL) framework to enhance the efficiency of vector field (VF) navigation for mobile robots. By integrating the idea of iterative learning control (ILC) control with VF, this framework utilizes historical data to enhance navigation efficiency significantly, reducing traversal time and expanding the applicability of IL to rapid navigation. Importantly, it has low-time complexity with O(n) per iteration, where n denotes the waypoints number, preventing the significant computational overhead caused by the increasing waypoints in existing methods, which often exceeds O(n^2), making it well-suited for real-time planning. Moreover, the approach is inherently model-free, leaning on historical data, thus enabling agile navigation with limited reliance on intricate model details. This article presents a comprehensive theoretical analysis of the stability, time optimality, time complexity, parameter insensitivity, robustness, and usage. Extensive simulations and experiments highlight its efficiency, promising a transformative impact on mobile robot navigation through the proposed IL.
PaperID: 70,
Authors: Kejia Ren, Gaotian Wang, Andrew S. Morgan, Lydia E. Kavraki, Kaiyu Hang
Affiliations: Department of Computer Science, Rice University, Houston, TX, USA; RAI Institute, Cambridge, MA, USA
Abstract: Nonprehensile actions, such as pushing, are crucial for addressing multiobject rearrangement problems. Many traditional methods generate robot-centric actions, which differ from intuitive human strategies and are typically inefficient. To this end, we adopt an object-centric planning paradigm and propose a unified framework for addressing a range of large-scale, physics-intensive nonprehensile rearrangement problems challenged by modeling inaccuracies and real-world uncertainties. By assuming that each object can actively move without being driven by robot interactions, our planner first computes desired object motions, which are then realized through robot actions generated online via a closed-loop pushing strategy. Through extensive experiments and in comparison with state-of-the-art baselines in both simulation and on a physical robot, we show that our object-centric planning framework can generate more intuitive and task-effective robot actions with significantly improved efficiency. In addition, we propose a benchmarking protocol to standardize and facilitate future research in nonprehensile rearrangement.
PaperID: 71,
Authors: David G. Black, Septimiu E. Salcudean
Affiliations: Department of Electrical and Computer Engineering, University of British Columbia, Vancouver, Canada
Abstract: Recent work introduced the concept of human teleoperation (HT), where the remote robot typically considered in conventional bilateral teleoperation is replaced by a novice person wearing a mixed-reality head-mounted display and tracking the motion of a virtual tool controlled by an expert. HT has advantages in cost, complexity, and patient acceptance for telemedicine in low-resource communities or remote locations. However, the stability, transparency, and performance of bilateral HT are unexplored. In this article, we, therefore, develop a mathematical model of the HT system using test data. We then analyze various control architectures with this model and implement them with the HT system, testing volunteer operators and a virtual fixture-based simulated patient to find the achievable performance, investigate stability, and determine the most promising teleoperation scheme in the presence of time delays. We show that instability in HT, while not destructive or dangerous, makes the system impossible to use. However, stable and transparent teleoperation is possible with small time delays (< \text200 ms) through three-channel teleoperation, or with large time delays through model-mediated teleoperation with local pose and force feedback for the novice.
PaperID: 72,
Authors: Rixin Wang, Shuopeng Wang, Jintao Ye, Ying Zhang, Lina Hao
Affiliations: School of Mechanical Engineering and Automation, Northeastern University, Shenyang, China
Abstract: Learning-based motion planning methods have shown significant promise in enhancing the efficiency of traditional algorithms. However, they often face performance degradation in novel environments with drastic scene changes due to the limited generalization ability of deep neural networks. This article introduces a confidence-driven motion planning network (CDMPNet), comprising a feature extraction autoencoder and a confidence-driven sampling network (CDSNet). The autoencoder compresses point clouds into latent vectors. The CDSNet is a closed-form continuous-time neural network, which predicts hyperparameters of an evidential distribution over the subsequent state’s mean and covariance for robot configuration sampling. We also present a CDMPNet-based neural planner and a CDMPNet-guided RRTConnect algorithm. Simulations and ablation studies are conducted on 2-D, 3-D, and 7-D planning tasks to validate the generalization ability of our method. Furthermore, we transfer the approach to a seven-degree-of-freedom Sawyer robotic arm to demonstrate the potential for real-world deployment.
PaperID: 73,
Authors: Qingwen Zhang, Ajinkya Khoche, Yi Yang, Li Ling, Sina Sharif Mansouri, Olov Andersson, Patric Jensfelt
Affiliations: Division of Robotics, Perception, and Learning, KTH Royal Institute of Technology, Stockholm, Sweden; Autonomous Transport Solutions Lab, Scania Group, Södertälje, Sweden
Abstract: Light detection and ranging (LiDAR) point cloud is essential for autonomous vehicles, but motion distortions from dynamic objects degrade the data quality. While previous work has considered distortions caused by ego motion, distortions caused by other moving objects remain largely overlooked, leading to errors in object shape and position. This distortion is particularly pronounced in high-speed environments, such as highways and in multi-LiDAR configurations, a common setup for heavy vehicles. To address this challenge, we introduce HiMo, a pipeline that repurposes scene flow estimation for nonego motion compensation, correcting the representation of dynamic objects in point clouds. During the development of HiMo, we observed that existing self-supervised scene flow estimators often produce degenerate or inconsistent estimates under high-speed distortion. We further propose SeFlow++, a real-time scene flow estimator that achieves state-of-the-art performance on both scene flow and motion compensation. Since well-established motion distortion metrics are absent in the literature, we introduce two evaluation metrics: compensation accuracy at a point level and shape similarity of objects. We validate HiMo through extensive experiments on Argoverse 2, ZOD and a newly collected real-world dataset featuring highway driving and multi-LiDAR-equipped heavy vehicles. Our findings show that HiMo improves the geometric consistency and visual fidelity of dynamic objects in LiDAR point clouds, benefiting downstream tasks, such as semantic segmentation and 3-D detection.
PaperID: 74,
Authors: Teng Li, Hyo-Sang Shin, Antonios Tsourdos
Affiliations: Centre for AI, Robotics and Space, FEAS, Cranfield University, Cranfield, U.K.; Korea Advanced Institute of Science and Technology (KAIST), Daejeon, Republic of Korea
Abstract: This article deals with large-scale decentralized task allocation problems for multiple heterogeneous robots. One of the grand challenges with decentralized task allocation problems is the NP-hardness for computation and communication. This article proposes a decentralized decreasing threshold task allocation (DTTA) algorithm that enables parallel allocation by leveraging a decreasing threshold to handle the NP-hardness. DTTA can release both computation and communication burdens for multiple robots in a decentralized network. In addition, DTTA provides a theoretical guarantee of the quality of the solution for maximizing submodular utility functions. Theoretical analysis indicates that DTTA can provide an optimality guarantee of (1-\epsilon)/2 with computation complexity of O(\min (r^2, \fracr\epsilon \ln \fracr\epsilon )) for each robot, where \epsilon is the parameter controlling the decreasing speed of the threshold, r is the number of tasks. To examine the performance of the proposed algorithm, we conduct numerical simulations based on a multitarget surveillance scenario. Simulation results demonstrate that DTTA delivers comparable solution quality significantly faster than state-of-the-art task allocation algorithms. Its advantages are particularly pronounced in large-scale missions with thousands of tasks and robots.
PaperID: 75,
Authors: Xinpan Meng, Long Cheng, Zhengwei Li
Affiliations: School of Artificial Intelligence, University of Chinese Academy of Sciences, Beijing, China
Abstract: Tactile and proximity sensing is essential for robotic tasks involving human–robot interaction and manipulation. However, existing dual-mode sensors often face challenges such as environmental interference, large sizes, and task-specific limitations. This study proposes a dual-mode photoelectric sensor that integrates tactile and proximity sensing. The tactile sensing mechanism is based on a variable optical path structure, while the proximity sensing relies on the surface light reflection. The sensor exhibits a high sensitivity (up to 1.12 V/N), compactness (4 mm thickness), and desirable stability with a drift of less than 1% over 8000 repetitive cycles under pressures ranging from 0 to 63 kPa. A general tactile-proximity servoing framework is also proposed for the dual-mode sensor array which enables tactile servoing, proximity servoing, and hybrid tactile-proximity servoing. Under this framework, parameters can be flexibly adjusted to adapt to different servoing tasks including position and orientation control of the robotic arm's end-effector. In more complex robotic tasks, a real-time fruit ripeness classification method is developed based on the proposed sensor. Using the proposed TPNet, the classification method can achieve an accuracy of 94.4% in a four-level tomato ripeness classification task during grasping.
PaperID: 76,
Authors: Julia Di, Zdravko Dugonjic, Will Fu, Tingfan Wu, Romeo Mercado, Kevin Sawyer, Victoria Rose Most, Gregg Kammerer, Stefanie Speidel, Richard E. Fan, Geoffrey A. Sonn, Mark R. Cutkosky, Mike Lambeta, Roberto Calandra
Affiliations: Stanford University, Stanford, CA, USA; LASR Lab, Technische Universität Dresden, Dresden, Germany; Meta, Menlo Park, CA, USA; School of Embedded and Composite AI (SECAI), Germany
Abstract: Vision-based tactile sensors have recently become popular due to their combination of low cost, very high spatial resolution, and ease of integration using widely available miniature cameras. The associated field of view and focal length, however, are difficult to package in a human-sized finger. In this article we employ optical fiber bundles to achieve a form factor that, at 15 mm diameter, is smaller than an average human fingertip. The electronics and camera are also located remotely, further reducing package size. The sensor achieves a spatial resolution of 0.22 mm and a minimum force resolution 5 mN for normal and shear contact forces. With these attributes, the DIGIT Pinki sensor is suitable for applications such as robotic and teleoperated digital palpation. We demonstrate its utility for palpation of the prostate gland and show that it can achieve clinically relevant discrimination of prostate stiffness for phantom and ex vivo tissue.
PaperID: 77,
Authors: Somayeh Hussaini, Michael Milford, Tobias Fischer
Affiliations: QUT Centre for Robotics, School of Electrical Engineering and Robotics, Queensland University of Technology, Brisbane, QLD, Australia
Abstract: In robotics, spiking neural networks (SNNs) are increasingly recognized for their largely unrealized potential energy efficiency and low latency particularly when implemented on neuromorphic hardware. This article highlights three advancements for SNNs in visual place recognition (VPR). First, we propose modular SNNs (Modular SNN), where each SNN represents a set of nonoverlapping geographically distinct places, enabling scalable networks for large environments. Second, we present ensembles of Modular SNNs, where multiple networks represent the same place, significantly enhancing accuracy compared to single-network models. Each of our Modular SNN modules is compact, comprising only 1500 neurons and 474k synapses, making them ideally suited for ensembling due to their small size. Finally, we investigate the role of sequence matching in SNN-based VPR, a technique where consecutive images are used to refine place recognition. We demonstrate competitive performance of our method on a range of datasets, including higher responsiveness to ensembling compared to conventional VPR techniques and higher R@1 improvements with sequence matching than VPR techniques with comparable baseline performance. Our contributions highlight the viability of SNNs for VPR, offering scalable and robust solutions, and paving the way for their application in various energy-sensitive robotic tasks.
PaperID: 78,
Authors: Yuan Yang, Aiguo Song, Lifeng Zhu, Baoguo Xu, Guangming Song, Yang Shi
Affiliations: State Key Laboratory of Digital Medical Engineering, Jiangsu Key Laboratory of Robot Sensing and Control, School of Instrument Science and Engineering, Southeast University, Nanjing, China; Department of Mechanical Engineering, University of Victoria, Victoria, BC, Canada
Abstract: This article proposes a distributed passivity-based bilateral teleoperation control for optimizing the velocity/force manipulability of the coordinated remote redundant manipulators during the task execution. Following the leader–follower paradigm, the control connects a local haptic device with a leader remote manipulator and coordinates all the leader and follower remote manipulators. The approach is novel in reconciling the potential conflicts between the pose synchronization task and the manipulability optimization task for the remote manipulators by two-layer auxiliary systems. The first layer decouples the pose synchronization constraints into separable position and orientation constraints, and the second layer optimizes the manipulability under the position and orientation constraints. The approach is robust by designing smooth controls for the manipulators without knowing their dynamic parameters. Finally, the control renders the bilateral teleoperator output strictly passive for stable physical interactions with the human user and the environment. Comparative experiments verify the effectiveness of the proposed control in the presence of time-varying communication delays.
PaperID: 79,
Authors: Fangcheng Zhu, Yunfan Ren, Longji Yin, Fanze Kong, Qingbo Liu, Ruize Xue, Wenyi Liu, Yixi Cai, Guozheng Lu, Haotian Li, Fu Zhang
Affiliations: Mechatronics and Robotic Systems Laboratory, Department of Mechanical Engineering, The University of Hong Kong, Hong Kong
Abstract: Aerial swarm systems possess immense potential in various aspects, such as cooperative exploration, target tracking, and search and rescue. Efficient accurate self- and mutual state estimation are the critical preconditions for completing these swarm tasks, which remain challenging research topics. This article proposes Swarm-LIO2, a fully decentralized, plug-and-play, computationally efficient, and bandwidth-efficient light detection and ranging (LiDAR)-inertial odometry for aerial swarm systems. Swarm-LIO2 uses a decentralized plug-and-play network as the communication infrastructure. Only bandwidth-efficient and low-dimensional information is exchanged, including identity, ego state, mutual observation measurements, and global extrinsic transformations. To support the plug and play of new teammate participants, Swarm-LIO2 detects potential teammate autonomous aerial vehicles (AAVs) and initializes the temporal offset and global extrinsic transformation all automatically. To enhance the initialization efficiency, novel reflectivity-based AAV detection, trajectory matching, and factor graph optimization methods are proposed. For state estimation, Swarm-LIO2 fuses LiDAR, inertial measurement units, and mutual observation measurements within an efficient error state iterated Kalman filter (ESIKF) framework, with careful compensation of temporal delay and modeling of measurements to enhance the accuracy and consistency. Moreover, the proposed ESIKF framework leverages the global extrinsic for ego state estimation in the case of LiDAR degeneration or refines the global extrinsic along with the ego state estimation otherwise. To enhance the scalability, Swarm-LIO2 introduces a novel marginalization method in the ESIKF, which prevents the growth of computational time with swarm size. Extensive simulation and real-world experiments demonstrate the broad adaptability to large-scale aerial swarm systems and complicated scenarios, including GPS-denied scenes and degenerated scenes for cameras or LiDARs. The experimental results showcase the centimeter-level localization accuracy, which outperforms other state-of-the-art LiDAR-inertial odometry for a single-AAV system. Furthermore, diverse applications demonstrate the potential of Swarm-LIO2 to serve as a reliable infrastructure for various aerial swarm missions.
PaperID: 80,
Authors: Jonathan Vorndamme, Alessandro Melone, Robin Jeanne Kirschner, Luis F. C. Figueredo, Sami Haddadin
Affiliations: Munich Institute of Robotics and Machine Intelligence (MIRMI), Technische Universität München (TUM), Munich, Germany
Abstract: Recent advances in control and planning allow for seamless physical human–robot interaction (pHRI). At the same time, novel challenges appear in orchestrating intelligent decision-making and ensuring safe control of robots. Particularly in scenarios involving unforeseen or unintended collisions, robots face the imperative of reacting judiciously to avert potential risks to humans, other robots, obstacles, or themselves. At the same time, they need to maintain focus on their primary task or be able to safely resume it. Collision detection and identification algorithms are now well established in industry, yet complex collision reflexes have not transitioned into industrial applications beyond basic stopping reactions. Despite the introduction of numerous advanced high-performance reflex controllers over the past decades, their real-world adoption has remained a challenge. This work establishes a systematic framework to address that gap. For this, the reflex control problem is defined, reflex behaviors are systematically classified and categorized, and relevant safety data is acquired following existing international standards. We argue that this foundational step is crucial for improving the safety and capabilities of robots in both complex industrial and domestic environments. We validate our approach within the system class of articulated manipulators through a state-of-the-art cooperative pick-and-place task, providing a blueprint for future implementations for other robot classes.
PaperID: 81,
Authors: Erfan Shahriari, Petr Svarný, Seyed Ali Baradaran Birjandi, Matej Hoffmann, Sami Haddadin
Affiliations: Chair of Robotics and Systems Intelligence, Munich Institute of Robotics and Machine Intelligence, Technical University of Munich, Munich, Germany; Department of Cybernetics, Faculty of Electrical Engineering, Czech Technical University in Prague, Prague, Czech Republic
Abstract: Robots have surpassed humans in terms of strength and precision, yet humans retain an unparalleled ability for decision-making in the face of unpredictable disturbances. This article aims to combine the strengths of both entities within a singular task: human motion guidance under strict geometric constraints, particularly adhering to predetermined paths. To tackle this challenge, a modular haptic guidance law is proposed that takes the human-applied wrench as an input. Using an auxiliary variable called phase, the generated desired motion is guaranteed to consistently adhere to the constraint path. It is demonstrated how the guidance policy can be generalized into physically interpretable terms, adjustable either prior to initiating the task or dynamically while the task is in progress. Additionally, an illustrative guidance adaptation policy is showcased that takes into account the human's manipulability. Leveraging passivity analysis, potential sources of instability are pinpointed, and subsequently, overall system stability is ensured by incorporating an augmented virtual energy tank. Lastly, a comprehensive set of experiments, including a 20-participant user study, explores various aspects of the approach in practice, encompassing both technical and usability considerations.
PaperID: 82,
Authors: Shaohong Zhong, Alessandro Albini, Perla Maiolino, Ingmar Posner
Affiliations: Oxford Robotics Institute, Department of Engineering Science, University of Oxford, Oxford, U.K.
Abstract: Recent advances in machine learning have driven a step-change in robot perception with modalities such as vision, where large amounts of training data are readily available or cheap to collect. However, in tactile perception, the relatively high cost of data collection still largely impedes the adoption of such data-driven learning solutions. In this article, we introduce TactGen, a novel, cross-modal framework to tackle this challenge. In particular, using a two-step data generation pipeline, we leverage easily accessible vision data to synthesise artificial tactile data for downstream classifier training. Specifically, we use readily collected video data of objects of interest to efficiently learn neural radiance field (NeRF) representations. The NeRF models are then used to render red–green–blue-depth (RGBD) images from any desired vantage points. In the second stage, the RGBD images are translated into corresponding tactile images typically generated by camera-based tactile sensors using a conditional generative adversarial network (cGAN). The cGAN model is itself trained with a large set of visuo-tactile images collected in simulation, and can be transferred into the real world without fine-tuning or additional data collection. We extensively validate this approach in the context of tactile object classification, showing that it effectively reduces data collection time by a factor of 20 while achieving similar performance to training a classifier on manually collected real data.
PaperID: 83,
Authors: Yichen Zhang, Xinyi Chen, Chen Feng, Boyu Zhou, Shaojie Shen
Affiliations: Department of Electronic and Computer Engineering, Hong Kong University of Science and Technology, Hong Kong, China; Southern University of Science and Technology, Shenzhen, China
Abstract: In this article, we introduce a novel Fast Autonomous expLoration framework using COverage path guidaNce (FALCON), which aims at setting a new performance benchmark in the field of autonomous aerial exploration. Despite recent advancements in the domain, existing exploration planners often suffer from inefficiencies, such as frequent revisitations of previously explored regions. FALCON effectively harnesses the full potential of online generated coverage paths in enhancing exploration efficiency. The framework begins with an incremental connectivity-aware space decomposition and connectivity graph construction, which facilitate efficient coverage path planning. Subsequently, a hierarchical planner generates a coverage path spanning the entire unexplored space, serving as a global guidance. Then, a local planner optimizes the frontier visitation order, minimizing traversal time while consciously incorporating the intention of the global guidance. Finally, minimum-time smooth and safe trajectories are produced to visit the frontier viewpoints. For fair and comprehensive benchmark experiments, we introduce a lightweight exploration planner evaluation environment that allows for comparing exploration planners across a variety of testing scenarios using an identical quadrotor simulator. In addition, an in-depth analysis and evaluation is conducted to highlight the significant performance advantages of FALCON in comparison with the state-of-the-art exploration planners based on objective criteria. Extensive ablation studies demonstrate the effectiveness of each component in the proposed framework. Real-world experiments conducted fully onboard further validate FALCON’s practical capability in complex and challenging environments. The source code of both the exploration planner FALCON and the exploration planner evaluation environment has been released to benefit the community.
PaperID: 84,
Authors: Anirvan Dutta, Etienne Burdet, Mohsen Kaboli
Affiliations: RoboTac Lab, BMW Group, Munich, Germany; Imperial College of Science, Technology and Medicine, London, U.K.
Abstract: Interactive exploration of unknown objects' properties, such as stiffness, mass, center of mass, friction coefficient, and shape, is crucial for autonomous robotic systems operating in unstructured environments. Precise identification of these properties is essential for stable and controlled object manipulation and for anticipating the outcomes of (prehensile or nonprehensile) manipulation actions, such as pushing, pulling, and lifting. Our study focuses on autonomously inferring the physical properties of a diverse set of homogeneous, heterogeneous, and articulated objects using a robotic system equipped with vision and tactile sensors. We propose a novel predictive perception framework to identify object properties by leveraging versatile exploratory actions: nonprehensile pushing and prehensile pulling. A key component of our framework is a novel active shape perception mechanism that seamlessly initiates exploration. In addition, our dual differentiable filtering with graph neural networks learns the object–robot interaction and enables consistent inference of indirectly observable, time-invariant object properties. Finally, we develop a N-step information gain approach to select the most informative actions for efficient learning and inference. Extensive real-robot experiments with planar objects show that our predictive perception framework outperforms state-of-the-art baselines and showcases it in three major applications for object tracking, goal-driven task, and environmental change detection.
PaperID: 85,
Authors: Linh Viet Nguyen, Khoi Thanh Nguyen, Van Anh Ho
Affiliations: Japan Advanced Institute of Science and Technology, Nomi, Japan
Abstract: In this article, we present a design concept, in which a monolithic soft body is incorporated with a vibration-driven mechanism, called Leafbot. We first report a morphological design of the robot's limbs that facilitates the forward locomotion of our vibration-driven model and enhances the capability of coping with sloped obstacles and irregular terrains. Second, the fabrication technique to achieve such a soft monolithic structure and limb morphology is fully addressed. Third, we clarify the locomotion of the Leafbot under high-frequency excitation via analytical and empirical methods in flat and even surface conditions. The maximum attained velocity in such a condition is 5 body length/ second. Finally, three model designs are constructed, each featuring a different limb pattern. We examine the terradynamics characteristics of three patterns in three pre-defined conditions, i.e., the success rate of overcoming the slope, semi-circular obstacles, and step-field terrains specialized by the rugosity factor. This proposed investigation aims to build a foundation for further terradynamics study of vibration-driven soft robots in a more complicated and confined environment, with potential applications in inspection tasks.
PaperID: 86,
Authors: Xusheng Hui, Jianjun Luo, Haonan You, Hao Sun
Affiliations: School of Astronautics, Northwestern Polytechnical University, Xi'an, China; Beijing Advanced Medical Technologies, Ltd. Inc., Beijing, China
Abstract: Noncontact manipulation in liquid environments holds significant applications in micro/nanofluidics, microassembly, micromanufacturing, and microrobotics. Achieving compatibility in manipulating both sedimented and floating objects, as well as independently and synergistically manipulating multiple targets, remains a significant challenge. Here, a noncontact manipulator is developed for both sedimented and floating objects using laser-induced thermocapillary convection. Various strategies are proposed based on the distinct responses of sedimented and floating objects. Predefined scanning and “checkpoint” methods facilitate accurate movements of individual and multiple particles, respectively. Ultrafast programmed scanning and laser multiplexing enable independent manipulation and high-throughput ordered distribution of multiple particles. At the air–liquid interface, “laser cage” and “laser wall” are proposed to serve as effective tools for manipulating floating objects, especially with vision-based closed-loop control. Methods and strategies here do not rely on specific features of targets, solvents, and substrates. Multiple examples, including complex path replication, maze traversal, and precise assembly and disassembly, are demonstrated to validate the feasibility of this manipulator. This work provides a versatile platform and a novel methodology for noncontact manipulation in liquid.
PaperID: 87,
Authors: Meng Ren, Wenhang Liu, Kun Song, Ling Shi, Zhenhua Xiong
Affiliations: School of Mechanical Engineering, State Key Laboratory of Mechanical System and Vibration, Shanghai Jiao Tong University, Shanghai, China; Department of Electronic and Computer Engineering, Hong Kong University of Science and Technology, Hong Kong
Abstract: The containment of multirobot systems (MRSs) has a wide range of applications. However, time delays in communication among robots introduce difficulties to the system to accomplish containment. In addition, the specific dynamics of robot models pose new nonlinear and nonholonomic challenges. To solve these problems, a containment control law is proposed first for double-integrator MRSs subject to nonuniform time-varying delays. In contrast to impractical uniform delays, nonuniform time-varying delays are considered more deeply from the perspective of the Laplacian matrix in this article. The stability is proved by the Lyapunov–Krasovskii function and linear matrix inequalities. The proposed control law is further refined into a dual-loop structure for multi-nonholonomic-mobile-robot systems, addressing the problem of nonholonomic constraints. Specifically, the first loop decouples the control inputs in a finite time, and then the nonholonomic robot models are regarded as linear models, which facilitates the proof of system stability. The effectiveness of the aforementioned two control laws is validated through simulations and experiments. Under these containment control laws, followers in the system reach the convex hull formed by leaders and meet the convergence objective despite the constraint of nonuniform time-varying delays.
PaperID: 88,
Authors: Kuan Xu, Yuefan Hao, Shenghai Yuan, Chen Wang, Lihua Xie
Affiliations: Centre for Advanced Robotics Technology Innovation (CARTIN), School of Electrical and Electronic Engineering, Nanyang Technological University, Singapore; Spatial AI and Robotics Lab, Computer Science and Engineering, University at Buffalo, Buffalo, NY, USA
Abstract: In this article, we present an efficient visual simultaneous localization and mapping (SLAM) system designed to tackle both short-term and long-term illumination challenges. Our system adopts a hybrid approach that combines deep learning techniques for feature detection and matching with traditional back-end optimization methods. Specifically, we propose a unified convolutional neural network that simultaneously extracts keypoints and structural lines. These features are then associated, matched, triangulated, and optimized in a coupled manner. In addition, we introduce a lightweight relocalization pipeline that reuses the built map, where keypoints, lines, and a structure graph are used to match the query frame with the map. To enhance the applicability of the proposed system to real-world robots, we deploy and accelerate the feature detection and matching networks using C++ and NVIDIA TensorRT. Extensive experiments conducted on various datasets demonstrate that our system outperforms other state-of-the-art visual SLAM systems in illumination-challenging environments. Efficiency evaluations show that our system can run at a rate of 73\,\mathrmHz on a PC and 40\,\mathrmHz on an embedded platform.
PaperID: 89,
Authors: Jules Sanchez, Jean-Emmanuel Deschaud, François Goulette
Affiliations: Centre of Robotics, Mines Paris - PSL, PSL University, Paris, France
Abstract: LiDAR semantic segmentation (LSS) for autonomous driving has been a growing field of interest in recent years. Datasets and methods have appeared and expanded very quickly, but methods have not been updated to exploit this new data availability and rely on the same classical datasets. Different ways of performing LSS training and inference can be divided into several subfields, which include the following: domain generalization, source-to-source segmentation, and pretraining. In this work, we aim to improve results in all of these subfields with the novel approach of multisource training. Multisource training relies on the availability of various datasets at training time. To overcome the common obstacles in multisource training, we introduce the coarse labels and call the newly created multisource dataset COLA. We propose three applications of this new dataset that display systematic improvement over single-source strategies: COLA-DG for domain generalization (+10% ), COLA-S2S for source-to-source segmentation (+5.3% ), and COLA-PT for pretraining (+12% ). We demonstrate that multisource approaches bring systematic improvement over single-source approaches.
PaperID: 90,
Authors: Dazhe Zhao, Renkun Wang, Sen Ding, Jiaze Shan, Xiao Guan, Zhaoyang Li, Jiaming Liang, Wenxi Gu, Bingpu Zhou, Iek Man Lei, Liwei Lin, Junwen Zhong
Affiliations: Department of Electromechanical Engineering, University of Macau, Macau, China; Institute of Applied Physics and Materials Engineering, University of Macau, Macau, China; Tencent Robotics X, Tencent, Shenzhen, China; Department of Mechanical Engineering, University of California at Berkeley, Berkeley, CA, USA
Abstract: High-speed and good trajectory controllability are two critical attributes of small artificial aquatic surface robots. Inspired by the moving mechanism of water striders, we herein propose insect-scale soft aquatic surface robots utilizing piezoelectric actuation coupled with asymmetric footpads. The aquatic surface robots move quickly without penetrating the water-air interface and utilize incoordinate propulsive force from asymmetric footpads to realize precise trajectory control. An ultrafast linear speed of 21.82 BL/s (24 cm/s) and a high angular speed of 303 °/s are achieved, which are advanced among small aquatic surface robots. We showcase agility and maneuverability by navigating through a water maze with a total route length of 88 cm in an actual driving time of 16.5 s. Moreover, proof-of-concept for search and rescue operations is demonstrated by using a robot to tow an on-water monitoring system to record a real-time video showing the “SOS” symbol. An untethered robot is also demonstrated to improve the practical potential. The design principles, operation mechanisms, and steering characteristics presented in this work provide fundamental guidelines for the development of future small aquatic surface robots.
PaperID: 91,
Authors: Sha Lu, Xuecheng Xu, Dongkun Zhang, Yuxuan Wu, Haojian Lu, Xieyuanli Chen, Rong Xiong, Yue Wang
Affiliations: Zhejiang University, Hangzhou, China; Shanghai Jiao Tong University, Shanghai, China; National University of Defense Technology, Changsha, China
Abstract: Global localization using onboard perception sensors, such as cameras and light detection and ranging (LiDAR) sensors, is crucial in autonomous driving and robotics applications when Global Positioning System (GPS) signals are unreliable. Most approaches achieve global localization by sequential place recognition (PR) and pose estimation (PE). Some methods train separate models for each task, while others employ a single model with dual heads, trained jointly with separate task-specific losses. However, the accuracy of localization heavily depends on the success of PR, which often fails in scenarios with significant changes in viewpoint or environmental appearance. Consequently, this renders the final PE of localization ineffective. To address this, we introduce a new paradigm, PR-by-PE localization, which bypasses the need for separate PR by directly deriving it from PE. We propose RING#, an end-to-end PR-by-PE localization network that operates in the bird's-eye-view (BEV) space, compatible with both vision and LiDAR sensors. RING# incorporates a novel design that learns two equivariant representations from BEV features, enabling globally convergent and computationally efficient PE. Comprehensive experiments on the north campus long-term vision and LiDAR (NCLT) and Oxford datasets show that RING# outperforms state-of-the-art methods in both vision and LiDAR modalities, validating the effectiveness of the proposed approach.
PaperID: 92,
Authors: Dexin Wang, Chunsheng Liu, Faliang Chang, Yichen Xu
Affiliations: School of Control Science and Engineering, Shandong University, Ji'nan, China
Abstract: Decision-making in robotics using denoising diffusion processes has increasingly become a hot research topic, but end-to-end policies perform poorly in tasks with rich contact and have limited interactivity. This article proposes Hierarchical Diffusion Policy (HDP), a new robot manipulation policy of using contact points to guide the generation of robot trajectories. The policy is divided into two layers: the high-level policy predicts the contact for the robot's next object manipulation based on 3-D information, while the low-level policy predicts the action sequence toward the high-level contact based on the latent variables of observation and contact. We represent both-level policies as conditional denoising diffusion processes, and combine behavioral cloning and Q-learning to optimize the low-level policy for accurately guiding actions towards contact. We benchmark Hierarchical Diffusion Policy across six different tasks and find that it significantly outperforms the existing state-of-the-art imitation learning method Diffusion Policy with an average improvement of 20.8% . We find that contact guidance yields significant improvements, including superior performance, greater interpretability, and stronger interactivity, especially on contact-rich tasks. To further unlock the potential of HDP, this article proposes a set of key technical contributions including one-shot gradient optimization, trajectory augmentation, and prompt guidance, which improve the policy's optimization efficiency, spatial awareness, and interactivity respectively. Finally, real-world experiments verify that HDP can handle both rigid and deformable objects.
PaperID: 93,
Authors: Yuchen Liu, Ruiqi Ni, Ahmed H. Qureshi
Affiliations: Department of Computer Science, Purdue University, West Lafayette, IN, USA
Abstract: Mapping and motion planning are two essential elements of robot intelligence that are interdependent in generating environment maps and navigating around obstacles. The existing mapping methods create maps that require computationally expensive motion planning tools to find a path solution. In this article, we propose a new mapping feature called arrival time fields, which is a solution to the Eikonal equation. The arrival time fields can directly guide the robot in navigating the given environments. Therefore, this article introduces a new approach called active neural time fields, which is a physics-informed neural framework that actively explores the unknown environment and maps its arrival time field on the fly for robot motion planning. Our method does not require any expert data for learning and uses neural networks to directly solve the Eikonal equation for arrival time field mapping and motion planning. We benchmark our approach against state-of-the-art mapping and motion planning methods and demonstrate its superior performance in both simulated and real-world environments with a differential drive robot and a six-degree-of-freedom robot manipulator.
PaperID: 94,
Authors: Claudio D. Pose, Juan Ignacio Giribet, Gabriel Torre
Affiliations: Laboratorio de Automática y Robótica, Facultad de Ingeniería, Universidad de Buenos Aires and CONICET - Universidad de San Andrés, Argentina; Laboratorio de Inteligencia Artificial y Robótica, Universidad de San Andrés and CONICET, Argentina; Laboratorio de Inteligencia Artificial y Robótica, Universidad de San Andrés and Instituto de Ingeniería Biomédica, Facultad de Ingeniería, Universidad de Buenos Aires, Argentina
Abstract: This manuscript details an architecture and training methodology for a data-driven framework aimed at detecting, identifying, and quantifying damage in the propeller blades of multirotor unmanned aerial vehicles. Real flight data was collected by substituting one propeller with a damaged counterpart, representing three distinct damage types of varying severity. This data was then used to train a composite model, which included both classifiers and neural networks, capable of accurately identifying the type of failure, estimating damage severity, and pinpointing the affected rotor. The data employed for this analysis were exclusively sourced from inertial measurements and control command inputs. This strategic choice ensures the adaptability of the proposed methodology across diverse multirotor vehicle platforms.
PaperID: 95,
Authors: Zehuan Yu, Zhijian Qiao, Wenyi Liu, Huan Yin, Shaojie Shen
Affiliations: Department of Electronic and Computer Engineering, The Hong Kong University of Science and Technology, Hong Kong, China; Department of Mechanism Engineering, The University of Hong Kong, Hong Kong, China
Abstract: Light detection and ranging (LiDAR) point cloud maps are extensively utilized on roads for robot navigation due to their high consistency. However, dense point clouds face challenges of high memory consumption and reduced maintainability for long-term operations. In this study, we introduce scalable and lightweight LiDAR mapping (SLIM), a scalable and lightweight mapping system for long-term LiDAR mapping in urban environments. The system begins by parameterizing structural point clouds into lines and planes. These lightweight and structural representations meet the requirements of map merging, pose graph optimization, and bundle adjustment, ensuring incremental management and local consistency. For long-term operations, a map-centric nonlinear factor recovery method is designed to sparsify poses while preserving mapping accuracy. We validate the SLIM system with multisession real-world LiDAR data from classical LiDAR mapping datasets, including KITTI, NCLT, HeLiPR, and M2DGR. The experiments demonstrate its capabilities in mapping accuracy, lightweightness, and scalability. Map reuse is also verified through map-based robot localization. Finally, with multisession LiDAR data, the SLIM system provides a globally consistent map with low memory consumption (~ 130 KB/km on KITTI).
PaperID: 96,
Authors: Iretiayo Akinola, Jie Xu, Jan Carius, Dieter Fox, Yashraj Narang
Affiliations: NVIDIA Corporation, Seattle, WA, USA
Abstract: For both humans and robots, the sense of touch, known as tactile sensing, is critical for performing contact-rich manipulation tasks. Three key challenges in robotic tactile sensing are interpreting sensor signals, generating sensor signals in novel scenarios, and learning sensor-based policies. For visuotactile sensors, interpretation has been facilitated by their close relationship with vision sensors (e.g., RGB cameras). However, generation is still difficult, as visuotactile sensors typically involve contact, deformation, illumination, and imaging, all of which are expensive to simulate; in turn, policy learning has been challenging, as simulation cannot be leveraged for large-scale data collection. We present TacSL (taxel), a library for GPU-based visuotactile sensor simulation and learning. TacSL can be used to simulate visuotactile images and extract contact-force distributions over 200× faster than the prior state-of-the-art, all within the widely used Isaac simulator. Furthermore, TacSL provides a learning toolkit containing multiple sensor models, contact-intensive training environments, and online/offline algorithms that can facilitate policy learning for sim-to-real applications. On the algorithmic side, we introduce a novel online reinforcement-learning algorithm called asymmetric actor-critic distillation, designed to effectively and efficiently learn tactile-based policies in simulation that can transfer to the real world. Finally, we demonstrate the utility of our library and algorithms by evaluating the benefits of distillation and multimodal sensing for contact-rich manipulation tasks, and most critically, performing sim-to-real transfer.
PaperID: 97,
Authors: Peng Yin, Shiqi Zhao, Jing Wang, Ruohai Ge, Jianmin Ji, Yeping Hu, Huaping Liu, Jianda Han
Affiliations: City University of Hong Kong, Hong Kong, SAR, China; University of Southern California, Los Angeles, CA, USA; University of Science and Technology of China, Hefei, China; Lawrence Livermore National Laboratory, Livermore, CA, USA; Tsinghua University, Beijing, China; Nankai University, Tianjin, China
Abstract: In this article, we introduce iLoc, an innovative visual localization system designed to enhance the autonomy and adaptability of robotic agents in long-term and large-scale applications. iLoc specializes in: 1) extracting stable and consistent descriptors for place recognition, unaffected by changes in viewpoint and illumination; 2) performing swift and precise global relocalization to establish a robot's position within a large and complex environment; and 3) generating real-time tracking trajectories aligned with reference maps, ensuring continual orientation within known spaces. Distinctively, iLoc incorporates a transformer-based learning module and an attention-enhanced recognition approach, enabling it to adapt to diverse environmental and viewpoint conditions. iLoc leverages a coarse-to-fine global feature matching technique for enhanced localization and integrates robust state estimation combining visual odometry and loop closures through local refinement and pose graph optimization. iLoc demonstrates remarkable proficiency in place recognition, achieving localization over distances of up to 2 km within 0.5 s with average accuracy at 1 m. It maintains stable localization accuracy, even under variable conditions. Its versatile design allows integration across various environments, significantly broadening the scope of universal localization capabilities in robotics. iLoc represents a substantial step forward in visual-based localization systems, delivering unparalleled speed and accuracy in place recognition. Its ability to adapt and respond to diverse environmental stimuli marks it as a crucial tool in advancing the field of robotic localization.
PaperID: 98,
Authors: Timothy Chen, Ola Shorinwa, Joseph Bruno, Aiden Swann, Javier Yu, Weijia Zeng, Keiko Nagami, Philip M. Dames, Mac Schwager
Affiliations: Stanford University, Stanford, CA, USA; Temple University, Philadelphia, PA, USA; University of California San Diego, San Diego, CA, USA
Abstract: We present Splat-Nav, a real-time robot navigation pipeline for Gaussian splatting (GSplat) scenes, a powerful new 3-D scene representation. Splat-Nav consists of two components: first, Splat-Plan, a safe planning module, and second, Splat-Loc, a robust vision-based pose estimation module. Splat-Plan builds a safe-by-construction polytope corridor through the map based on mathematically rigorous collision constraints and then constructs a Bézier curve trajectory through this corridor. Splat-Loc provides real-time recursive state estimates given only an RGB feed from an on-board camera, leveraging the point-cloud representation inherent in GSplat scenes. Working together, these modules give robots the ability to recursively replan smooth and safe trajectories to goal locations. Goals can be specified with position coordinates, or with language commands by using a semantic GSplat. We demonstrate improved safety compared to point cloud-based methods in extensive simulation experiments. In a total of 126 hardware flights, we demonstrate equivalent safety and speed compared to motion capture and visual odometry, but without a manual frame alignment required by those methods. We show online replanning at more than 2 Hz and pose estimation at about 25 Hz, an order of magnitude faster than neural radiance field-based navigation methods, thereby enabling real-time navigation.
PaperID: 99,
Authors: Shida Xu, Kaicheng Zhang, Sen Wang
Affiliations: Department of Electrical and Electronic Engineering and I-X, Imperial College London, London, U.K.
Abstract: Underwater environments pose significant challenges for visual simultaneous localization and mapping (SLAM) systems due to limited visibility, inadequate illumination, and sporadic loss of structural features in images. Addressing these challenges, this article introduces a novel, tightly coupled acoustic-visual-inertial SLAM approach, termed AQUA-SLAM, to fuse a Doppler velocity log (DVL), a stereo camera, and an inertial measurement unit (IMU) within a graph optimization framework. Moreover, we propose an efficient sensor calibration technique, encompassing the multisensor extrinsic calibration (among the DVL, camera, and IMU) and the DVL transducer misalignment calibration, with a fast linear approximation procedure for real-time online execution. The proposed methods are extensively evaluated in a tank environment with ground truth, and validated for offshore applications in the North Sea. The results demonstrate that our method surpasses current state-of-the-art underwater and visual-inertial SLAM systems in terms of localization accuracy and robustness. The proposed system will be made open-source for the community.
PaperID: 100,
Authors: Ruihua Han, Shuai Wang, Shuaijun Wang, Zeqing Zhang, Jianjun Chen, Shijie Lin, Chengyang Li, Cheng-Zhong Xu, Yonina C. Eldar, Qi Hao, Jia Pan
Affiliations: Department of Computer Science, University of Hong Kong, Hong Kong; Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences, Shenzhen, China; Department of Computer Science and Engineering, Southern University of Science and Technology, Shenzhen, China; Robotics and Autonomous Systems Thrust, Hong Kong University of Science and Technology Guangzhou, Guangzhou, China; IOTSC, University of Macau, Macau, China; Weizmann Institute of Science, Rehovot, Israel; Sifakis Research Institute for Trustworthy Autonomous Systems, Southern University of Science and Technology, Shenzhen, China
Abstract: Navigating a nonholonomic robot in a cluttered, unknown environment requires accurate perception and precise motion control for real-time collision avoidance. This article presents neural proximal alternating-minimization network (NeuPAN): a real-time, highly accurate, map-free, easy-to-deploy, and environment-invariant robot motion planner. Leveraging a tightly coupled perception-to-control framework, NeuPAN has two key innovations compared to existing approaches: first, it directly maps raw point cloud data to a latent distance feature space for collision-free motion generation, avoiding error propagation from the perception to control pipeline; second, it is interpretable from an end-to-end model-based learning perspective. The crux of NeuPAN is solving an end-to-end mathematical model with numerous point-level constraints using a plug-and-play proximal alternating-minimization network, incorporating neurons in the loop. This allows NeuPAN to generate real-time, physically interpretable motions. It seamlessly integrates data and knowledge engines, and its network parameters can be fine-tuned via backpropagation. We evaluate NeuPAN on a ground mobile robot, a wheel-legged robot, and an autonomous vehicle, in extensive simulated and real-world environments. Results demonstrate that NeuPAN outperforms existing baselines in terms of accuracy, efficiency, robustness, and generalization capabilities across various environments, including the cluttered sandbox, office, corridor, and parking lot. We show that NeuPAN works well in unknown and unstructured environments with arbitrarily shaped objects, transforming impassable paths into passable ones.
PaperID: 101,
Authors: Tommaso Belvedere, Marco Cognetti, Giuseppe Oriolo, Paolo Robuffo Giordano
Affiliations: CNRS, Inria, IRISA, Campus de Beaulieu, Univ Rennes, Rennes Cedex, France; LAAS-CNRS, Université de Toulouse, CNRS, UPS, Toulouse cedex , France; Dipartimento di Ingegneria Informatica, Automatica e Gestionale, Sapienza Università di Roma, Roma, Italy
Abstract: This article introduces a computationally efficient robust model predictive control (MPC) scheme for controlling nonlinear systems affected by parametric uncertainties in their models. The approach leverages the recent notion of closed-loop state sensitivity and the associated ellipsoidal tubes of perturbed trajectories for taking into account online time-varying restrictions on state and input constraints. This makes the MPC controller “aware” of potential additional requirements needed to cope with parametric uncertainty, thus significantly improving the tracking performance and success rates during navigation in constrained environments. One key contribution lies in the introduction of a computationally efficient robust MPC formulation with a comparable computational complexity to a standard MPC (i.e., an MPC not explicitly dealing with parametric uncertainty). An extensive simulation campaign is presented to demonstrate the effectiveness of the proposed approach in handling parametric uncertainties and enhancing task performance, safety, and overall robustness. Furthermore, we also provide an experimental validation that shows the feasibility of the approach in real-world conditions and corroborates the statistical findings of the simulation campaign. The versatility and efficiency of the proposed method make it therefore a valuable tool for real-time control of robots subject to nonnegligible uncertainty in their models.
PaperID: 102,
Authors: Christopher J. Ford, Haoran Li, Manuel G. Catalano, Matteo Bianchi, Efi Psomopoulou, Nathan F. Lepora
Affiliations: Department of Engineering Mathematics and Bristol Robotics Laboratory, University of Bristol, Bristol, U.K.; Department of Soft Robotics for Human Cooperation and Rehabilitation, Istituto Italiano di Tecnologia (IIT), Genova, Italy; Department of Information Engineering and the Research Center “E. Piaggio,”, University of Pisa, Pisa, Italy
Abstract: This article presents a shear-based control scheme for grasping and manipulating delicate objects with a Pisa/IIT anthropomorphic SoftHand equipped with soft biomimetic tactile sensors on all five fingertips. These “microTac” tactile sensors are miniature versions of the TacTip vision-based tactile sensor, and can extract precise contact geometry and force information at each fingertip for use as feedback into a controller to modulate the grasp while a held object is manipulated. Using a parallel processing pipeline, we asynchronously capture tactile images and predict contact pose and force from multiple tactile sensors. Consistent pose and force models across all sensors are developed using supervised deep learning with transfer learning techniques. We then develop a grasp control framework that uses contact force feedback from all fingertip sensors simultaneously, allowing the hand to safely handle delicate objects even under external disturbances. This control framework is applied to several grasp-manipulation experiments: First, retaining a flexible cup in a grasp without crushing it under changes in object weight; Second, a pouring task where the center of mass of the cup changes dynamically; and third, a tactile-driven leader-follower task where a human guides a held object. These manipulation tasks demonstrate more human-like dexterity with underactuated robotic hands by using fast reflexive control from tactile sensing.
PaperID: 103,
Authors: Zijian An, Lifeng Zhou
Affiliations: Department of Electrical and Computer Engineering, Drexel University, Philadelphia, PA, USA
Abstract: In this article, we study the problem of game-theoretic robot allocation where two players strategically allocate robots to compete for multiple sites of interest. Robots possess offensive or defensive capabilities to interfere and weaken their opponents to take over a competing site. This problem belongs to the conventional an acronym colonel blotto game (CBG). Considering the robots' heterogeneous capabilities and environmental factors, we generalize the conventional Blotto game by incorporating heterogeneous robot types and graph constraints that capture the robot transitions between sites. Then, we employ the double oracle algorithm (DOA) to solve for the Nash equilibrium of the generalized Blotto game. Particularly, for cyclic-dominance-heterogeneous (CDH) robots that inhibit each other, we define a new transformation rule between any two robot types. Building on the transformation, we design a novel utility function to measure the game's outcome quantitatively. Moreover, we rigorously prove the correctness of the designed utility function. Finally, we conduct extensive simulations to demonstrate the effectiveness of DOA on computing Nash equilibrium for homogeneous, linear heterogeneous, and CDH robot allocation on graphs.
PaperID: 104,
Authors: Michael Amir, Alfred M. Bruckstein
Affiliations: University of Cambridge, U.K.; Technion - Israel Institute of Technology, Haifa, Israel
Abstract: We investigate the algorithmic problem of uniformly dispersing a swarm of robots in an unknown, grid-like environment. In this setting, our goal is to study the relationships between performance metrics and robot capabilities. We introduce a formal model comparing dispersion algorithms based on makespan, traveled distance, energy consumption, sensing, communication, and memory. Using this framework, we classify uniform dispersion algorithms according to their capability requirements and performance. We prove that while makespan and travel can be minimized in all environments, energy cannot, if the swarm’s sensing range is bounded. In contrast, we show that energy can be minimized by “ant-like” robots in synchronous settings and asymptotically minimized in asynchronous settings, provided the environment is topologically simply connected, by using our “find-corner depth-first search” (FCDFS) algorithm. Our theoretical and experimental results show that FCDFS significantly outperforms known algorithms. Our findings reveal key limitations in designing swarm robotics systems for unknown environments, emphasizing the role of topology in energy-efficient dispersion.
PaperID: 105,
Authors: Connor Holmes, Frederike Dümbgen, Timothy D. Barfoot
Affiliations: University of Toronto Robotics Institute, ON, Canada; Inria, École Normale Supérieure, PSL University, Paris, France
Abstract: A recent set of techniques in the robotics community, known as certifiably correct methods, frames robotics problems as polynomial optimization problems and applies convex, semidefinite programming (SDP) relaxations to either find or certify their global optima. In parallel, differentiable optimization allows optimization problems to be embedded into end-to-end learning frameworks and has received considerable attention in the robotics community. In this article, we consider the ill effect of convergence to spurious local minima in the context of learning frameworks that use differentiable optimization. We present SDPRLayers, an approach that seeks to address this issue by combining convex relaxations with implicit differentiation techniques to provide certifiably correct solutions and gradients throughout the training process. We provide theoretical results that outline conditions for the correctness of these gradients and provide efficient means for their computation. Our approach is first applied to two simple-but-demonstrative simulated examples, which expose the potential pitfalls of reliance on local optimization in existing, state-of-the-art, differentiable optimization methods. We then apply our method in a real-world application: we train a deep neural network to detect image keypoints for robot localization in challenging lighting conditions. We provide our open-source, PyTorch implementation of SDPRLayers and our differentiable localization pipeline.
PaperID: 106,
Authors: Giovanni Franzese, Ravi Prakash, Cosimo Della Santina, Jens Kober
Affiliations: Cognitive Robotics, Delft University of Technology, Delft, The Netherlands; Cyber Physical Systems, Indian Institute of Science Bangalore, Bangalore, India
Abstract: Learning from Interactive Demonstrations has revolutionized the way nonexpert humans teach robots. It is enough to kinesthetically move the robot around to teach pick-and-place, dressing, or cleaning policies. However, the main challenge is correctly generalizing to novel situations, e.g., different surfaces to clean or different arm postures to dress. This article proposes a novel task parameterization and generalization to transport the original robot policy, i.e., position, velocity, orientation, and stiffness. Unlike the state of the art, only a set of keypoints is tracked during the demonstration and the execution, e.g., a point cloud of the surface to clean. We then propose to fit a nonlinear transformation that would deform the space and then the original policy using the paired source and target point sets. The use of function approximators like Gaussian Processes allows us to generalize, or transport, the policy from every space location while estimating the uncertainty of the resulting policy due to the limited task keypoints and the reduced number of demonstrations. We compare the algorithm’s performance with state-of-the-art task parameterization alternatives and analyze the effect of different function approximators. We also validated the algorithm on robot manipulation tasks, i.e., different posture arm dressing, different location product reshelving, and different shape surface cleaning.
PaperID: 107,
Authors: Zhan Gao, Guang Yang, Amanda Prorok
Affiliations: Department of Computer Science and Technology, University of Cambridge, Cambridge, U.K.
Abstract: This work views the multiagent system and its surrounding environment as a coevolving system, where the behavior of one affects the other. The goal is to take both agent actions and environment configurations as decision variables, and optimize these two components in a coordinated manner to improve some measure of interest. Toward this end, we consider the problem of decentralized multiagent navigation in a cluttered environment, where we assume that the layout of the environment is reconfigurable. By introducing two subobjectives—multiagent navigation and environment optimization—we propose an agent-environment co-optimization problem and develop a coordinated algorithm that alternates between these subobjectives to search for an optimal synthesis of agent actions and environment configurations; ultimately, improving the navigation performance. Due to the challenge of explicitly modeling the relation between the agents, the environment and their performance therein, we leverage policy gradient to formulate a model-free learning mechanism within the coordinated framework. A formal convergence analysis shows that our coordinated algorithm tracks the local minimum solution of an associated time-varying nonconvex optimization problem. Experiments corroborate theoretical findings and show the benefits of co-optimization. Interestingly, the results also indicate that optimized environments can offer structural guidance to deconflict agents in motion.
PaperID: 108,
Authors: Crystal E. Winston, Hojung Choi, Rianna M. Jitosho, Zhenishbek Zhakypov, Jasmin E. Palmer, Mark R. Cutkosky, Allison M. Okamura
Affiliations: Department of Mechanical Engineering, Stanford University, Stanford, CA, USA
Abstract: Skin deformation haptic devices worn on the finger pad provide realistic touch feedback during interactions with virtual objects. Two primary challenges in creating such devices are: first, making a multidegree-of-freedom device (DoF) that is small and lightweight so it does not encumber the wearer and second, providing accurate control of forces displayed to the finger pad. This work presents a 4-DoF finger pad haptic device, called Fourigami, that addresses these challenges. We address the first challenge using origami manufacturing methods and pneumatic actuation to fabricate a 25 g prototype that displays normal, shear, and twist and can be easily worn on the finger pad. We address the second challenge using a low-profile, 6-DoF, force/torque sensor to control forces displayed to the finger. Fourigami has a bandwidth ranging from 2 to 4 Hz depending on direction, and when acting on a human finger, it exerts forces ranging from \pm 1.0 N in shear, 4.2 N in normal, and \pm 4.2 N \cdot mm of twist. Finally, we demonstrate the device’s efficacy when rendering haptic feedback to a user tracking a sinusoidal trajectory and a trajectory representing interactions with a virtual object.
PaperID: 109,
Authors: Zhixiang Wang, Xudong Li, Yizhai Zhang, Fan Zhang, Panfeng Huang
Affiliations: Research Center for Intelligent Robotics, School of Astronautics, Northwestern Polytechnical University, Xi’an, China
Abstract: Event cameras, when combined with inertial sensors, show significant potential for motion estimation in challenging scenarios, such as high-speed maneuvers and low-light environments. While numerous methods exist for producing such estimations, most boil down to solving a synchronous discrete-time fusion problem. However, the asynchronous nature of event cameras and their unique fusion mechanism with inertial sensors remain underexplored. In this article, we introduce a monocular event-inertial odometry method called asynchronous event-inertial odometry (AsynEIO), designed to fuse asynchronous event and inertial data within a unified Gaussian process (GP) regression framework. Our approach incorporates an event-driven front-end that tracks feature trajectories directly from raw event streams at a high temporal resolution. These tracked feature trajectories, along with various inertial factors, are integrated into the same GP regression framework to enable asynchronous fusion. With deriving analytical residual Jacobians and noise models, our method constructs a factor graph that is iteratively optimized and pruned using a sliding-window optimizer. Comparative assessments highlight the performance of different inertial fusion strategies, suggesting optimal choices for varying conditions. Experimental results on both public datasets and our own event-inertial sequences indicate that AsynEIO outperforms existing methods, especially in high-speed and low-illumination scenarios.
PaperID: 110,
Authors: Yusuke Tanaka, Yuki Shirai, Alexander Schperberg, Xuan Lin, Dennis W. Hong
Affiliations: Department of Mechanical and Aerospace Engineering, University of California, Los Angeles, CA, USA
Abstract: This article presents Spine-enhanced Climbing Autonomous Limbed Exploration Robot (SCALER), a versatile free-climbing multilimbed robot that is designed to achieve tightly coupled simultaneous locomotion and dexterous grasping. While existing quadrupedal-limbed robots have demonstrated impressive dexterous capabilities, achieving a balance between power-demanding locomotion and precise grasping remains a critical challenge. We design a torso mechanism and a parallel–serial limb to meet the conflicting requirements that pose unique challenges in hardware design. SCALER employs underactuated two-fingered GOAT grippers that can mechanically adapt and offer seven modes of grasping, enabling SCALER to traverse extreme terrains with multimodal grasping strategies. We study the whole-body approach, where SCALER utilizes its body and limbs to generate additional forces for stable grasping in various environments, thereby further enhancing its versatility. Furthermore, we improve the GOAT gripper actuation speed to realize more dynamic climbing in a closed-loop control fashion. With these proposed technologies, SCALER can traverse vertical, overhanging, upside-down, slippery terrains and bouldering walls with nonconvex-shaped climbing holds under the Earth’s gravity.
PaperID: 111,
Authors: Zhenwei Zhang, Yuhao Zhang, Xingwei Zhao, Bo Tao, Han Ding
Affiliations: State Key Laboratory of Intelligent Manufacturing Equipment and Technology, Department of Mechanical Science and Engineering, Huazhong University of Science and Technology, Wuhan, China
Abstract: Multirobot coordination in shared workspaces is prone to deadlocks, which can compromise operational capabilities and task efficiency. Accurately determining the timing and spatial locations of deadlocks is essential for effective resolution, yet remains challenging due to dynamic robot interactions and growing system complexity. To this end, a distributed deadlock-aware control framework is proposed for robots to detect and avoid deadlocks while maintaining safe task execution. First, deadlocks are characterized by analyzing undesired equilibria in robot dynamics under safety constraints imposed by multiple stacked control barrier functions (CBFs). Our analysis reveals two critical properties: 1) deadlocks occur at intersections of all active CBF boundaries; and 2) deadlocks arise when robot stabilizing force are confined within the conical hull formed by active safety forces. These theoretical insights underpin a new detection method that identifies potential deadlocks from conflicts between safety requirements and task objectives. Furthermore, a reactive deadlock avoidance method is designed to help robots escape and prevent entry into potential deadlock regions by adaptively modulating the stabilizing force. A generalized workflow is established to systematically address deadlocks across various multirobot tasks. Simulation and hardware experiments are conducted on robots collaborating in dense environments to validate the framework’s effectiveness in preventing task failures caused by deadlocks.
PaperID: 112,
Authors: Peter Amorese, Shohei Wakayama, Nisar R. Ahmed, Morteza Lahijanian
Affiliations: University of Colorado Boulder, Boulder, CO, USA
Abstract: When a robot autonomously performs a complex task, it frequently must balance competing objectives while maintaining safety. This becomes more difficult in uncertain environments with stochastic outcomes. Enhancing transparency in the robot’s behavior and aligning with user preferences are also crucial. This article introduces a novel framework for multiobjective reinforcement learning that ensures safe task execution, optimizes tradeoffs between objectives, and adheres to user preferences. The framework has two main layers: a multiobjective task planner and a high-level selector. The planning layer generates a set of optimal tradeoff plans that guarantee satisfaction of a temporal logic task. The selector uses active inference to decide which generated plan best complies with user preferences and aids learning. Operating iteratively, the framework updates a parameterized learning model based on collected data. Case studies and benchmarks on both manipulation and mobile robots show that our framework outperforms other methods and (i) learns multiple optimal tradeoffs, (ii) adheres to a user preference, and (iii) allows the user to adjust the balance between (i) and (ii).
PaperID: 113,
Authors: Giovanni Cioffi, Leonard Bauersfeld, Davide Scaramuzza
Affiliations: Robotics and Perception Group, Department of Informatics, University of Zurich, Zürich, Switzerland
Abstract: Visual-inertial odometry (VIO) is widely used for state estimation in autonomous micro aerial vehicles using onboard sensors. Current methods improve VIO by incorporating a model of the translational vehicle dynamics, yet their performance degrades when faced with low-accuracy vehicle models or continuous external disturbances, like wind. Additionally, incorporating rotational dynamics in these models is computationally intractable when they are deployed in online applications, e.g., in a closed-loop control system. We present HDVIO2.0, which models full 6-DoF, translational and rotational, vehicle dynamics and tightly incorporates them into a VIO system with minimal impact on the runtime. HDVIO2.0 builds upon the previous work, HDVIO, and addresses these challenges through a hybrid dynamics model combining a point-mass vehicle model with a learning-based component, with access to control commands and inertial measurement unit (IMU) history, to capture complex aerodynamic effects. The key idea behind modeling the rotational dynamics is to represent them with continuous-time functions. HDVIO2.0 leverages the divergence between the actual motion and the predicted motion from the hybrid dynamics model to estimate external forces as well as the robot state. Our system surpasses the performance of state-of-the-art methods in experiments using public and new drone dynamics datasets, as well as real-world flights in winds up to 25 km/h. Unlike existing approaches, we also show that accurate vehicle dynamics predictions are achievable without precise knowledge of the vehicle state.
PaperID: 114,
Authors: Jiahao Liang, Yuanzhe Wang, Guohao Peng, Zhenyu Wu, Danwei Wang
Affiliations: School of Electrical and Electronic Engineering, Nanyang Technological University, Singapore; School of Control Science and Engineering, Shandong University, Jinan, China
Abstract: Curb following is a critical technology for autonomous road sweeping vehicles. However, existing solutions face two primary challenges: 1) unreliable curb detection; and 2) inefficient motion generation. Unreliable curb detection stems from the wide variability in curb dimensions and types, as well as interference from roadside features, such as vegetation and infrastructure. Inefficient motion generation occurs when existing methods prioritize tracking accuracy while neglecting task completion efficiency, leading to prolonged operation times. To address these challenges, we propose Curb-Tracker, an integrated curb-following system designed for autonomous vehicles operating in diverse road environments. First, we develop a robust and adaptive curb detection algorithm that leverages a 2.5-D elevation map of the local environment and dynamically adjusts key parameters online to ensure reliable detection across varying scenarios. Second, to achieve accurate and efficient curb-aligned motion generation, we leverage model predictive contouring control as a tailored framework specifically designed for the curb-following task to generate an optimal control sequence for the vehicle to maintain a specified lateral offset from the curb while maximizing travel progress along it. The proposed system has been implemented on a Hunter 2.0, a front-wheel Ackerman-steering mobile robot, and has been validated through extensive experiments in both Gazebo simulation and real-world environments. Experimental results demonstrate the effectiveness, adaptability, and robustness of the proposed system across a wide range of road scenarios.
PaperID: 115,
Authors: Hanying Zhao, Lingwei Xu, Yi Li, Feiyang Wen, Haoran Gao, Changwu Liu, Jincheng Yu, Yu Wang, Yuan Shen
Affiliations: Department of Electronic Engineering, Tsinghua University, Beijing, China; Department of Electronic Engineering, Beijing National Research Center for Information Science and Technology, Tsinghua University, Beijing, China
Abstract: In environments where robots operate with limited global navigation satellite system accessibility, ultra-wideband (UWB) localization technology is a popular auxiliary solution to assist visual–inertial odometry systems. However, current UWB approaches lack 3-D pairwise localization capability and suffer from rapidly declining localization update rates as the network scales, limiting their effectiveness for swarm robotic applications. This article presents a novel UWB sensor that enables 3-D pairwise localization and a localization scheme that can deliver robust, scalable, and accurate position awareness for multi-robot systems. Our approach begins with calibrating intrinsic UWB errors from hardware deviations and propagation effects, yielding high-accuracy distance and direction measurements. Using these measurements, we perform distributed relative localization through inter- and intra-node cooperation by integrating UWB and inertial measurement unit data. To enable swarm-scale operation, our platform implements the signal-multiplexing network ranging protocol to maximize update rates and network capacity. Experimental results show that our approach achieves centimeter-level localization accuracy at high update rates (100 Hz with UWB only), validating its robustness, scalability, and accuracy for robotic applications.
PaperID: 116,
Authors: Walker Gosrich, Saurav Agarwal, Kashish Garg, Siddharth Mayya, Matthew Malencia, Mark Yim, Vijay Kumar
Affiliations: GRASP Laboratory, University of Pennsylvania, Philadelphia, PA, USA; Aquatic Labs, Cambridge, MA, USA; Zipline, San Francisco, CA, USA
Abstract: We propose a new formulation for the multirobot task allocation problem that incorporates 1) complex precedence relationships between tasks, 2) efficient intratask coordination, and 3) cooperation through the formation of robot coalitions. A task graph specifies the tasks and their relationships, and a set of reward functions models the effects of coalition size and preceding task performance. Maximizing task rewards is NP-hard; hence, we propose network flow-based algorithms to approximate solutions efficiently. A novel online algorithm performs iterative reallocation, providing robustness to task failures and model inaccuracies to achieve higher performance than offline approaches. We comprehensively evaluate the algorithms in a testbed with random missions and reward functions and compare them to a mixed-integer solver and a greedy heuristic. In addition, we validate the overall approach in an advanced simulator, modeling reward functions based on realistic physical phenomena and executing the tasks with realistic robot dynamics. Results establish efficacy in modeling complex missions and efficiency in generating high-fidelity task plans while leveraging task relationships.
PaperID: 117,
Authors: Ke Wu, Zicheng Zhang, Muer Tie, Ziqing Ai, Zhongxue Gan, Wenchao Ding
Affiliations: College of Intelligent Robotics and Advanced Manufacturing, Fudan University, Shanghai, China; College of Computer Science and Artificial Intelligence, Fudan University, Shanghai, China
Abstract: VINGS-Mono is a monocular (inertial) Gaussian splatting (GS) SLAM framework designed for large scenes. The framework comprises four main components: visual-inertial odometry (VIO) front end, 2-D Gaussian map, novel view synthesis (NVS) loop closure, and dynamic eraser. In the VIO front end, RGB frames are processed through dense bundle adjustment and uncertainty estimation to extract scene geometry and poses. Based on this output, the mapping module incrementally constructs and maintains a 2-D Gaussian map. Key components of the 2-D Gaussian map include a sample-based rasterizer, score manager, and pose refinement, which collectively improve mapping speed and localization accuracy. This enables the SLAM system to handle large-scale urban environments with up to 50 million Gaussian ellipsoids. To ensure global consistency in large-scale scenes, we design a loop-closure module, which innovatively leverages the NVS capabilities of GS for loop-closure detection and correction of the Gaussian map. In addition, we propose a dynamic eraser to address the inevitable presence of dynamic objects in real-world outdoor scenes. Extensive evaluations in indoor and outdoor environments demonstrate that our approach achieves localization performance on par with VIO while surpassing recent GS/NeRF SLAM methods. It also significantly outperforms all existing methods in terms of mapping and rendering quality. Furthermore, we developed a mobile app and verified that our framework can generate high-quality Gaussian maps in real time using only a smartphone camera and a low-frequency IMU sensor. To the best of our knowledge, VINGS-Mono is the first monocular Gaussian SLAM method capable of operating in outdoor environments and supporting kilometer-scale large scenes.
PaperID: 118,
Authors: Zhanwei Wang, Huaijin Chen, Syeda Shadab Zehra Zaidi, Ellen Roels, Hendrik Cools, Bram Vanderborght, Seppe Terryn
Affiliations: Brubotics, Vrije Universiteit Brussel and IMEC, Elsene, Belgium; Biorobotics Institute, Scuola Superiore Sant’Anna (SSSA), Pisa, Italy
Abstract: While vacuum-based bending actuation offers benefits such as safety and compactness in soft robotics, it is often overlooked due to its limited actuation pressure, which restricts both bending angle and force output. This study presents a crease-free, origami-inspired vacuum bending actuator that advances both state-of-the-art vacuum bending actuators and traditional origami deformation principles by introducing orderly self folding through optimized stiffness distribution. Achieved through finite element method, this design provides several advantages: 1) Self-folding allows for high bending angles (up to 138^\circ ) in a compact form. 2) The crease-free design facilitates 3-D printing from a single soft material using a consumer-level fused filament fabrication printer, specifically thermoplastic polyurethane with a Shore hardness of 60 A, potentially higher flexibility and durability. 3) The compact configuration enables modular design, supporting reconfiguration as demonstrated in adaptable locomotion soft robots. 4) The large bending angles allow the actuator to wrap around objects, offering extensive contact compared to other designs. This capability, combined with its vacuum-driven mechanism, enables synergy with self-closing suction cups in an octopus-like vacuum gripper, providing large versatility and grasping force for handling a wide range of objects, from small, irregular shapes to larger, flat items.
PaperID: 119,
Authors: Saurav Agarwal, Ramya Muthukrishnan, Walker Gosrich, Vijay Kumar, Alejandro Ribeiro
Affiliations: GRASP Laboratory, University of Pennsylvania, Philadelphia, PA, USA; CSAIL, Massachusetts Institute of Technology, Cambridge, MA, USA; Department of Electrical and Systems Engineering, University of Pennsylvania, Philadelphia, PA, USA
Abstract: Coverage control is the problem of navigating a robot swarm to collaboratively monitor features or a phenomenon of interest not known a priori. The problem is challenging in decentralized settings with robots that have limited communication and sensing capabilities. We propose a learnable perception-action-communication (LPAC) architecture for the problem, wherein a convolutional neural network (CNN) processes localized perception; a graph neural network (GNN) facilitates robot communications; finally, a shallow multilayer perceptron computes robot actions. The GNN enables collaboration in the robot swarm by computing what information to communicate with nearby robots and how to incorporate received information. Evaluations show that the LPAC models—trained using imitation learning—outperform standard decentralized and centralized coverage control algorithms. The learned policy generalizes to environments different from the training dataset, transfers to larger environments with more robots, and is robust to noisy position estimates. The results indicate the suitability of LPAC architectures for decentralized navigation in robot swarms to achieve collaborative behavior.
PaperID: 120,
Authors: Yongbo Chen, Yanhao Zhang, Shaifali Parashar, Liang Zhao, Shoudong Huang
Affiliations: Robotics Institute, University of Technology Sydney, Ultimo, NSW, Australia; Institut National des Sciences Appliquées de Lyon (LIRIS, INSA-Lyon), Villeurbanne, France; School of Informatics, University of Edinburgh, Edinburgh, U.K.
Abstract: Nonrigid structure-from-motion (NRSfM), a promising technique for addressing the mapping challenges in monocular visual deformable simultaneous localization and mapping, has attracted growing attention. We introduce a novel method, called Con-NRSfM, for NRSfM under conformal deformations, encompassing isometric deformations as a subset. Our approach performs point-wise reconstruction using 2-D selected image warps optimized through a graph-based framework. Unlike existing methods that rely on strict assumptions, such as locally planar surfaces or locally linear deformations, and fail to recover the conformal scale, our method eliminates these constraints and accurately computes the local conformal scale. In addition, our framework decouples constraints on depth and conformal scale, which are inseparable in other approaches, enabling more precise depth estimation. To address the sensitivity of the formulated problem, we employ a parallel separable iterative optimization strategy. Furthermore, a self-supervised learning framework, utilizing an encoder–decoder network, is incorporated to generate dense 3-D point clouds with texture. Simulation and experimental results using both synthetic and real datasets demonstrate that our method surpasses existing approaches in terms of reconstruction accuracy and robustness.
PaperID: 121,
Authors: Kun Cao, Xinhang Xu, Wanxin Jin, Karl Henrik Johansson, Lihua Xie
Affiliations: Department of Control Science and Engineering, College of Electronics and Information Engineering, Tongji University, Shanghai, China; School of Electrical and Electronic Engineering, Nanyang Technological University, Singapore; School for Engineering of Matter, Transport, and Energy, Arizona State University, Tempe, AZ, USA; Division of Decision and Control Systems, School of Electrical Engineering and Computer Science, KTH Royal Institute of Technology, Stockholm, Sweden
Abstract: A differential dynamic programming (DDP)-based framework for inverse reinforcement learning (IRL) is introduced to recover the parameters in the cost function, system dynamics, and constraints from demonstrations. Different from existing work, where DDP was usually used for the inner forward problem, our proposed framework uses it to efficiently compute the gradient required in the outer inverse problem with equality and inequality constraints. The equivalence between the proposed and existing methods based on Pontryagin's maximum principle (PMP) is established. More importantly, using this DDP-based IRL with an open-loop loss function, a closed-loop IRL framework is presented. In this framework, a loss function is proposed to capture the closed-loop nature of demonstrations. It is shown to be better than the commonly used open-loop loss function. We show that the closed-loop IRL framework reduces to a constrained inverse optimal control problem under certain assumptions. Under these assumptions and a rank condition, it is proven that the learning parameters can be recovered from the demonstration data. The proposed framework is extensively evaluated through four numerical robot examples and one real-world quadrotor system. The experiments validate the theoretical results and illustrate the practical relevance of the approach.
PaperID: 122,
Authors: Xinyu Gao, Weiwei Shang, Bin Zhang
Affiliations: Department of Automation, University of Science and Technology of China, Hefei, China
Abstract: In this article, a disturbance observer-based model predictive control (DOB-based MPC) strategy is proposed for the trajectory tracking of cable-driven parallel robots (CDPRs). The original nonlinear optimization problem of the MPC explicitly handles the positive bounded constraints of cable tensions and is transformed into a quadratic problem (QP) based on the desired trajectory and offline workspace analysis. Additionally, the uncertainties and external disturbances in the system are considered and derived as the lumped disturbance. Then, a nonlinear DOB is used to estimate the lumped disturbance, and accordingly the estimation is used to enhance the prediction model of the MPC. The control input of the proposed MPC strategy is redesigned to incorporate an auxiliary controller. The estimation error of the DOB and the time-varying characteristics of the disturbance are leveraged to tighten the constraints of the QP less conservatively and generate feasible tubes. Such tube techniques guarantee the recursive feasibility and the input-to-state stability of the proposed MPC strategy. Furthermore, the whole algorithm for deployment including an online constraint updating method is developed. Both simulations and experiments are carried out thoroughly, showing that the DOB-based MPC can effectively improve trajectory tracking accuracy and ensure that the cable tensions satisfy the constraints in the case of unknown disturbances and model uncertainties.
PaperID: 123,
Authors: Wei Zhang, Qing Cheng, David Skuddis, Niclas Zeller, Daniel Cremers, Norbert Haala
Affiliations: Institute for Photogrammetry and Geoinformatics, University of Stuttgart, Stuttgart, Germany; Technical University of Munich, Munich, Germany; Karlsruhe University of Applied Sciences, Karlsruhe, Germany
Abstract: We present HI-SLAM2, a geometry-aware Gaussian SLAM system that achieves fast and accurate monocular scene reconstruction using only RGB input. Existing neural SLAM or 3DGS-based SLAM methods often tradeoff between rendering quality and geometry accuracy, our research demonstrates that both can be achieved simultaneously with RGB input alone. The key idea of our approach is to enhance the ability for geometry estimation by combining easy-to-obtain monocular priors with learning-based dense SLAM, and then using 3-D Gaussian splatting as our core map representation to efficiently model the scene. Upon loop closure, our method ensures on-the-fly global consistency through efficient pose graph bundle adjustment and instant map updates by explicitly deforming the 3-D Gaussian units based on anchored keyframe updates. Furthermore, we introduce a grid-based scale alignment strategy to maintain improved scale consistency in prior depths for finer depth details. Through extensive experiments on Replica, ScanNet, Waymo Open, ETH3D SLAM and ScanNet++ datasets, we demonstrate significant improvements over existing neural SLAM methods and even surpass RGB-D-based methods in both reconstruction and rendering quality.
PaperID: 124,
Authors: Xu Liu, Jiuzhou Lei, Ankit Prabhu, Yuezhan Tao, Igor Spasojevic, Pratik Chaudhari, Nikolay Atanasov, Vijay Kumar
Affiliations: Microsoft, Redmond, WA, USA; Texas A&M University, College Station, TX, USA; GRASP Laboratory, University of Pennsylvania, Philadelphia, PA, USA; University of California, Riverside, CA, USA; Department of Electrical and Computer Engineering, University of California San Diego, La Jolla, CA, USA
Abstract: This article develops a real-time decentralized metric-semantic simultaneous localization and mapping (SLAM) algorithm that enables a heterogeneous robot team to collaboratively construct object-based metric-semantic maps. The proposed framework integrates a data-driven front-end, for instance, segmentation from either RGBD cameras or light detection and ranging (LiDAR) and a custom back-end for optimizing robot trajectories and object landmarks in the map. To allow multiple robots to merge their information, we design semantics-driven place recognition algorithms that leverage the informativeness and viewpoint invariance of the object-level metric-semantic map for inter-robot loop closure detection. A communication module is designed to track each robot's observations and those of other robots whenever communication links are available. The framework supports real-time, decentralized operation onboard the robots and has been integrated with three types of aerial and ground platforms. We validate its effectiveness through experiments in both indoor and outdoor environments, as well as benchmarks on public datasets and comparisons with existing methods. The framework is open-sourced and suitable for both single-agent and multirobot real-time metric-semantic SLAM applications.
PaperID: 125,
Authors: Andrey Zhitnikov, Vadim Indelman
Affiliations: Technion Autonomous Systems Program (TASP), Technion—Israel Institute of Technology, Haifa, Israel; Stephen B. Klein Faculty of Aerospace Engineering, Technion—Israel Institute of Technology, Haifa, Israel
Abstract: Taking into account future risk is essential for an autonomously operating robot to find online not only the best but also a safe action to execute. In this article, we build upon the recently introduced formulation of probabilistic belief-dependent constraints. In our methodology, safety can be materialized with any general belief-dependent operator we call payoff. We present an anytime approach employing the Monte Carlo Tree Search (MCTS) method in continuous domains in terms of states, actions, and observations and general-belief-dependent reward and payoff operators. Unlike previous approaches, our method ensures safety anytime with respect to the currently expanded search tree without relying on the convergence of the search. We prove convergence in probability with an exponential rate of a version of our algorithms and study proposed techniques via extensive simulations. Even with a tiny number of tree queries, the best action found by our approach is much safer than the baseline. Moreover, our approach constantly yields better than the baseline action in terms of objective function. This is because we revise the values and statistics maintained in the search tree and remove from them the contribution of the pruned actions. We rigorously show that our cleaning routine is necessary. Without it, at the limit of convergence of MCTS, an infinite amount of sampled dangerous actions can be detrimental to the objective function.
PaperID: 126,
Authors: Giuseppe Milazzo, Simon Lemerle, Giorgio Grioli, Antonio Bicchi, Manuel G. Catalano
Affiliations: Soft Robotics for Human Cooperation and Rehabilitation, Istituto Italiano di Tecnologia, Genova, Italy
Abstract: Intuitively, prostheses with user-controllable stiffness could mimic the intrinsic behavior of the human musculoskeletal system, promoting safe and natural interactions and task adaptability in real-world scenarios. However, prosthetic design often disregards compliance because of the additional complexity, weight, and needed control channels. This article focuses on designing a variable stiffness actuator (VSA) with weight, size, and performance compatible with prosthetic applications, addressing its implementation for the elbow joint. While a direct biomimetic approach suggests adopting an agonist-antagonist (AA) layout to replicate the biceps and triceps brachii with elastic actuation, this solution is not optimal to accommodate the varied morphologies of residual limbs. Instead, we employed the AA layout to craft an elbow prosthesis fully contained in the user's forearm, catering to individuals with distal transhumeral amputations. In addition, we introduce a variant of this design where the two motors are split in the upper arm and forearm to distribute mass and volume more evenly along the bionic limb, enhancing comfort for patients with more proximal amputation levels. We characterize and validate our approach, demonstrating that both architectures meet the target requirements for an elbow prosthesis. The system attains the desired 120^\circ range of motion, achieves the target stiffness range of [2, 60] N \cdot m/rad, and can actively lift up to 3 kg. Our novel design reduces weight by up to 50% compared to existing VSAs for elbow prostheses while achieving performance comparable to the state of the art. Case studies suggest that passive and variable compliance could enable robust and safe interactions and task adaptability in the real world.
PaperID: 127,
Authors: Elena-Sorina Lupu, Fengze Xie, James A. Preiss, Jedidiah Alindogan, Matthew Anderson, Soon-Jo Chung
Affiliations: Graduate Aerospace Laboratories, California Institute of Technology, Pasadena, CA, USA; Department of Computing and Mathematical Sciences, California Institute of Technology, Pasadena, CA, USA; Department of Computer Science, University of California, Santa Barbara, Santa Barbara, CA, USA
Abstract: Control of off-road vehicles is challenging due to the complex dynamic interactions with the terrain. Accurate modeling of these interactions is important to optimize driving performance, but the relevant physical phenomena, such as slip, are too complex to model from first principles. Therefore, we present an offline meta-learning algorithm to construct a rapidly-tunable model of residual dynamics and disturbances. Our model processes terrain images into features using a visual foundation model (VFM), then maps these features and the vehicle state to an estimate of the current actuation matrix using a deep neural network (DNN). We then combine this model with composite adaptive control to modify the last layer of the DNN in real time, accounting for the remaining terrain interactions not captured during offline training. We provide mathematical guarantees of stability and robustness for our controller, and demonstrate the effectiveness of our method through simulations and hardware experiments with a tracked vehicle and a car-like robot. We evaluate our method outdoors on different slopes with varying slippage and actuator degradation disturbances, and compare against an adaptive controller that does not use the VFM terrain features. We show significant improvement over the baseline in both hardware experimentation and simulation.
PaperID: 128,
Authors: Da Sun, Qianfang Liao
Affiliations: School of Computer Science and Technology, University of Science and Technology of China, Hefei, China
Abstract: Efficiently coordinating multiple robotic arms is vital for secure and optimal operation in a shared workspace. This requires not only successful task completion but also minimizing collision risks from overlapping movements. Introducing real-time motion modulation adds an extra layer of challenge to this coordination task. In this article, we introduce a novel framework for real-time multiarm coordination, offering two main contributions: First, based on fuzzy model-based movement primitives, we propose a method for real-time trajectory modulation by learning from single demonstrations. This capability allows robots to modulate their motions online to reach arbitrary new desired places smoothly without necessitating extra demonstrations from users. Second, our framework incorporates a real-time multiarm coordination strategy that seamlessly integrates the trajectory modulation method with an extended reactive approach. This strategy empowers multiple robotic arms operating within a shared workspace to dynamically regulate their movements and execute tasks simultaneously in a human-desired manner while reactively avoiding mutual collisions. In the experiments, we utilize a group of robotic arms working in a shared workspace to validate the effectiveness of our framework and to make comparisons with state-of-the-art methods.
PaperID: 129,
Authors: Rilun Xia, Dongming Wang, Chenqi Mou
Affiliations: LMIB-School of Mathematical Sciences, Beihang University, Beijing, China; LMIB-School of Artificial Intelligence, Beihang University, Beijing, China
Abstract: The problem of collision detection plays an important role in many fields of science and engineering. This article presents a collision detection method for general convex objects bounded by pieces of implicit surfaces. There are two key ideas that underlie our method: one is the introduction of a new kind of pseudodistance, called the \delta-distance, for implicitly represented convex objects which has the desired properties of convexity and square differentiability; the other is the use of \delta-distance functions to construct a virtual potential field in the real space, so that the problem of collision detection can be reduced to a problem of unconstrained convex optimization. The method is extended and applied to detect whether two objects collide when they are moving continuously along linearly translational trajectories, which is a special case of one of the continuous collision detection subproblems. We have implemented collision detection algorithms in C++ and conducted a large number of experiments, with test examples involving objects modeled by planar, quadric, superquadric, superellipsoidal, and hyperquadric surfaces, as well as pieces of them, in both stationary and linearly translational moving states. The experimental results show that our method has good performance and it is computationally efficient and widely applicable.
PaperID: 130,
Authors: Jakub Rozlivek, Alessandro Roncone, Ugo Pattacini, Matej Hoffmann
Affiliations: Department of Cybernetics, Faculty of Electrical Engineering, Czech Technical University in Prague, Prague, Czech Republic; Human Interaction and RObotics (HIRO), Department of Computer Science, University of Colorado Boulder, Boulder, CO, USA; iCub Tech, Istituto Italiano di Tecnologia, Genova, Italy
Abstract: For safe and effective operation of humanoid robots in human-populated environments, the problem of commanding a large number of degrees of freedom (DoFs) while simultaneously considering dynamic obstacles and human proximity has still not been solved. In this article, we present a new reactive motion controller that commands two arms of a humanoid robot and three torso joints (17 DoF in total). We formulate a quadratic program that seeks joint velocity commands respecting multiple constraints while minimizing the magnitude of the velocities. We introduce a new unified treatment of obstacles that dynamically maps visual and proximity (precollision) and tactile (postcollision) obstacles as additional constraints to the motion controller, in a distributed fashion over the surface of the upper body of the iCub robot (with 2000 pressure-sensitive receptors). This results in a bioinspired controller that: first, gives rise to a robot with whole-body visuo-tactile awareness, resembling peripersonal space representations, and, second, produces human-like minimum jerk movement profiles. The controller was extensively experimentally validated, including a physical human–robot interaction scenario.
PaperID: 131,
Authors: Dinmukhammed Mukashev, Saltanat Seitzhan, Jabrail Chumakov, Soibkhon Khajikhanov, Madina Yergibay, Nurlan Zhaniyar, Rustam Chibar, Ayan Mazhitov, Matteo Rubagotti, Zhanat Kappassov
Affiliations: Institute of Smart Systems and Artificial Intelligence, Nazarbayev University, Astana, Kazakhstan; Department of Robotics, School of Engineering and Digital Sciences, Nazarbayev University, Astana, Kazakhstan
Abstract: The prompt and robust detection of tactile information is a relevant and challenging problem, and a considerable research effort is, thus, being put into innovative transduction methods for tactile sensors. In this article, we investigate the possibility of using event-based cameras to sense contact forces applied to objects by a robot end effector. The proposed optical tactile sensor incorporates a soft hemispherical pad made of silicone rubber with imprinted markers and a pulsewidth modulation light source that emits optical pulses, allowing robust detection of the markers to track deformations of the pad. To test the effectiveness of our sensor, experiments were carried out attaching it to a teleoperated robot arm to finely control it when out of the user's field of view, as accurately as if the user could see it. In the experiments, an augmented reality display and a haptic device were used to convey the force detected by the event-based tactile sensor back to the human operator. The experiments included a practical application of a soft tissue puncturing tool, and psychophysical test results from ten participants were recorded, to validate efficacy and usability of the system.
PaperID: 132,
Authors: Martin Vonheim Larsen, Kim Mathiassen
Affiliations: Defence Systems Division, Norwegian Defence Research Establishment, Kjeller, Norway
Abstract: In this article, we present a novel method for self-calibrating a pan–tilt–zoom (PTZ) camera system model, specifically suited for long-range multitarget tracking with maneuvering low-cost PTZ cameras. Traditionally, such camera systems cannot provide accurate mappings from pixels to directions in the platform frame due to imprecise pan/tilt measurements or lacking synchronization between the pan/tilt unit and the video stream. Using a direction-only bundle adjustment (BA) incorporating pan/tilt measurements, we calibrate camera intrinsics, rolling shutter characteristics, and pan/tilt mechanics, and obtain clock synchronization between the video stream and pan/tilt telemetry. We call the resulting method pan/tilt camera extrinsic and intrinsic estimation (PTCEE). In a thorough simulation study, we show that the proposed estimation scheme identifies model parameters with subpixel precision across a wide range of camera setups. Leveraging the map of landmarks from the BA, we propose a method for estimating camera orientation in real-time, and demonstrate pixel-level mapping precision on real-world data. Through the proposed calibration and orientation schemes, PTCEE enables high-precision target tracking during camera maneuvers in many low-cost systems, which was previously reserved for high-end systems with specialized hardware.
PaperID: 133,
Authors: Cheng Chen, Hongliang Ren, Hongqiang Wang
Affiliations: Shenzhen Key Laboratory of Intelligent Robotics and Flexible Manufacturing Systems, Southern University of Science and Technology, Shenzhen, China; Department of Electronic Engineering, Shun Hing Institute of Advanced Engineering, The Chinese University of Hong Kong, Hong Kong, SAR, China
Abstract: Various variable stiffness mechanisms have been developed to bestow new capabilities for the robotics community by changing the mechanical behaviors of robots. However, variable stiffness is limited in actuation, response speed, stiffness ratio, and, most importantly, modeling. This article proposes hybrid actuated laminar jamming to outperform individual actuated variable stiffness mechanisms. An analytical model for multilayer laminar jamming that accurately characterizes mechanical behaviors in experiments is first built. Comprehensive parametrical analysis based on this model serves as design guidelines for performance improvements of laminar jamming. Feedforward control further proves the validity of the proposed model and exhibits good controllability, showing response speed as fast as 5 ms. The synergy between electroadhesion and vacuum actuation significantly enhances overall performance, resulting in far greater effects than individual contributions. For instance, the proposed device generates a high stiffness that is almost impossible for individual vacuum or electroadhesion. Moreover, vacuuming increases 23% of the breakdown voltage, which leads to a larger electroadhesion force and, hence, a higher stiffness.
PaperID: 134,
Authors: Trevor Ablett, Oliver Limoyo, Adam Sigal, Affan Jilani, Jonathan Kelly, Kaleem Siddiqi, Francois Robert Hogan, Gregory Dudek
Affiliations: Samsung AI Centre, Montreal, QC, Canada; Space and Terrestrial Autonomous Robotics Systems (STARS) Laboratory and the Robotics Institute (RI), University of Toronto, Toronto, ON, Canada
Abstract: Contact-rich tasks continue to present many challenges for robotic manipulation. In this work, we leverage a multimodal visuotactile sensor within the framework of imitation learning (IL) to perform contact-rich tasks that involve relative motion (e.g., slipping and sliding) between the end-effector and the manipulated object. We introduce two algorithmic contributions, tactile force matching and learned mode switching, as complimentary methods for improving IL. Tactile force matching enhances kinesthetic teaching by reading approximate forces during the demonstration and generating an adapted robot trajectory that recreates the recorded forces. Learned mode switching uses IL to couple visual and tactile sensor modes with the learned motion policy, simplifying the transition from reaching to contacting. We perform robotic manipulation experiments on four door-opening tasks with a variety of observation and algorithm configurations to study the utility of multimodal visuotactile sensing and our proposed improvements. Our results show that the inclusion of force matching raises average policy success rates by 62.5%, visuotactile mode switching by 30.3%, and visuotactile data as a policy input by 42.5%, emphasizing the value of see-through tactile sensing for IL, both for data collection to allow force matching, and for policy execution to enable accurate task feedback.
PaperID: 135,
Authors: Elena Merlo, Marta Lagomarsino, Edoardo Lamon, Arash Ajoudani
Affiliations: Human-Robot Interfaces and Interaction, Istituto Italiano di Tecnologia, Genoa, Italy; Department of Information Engineering and Computer Science, University of Trento, Trento, Italy
Abstract: Observational learning is a promising approach to enable people without expertise in programming to transfer skills to robots in a user-friendly manner, since it mirrors how humans learn new behaviors by observing others. Many existing methods focus on instructing robots to mimic human trajectories, but motion-level strategies often pose challenges in skills generalization across diverse environments. This article proposes a novel framework that allows robots to achieve a higher-level understanding of human-demonstrated manual tasks recorded in RGB videos. By recognizing the task structure and goals, robots generalize what observed to unseen scenarios. We found our task representation on Shannon's Information Theory (IT), which is applied for the first time to manual tasks. IT helps extract the active scene elements and quantify the information shared between hands and objects. We exploit scene graph properties to encode the extracted interaction features in a compact structure and segment the demonstration into blocks, streamlining the generation of behavior trees for robot replicas. Experiments validated the effectiveness of IT to automatically generate robot execution plans from a single human demonstration. In addition, we provide HANDSOME, an open-source dataset of HAND Skills demOnstrated by Multi-subjEcts, to promote further research and evaluation in this field.
PaperID: 136,
Authors: Lijun Han, Yiming Liu, Hesheng Wang
Affiliations: Department of Automation, Key Laboratory of System Control and Information Processing of Ministry of Education, State Key Laboratory of Avionics Integration and Aviation System-of-Systems Synthesis, Shanghai Jiao Tong University, Shanghai, China
Abstract: Pushing is an essential nonprehensile manipulation for robots to achieve complex tasks. Until now, object rigidity remains one of the common assumptions in robotic pushing. To endow robots with the advanced capability of pushing deformable objects, we propose a mathematical model and control method for the planar pushing of deformable objects. Given the robotic end-effector velocity or position input, the model predicts the motion and deformation of the pushed object, which is developed based on the quasi-static finite element analysis with reasonable simplification, considering the contact conditions of nodes with both the operator and the contact surface. By combining the designed model to estimate the state of the object and interactions with the environment, we further propose a method based on model predictive control to realize the pushing control. With a specialized simplified model to accelerate prediction, the controller is solved by iterative linear quadratic regulator with a dynamic weight, which balances the object motion and pushing area adjustment. The accuracy and efficiency of the proposed deformable model are validated by comparing the theoretical results with the experimental ones under different conditions, and the controller is verified by simulation and experiments.
PaperID: 137,
Authors: Yanghong Li, Li Zheng, Yahao Wang, Erbao Dong, Shiwu Zhang
Affiliations: Humanoid Robotics Institute, State Key Laboratory of Precision and Intelligent Chemistry, CAS Key Laboratory of Mechanical Behavior and Design of Materials, Department of Precision Machinery and Precision Instrumentation, University of Science and Technology of China, Hefei, China; Humanoid Robotics Institute, CAS Key Laboratory of Mechanical Behavior and Design of Materials, Department of Precision Machinery and Precision Instrumentation, University of Science and Technology of China, Hefei, China
Abstract: Aiming at the robust force tracking challenge for robots in continuous contact with uncertain environments, a novel adaptive variable impedance control policy based on deep reinforcement learning (DRL) is proposed in this article. The policy includes a neural network feedforward controller and a variable impedance feedback controller. Based on the DRL algorithm, the iterative network feedforward controller explores and prelearns the optimal policy for impedance tuning in simulation scenarios with randomly generated terrain. The converged results are then used as feedforward inputs in the variable impedance feedback controller to improve the force-tracking performance of the robot during contact. A simplified dynamic contact model between the robot and the uncertain environment called the “couch model,” which satisfies the Lipschiz continuity condition, is developed to provide boundary conditions for the safe transfer of capabilities learned in simulation to real robots. Unlike the exhaustive example that relies on the completeness of the learning samples, this article gives theoretical proofs of the stability and convergence of the proposed control policy via Lyapunov’s theorem and contraction mapping principle. The control method proposed in this article is more interpretable and shows higher sample utilization efficiency and generalization ability in simulations and experiments.
PaperID: 138,
Authors: Zuoxue Wang, Pei Jiang, Xiao-Bin Li, Huajun Cao, Xi Vincent Wang, Xiangfei Li, Min Cheng
Affiliations: College of Mechanical and Vehicle Engineering, Chongqing University, Chongqing, China; Department of Production Engineering, KTH Royal Institute of Technology, Stockholm, Sweden; State Key Laboratory of Intelligent Manufacturing Equipment and Technology, School of Mechanical Science and Engineering, Huazhong University of Science and Technology, Wuhan, China
Abstract: Industrial robots (IRs) have considerable energy-saving potential due to their vast application scale and wide range of applications. Although substantial work on the energy consumption (EC) optimization of IRs has emerged, most optimization approaches require prior knowledge of the IRs' dynamic characteristics and the electro-mechanical parameters of their drive systems, which are typically not provided by IR manufacturers. Therefore, this article proposes an EC modeling and optimization method based on the time-scaling technique and custom identification experimental data without joint torque information. Specifically, this article develops an energy characteristic parameter submodel (ECPSM) to formulate the EC resulting from configuration transitions. In addition, theoretical proof demonstrates that all coefficients in the proposed ECPSM can be identified based on the data of a finite number of identification experiments. Building upon the proposed EC model, a bidirectional dynamic programming (BDP) algorithm optimizes the IR's trajectory for energy-saving, while utilizing parallel processing significantly reduces the time required for the optimization process. Experimental results on the KUKA KR60-3 demonstrate that the proposed method achieves an average relative error of 1.59% for predicting the EC of linear scaling trajectories and 6.19% for nonlinear scaled trajectories. Moreover, the BDP-based optimization method dramatically reduces the computational time required to obtain the optimal scaling trajectory and its EC.
PaperID: 139,
Authors: Xinglong Zhang, Wei Pan, Cong Li, Xin Xu, Xiangke Wang, Ronghua Zhang, Dewen Hu
Affiliations: College of Intelligence Science and Technology, National University of Defense Technology, Changsha, China; Department of Computer Science, The University of Manchester, Manchester, U.K.
Abstract: Distributed model predictive control (DMPC) is promising in achieving optimal cooperative control in multirobot systems (MRS). However, real-time DMPC implementation relies on numerical optimization tools to periodically calculate local control sequences online. This process is computationally demanding and lacks scalability for large-scale, nonlinear MRS. This article proposes a novel distributed learning-based predictive control framework for scalable multirobot control. Unlike conventional DMPC methods that calculate open-loop control sequences, our approach centers around a computationally fast and efficient distributed policy learning algorithm that generates explicit closed-loop DMPC policies for MRS without using numerical solvers. The policy learning is executed incrementally and forward in time in each prediction interval through an online distributed actor–critic implementation. The control policies are successively updated in a receding-horizon manner, enabling fast and efficient policy learning with the closed-loop stability guarantee. The learned control policies could be deployed online to MRS with varying robot scales, enhancing scalability and transferability for large-scale MRS. Furthermore, we extend our methodology to address the multirobot safe learning challenge through a force field-inspired policy learning approach. We validate our approach's effectiveness, scalability, and efficiency through extensive experiments on cooperative tasks of large-scale wheeled robots and multirotor drones. Our results demonstrate the rapid learning and deployment of DMPC policies for MRS with scales up to 10 000 units.
PaperID: 140,
Authors: Shuolong Chen, Xingxing Li, Shengyu Li, Yuxuan Zhou, Xiaoteng Yang
Affiliations: School of Geodesy and Geomatics (SGG), Wuhan University, Wuhan, China
Abstract: The integrated inertial system, typically integrating an IMU and an exteroceptive sensor, such as radar, light detection and ranging (LiDAR), and camera, has been widely accepted and applied in modern robotic applications for ego-motion estimation, motion control, or autonomous exploration. To improve system accuracy, robustness, and further usability, both multiple and various sensors are generally resiliently integrated, which benefits the system performance regarding failure tolerance, perception capability, and environment compatibility. For such systems, accurate and consistent spatiotemporal calibration is required to maintain a unique spatiotemporal framework for multisensor fusion. Considering that most existing calibration methods first, are generally oriented to specific integrated inertial systems, second, often focus on spatial-only determination, and third, usually require artificial targets, lacking convenience and usability, we propose iKalibr: a unified targetless spatiotemporal calibration framework for resilient integrated inertial systems, which overcomes the above issues, and enables both accurate and consistent calibration. Altogether four commonly employed sensors are supported in iKalibr currently, namely, IMU, radar, LiDAR, and camera. The proposed method starts with a rigorous and efficient dynamic initialization, where all parameters in the estimator would be accurately recovered. Subsequently, several continuous-time batch optimizations are conducted to refine the initialized parameters toward better states. Sufficient real-world experiments were conducted to verify the feasibility and evaluate the calibration performance of iKalibr. The results demonstrate that iKalibr can achieve accurate resilient spatiotemporal calibration.
PaperID: 141,
Authors: Lanxiang Zheng, Mingxin Wei, Ruidong Mei, Kai Xu, Junlong Huang, Hui Cheng
Affiliations: School of Computer Science and Engineering, Sun Yat-Sen University, Guangzhou, China; School of Artificial Intelligence, Sun Yat-Sen University, Zhuhai, China; School of Systems Science and Engineering, Sun Yat-Sen University, Guangzhou, China; School of Intelligent Systems Engineering, Sun Yat-Sen University, Shenzhen, China
Abstract: The article presents an air-assisted ground robotic autonomous exploration framework, which leverages the high mobility and wide aerial perspective of unmanned aerial vehicles (UAVs) to assist unmanned ground vehicles (UGVs) in detailed exploration, enhancing exploration efficiency and improving the quality of point cloud collection in regions of interest in large-scale, unknown environments. In this framework, the UAV, equipped with an onboard RGB camera, rapidly surveys large unknown areas and generates a bird's eye view (BEV) to identify critical zones for UGV exploration. With prior information about the unexplored area's outline from the real-time shared BEV, the UGV can carry out more efficient and informed exploration from a global perspective. To maximize the utility of this prior information and optimize point cloud collection, a hierarchical exploration strategy and an attention mechanism are incorporated to guide the UGV's focus toward areas requiring detailed mapping, rather than broad, featureless regions. Real-world experiments validate the effectiveness of the framework, demonstrating significant improvements in exploration efficiency and point cloud collection compared to state-of-the-art methods. The results further show that even with a coarse BEV, the UGV's exploration efficiency is greatly enhanced.
PaperID: 142,
Authors: Anqing Duan, Wanli Liuchen, Jinsong Wu, Raffaello Camoriano, Lorenzo Rosasco, David Navarro-Alarcon
Affiliations: Mohamed Bin Zayed University of Artificial Intelligence (MBZUAI), Abu Dhabi, UAE; Department of Mechanical Engineering, The Hong Kong Polytechnic University, Hong Kong; Visual and Multimodal Applied Learning Laboratory, DAUIN, Politecnico di Torino, Turin, Italy; Laboratory for Computational and Statistical Learning (LCSL), Machine Learning Genoa Center (MaLGa), and the Dipartimento di Informatica, Bioingegneria, Robotica e Ingegneria dei Sistemi (DIBRIS), University of Genova, Genoa, Italy
Abstract: The increasing deployment of robots has significantly enhanced the automation levels across a wide and diverse range of industries. This article investigates the automation challenges of laser-based dermatology procedures in the beauty industry. This group of related manipulation tasks involves delivering energy from a cosmetic laser onto the skin with repetitive patterns. To automate this procedure, we propose to use a robotic manipulator and endow it with the dexterity of a skilled dermatology practitioner through a learning-from-demonstration framework. To ensure that the cosmetic laser can properly deliver the energy onto the skin surface of an individual, we develop a novel structured prediction-based imitation learning algorithm with the merit of handling geometric constraints. Notably, our proposed algorithm effectively tackles the imitation challenges associated with quasi-periodic motions, a common feature of many laser-based cosmetic tasks. The conducted real-world experiments illustrate the performance of our robotic beautician in mimicking realistic dermatological procedures. Our new method is shown to not only replicate the rhythmic movements from the provided demonstrations but also to adapt the acquired skills to previously unseen scenarios and subjects.
PaperID: 143,
Authors: Joaquim Ortiz de Haro, Wolfgang Hönig, Valentin N. Hartmann, Marc Toussaint
Affiliations: Machines in Motion Laboratory, New York University, New York, NY, USA; Technical University Berlin, Berlin, Germany
Abstract: Motion planning for robotic systems with complex dynamics is a challenging problem. While recent sampling-based algorithms achieve asymptotic optimality by propagating random control inputs, their empirical convergence rate is often poor, especially in high-dimensional systems such as multirotors. An alternative approach is to first plan with a simplified geometric model and then use trajectory optimization to follow the reference path while accounting for the true dynamics. However, this approach may fail to produce a valid trajectory if the initial guess is not close to a dynamically feasible trajectory. In this article, we present Iterative Discontinuity Bounded A (iDb-A), a novel kinodynamic motion planner that combines search and optimization iteratively. The search step utilizes a finite set of short trajectories (motion primitives) that are interconnected while allowing for a bounded discontinuity between them. The optimization step locally repairs the discontinuities with trajectory optimization. By progressively reducing the allowed discontinuity and incorporating more motion primitives, our algorithm achieves asymptotic optimality with excellent any-time performance. We provide a benchmark of 43 problems across eight different dynamical systems, including different versions of unicycles and multirotors. Compared to state-of-the-art methods, iDb-A consistently solves more problem instances and finds lower-cost solutions more rapidly.
PaperID: 144,
Authors: Xinliang Guo, Zheyu Liu, Vincent Crocher, Ying Tan, Denny Oetomo, Arno H. A. Stienen
Affiliations: Human Robotics Laboratory, Department of Mechanical Engineering, The University of Melbourne, Parkville, VIC, Australia; Department of BioMechanical Engineering, Faculty of Mechanical Engineering, The Delft University of Technology, Delft, The Netherlands
Abstract: Haptic interaction is critical in physical human–robot Interaction (pHRI), given its wide applications in manufacturing, medical and healthcare, and various industry tasks. A stable haptic interface is always needed while the human operator interacts with the robot. Passivity-based approaches have been widely utilized in the control design as a sufficient condition for stability. However, it is a conservative approach which therefore sacrifices performance to maintain stability. This article proposes a novel concept to characterize an ultimately passive system, which can achieve the boundedness of the energy in the steady-state. A so-called ultimately passive controller (UPC) is then proposed. This algorithm switches the system between a nominal mode for keeping desired performance and a conservative mode when needed to remain stable. An experimental evaluation on two robotic systems, one admittance-based and one impedance-based, demonstrates the potential interest of the proposed framework compared to existing approaches. The results demonstrate the possibility of UPC in finding a more aggressive tradeoff between haptic performance and system stability, while still providing a stability guarantee.
PaperID: 145,
Authors: Caitlin Freeman, Arun Niddish Mahendran, Vishesh Vikas
Affiliations: Agile Robotics Lab, University of Alabama, Tuscaloosa, AL, USA
Abstract: Locomotion gaits are fundamental for control of soft terrestrial robots. However, synthesis of these gaits is challenging due to modeling of robot-environment interaction and lack of a mathematical framework. This work presents an environment-centric, data-driven, and fault-tolerant probabilistic model-free control framework that allows for soft multilimb robots to learn from their environment and synthesize diverse sets of locomotion gaits for realizing open-loop control. Here, discretization of factors dominating robot-environment interactions enables an environment-specific graphical representation where the edges encode experimental locomotion data corresponding to the robot motion primitives. In this graph, locomotion gaits are defined as simple cycles that are transformation invariant, i.e., the locomotion is independent of the starting vertex of these periodic cycles. Gait synthesis, the problem of finding optimal locomotion gaits for a given substrate, is formulated as binary integer linear programming problems with a linearized cost function, linear constraints, and iterative simple cycle detection. Experimentally, gaits are synthesized for varying robot-environment interactions. Variables include robot morphology—three-limb and four-limb robots, TerreSoRo-III and TerreSoRo-IV; substrate—rubber mat, whiteboard and carpet; and actuator functionality—simulated loss of robot limb actuation. On an average, gait synthesis improves the translation and rotation speeds by 82% and 97%, respectively. The results highlight that data-driven methods are vital to soft robot locomotion control due to complex robot-environment interactions and simulation-to-reality gaps, particularly when biological analogues are unavailable.
PaperID: 146,
Authors: Charles Dawson, Anjali Parashar, Chuchu Fan
Affiliations: Department of Aeronautics and Astronautics, MIT, Cambridge, MA, USA; Department of Mechanical Engineering, MIT, Cambridge, MA, USA
Abstract: Before deploying autonomous systems in safety-critical applications, we must be able to understand and verify the safety of these systems. For cases where the risk or cost of real-world testing is prohibitive, we propose a simulation-based framework for 1) predicting ways in which an autonomous system is likely to fail and 2) automatically adjusting the system's design and control policy to preemptively mitigate those failures. Existing tools for failure prediction struggle to search over high-dimensional environmental parameters, cannot efficiently handle end-to-end testing for systems with vision in the loop, and provide little guidance on how to mitigate failures once they are discovered. We approach this problem through the lens of approximate Bayesian inference, using differentiable simulation and rendering for efficient failure case prediction and repair (and providing a gradient-free version of our algorithm for cases where a differentiable simulator is not available). We include a theoretical and empirical evaluation of the tradeoffs between gradient-based and gradient-free methods, applying our approach to a range of robotics and control problems, including optimizing search patterns for robot swarms, UAV formation control, and robust network control. Compared to optimization-based falsification methods, our method predicts a more diverse, representative set of failure modes, and we find that our use of differentiable simulation yields solutions that have up to 10x lower cost and requires up to 2x fewer iterations to converge relative to gradient-free techniques. In hardware experiments, we find that repairing control policies using our method leads to a 5x robustness improvement.
PaperID: 147,
Authors: Antoine N. André, Fabio Morbidi, Guillaume Caron
Affiliations: CNRS- AIST Joint Robotics Laboratory (JRL), IRL, Tsukuba, Japan; MIS laboratory, University of Picardie Jules Verne, Amiens, France; CNRS-AIST Joint Robotics Laboratory (JRL), IRL, Tsukuba, Japan
Abstract: In this article, we present a new spherical image representation, called uniform spherical mapping of omnidirectional images (UniphorM), and show its strong potential in robotic vision. UniphorM provides an accurate and distortion-free representation of a 360-degree image, by relying on multiple subdivisions of an icosahedron and its associated Voronoi diagrams. The geometric mapping procedure is described in detail, and the tradeoff between pixel accuracy and computational complexity is investigated. To demonstrate the benefits of UniphorM in real-world problems, we applied it to direct visual attitude estimation and visual place recognition (VPR), by considering dual-fisheye images captured by a camera mounted on multiple robotic platforms. In the experiments, we measured the impact of the number of subdivision levels of the icosahedron on the attitude estimation error, time efficiency, and size of convergence domain of an existing visual gyroscope, using UniphorM and three competing mapping algorithms. A similar evaluation procedure was carried out for VPR. Finally, two new omnidirectional image datasets, one recorded with a hexacopter, called SVMIS+, the other based on the Mapillary platform, have been created and released for the entire research community.
PaperID: 148,
Authors: Marta Colakovic-Benceric, Juraj Persic, Ivan Markovic, Ivan Petrovic
Affiliations: Faculty of Electrical Engineering and Computing,Laboratory for Autonomous Systems and Mobile Robotics, University of Zagreb, Zagreb, Croatia; Calirad, Zagreb, Croatia
Abstract: The operational reliability of an autonomous robot depends crucially on extrinsic sensor calibration as a prerequisite for precise and accurate data fusion. Exploring the calibration of unscaled sensors (e.g., monocular cameras) and the effective utilization of uncertainties are difficult and often overlooked. The development of a solution for the simultaneous calibration of hand-eye sensors and scale estimation based on the Gauss–Helmert model aims to utilize the valuable information contained in the uncertainty of odometry. In this work, we propose a versatile and robust solution for batch calibration based on the analytical on-manifold approach for estimation. The versatility of our method is demonstrated by its ability to calibrate multiple unscaled and metric-scaled sensors while dealing with odometry failures and reinitializations. Importantly, all estimated parameters are provided with their corresponding uncertainties. The validation of our method and its comparison with five competing state-of-the-art calibration methods in both simulations and real-world experiments show its superior accuracy, with particularly promising results observed in high-noise scenarios.
PaperID: 149,
Authors: Aramis Augusto Bonzini, Lucia Seminara, Simone Macciò, Alessandro Carfì, Lorenzo Jamone
Affiliations: Advanced Robotics at Queen Mary (ARQ), School of Engineering and Materials Science, Queen Mary University of London, London, U.K.; Department of Naval, Electrical, Electronic, and Telecommunications Engineering (DITEN), University of Genoa, Genova, Italy; Department of Engineering (DIBRIS), University of Genoa, Genova, Italy
Abstract: Haptic robotic exploration aims to control the movements of a robot with the objective of touching an object and retrieving physical information about it. In this work, we present an innovative exploration strategy to simultaneously detect symmetries in a 3-D object and use this information to enhance shape estimation. This is achieved by leveraging a novel formulation of Gaussian process models that allows the modeling of symmetric surfaces. Our procedure does not assume any prior knowledge about the object, neither about its shape nor about the presence and type of symmetry, necessitating only an approximate estimate of the size and boundaries (bounding box). We report experimental results both in simulation and in the real world, showing that using symmetric models leads to a reduction in shape estimation error, exploration time, and in the number of physical contacts performed by a robot when exploring objects that have symmetries.
PaperID: 150,
Authors: Yanggang Feng, Xingyu Hu, Yuebing Li, Ke Ma, Jiaxin Ren, Zhihao Zhou, Fuzhen Yuan, Yan Huang, Liu Wang, Qining Wang, Wuxiang Zhang, Xilun Ding
Affiliations: School of Mechanical Engineering and Automation, Beihang University, Beijing, China; College of Engineering, Peking University, Beijing, China; Department of Sports Medicine, Peking University Third Hospital, Institute of Sports Medicine of Peking University, Beijing, China; School of Mechatronical Engineering, Beijing Institute of Technology, Beijing, China; Department of Modern Mechanics, University of Science and Technology of China, Hefei, China
Abstract: Knee pain is prevalent in over 20% of the population, limiting the mobility of those affected. In turn, isokinetic dynamometers and robots have been used to facilitate rehabilitation for those still capable of ambulation. However, there are at most only a few wearable robots capable of delivering isokinetic training for bedridden patients. Here, we developed a wearable robot that provides bedside isokinetic training by utilizing a variable stiffness actuator and dynamic energy regeneration. The efficacy of this device was validated in a study involving six subjects with debilitating knee injuries. During two courses of rehabilitation over a total of three weeks, the average peak torque, average torque, and average work produced by their affected knees increased significantly by 81.0%, 101.4%, and 117.6%, respectively. Furthermore, the device's energy regeneration features were found capable of extending its operating time to 198 days under normal usage, representing a 57.8% increase over the same device without regeneration. These results suggest potential methodologies for delivering isokinetic joint rehabilitation to bedridden patients in areas with limited infrastructure.
PaperID: 151,
Authors: Yijun Yuan, Michael Bleier, Andreas Nüchter
Affiliations: Tsinghua University, Beijing, China; Julius-Maximilians-Universität Würzburg, Würzburg, Germany
Abstract: In this article, we present SceneFactory, a workflow-centric and unified framework for incremental scene modeling that conveniently supports a wide range of applications, such as (unposed and/or uncalibrated) multiview depth estimation, LiDAR completion, (dense) RGB-D/RGB-LiDAR (RGB-L)/Mono/Depth-only reconstruction, and simultaneous localization and mapping (SLAM). The workflow-centric design uses multiple blocks as the basis for constructing different production lines. The supported applications, i.e., productions avoid redundancy in their designs. Thus, the focus is placed on each block itself for independent expansion. To support all input combinations, our implementation consists of four building blocks that form SceneFactory: first, tracking, second, flexion, third, depth estimation, and fourth, scene reconstruction. The tracking block is based on Mono SLAM and is extended to support RGB-D and RGB-L inputs. Flexion is used to convert the depth image (untrackable) into a trackable image. For general-purpose depth estimation, we propose an unposed and uncalibrated multiview depth estimation model (U^2-MVD) to estimate dense geometry. U^2-MVD exploits dense bundle adjustment to solve for poses, intrinsics, and inverse depth. A semantic-aware ScaleCov step is then introduced to complete the multiview depth. Relying on U^2-MVD, SceneFactory both supports user-friendly 3-D creation (with just images) and bridges the applications of Dense RGB-D and Dense Mono. For high-quality surface and color reconstruction, we propose dual-purpose multiresolutional neural points for the first surface accessible surface color field design, where we introduce improved point rasterization for point cloud-based surface query. We implement and experiment with SceneFactory to demonstrate its broad applicability and high flexibility. Its quality also competes or exceeds the tightly-coupled state of the art approaches in all tasks.
PaperID: 152,
Authors: Tiandong Zhang, Rui Wang, Qiyuan Cao, Shaowei Cui, Gang Zheng, Shuo Wang
Affiliations: State Key Laboratory of Multimodal Artificial Intelligence Systems, Institute of Automation, Chinese Academy of Sciences, Beijing, China; Centrale Lille, CRIStAL - Centre de Recherche en Informatique Signal et Automatique de Lille, University of Lille, Lille, France
Abstract: This article presents a novel vision-based artificial lateral line (ALL) sensor, FlowSight, enhancing the perception capabilities of underwater robots. Through an autonomous vision system, FlowSight allows for simultaneous sensing the speed and direction of local water flow without relying on external auxiliary equipment. Inspired by the lateral line neuromast of fish, a flexible bionic tentacle is designed to sense water flow. Deformation and motion characteristics of the tentacle are modeled and analyzed using bidirectional fluid-structure interaction (FSI) simulation. Upon contact with water flow, the tentacle converts water flow information into elastic deformation information, which is captured and processed into an image sequence by the autonomous vision system. Subsequently, a water flow perception method based on deep neural networks is proposed to estimate the flow speed and direction from the captured image sequence. The perception network is trained and tested using data collected from practical experiments conducted in a controllable swim tunnel. Finally, the FlowSight sensor is integrated into the bionic underwater robot RoboDact, and a closed-loop motion control experiment based on water flow perception is conducted. Experiments conducted in the swim tunnel and water pool demonstrate the feasibility and effectiveness of FlowSight sensor and the water flow perception method.
PaperID: 153,
Authors: Bing Chen, Xiang Ni, Lei Zhou, Bin Zi, Eric Li, Dan Zhang
Affiliations: School of Mechanical Engineering, Hefei University of Technology, Hefei, China; School of Mechano-Electronic Engineering, Xidian University, Xi'an, China; School of Computing, Engineering & Digital Technologies, Teesside University, Middlesbrough, U.K.; Faculty of Engineering, The Hong Kong Polytechnic University, Hong Kong
Abstract: Frequent and high-load manual material handling (MMH) tasks often cause back injuries to the workers, and back-support exoskeletons are developed for individuals with MMH tasks. However, these exoskeletons usually cannot adapt well to the movements of the wearer's spine. This article introduces a new bioinspired five degree of freedom (DOF) origami, and via mechanical design, a unique rigid-flexible coupled bioinspired origami mechanism is proposed. This origami mechanism is compact and lightweight, and it has stable kinematic behaviors. With the designed origami mechanisms, a novel active origami-based robotic spine assistive exoskeleton (OSAE) is developed to assist individuals with MMH tasks during the symmetric and asymmetric lifting. The OSAE is actuated by a cable-driven module through an underactuated spine module that consists of seven origami mechanisms. With the designed spine module, the OSAE can adapt well to the wearer's spine motions during MMH tasks. Modeling of the five-DOF origami is described, and an adaptive control strategy is proposed for the exoskeleton to adapt to different lifting methods and objects with different weights. The experimental results demonstrate the effectiveness of the proposed OSAE. During the symmetric lifting of a 10-kg object, a reduction of 41.28% of the average muscle activity of the wearer's lumbar erector spinae muscle (LES) is observed, and reductions of 30.15% and 39.54% of the average muscle activities of the wearer's left and right LES are observed, respectively, during the asymmetric lifting of a 10-kg object.
PaperID: 154,
Authors: Jingtao Tang, Zining Mao, Hang Ma
Affiliations: School of Computing Science, Simon Fraser University, Burnaby, BC, Canada
Abstract: In this article, we study multirobot coverage path planning (MCPP) on a four-neighbor 2-D grid G, which aims to compute paths for multiple robots to cover all cells of G. Traditional approaches are limited as they first compute coverage trees on a quadrant coarsened grid \mathcal H and then employ the spanning tree coverage (STC) paradigm to generate paths on G, making them inapplicable to grids with partially obstructed 2 × 2 blocks. To address this limitation, we reformulate the problem directly on G, revolutionizing grid-based MCPP solving and establishing new NP-hardness results. We introduce extended STC (ESTC), a novel paradigm that extends STC to ensure complete coverage with bounded suboptimality, even when \mathcal H includes partially obstructed blocks. Furthermore, we present LS-MCPP, a new algorithmic framework that integrates ESTC with three novel types of neighborhood operators within a local search strategy to optimize coverage paths directly on G. Unlike prior grid-based MCPP work, our approach also incorporates a versatile postprocessing procedure that applies multiagent path finding (MAPF) techniques to MCPP for the first time, enabling a fusion of these two important fields in multirobot coordination. This procedure effectively resolves inter-robot conflicts and accommodates turning costs by solving an MAPF variant, making our MCPP solutions more practical for real-world applications. Extensive experiments demonstrate that our approach significantly improves solution quality and efficiency, managing up to 100 robots on grids as large as \text256 × \text256 within minutes of runtime. Validation with physical robots confirms the feasibility of our solutions under real-world conditions.
PaperID: 155,
Authors: Puze Liu, Haitham Bou-Ammar, Jan Peters, Davide Tateo
Affiliations: Intelligent Autonomous Systems Group, Technical University of Darmstadt, Darmstadt, Germany; Huawei R&D London, Cambridge, U.K.
Abstract: Integrating learning-based techniques, especially reinforcement learning, into robotics is promising for solving complex problems in unstructured environments. Most of the existing approaches rely on training in carefully calibrated simulators before being deployed on real robots, often without real-world fine-tuning. While effective in controlled settings, this framework falls short in applications where precise simulation is unavailable or the environment is too complex to model. Instead, on-robot learning, which learns by interacting directly with the real world, offers a promising alternative. One major problem for on-robot reinforcement learning is ensuring safety, as uncontrolled exploration can cause catastrophic damage to the robot or the environment. Indeed, safety specifications, often represented as constraints, can be complex and nonlinear, making safety challenging to guarantee in learning systems. In this article, we show how we can impose complex safety constraints on learning-based robotics systems in a principled manner, both from theoretical and practical points of view. Our approach is based on the concept of the constraint manifold, representing the set of safe robot configurations. Exploiting differential geometry techniques, i.e., the tangent space, we can construct a safe action space, allowing learning agents to sample arbitrary actions while ensuring safety. We demonstrate the method's effectiveness in a real-world robot air hockey task, showing that our method can handle high-dimensional tasks with complex constraints.
PaperID: 156,
Authors: Zirui Xu, Sandilya Sai Garimella, Vasileios Tzoumas
Affiliations: Department of Aerospace Engineering, University of Michigan, Ann Arbor, MI, USA; Department of Robotics, University of Michigan, Ann Arbor, MI, USA
Abstract: In this article, we provide a communication- and computation-efficient method for distributed submodular optimization in robot mesh networks. Submodularity is a property of diminishing returns that arises in active information gathering such as mapping, surveillance, and target tracking. Our method, resource-aware distributed greedy (RAG), introduces a new distributed optimization paradigm that enables scalable and near-optimal action coordination. To this end, RAG requires each robot to make decisions based only on information received from and about their neighbors. In contrast, the current paradigms allow the relay of information about all robots across the network. As a result, RAG’s decision-time scales linearly with the network size, while state-of-the-art near-optimal submodular optimization algorithms scale cubically. We also characterize how the designed mesh-network topology affects RAG’s approximation performance. Our analysis implies that sparser networks favor scalability without proportionally compromising approximation performance: while RAG’s decision-time scales linearly with network size, the gain in approximation performance scales sublinearly. We demonstrate RAG’s performance in simulated scenarios of area detection with up to 45 robots, simulating realistic robot-to-robot (r2r) communication speeds such as the 0.25 Mb/s speed of the Digi XBee 3 Zigbee 3.0. In the simulations, RAG enables real-time planning, up to three orders of magnitude faster than competitive near-optimal algorithms, while also achieving superior mean coverage performance. To enable the simulations, we extend the high-fidelity and photo-realistic simulator AirSim by integrating a scalable collaborative autonomy pipeline to tens of robots and simulating r2r communication delays.
PaperID: 157,
Authors: Xingxu Li, Yiheng Han, Nan Ma, Yong-Jin Liu, Jia Pan, Shun Yang, Siyi Zheng
Affiliations: School of Information Science and Technology, Beijing University of Technology, Beijing, China; Department of Computer Science and Technology, Tsinghua University, Beijing, China; Department of Computer Science, University of Hong Kong, Hong Kong, SAR, China; Beijing AIForce Technology Company Ltd., Beijing, China
Abstract: Using robots for tomato truss harvesting represents a promising approach to agricultural production. However, incomplete acquisition of perception information and clumsy operations often results in low harvest success rates or crop damage. To addressthis issue, we designed a new method for tomato truss perception, an autonomous harvesting method, and a novel circular rotary cutting end-effector. The robot performs object detection and keypoint detection on tomato trusses using the proposed top–down fusion network, making decisions on suitable targets for harvesting based on phenotyping and pose estimation. The designed end-effector moves gradually from the bottom up to wrap around the tomato truss, cutting the peduncle to complete the harvest. Experiments conducted in real-world scenarios for robotic perception and autonomous harvesting of tomato trusses show that the proposed method increases accuracy by up to 11.42% and 22.29% for complete and limited dataset conditions, compared to baseline models. Furthermore, we have implemented an automatic tomato harvesting system based on TDFNet, which reaches an average harvest success rate of 89.58% in the greenhouse.
PaperID: 158,
Authors: Jialiang Hou, Xin Zhou, Neng Pan, Ang Li, Yuxiang Guan, Chao Xu, Zhongxue Gan, Fei Gao
Affiliations: Academy for Engineering and Technology, Fudan University, Shanghai, China; Institute of Cyber-Systems and Control, Zhejiang University, Hangzhou, China; Huzhou Institute of Zhejiang University, Huzhou, China
Abstract: Achieving large-scale aerial swarms is challenging due to the inherent contradictions in balancing computational efficiency and scalability. This article introduces primitive-swarm, an ultra-lightweight and scalable planner designed specifically for large-scale autonomous aerial swarms. The proposed approach adopts a decentralized and asynchronous replanning strategy. Within it is a novel motion primitive library consisting of time-optimal and dynamically feasible trajectories. They are generated utilizing a novel time-optimal path parameterization algorithm based on reachability analysis. Then, a rapid collision checking mechanism is developed by associating the motion primitives with the discrete surrounding space according to conflicts. By considering both spatial and temporal conflicts, the mechanism handles robot-obstacle and robot–robot collisions simultaneously. Then, during a replanning process, each robot selects the safe and minimum cost trajectory from the library based on user-defined requirements. Both the time-optimal motion primitive library and the occupancy information are computed offline, turning a time-consuming optimization problem into a linear-complexity selection problem. This enables the planner to comprehensively explore the nonconvex, discontinuous 3-D safe space filled with numerous obstacles and robots, effectively identifying the best hidden path. Benchmark comparisons demonstrate that our method achieves the shortest flight time and traveled distance with a computation time of less than 1 ms in dense environments. Super large-scale swarm simulations, involving up to 1000 robots, running in real time, verify the scalability of our method. Real-world experiments validate the feasibility and robustness of our approach. The code will be released to foster community collaboration.
PaperID: 159,
Authors: Ruoyu Lin, Soobum Kim, Magnus Egerstedt
Affiliations: Department of Electrical Engineering and Computer Science, University of California, Irvine, CA, USA; School of Interactive Computing, Georgia Institute of Technology, Atlanta, GA, USA
Abstract: Inspired by common features found in collaborative behaviors in nature, we investigate a general collaborative pursuit framework enabling heterogeneous multi-robot systems to adapt to dynamic environments and diverse tasks. A class of augmented Fokker–Planck equations is formulated to characterize dynamic environmental conditions, and the resulting time-varying density functions drive a novel coverage-based controller, with provable stability properties, for the participating robots to perform tasks in real time. The developed framework is decentralized and incorporates heterogeneity among different robots in task suitability, relative performance in a specific task, and safe operating regions. To demonstrate its adaptivity and effectiveness, the framework is implemented across four experimental applications ranging from multi-robot coordination to collaboration, namely forest firefighting, pursuit–evasion, monitoring of various environmental phenomena, and phoretic interactions.
PaperID: 160,
Authors: Shohei Wakayama, Alberto Candela, Paul O. Hayne, Nisar R. Ahmed
Affiliations: Smead Aerospace Engineering Sciences Department, University of Colorado Boulder, Boulder, CO, USA; Jet Propulsion Laboratory, California Institute of Technology, Pasadena, CA, USA; Astrophysical and Planetary Sciences Department, University of Colorado Boulder, CO, USA
Abstract: Autonomous selection of optimal options for data collection from multiple alternatives is challenging in uncertain environments. When secondary information about options is accessible, such problems can be framed as contextual multiarmed bandits (CMABs). Neuroinspired active inference (AIF) has gained interest for its ability to balance exploration and exploitation using the expected free energy objective function. Unlike previous studies that showed the effectiveness of AIF-based strategy for CMABs using synthetic data, this study aims to apply AIF to realistic scenarios, using a simulated mineralogical survey site selection problem. Hyperspectral data from the next generation airborne visible–infrared imaging spectrometer at Cuprite, Nevada, serves as contextual information for predicting outcome probabilities, while geologists’ mineral labels represent outcomes. Monte Carlo simulations assess the robustness of AIF against changing expert preferences. Results show AIF requires fewer iterations than standard bandit approaches with real-world noisy and biased data, and performs better when outcome preferences vary online by adapting the selection strategy to align with expert shifts.
PaperID: 161,
Authors: Biao Wu, Chaoyi Huang, Xiangru Li, Jiahao Xu, Sicong Liu, James Lam, Zheng Wang, Jian S. Dai
Affiliations: Department of Mechanical and Energy Engineering, Shenzhen Key Laboratory of Intelligent Robotics and Flexible Manufacturing Systems, Southern University of Science and Technology, Shenzhen, China; Department of Mechanical Engineering, The University of Hong Kong, Pokfulam, Hong Kong SAR; Department of Mechanical and Aerospace Engineering, The Hong Kong University of Science and Technology, Kowloon, Hong Kong SAR; Sino-German College of Intelligent Manufacturing, Shenzhen Technology University, Pingshan, China; Wisson Robotics, Shenzhen, China
Abstract: With the vast demand in marine development, robotic fish show promising potential in underwater exploration for their high-performance propulsion ability. However, fish-inspired robots are yet to utilize the structural flexibility of rhythmic actuation such as bony fish (Osteichthyes). The Body and Caudal Fin (BCF) locomotion in fish optimizes the use of muscle power and body flexibility by synchronizing muscle activation with the undulating-oscillatory tail-flapping, such as Thunniform, while robotic fish are primarily designed as motion trackers rather than as efficient swimmers. In this article, we propose a power allocation strategy (PAS) that imitates muscle rhythmic actuation, which increases the flapping amplitude by the coupling of the peduncle motion and the tail deformation. Inspired by this peduncle-tail mechanism, we developed a direct-drive fish robot (DDRFishBot). The DDRFishBot is enhanced by our developed PAS in tail-elastic potential energy release by 228%, in propulsion by 45.6%, and in efficiency coefficient by 16.3%. This study establishes the performance enhancement principle of exploiting tail flexibility through a simple scotch yoke mechanism, expanding the performance space of fish-inspired tail-flapping swimming robot.
PaperID: 162,
Authors: Shengyin Wang, Matteo Leonetti, Mehmet Remzi Dogar
Affiliations: School of Computer Science, University of Leeds, Leeds, U.K.; Department of Informatics, King’s College London, London, U.K.
Abstract: Motion planning for deformable object manipulation has been a challenge for a long time in robotics due to its high computational cost. In this work, we propose to mitigate this cost by limiting the number of picking points on a deformable object within the action space and simplifying the dynamics model. We do this first by identifying a minimal geometric model that closely approximates the original model at the goal state; specifically, we implement this general approach for 1-D linear deformable objects (e.g., ropes) using a piece-wise line-fitted model, and for 2-D surface deformable objects (e.g., cloth) using a mesh-simplified model. Then a small number of key particles are extracted as the pickable points in the action space which are sufficient to represent and reach the given goal. In addition, a simplified dynamics model is constructed based on the simplified geometric model, containing much fewer particles and thus being much faster to simulate than the original dynamics model, albeit with some loss of precision. We further refine this model iteratively by adding more details from the actually achieved final state of the original model until a satisfactory trajectory is generated. Extensive simulation experiments are conducted on a set of representative tasks for ropes and cloth, which show a significant decrease in time cost while achieving similar or better trajectory costs. Finally, we establish a closed-loop system of perception, planning, and control with a real robot for cloth folding, which validates the effectiveness of our proposed method.
PaperID: 163,
Authors: Zhanteng Xie, Philip M. Dames
Affiliations: Department of Mechanical Engineering, Temple University, Philadelphia, PA, USA
Abstract: This article presents a family of Stochastic Cartographic Occupancy Prediction Engines that enable mobile robots to predict the future states of complex dynamic environments. They do this by accounting for the motion of the robot itself, the motion of dynamic objects, and the geometry of static objects in the scene, and they generate a range of possible future states of the environment. These prediction engines are software-optimized for real-time performance for navigation in crowded dynamic scenes, achieving up to 89 times faster inference speed and 8 times less memory usage than other state-of-the-art engines. Three simulated and real-world datasets collected by different robot models are used to demonstrate that these proposed prediction algorithms are able to achieve more accurate and robust stochastic prediction performance than other algorithms. Furthermore, a series of simulation and hardware navigation experiments demonstrate that the proposed predictive uncertainty-aware navigation framework with these stochastic prediction engines is able to improve the safe navigation performance of current state-of-the-art model- and learning-based control policies.
PaperID: 164,
Authors: Mingrui Liu, Xingxing Zuo, Renlang Huang, Minglei Zhao, Jiming Chen, Liang Li
Affiliations: College of Control Science and Engineering, Zhejiang University, Hangzhou, China; Department of Robotics, Mohamed Bin Zayed University of Artificial Intelligence, Abu Dhabi, UAE
Abstract: This work presents a visual odometry (VO) system that leverages image edge features. Edges are spatially expressive cues commonly present across diverse environments, offering rich textural and structural information. However, existing edge-based VO methods often fail to fully exploit this potential. To this end, we introduce a novel feature representation termed organized edges, which transforms disjoint edge pixels into sequentialized clusters, enabling more effective retention and utilization of the underlying textural and structural information. Another nice property of this formulation is that organized edges can perform edge-level association across multiple frames, enabling the establishment of a covisibility graph. To achieve precise and efficient pose estimation, we propose a range of particularly designed tracking and joint optimization methods based on the characteristics of organized edges. For tracking, we formulate edge-wise rather than pixel-wise residuals to achieve robust and accurate interframe registration. For joint optimization, we introduce a novel shape-preserving edge-fitting method and an organized edge-based bundle adjustment (BA) approach, which decomposes the traditional BA problem into fitting and registration to preserve the structural integrity. Based on these novel techniques, we develop a complete VO system that exclusively employs organized edge features, achieving efficient tracking and precise local mapping. Extensive experiments demonstrate its accuracy and robustness in indoor environments, outperforming or achieving comparable performance to state-of-the-art methods.
PaperID: 165,
Authors: Michele Antonazzi, Matteo Alberti, Alex Bassot, Matteo Luperto, Nicola Basilico
Affiliations: Department of Computer Science, University of Milan, Milano, Italy
Abstract: Cloud robotics allows low-power robots to perform computationally intensive inference tasks by offloading them to the cloud, raising privacy concerns when transmitting sensitive images. Although end-to-end encryption secures data in transit, it does not prevent misuse by inquisitive third-party services since data must be decrypted for processing. This article tackles these privacy issues in cloud-based object detection tasks for service robots. We propose a cotrained encoder-decoder architecture that retains only task-specific features while obfuscating sensitive information, utilizing a novel weak loss mechanism with proposal selection for privacy preservation. A theoretical analysis of the problem is provided, along with an evaluation of the tradeoff between detection accuracy and privacy preservation through extensive experiments on public datasets and a real robot.
PaperID: 166,
Authors: Chunyu Li, Mengfan He, Chao Chen, Jiacheng Liu, Xu Lyu, Guoquan Huang, Ziyang Meng
Affiliations: Department of Precision Instrument, Tsinghua University, Beijing, China; Autonomous Aerial Vehicle Lab, Meituan, Beijing, China; Department of Mechanical Engineering, Department of Computer and Information Sciences, University of Delaware, Newark, DE, USA
Abstract: This article introduces GeoVINS, a vision-based navigation framework designed for large-scale global state estimation. By utilizing geographic information from satellite orthoimagery, GeoVINS tackles scale ambiguity and accumulative drift problems inherent in standard visual-inertial simultaneous localization and mapping systems, and provides accurate, robust, and real-time global localization. In particular, to address the challenge of memory explosion issue for large-scale localization, we propose a novel aerial “classify-then-retrieve” aerial visual place recognition (VPR) approach, where geographic locations can be efficiently identified and memory usage can be reduced by 3 to 4 orders of magnitude compared to classical retrieval-based approaches. In addition, a hierarchical geographic data association scheme, enhanced by state-of-the-art deep learning-based feature matching, guarantees high efficiency and robustness against variations in appearance and viewpoint. Relying on the obtained three-dimensional (3-D) geographic information, GeoVINS achieves efficient global state initialization and precise motion tracking. To address GPU limitations in embedded devices, a collaborative CPU–GPU utilization approach is proposed, seamlessly integrating asynchronous global information to eliminate accumulative errors. Relying solely on satellite imagery and without requiring any other prior information, GeoVINS achieves rapid place recognition in unseen environments of city-scale (e.g., 2500 \textkm^2) on an embedded computing device with an inference time of 43 \,\mathrmm\mathrms, and performs state estimation at a frequency of 25 \,\mathrmHz. The system enables autonomous aerial vehicle navigation as an alternative to Global Navigation Satellite System.
PaperID: 167,
Authors: Yuanfei Lin, Zekun Xing, Xuyuan Han, Matthias Althoff
Affiliations: Department of Computer Engineering, Technical University of Munich, Garching, Germany
Abstract: Complying with traffic rules is challenging for automated vehicles, as numerous rules need to be considered simultaneously. If a planned trajectory violates traffic rules, it is common to replan a new trajectory from scratch. We instead propose a trajectory repair technique to save computation time. By coupling satisfiability modulo theories with set-based reachability analysis, we determine if and in what manner the initial trajectory can be repaired. Experiments in high-fidelity simulators and in the real world demonstrate the benefits of our proposed approach in various scenarios. Even in complex environments with intricate rules, we efficiently and reliably repair rule-violating trajectories, enabling automated vehicles to swiftly resume legally safe operation in real time.
PaperID: 168,
Authors: Bin-Bin Hu, Weijia Yao, Yanxin Zhou, Henglai Wei, Chen Lv
Affiliations: School of Mechanical and Aerospace Engineering, Nanyang Technological University, Singapore; School of Robotics, Hunan University, Hunan, China; School of Transportation Science and Engineering, Beihang University, Beijing, China
Abstract: Reducing undesirable path crossings among trajectories of different robots is vital in multirobot navigation missions, which not only reduces detours and conflict scenarios, but also enhances navigation efficiency and boosts productivity. Despite recent progress in multirobot path-crossing-minimal (MPCM) navigation, the majority of approaches depend on the minimal squared-distance reassignment of suitable desired points to robots directly. However, if obstacles occupy the passing space, calculating the actual robot-point distances becomes complex or intractable, which may render the MPCM navigation in obstacle environments inefficient or even infeasible. In this article, the concurrent-allocation task execution (CATE) algorithm is presented to address this problem (i.e., MPCM navigation in obstacle environments). First, the path-crossing-related elements in terms of, first, robot allocation, second, desired-point convergence, and first, collision and obstacle avoidance are encoded into integer and control barrier function (CBF) constraints. Then, the proposed constraints are used in an online constrained optimization framework, which implicitly yet effectively minimizes the possible path crossings and trajectory length in obstacle environments by minimizing the desired point allocation cost and slack variables in CBF constraints simultaneously. In this way, the MPCM navigation in obstacle environments can be achieved with flexible spatial orderings. Note that the feasibility of solutions and the asymptotic convergence property of the proposed CATE algorithm in obstacle environments are both guaranteed, and the calculation burden is also reduced by concurrently calculating the optimal allocation and the control input directly without the path planning process. Finally, extensive simulations and experiments are conducted to validate that the CATE algorithm, first, outperforms the existing state-of-the-art baselines in terms of feasibility and efficiency in obstacle environments, second, is effective in environments with dynamic obstacles and is adaptable for performing various navigation tasks in 2-D and 3-D, third, demonstrates its efficacy and practicality by 2-D experiments with a multi-autonomous mobile robot (AMR) onboard navigation system, and, first, provides a possible solution to evade deadlocks and pass through a narrow gap.
PaperID: 169,
Authors: Alvaro Calvo, Jesús Capitán
Affiliations: Multirobot and Control Systems Group, Universidad de Sevilla, Seville, Spain
Abstract: In this article, we present a framework for multirobot task allocation (MRTA) in heterogeneous teams performing long-endurance missions in dynamic scenarios. Given the limited battery of robots, especially for aerial vehicles, we allow for robot recharges and the possibility of fragmenting and/or relaying certain tasks. We also address tasks that must be performed by a coalition of robots in a coordinated manner. Given these features, we introduce a new class of heterogeneous MRTA problems, which we analyze theoretically and optimally formulate as a mixed-integer linear program (MILP). We then contribute a heuristic algorithm to compute approximate solutions and integrate it into a mission planning and execution architecture capable of reacting to unexpected events by repairing or recomputing plans online. Our experimental results show the relevance of our newly formulated problem in a realistic use case for inspection with aerial robots. We assess the performance of our heuristic solver in comparison with other variants and with exact optimal solutions in small-scale scenarios. In addition, we evaluate the ability of our replanning framework to repair plans online.
PaperID: 170,
Authors: Xiaoshan Bai, Baode Li, Jian-Qiang Li, Zongze Wu, Weidong Zhang, Shuzhi Sam Ge
Affiliations: College of Mechatronics and Control Engineering, Shenzhen University, Shenzhen, China; School of Artificial Intelligence, National Engineering Laboratory for Big Data System Computing Technology, Shenzhen University, Shenzhen, China; College of Mechatronics and Control Engineering, Guangdong Laboratory of Artificial Intelligence and Digital Economy (SZ), Shenzhen University, Shenzhen, China; Department of Automation, Shanghai Jiao Tong University, Shanghai, China; Department of Electrical and Computer Engineering, National University of Singapore, Singapore
Abstract: As the demand for efficient parcel delivery continues to grow in the logistics industry, optimizing multirobot task assignment has become crucial for enhancing overall delivery performance. This article addresses the precedence-constrained multitruck multidrone package delivery task assignment problem, where each truck coordinates with a drone to serve multiple dispersed customers under precedence constraints that specify the required order of service. While trucks deliver packages to designated customers, drones can simultaneously serve other customers, subject to their limited flight endurance and payload capacity. To tackle this challenge, a three-phase heuristic algorithm is proposed to minimize the total delivery time required to serve the last customer while ensuring all precedence constraints are satisfied. In the first phase, an extended minimum marginal cost algorithm is applied to quickly construct truck-only routes that comply with precedence constraints. In the second phase, a splitting algorithm combined with an endurance checking procedure is employed to generate hybrid truck–drone routes considering drone limitations. In the final phase, a variable neighborhood descent approach is introduced to further improve the solution by strategically perturbing the truck-only routes. Extensive simulations and experiments demonstrate that the proposed three-phase heuristic algorithm consistently achieves higher-quality solutions with reduced computation time compared with the widely used adaptive large neighborhood search method.
PaperID: 171,
Authors: Yinan Deng, Yufeng Yue, Jianyu Dou, Jingyu Zhao, Jiahui Wang, Yujie Tang, Yi Yang, Mengyin Fu
Affiliations: School of Automation, Beijing Institute of Technology, Beijing, China
Abstract: Robotic systems demand accurate and comprehensive 3-D environment perception, requiring simultaneous capture of photorealistic appearance (optical), precise layout shape (geometric), and open-vocabulary scene understanding (semantic). Existing methods typically achieve only partial fulfillment of these requirements while exhibiting optical blurring, geometric irregularities, and semantic ambiguities. To address these challenges, we propose OmniMap. Overall, OmniMap represents the first online mapping framework that simultaneously captures optical, geometric, and semantic scene attributes while maintaining real-time performance and model compactness. At the architectural level, OmniMap employs a tightly coupled 3DGS–Voxel hybrid representation that combines fine-grained modeling with structural stability. At the implementation level, OmniMap identifies key challenges across different modalities and introduces several innovations: adaptive camera modeling for motion blur and exposure compensation, hybrid incremental representation with normal constraints, and probabilistic fusion for robust instance-level understanding. Extensive experiments show OmniMap’s superior performance in rendering fidelity, geometric accuracy, and zero-shot semantic segmentation compared to state-of-the-art methods across diverse scenes. The framework’s versatility is further evidenced through a variety of downstream applications, including multidomain scene Q&A, interactive editing, perception-guided manipulation, and map-assisted navigation.
PaperID: 172,
Authors: Yuxiang Peng, Chuchu Chen, Kejian Wu, Guoquan Huang
Affiliations: Department of Mechanical Engineering, University of Delaware, Newark, DE, USA; Department of Department of Mechanical and Aerospace Engineering, George Washington University, Washington, DC, USA; XREAL Inc., Beijing, China
Abstract: In this article, we develop and open-source, for the first time, a robust and efficient square-root filter (SRF)-based visual–inertial navigation system (VINS), termed \sqrt\textVINS, which is ultra-fast, numerically stable, and capable of dynamic initialization even under extreme conditions (i.e., extremely small time window). Despite recent advancements in VINS, resource constraints and numerical instability on embedded (robotic) systems with limited precision remain critical challenges. A square-root covariance-based filter offers a promising solution by providing numerical stability, efficient memory usage, and guaranteed positive semidefiniteness. However, canonical SRFs suffer from inefficiencies caused by disruptions in the triangular structure of the covariance matrix during updates. The proposed method significantly improves VINS efficiency with a novel Cholesky decomposition (LLT)-based SRF update, by fully exploiting the system structure and the SRF to preserve the upper triangular structure of square-root covariance. Moreover, we design a fast, robust, and dynamic initialization method, which first quickly recovers the minimal states without triangulating 3D features and then efficiently performs iterative SRF update to refine the full states, enabling seamless VINS operation even in challenging scenarios. The proposed LLT-based SRF is extensively verified through numerical studies, demonstrating superior numerical stability under challenging conditions and achieving robust efficient performance on 32-b single-precision floats, operating at twice the speed of state-of-the-art methods. Our initialization method, tested on both mobile workstations and Jetson Nano computers achieving a high success rate of initialization even within a 100-ms window under minimal conditions. Finally, the proposed \sqrt\textVINS is extensively validated across diverse scenarios, demonstrating strong efficiency, robustness, and reliability.
PaperID: 173,
Authors: Qinghua Yu, Mengjie Zhang, Chengru Jiang, Guoying Gu, Dong Wang
Affiliations: State Key Laboratory of Mechanical System and Vibration, School of Mechanical Engineering, Shanghai Key Laboratory of Intelligent Robotics, Shanghai Jiao Tong University, Shanghai, China
Abstract: PneuNet, consists of a series of interconnected chambers embedded within a soft elastomer material, can exhibit diverse deformations. 3-D printing allows for precise control over both material combinations and geometrical configurations, enabling the fabrication of PneuNets with complicated structures and multifunctionality. However, the increased freedom in material and structures introduced by 3-D printing also presents significant challenges for modeling and design, including material nonlinearities, complex cross-sections, and varying initial curvatures. In this work, we develop 3-D-printed PneuNets with varying initial curvatures and cross-sections demonstrating finite deformation with multiple complete turns. To model the helical shape, we establish a general nonlinear framework based on the minimum potential energy method. The model is validated by PneuNets with various material combinations and geometrical configurations across a range of constitutive models, including Mooney–Rivlin, Ogden, Neo–Hookean, and Yeoh models. Results show that the nonlinear model, especially the Mooney–Rivlin model, accurately captures the deformation without any fitting parameters, achieving an R^2 value of 0.975, compared to 0.017 for the linear model. Based on the validated model, PneuNets are inverse-designed to achieve desired spatial deformations. Their dynamic responses and payload capacities are also evaluated. We design a 3-D-printed octopus with tentacles composed of PneuNets, capable of mimicking the grasping and movement of a real octopus. In addition, we demonstrate the multifunctional capabilities such as fluid transition and sensing. This study lays a solid foundation for the design and application of 3-D-printed PneuNets.
PaperID: 174,
Authors: Lepeng Chen, Rongxin Cui, Weisheng Yan, Chenguang Yang, Zhijun Li, Hui Xu, Haitao Yu
Affiliations: School of Marine Science and Technology, Northwestern Polytechnical University, Xi'an, China; Department of Computer Science, University of Liverpool, Liverpool, U.K.; School of Mechanical Engineering, and Translational Research Center, Shanghai Yang Zhi Rehabilitation Hospital (Shanghai Sunshine Rehabilitation Center), Tongji University, Shanghai, China
Abstract: The stability criterion is critical for the design of legged robots' motion planning and control algorithms. If these algorithms cannot theoretically ensure legged robots' stability, we need many trials to identify suitable parameters for stable locomotion. However, most existing stability criteria are tailored to robots driven solely by legs and cannot be applied to thruster-assisted legged robots. Here, we propose a stability criterion for a thruster-assisted underwater hexapod robot by finding maximum and minimum allowable thruster forces and comparing them with the current thrusts to check its stability. On this basis, we propose a method to increase the robot's stability margin by adjusting the value of thrusts. This process is called stability enhancement. The criterion uses the optimization method to transform multiple variables such as attitude, velocity, acceleration of the robot body, and the angle and angular velocity of leg joints into one kind of variable (thrust) to judge the stability directly. In addition, the stability enhancement method is straightforward to implement because it only needs to adjust the thrusts. These provide insights into how multiclass forces such as inertia force, fluid force, thrust, gravity, and buoyancy affect the robot's stability.
PaperID: 175,
Authors: Emanuele Aucone, Carmelo Sferrazza, Manuel Gregor, Raffaello D'Andrea, Stefano Mintchev
Affiliations: Environmental Robotics Laboratory, Department of Environmental Systems Science, ETH Zürich, Zürich, Switzerland; Robot Learning Lab, UC Berkeley, Berkeley, CA, USA; Institute for Dynamic Systems and Control, Department of Mechanical and Process Engineering, ETH Zürich, Zürich, Switzerland
Abstract: Distributed tactile sensing for multiforce detection is crucial for various aerial robot interaction tasks. However, current contact sensing solutions on drones only exploit single end-effector sensors and cannot provide distributed multicontact sensing. Designed to be easily mounted at the bottom of a drone, we propose an optical tactile sensor that features a large and curved soft-sensing surface, a hollow structure and a new illumination system. Even when spaced only 2 cm apart, multiple contacts can be detected simultaneously using our software pipeline, which provides real-world quantities of 3-D contact locations (mm) and 3-D force vectors (N), with an accuracy of 1.5 mm and 0.17 N, respectively. We demonstrate the sensor's applicability and reliability onboard and in real time with two demos related to, first, the estimation of the compliance of different perches and subsequent realignment and landing on the stiffer one, and second, the mapping of sparse obstacles. The implementation of our distributed tactile sensor represents a significant step toward attaining the full potential of drones as versatile robots capable of interacting with and navigating within complex environments.
PaperID: 176,
Authors: Abdolreza Taheri, Amy Rankka, Pelle Gustafsson, Joni Pajarinen, Reza Ghabcheloo
Affiliations: Motion Technology R&D, HIAB, Hudiksvall, Sweden; Department of Electrical Engineering and Automation, Aalto University, Aalto, Finland; Faculty of Engineering and Natural Sciences, Tampere University, Tampere, Finland
Abstract: Loader cranes with multiple actuated joints are complex systems to be operated by humans. Development of advanced assistance functions, such as end-effector velocity control in Cartesian space allows for utilizing the machine to its full speed and potential, wherein actuator limits, load balance, and singularities, as well as other complicated effects are handled by the automated function. To this end, this article provides a reinforcement learning-based policy optimization workflow for training and evaluating controllers using large-scale, parallelized invocations of forward kinematics. Monte Carlo evaluations of the closed-loop model are performed to inspect the stability and performance in the whole operational envelope of the loader crane for safe deployment on real machines. Our approach does not require any explicit inverse-kinematics model and is free from complex or hard-coded actuator limits or objectives. Results of simulations and experiments on a real loader crane are provided to showcase the performance of our approach in comparison to Jacobian inverse-based methods.
PaperID: 177,
Authors: Yizhao Wang, Weibang Bai, Zhuangzhuang Zhang, Haili Wang, Han Sun, Qixin Cao
Affiliations: State Key Laboratory of Mechanical System and Vibration, School of Mechanical Engineering, Shanghai Jiao Tong University, Shanghai, China; ShanghaiTech Automation and Robotics Center, School of Information Science and Technology, ShanghaiTech University, Shanghai, China
Abstract: Existing visual simultaneous localization and mapping (SLAM) systems struggle with loop closure under significant viewpoint variations, such as revisiting the same place orthogonally or oppositely. This limitation primarily stems from the lack of viewpoint invariance in both macroscopic place definition and microscopic feature description. Based on the crucial insight that geometric information is more viewpoint-invariant than visual information, we create map segments via leveraging the position and coplanarity distribution of map elements within the map constructed by monocular SLAM to overcome the limitations of frame-based place definition and extract perspective-invariant ORB (PRIOR) features via accounting for local surface perspective distortions to enhance perspective invariance without the need for costly perspective invariance estimation or additional depth data. We further utilize map segments and PRIOR features to hierarchically detect and correct loop closures with a coarse-to-fine geometric consistency check. We integrate our novelties into the prevalent SLAM framework, thereby proposing PRIOR-SLAM, which achieves state-of-the-art performance in feature matching and retrieval, visual place recognition, and loop closure under large viewpoint changes.
PaperID: 178,
Authors: Guanrui Li, Xinyang Liu, Giuseppe Loianno
Affiliations: New York University, New York, NY, USA
Abstract: Human–robot interaction will play an essential role in various industries and daily tasks, enabling robots to effectively collaborate with humans and reduce physical workload. Most existing approaches for physical human–robot interaction focus on collaboration between a human and a single ground or aerial robot. In recent years, very little progress has been made in this research area when considering multiple aerial robots, which offer increased versatility and mobility. This article presents a novel approach for physical human–robot collaborative transportation and manipulation of a cable-suspended payload with multiple aerial robots. The proposed method enables smooth and intuitive interaction between the transported objects and a human worker. We address the inter-robots and inter-robot–human separation during the operations by exploiting the internal redundancy of the multirobot transportation system. The key elements of our approach are, first, a collaborative payload external wrench estimator that does not rely on any force sensor; second, a 6-D admittance controller for human–aerial–robot collaborative transportation and manipulation; third, a human-aware force distribution that exploits the internal system redundancy to guarantee the execution of additional tasks such as inter-human–robot separation without compromising the payload trajectory tracking or interaction quality. We validate our approach through extensive simulation and real-world experiments. These include scenarios where the robot team assists the human in transporting and manipulating a load, or where the human helps the robot team navigate the environment. We experimentally demonstrate for the first time, to the best of authors' knowledge that our approach enables a quadrotor team to physically collaborate with a human in manipulating a payload in all 6 degrees of freedom in collaborative human–robot transportation and manipulation tasks.
PaperID: 179,
Authors: He Li, Patrick M. Wensing
Affiliations: Department of Aerospace and Mechanical Engineering, University of Notre Dame, Notre Dame, IN, USA
Abstract: This work introduces an optimization-based planning and control framework for real-time synthesis of whole-body motions for legged robots. At the core of the proposed framework is a cascaded-fidelity model predictive controller (Cafe-Mpc). Cafe-Mpc strategically relaxes the planning problem along the prediction horizon (i.e., with descending model fidelity, increasingly coarse time steps, and relaxed constraints) for computational and performance gains. This problem is numerically solved with an efficient customized multiple-shooting iLQR solver that is tailored for hybrid systems. The action-value function from Cafe-Mpc is then used as the basis for a new value-function-based whole-body control (VWBC) technique that avoids additional tuning. In this respect, the proposed framework unifies whole-body MPC and more conventional whole-body quadratic programming, which have been treated as separate components in previous works. We study the effects of the cascaded relaxations in Cafe-Mpc on the tracking performance and required computation time. We also show that Cafe-Mpc, if configured appropriately, advances the performance of whole-body MPC without necessarily increasing computational cost. Furthermore, we show the superior performance of VWBC over a conventional Riccati feedback controller in terms of constraint handling. The proposed framework enables accomplishing a gymnastic-style running barrel roll for the first time on quadruped hardware, where Cafe-Mpc runs at 50 Hz, and the solver spends on average 5.3 ms per iteration. Results are demonstrated in the accompanying video.
PaperID: 180,
Authors: Alec Farid, Sushant Veer, Divyanshu Pachisia, Anirudha Majumdar
Affiliations: Princeton University, Princeton, NJ, USA
Abstract: Our goal is to perform out-of-distribution (OOD) detection, i.e., to detect when a robot is operating in environments drawn from a different distribution than the ones used to train the robot. We leverage probably approximately correct-Bayes theory to train a policy with a guaranteed bound on performance on the training distribution. Our idea for OOD detection relies on the following intuition: violation of the performance bound on test environments provides evidence that the robot is operating OOD. We formalize this via statistical techniques based on p-values and concentration inequalities. The approach provides guaranteed confidence bounds on OOD detection including bounds on both the false-positive and false-negative rates of the detector and is task-driven and only sensitive to changes that impact the robot's performance. We demonstrate our approach in simulation and hardware for a grasping task using objects with unfamiliar shapes or poses and a drone performing vision-based obstacle avoidance in environments with wind disturbances and varied obstacle densities. Our examples demonstrate that we can perform task-driven OOD detection within just a handful of trials.
PaperID: 181,
Authors: Andrew G. Curtis, Mark Yim, Michael Rubenstein
Affiliations: Center for Robotics and Biosystems, McCormick School of Engineering, Northwestern University, Evanston, IL, USA; General Robotics, Automation, Sensing and Perception (GRASP) Lab, University of Pennsylvania, Philadelphia, PA, USA
Abstract: Despite their growing popularity, swarms of robots remain limited by the operating time of each individual. We present algorithms that allow a human to sculpt a swarm of robots into a shape that persists in space perpetually, independent of onboard energy constraints, such as batteries. Robots generate a path through a shape such that robots cycle in and out of the shape. Robots inside the shape react to human initiated changes and adapt the path through the shape accordingly. Robots outside the shape recharge and return to the shape so that the shape can persist indefinitely. The presented algorithms communicate shape changes throughout the swarm using message passing and robot motion. These algorithms enable the swarm to persist through any arbitrary changes to the shape. We describe these algorithms in detail and present their performance in simulation and on a swarm of mobile robots. The result is a swarm behavior more suitable for extended duration, dynamic shape-based tasks in applications, such as entertainment, agriculture, and emergency response.
PaperID: 182,
Authors: Keenan Burnett, Angela P. Schoellig, Timothy D. Barfoot
Affiliations: University of Toronto, Institute for Aerospace Studies, Toronto, ON, Canada; Technical University of Munich, Munich, Germany
Abstract: In this work, we demonstrate continuous-time radar-inertial and lidar-inertial odometry using a Gaussian process motion prior. Using a sparse prior, we demonstrate improved computational complexity during preintegration and interpolation. We use a white-noise-on-acceleration motion prior and treat the gyroscope as a direct measurement of the state while preintegrating accelerometer measurements to form relative velocity factors. Our odometry is implemented using sliding-window batch trajectory estimation. To our knowledge, our work is the first to demonstrate radar-inertial odometry with a spinning mechanical radar using both gyroscope and accelerometer measurements. We improve the performance of our radar odometry by 43% by incorporating an inertial measurement unit. Our approach is efficient and we demonstrate real-time performance. Code for this article can be found at: https://github.com/utiasASRL/steam_icp.
PaperID: 183,
Authors: Ihab S. Mohamed, Junhong Xu, Gaurav S. Sukhatme, Lantao Liu
Affiliations: Luddy School of Informatics, Computing, and Engineering, Indiana University, Bloomington, IN, USA; Department of Computer Science, University of Southern California, Los Angeles, CA, USA
Abstract: The classical model predictive path integral (MPPI) control framework, while effective in many applications, lacks reliable safety features due to its reliance on a risk-neutral trajectory evaluation technique, which can present challenges for safety-critical applications such as autonomous driving. Furthermore, when the majority of MPPI sampled trajectories concentrate in high-cost regions, it may generate an infeasible control sequence. To address this challenge, we propose the U-MPPI control strategy, a novel methodology that can effectively manage system uncertainties while integrating a more efficient trajectory sampling strategy. The core concept is to leverage the unscented transform (UT) to propagate not only the mean but also the covariance of the system dynamics, going beyond the traditional MPPI method. As a result, it introduces a novel and more efficient trajectory sampling strategy, significantly enhancing state-space exploration and ultimately reducing the risk of being trapped in local minima. Furthermore, by leveraging the uncertainty information provided by UT, we incorporate a risk-sensitive cost function that explicitly accounts for risk or uncertainty throughout the trajectory evaluation process, resulting in a more resilient control system capable of handling uncertain conditions. By conducting extensive simulations of 2-D aggressive autonomous navigation in both known and unknown cluttered environments, we verify the efficiency and robustness of our proposed U-MPPI control strategy compared to the baseline MPPI. We further validate the practicality of U-MPPI through real-world demonstrations in unknown cluttered environments, showcasing its superior ability to incorporate both the UT and local costmap into the optimization problem without introducing additional complexity.
PaperID: 184,
Authors: Hang Shi, Yali Meng, Wenlong Cui, Meng Rao, Shuting Wang, Yangmin Xie
Affiliations: Department of Automation, Shanghai University, Shanghai, China; Shanghai Key Laboratory of Intelligent Manufacturing and Robotics, Shanghai University, Shanghai, China
Abstract: This study draws inspiration from the locomotion and adaptability of aquatic snakes to develop an innovative soft-bodied, hydraulic-driven untethered underwater snake robot “BaiLong.” The robot consists of a segmented soft structure and embeds actuation, control, and power modules in the head. Featuring the self-shape perception capability, it leverages an online iterative learning control method to effectively mitigate body shape deformation errors and attain precise gait movements. As a result, the soft robot has achieved movements emulating the serpentine motion of real snakes with locomotion consistency equivalent to rigid robots. Extensive experiments in both artificial and natural aquatic environments have presented improved swimming speed among soft snakes with promising turning agility, and revealed the gait parameter influence on the linear velocity described by a near-constant Strouhal number. The reported investigation sufficiently demonstrates the swimming feasibility and performance of underwater soft snake robots and significantly advances their capabilities for long-range applications.
PaperID: 185,
Authors: Sijia Liu, Chunbao Liu, Guowu Wei, Luquan Ren, Lei Ren
Affiliations: School of Mechanical and Aerospace Engineering, Jilin University, Changchun, China; School of Science, Engineering, and Environment, University of Salford, Salford, U.K.; Key Laboratory of Bionic Engineering, Ministry of Education, Jilin University, Changchun, China
Abstract: This article explores a hydraulically powered double-joint soft robotic fish called HyperTuna and a set of locomotion optimization methods. HyperTuna has an innovative, highly efficient actuation structure that includes a four-cylinder piston pump and a double-joint soft actuator with self-sensing. We conducted deformation analysis on the actuator and established a finite element model to predict its performance. A closed-loop strategy combining a central pattern generator controller and a proportional–integral–derivative controller was developed to control the swimming posture accurately. Next, a dynamic model for the robotic fish was established considering the soft actuator, and the model parameters were identified via data-driven methods. Then, a particle swarm optimization algorithm was adopted to optimize the control parameters and improve the locomotion performance. Experimental results showed that the maximum speed increased by 3.6% and the cost of transport (\textCOT) decreased by up to 13.9% at 0.4 m/s after optimization. The proposed robotic fish achieved a maximum speed of 1.12 BL/s and a minimum \textCOT of 12.1 J/(kg·m), which are outstanding relative to those of similar soft robotic fish. Finally, HyperTuna completed turning and diving–floating movements and long-distance continuous swimming in open water, which confirmed its potential for practical application.
PaperID: 186,
Authors: Xuan Liu, Cagdas D. Onal, Jie Fu
Affiliations: School of Aeronautics and Astronautics, Zhejiang University, Hangzhou, China; Robotics Engineering Department, Worcester Polytechnic Institute, Worcester, MA, USA; Department of Electrical and Computer Engineering, University of Florida, Gainesville, FL, USA
Abstract: Contact-awareness poses a significant challenge in the locomotion control of soft snake robots. This article is to develop bioinspired contact-aware locomotion controllers, grounded in a novel theory pertaining to the feedback mechanism of the Matsuoka oscillator. This mechanism enables the Matsuoka central pattern generator (CPG) system to function analogously to a “spinal cord” in the entire contact-aware control framework. Specifically, it concurrently integrates stimuli, such as tonic input signals originating from the “brain” (a goal-tracking locomotion controller) and sensory feedback signals from the “reflex arc” (the contact reactive controller), for generating different types of rhythmic signals to orchestrate the movement of the soft snake robot traversing through densely populated obstacles and even narrow aisles. Within the “reflex arc” design, we have designed two distinct types of contact reactive controllers: 1) a reinforcement learning-based sensor regulator that learns to modulate the sensory feedback inputs of the CPG system, and 2) a local reflexive controller that establishes a direct connection between sensor readings and the CPG's feedback inputs, adhering to a specific topological configuration. These two reactive controllers, when combined with the goal-tracking locomotion controller and the Matsuoka CPG system, facilitate the implementation of two contact-aware locomotion control schemes. Both control schemes have been rigorous tested and evaluated in both simulated and real-world soft snake robots, demonstrating commendable performance in contact-aware locomotion tasks. These experimental outcomes further validate the benefits of the modified Matsuoka CPG system, augmented by a novel sensory feedback mechanism, for the design of bioinspired robot controllers.
PaperID: 187,
Authors: Giovanni Soleti, Paolo Roberto Massenio, Julian Kunze, Gianluca Rizzello
Affiliations: Department of Systems Engineering, Saarland University, Saarbrücken, Germany; Department of Electrical and Information Engineering, Polytechnic University of Bari, Bari, Italy
Abstract: Achieving accurate closed-loop position control of soft robots remains an ongoing research problem, due to the challenges posed by underactuation, elastic nonlinearities, and material creep. Although soft driving technologies relying on tendons and smart material transducers (e.g., dielectric elastomers, shape memory alloys) offer more ease of controllability compared to pneumatics, the corresponding controller design problem becomes even more challenging because of additional nonlinear effects. Those include a configuration-dependent actuation matrix, that stems from the kinematics of the actuation, and control input saturation, which is especially critical for smart material actuators. In this article, we investigate for the first time the closed-loop position control of a soft-robotic system driven by dielectric elastomer actuators. The objective is to regulate the robot state to a constant setpoint, accounting for the effects of open-loop instability, underactuation, control input saturation, and constant external disturbances. To achieve this goal, we propose a model-based feedback scheme, which combines a stabilizing energy-shaping controller with a robustifying PI-like law. After presenting the general theory, a linear matrix inequalities algorithm is proposed to practically address the controller design in spite of strong model nonlinearities. Experimental validation conducted on a prototype of the soft-robotic system confirms the effectiveness of the proposed control approach.
PaperID: 188,
Authors: Fei Li, Hao Yang, Guoying Gu, Yongqing Wang, Haijun Peng
Affiliations: School of Mechanics and Aerospace Engineering, State Key Laboratory of Structural Analysis, Optimization and CAE Software for Industrial Equipment, Dalian University of Technology, Dalian, China; State Key Laboratory of Mechanical System and Vibration, School of Mechanical Engineering, Shanghai Jiao Tong University, Shanghai, China; School of Mechanical Engineering, Dalian University of Technology, Dalian, China
Abstract: Trajectory tracking control of flexible continuum robots is challenging due to their inherent compliance and high nonlinearity. Many related works exclude the control of the end's orientation, i.e., only the end's position is considered. In this article, a differential-algebraic equations (DAEs) model-based instantaneous optimal control (IOC) framework for the end's position and orientation cooperative tracking of a cable-driven tensegrity continuum robot (TCR) is developed. Based on the tensegrity concept, a TCR is designed first as the control object, which can achieve multimode deformations such as bending, scoliosis, contraction, and the S- or J-shape. Then, the actuation of cables is introduced as the system kinematic constraints from the view of multibody dynamics so that a control-oriented model of the TCR can be built by DAEs. Subsequently, the original continuous trajectory tracking problem is approximated for a series of IOC problems at each discrete time slot. Finally, considering the constraints of control input saturation, a linear complementarity problem was derived for solving these IOC problems. The method provides an easy-to-implement and unified framework for addressing the trajectory tracking control issues of cable-driven continuum robots, which can improve the control performance of the position-only tracking controllers and exploit the TCR's advantages to handle more application scenarios. The advanced performance and potential applications of the proposed controller have been evaluated via several numerical simulations and experiments on the TCR prototype.
PaperID: 189,
Authors: Alireza Rezazadeh, Houjian Yu, Karthik Desingh, Changhyun Choi
Affiliations: Department of Electrical and Computer Engineering, University of Minnesota, Minneapolis, MN, USA; Department of Computer Science and Engineering, University of Minnesota, Minneapolis, MN, USA
Abstract: Learning multiobject dynamics purely from visual data is challenging due to the need for robust object representations that can be learned through robot interactions. In previous work (Rezazadeh et al., 2023), we introduced two novel architectures: SlotTransport for discovering object-centric representations from singleview RGB images, referred to as slots, and SlotGNN for predicting scene dynamics from singleview RGB images and robot interactions using the discovered slots. This article introduces InvSlotGNN, a novel framework for learning multiview slot discovery and dynamics that are invariant to the camera viewpoint. First, we demonstrate that SlotTransport can be trained on multiview data such that a single model discovers temporally aligned, object-centric representations from a wide range of different camera angles. These slots bind to objects from various viewpoints, even under occlusion or absence. Next, we introduce InvSlotGNN, an extension of SlotGNN, that learns multiobject dynamics invariant to the camera angle and predicts the future state from observations taken by uncalibrated cameras. InvSlotGNN learns a graph representation of the scene using the slots from SlotTransport and performs relational and spatial reasoning to predict the future state of the scene for arbitrary viewpoints, conditioned on robot actions. We demonstrate the effectiveness of SlotTransport in learning multiview object-centric features that accurately encode visual and positional information. Furthermore, we highlight the accuracy of InvSlotGNN in downstream robotic tasks, including long-horizon prediction and multiobject rearrangement. Finally, with minimal real data, our framework robustly predicts slots and their dynamics in real-world multiview scenarios.
PaperID: 190,
Authors: Tianxiao Gao, Mingle Zhao, Cheng-Zhong Xu, Hui Kong
Affiliations: State Key Laboratory of Internet of Things for Smart City (SKL-IOTSC), Faculty of Science and Technology, University of Macau, Macao, China
Abstract: Accurate and robust state estimation at nighttime is essential for autonomous robotic navigation to achieve nocturnal or round-the-clock tasks. An intuitive question arises: can low-cost standard cameras be exploited for nocturnal state estimation? Regrettably, most existing visual methods may fail under adverse illumination conditions, even with active lighting or image enhancement. A pivotal insight, however, is that streetlights in most urban scenarios act as stable and salient prior visual cues at night, reminiscent of stars in deep space aiding spacecraft voyage in interstellar navigation. Inspired by this, we propose Night-Voyager, an object-level nocturnal vision-aided state estimation framework that leverages prior object maps and keypoints for versatile localization. We also find that the primary limitation of conventional visual methods under poor lighting conditions stems from the reliance on pixel-level metrics. In contrast, metric-agnostic, nonpixel-level object detection serves as a bridge between pixel-level and object-level spaces, enabling effective propagation and utilization of object map information within the system. Night-Voyager begins with a fast initialization to solve the global localization problem. By employing an effective two-stage cross-modal data association, the system delivers globally consistent state updates using map-based observations. To address the challenge of significant uncertainties in visual observations at night, a novel matrix Lie group formulation and a feature-decoupled multistate invariant filter are introduced, ensuring consistent and efficient estimation. Through comprehensive experiments in both simulation and diverse real-world scenarios (spanning approximately 12.3 km), Night-Voyager showcases its efficacy, robustness, and efficiency, filling a critical gap in nocturnal vision-aided state estimation.
PaperID: 191,
Authors: Shihao Cheng, Curt A. Laubscher, T. Kevin Best, Robert D. Gregg
Affiliations: Department of Robotics, University of Michigan, Ann Arbor, MI, USA
Abstract: For powered prosthetic legs to be viable in everyday situations, they require an activity classification system that is not only accurate but also straightforward to understand and use. However, incorporating the numerous activity modes in real-world ambulation often requires high-dimensional feature spaces and restrictions on the leg leading each transition. This article addresses these challenges by delegating sit/stand transitions and variable-incline walking to the mid-level controller, effectively reducing the classification space to four states with easily distinguishable features. We implement simple heuristic rules for both prosthetic-led and intact-led (i.e., ambilateral) transitions, using lower limb kinematic features, ground contact and inclination, and environmental distance from an ultrasonic sensor. Two transfemoral amputee subjects using a powered knee-ankle prosthesis demonstrated an ambilateral transition accuracy of 99.2% under both self-paced and rapid-paced/fatiguing conditions, with a 100% recovery rate due to backup logic or user-cued resets. The incline estimator enabled the prosthesis to continuously adapt between level and inclined surfaces without explicit classification. These results and an outdoor multiterrain demonstration indicate that simple and straightforward transition logic can enable powered prosthetic legs to be used reliably across a broad array of daily activities.
PaperID: 192,
Authors: Hongpeng Wang, Zhongzhi Cao, Yue Fei, Peizhao Wang, Yaojing Li, Chuanyu Sun, Ming He, Jianda Han
Affiliations: College of Artificial Intelligence, Nankai University, Tianjin, China; College of Electronic Information and Optical Engineering, Nankai University, Tianjin, China
Abstract: Autonomous, accurate, and dynamic 3-D reconstruction for wide-area environments is crucial for unmanned aerial vehicle monitoring and rescue tasks, however, when conducted in an unknown complex terrain, the reconstruction result obtained from a single flight suffers poor quality. In this article, we present an Active Iterative Optimization framework for trajectory planning and visual reconstruction. Firstly, the trajectory is planned under the photogrammetric constraints based on rough terrain. Due to the visual field deviation caused by pose error during actual flight, the view loss evaluation is established and keyframes are selected to conduct 3-D reconstruction. A comprehensive metric is designed to quantitatively evaluate reconstruction effect without ground truth. The point cloud is then rasterized and divided into normal or low-scoring region according to the evaluation metric. In the next iteration, trajectory is replanned in low-scoring region to purposefully optimize the point cloud of local area. Thus the reconstruction result can be iteratively optimized. We validated the effectiveness of the proposed framework in simulation and physical experiments.
PaperID: 193,
Authors: Nicola Piccinelli, Riccardo Muradore
Affiliations: Department of Engineering for Innovation Medicine, University of Verona, Verona, Italy
Abstract: Bilateral teleoperation systems are often used in safety–critical scenarios where human operators may interact with the environment remotely, as in robotic-assisted surgery or nuclear plant maintenance. Teleoperation's stability and transparency are the two most important properties to be satisfied, but they cannot be optimized independently since they are in contrast. This article presents a passive linear MPC control scheme to implement bilateral teleoperation that optimizes the tradeoff between stability and transparency (a.k.a. performance). First, we introduce a linear virtual energy tank with a novel energy-sharing policy, allowing us to define a passive linear model predictive control (MPC). Second, we provide conditions to guarantee the stability of the nonlinear closed-loop system. We validate the proposed approach in a teleoperation scheme using two 7-degree of freedom manipulators while performing an assembly task. This novel passivity-based bilateral teleoperation using linear MPC and linearized energy tank reduces the computational effort of existing passive nonlinear MPC controllers.
PaperID: 194,
Authors: Ibrahim Ibrahim, Wilm Decré, Jan Swevers
Affiliations: MECO Research Team, Department of Mechanical Engineering, KU Leuven, Leuven, Belgium
Abstract: In this study, we present a simple and intuitive method for accelerating optimal Reeds–Shepp path computation. Our approach uses geometrical reasoning to analyze the behavior of optimal paths, resulting in a new partitioning of the state space and a further reduction in the minimal set of viable paths. We revisit and reimplement classic methodologies from literature, which lack contemporary open-source implementations, to serve as benchmarks for evaluating our method. In addition, we address the underspecified Reeds–Shepp planning problem where the final orientation is unspecified. We perform exhaustive experiments to validate our solutions. Compared to the modern C++ implementation of the original Reeds–Shepp solution in the Open Motion Planning Library, our method demonstrates a 15× speedup, while classic methods achieve a 5.79× speedup. Both approaches exhibit machine-precision differences in path lengths compared to the original solution. We release our proposed C++ implementations for both the accelerated and underspecified Reeds–Shepp problems as open-source code.
PaperID: 195,
Authors: Bo Pang, Deming Zhai, Jianan Zhen, Long Wang, Xianming Liu
Affiliations: School of Computer Science and Technology, Harbin Institute of Technology, Harbin, China; SenseTime Ltd., Hangzhou, China
Abstract: Aligning a point cloud to a fixed 3-D model is a crucial task in many applications, such as 6-D pose estimation for robotic grasping. Typically, an initial pose is estimated by analyzing both the point cloud and the 3-D model, after which the iterative closest point (ICP) algorithm is used to refine the pose, reducing large errors and improving accuracy. In this article, we propose an accurate and efficient alternative to the ICP. Our method encodes the fixed 3-D model into an implicit neural network, which is trained offline as a one-time process in just a few minutes, requiring only the CAD model of the object. The network takes the point cloud and pose as inputs and outputs the signed distance field (SDF) value. By minimizing the absolute SDF value with the fixed point cloud and network weights, while optimizing the pose, we obtain the final precise alignment. The key advantage of our method is that it eliminates the need to explicitly establish one-to-one correspondences between the point cloud and the 3-D model, a necessary step in the ICP and its variants. This enables our framework to avoid local optima and makes it more robust to challenging conditions such as large initial pose gaps, noisy data, variations in scale, occlusions, and reflections. Furthermore, the end-to-end network of our framework offers significant runtime efficiency. We validate the superior performance of our approach through extensive comparisons with various ICP variants on both synthetic and real-world datasets.
PaperID: 196,
Authors: Seth Stewart, Joseph Pawelski, Steve Ward, Andrew J. Petruska
Affiliations: Department of Mechanical Engineering, Colorado School of Mines, Golden, CO, USA; CisLunar Industries, Denver, CO, USA
Abstract: Control of objects using remotely generated magnetic fields has established itself as a viable option for 3-D position control, though the objects being manipulated to date have largely been limited to soft and hard-magnetic objects that react to a static magnetic field. This limits the application to a small subset of materials. This work presents the first analytically derived model for 3-D position control of any electrically conductive material subject to a time-varying magnetic field. By leveraging the induced eddy current and subsequent induced dipole, this model shows that conductive materials behave equivalently to diamagnetic materials and are, therefore, not subject to the limitations of the Earnshaw’s theorem, making stable, open-loop levitation possible. This is demonstrated by open-loop position control of a semibuoyant aluminum sphere.
PaperID: 197,
Authors: Pengbo Huang, Zhijun Li, Mengchu Zhou, Guoxin Li, Yang Song, Rongxin Cui
Affiliations: School of Mechanical Engineering, Translational Research Center, Shanghai Yangzhi Rehabilitation Hospital (Shanghai Sunshine Rehabilitation Center), Tongji University, Shanghai, China; School of Information and Electronic Engineering, Zhejiang Gongshang University, Hangzhou, China; School of Electronic and Information Engineering, Tongji University, Shanghai, China; School of Marine Science and Technology, Northwestern Polytechnical University, Xi’an, China
Abstract: This article presents a hybrid long short-term motor (HLSM) optimization and control approach for a walking exoskeleton. It consists of long-term global optimization, short-term local optimization, human-in-the-loop trajectory adaptation, and hybrid cerebellar model articulation controller (HCMAC). In the long-term global optimization, a graphic spiking neural network (SNN) is utilized for an optimal global path. Along the path, the short-term motor optimization includes footstep optimization and obtains a sequence of footsteps. While in response to the unexpected obstacles along the footstep sequence, a human-in-the-loop planning strategy is designed by a virtual impedance model between the centers of mass (COMs) of the human and the exoskeleton, regulating the COM of the exoskeleton and generating footstep adaptation of the exoskeleton such that the exoskeleton can avoid obstacles and maintain its original global trajectory. Moreover, considering the unmodeled dynamics, we propose an HCMAC based on an integral Lyapunov function, which is exploited to counteract the system’s nonlinear uncertainties, external disturbances, and reduces a relatively high computational cost. We validate the effectiveness of the HLSM planner and controller in a practical indoor setting. The results demonstrate the effectiveness of HLSM planning and control in a real scenario for a walking exoskeleton.
PaperID: 198,
Authors: Yuwen Zhao, Jiaqi Zhu, Jie Zhang, Siyuan Zhang, Maosen Shao, Zhiping Chai, Yimu Liu, Jianing Wu, Zhigang Wu, Jinxiu Zhang
Affiliations: School of Aeronautics and Astronautics, Shenzhen Campus of Sun Yat-sen University, Shenzhen, China; State Key Laboratory of Intelligent Manufacturing Equipment and Technology, Huazhong University of Science and Technology, Wuhan, China; State Key Laboratory of Structural Analysis, Optimization and CAE Software for Industrial Equipment, Department of Engineering Mechanics, Dalian University of Technology, Dalian, China; School of Advanced Manufacturing, Shenzhen Campus of Sun Yat-sen University, Shenzhen, China
Abstract: Multimodal grasping has emerged as a promising strategy to enhance the grasping diversity of grippers in response to the rapid expansion of application scenarios. Among various designs, the pinch-suction hybrid mechanism and the soft-rigid hybrid structure have proved to be two practical strategies to achieve multimodality. However, existing research on these two strategies still lacks simple and effective collaborative mechanisms to fully leverage the advantages of each mode while ensuring mutual noninterference. In this article, we propose a pinch-suction and soft-rigid hybrid multimodal gripper (HMG), integrating four operating modes into a compact structure. Two simple and effective collaborative mechanisms are introduced to coordinate between pinch and suction operation and between soft and rigid components, respectively. Through the collaboration of different modes, the HMG exhibits a competitive grasping diversity across four aspects, including weight (from 0.2 g to 10 kg), fragility (from jelly to aluminum profile), size scale (from 0.46 mm to 0.55 m), and shape (from poorly pinchable to poorly suckable). We further demonstrate its adaptability and robustness in handling irregular-shaped objects, and its proficiency in executing complex real-world manipulation tasks, underwater operations, and closed-loop grasping. Its enhanced grasping diversity is poised to accelerate diverse applications in daily life, industrial settings, and underwater scenarios.
PaperID: 199,
Authors: Zisong Xu, Rafael Papallas, Jaina Modisett, Markus Billeter, Mehmet Remzi Dogar
Affiliations: School of Computer Science, University of Leeds, Leeds, U.K.
Abstract: This article introduces a method for 6-D pose tracking and control of multiple objects during nonprehensile manipulation by a robot. The tracking system estimates objects’ poses by integrating physics predictions, derived from robotic joint state information, with visual inputs from an RGB-D camera. Specifically, the methodology is based on particle filtering, which fuses control information from the robot as an input for each particle movement and with real-time camera observations to track the pose of objects. Comparative analyses reveal that this physics-based approach substantially improves pose tracking accuracy over baseline methods that rely solely on visual data, particularly during manipulation in clutter, where occlusions are a frequent problem. The tracking system is integrated with a model predictive control approach which shows that the probabilistic nature of our tracking system can help robust manipulation planning and control of multiple objects in clutter, even under heavy occlusions.
PaperID: 200,
Authors: Chaofan Zhang, Shaowei Cui, Jingyi Hu, Tianyu Jiang, Tiandong Zhang, Rui Wang, Shuo Wang
Affiliations: Institute of Automation, Chinese Academy of Sciences, Beijing, China
Abstract: Visuotactile sensors have been shown to provide rich contact information for robots. However, how to build a high-fidelity visuotactile simulator that supports multimode tactile imprints and various sensor configurations (such as coating patterns) remains a challenging problem. In this article, we present TacFlex, an efficient and flexible simulator for visuotactile sensors, which physically simulates the elastomer deformation using finite element methods, and focuses on linking the deformed elastomer mesh to diverse tactile imprints, including tactile images with arbitrary coating patterns and tactile 3-D point clouds. We further propose a ray tracing-based rectification method to deal with multimedium refraction effects to make the simulated tactile images more realistic. Extensive qualitative and quantitative experiments are conducted to demonstrate the effectiveness of TacFlex on several visuotactile sensors. Furthermore, we explore the Sim2Real performance of different tactile imprints provided by TacFlex in tactile perception and manipulation tasks, such as cylindrical object pose estimation and peg-in-hole. The perception/policy models trained in simulation are successfully deployed in the real world. Finally, we present the outlook on the potential of TacFlex in visuotactile manipulation learning. The TacFlex simulator is open-sourced to the community (https://sites.google.com/view/tacflex/).
PaperID: 201,
Authors: Pit Henrich, Franziska Mathis-Ullrich, Paul Maria Scheikl
Affiliations: Department Artificial Intelligence in Biomedical Engineering, Friedrich-Alexander-University Erlangen-Nürnberg, Erlangen, Germany
Abstract: Accurately determining the shape of deformable objects and the location of their internal structures is crucial for medical tasks that require precise targeting, such as robotic biopsies. We introduce a method for accurate low-latency understanding of deformable objects (LUDO). LUDO reconstructs objects in their deformed state, including their internal structures, from a single-view point cloud observation in under 30 ms using occupancy networks. LUDO provides uncertainty estimates for its predictions. In addition, it provides explainability by highlighting key features in its input observations. Both uncertainty and explainability are important for safety-critical applications, such as surgery. We evaluate LUDO in real-world robotic experiments, achieving a success rate of 98.9% for puncturing various regions of interest (ROIs) inside deformable objects. We compare LUDO to a popular baseline and show its superior ROI localization accuracy, training time, and memory requirements. LUDO demonstrates the potential to interact with deformable objects without the need for deformable registration methods.
PaperID: 202,
Authors: Hongyu Zhou, Vasileios Tzoumas
Affiliations: Department of Aerospace Engineering, University of Michigan, Ann Arbor, MI, USA
Abstract: We provide an algorithm for the simultaneous system identification and model predictive control of nonlinear systems. The algorithm has finite-time near-optimality guarantees and asymptotically converges to the optimal (noncausal) controller. Particularly, the algorithm enjoys sublinear dynamic regret, defined herein as the suboptimality against an optimal clairvoyant controller that knows how the unknown disturbances and system dynamics will adapt to its actions. The algorithm is self-supervised and applies to control-affine systems with unknown dynamics and disturbances that can be expressed in reproducing kernel Hilbert spaces. Such spaces can model external disturbances and modeling errors that can even be adaptive to the system’s state and control input. For example, they can model wind and wave disturbances to aerial and marine vehicles, or inaccurate model parameters such as inertia of mechanical systems. We are motivated by the future of autonomy where robots will autonomously perform complex tasks despite real-world unknown disturbances such as wind gusts. The algorithm first generates random Fourier features that are used to approximate the unknown dynamics or disturbances. Then, it employs model predictive control based on the current learned model of the unknown dynamics (or disturbances). The model of the unknown dynamics is updated online using least squares based on the data collected while controlling the system. We validate our algorithm in both hardware experiments and physics-based simulations. The simulations include a cart-pole aiming to maintain the pole upright despite inaccurate model parameters and a quadrotor aiming to track reference trajectories despite unmodeled aerodynamic drag effects. The hardware experiments include a quadrotor aiming to track a circular trajectory despite unmodeled aerodynamic drag effects, ground effects, and wind disturbances.
PaperID: 203,
Authors: Taoran Jiang, Yixuan Guan, Liqian Ma, Jing Xu, Jiaojiao Meng, Weihang Chen, Zecui Zeng, Lusong Li, Dan Wu, Rui Chen
Affiliations: Department of Mechanical Engineering, Tsinghua University, Beijing, China; JD Explore Academy, Beijing, China
Abstract: Articulated objects are ubiquitous in daily life. In this article, we present DexSim2Real^\mathbf2, a novel framework for goal-conditioned articulated object manipulation. The core of our framework is constructing an explicit world model of unseen articulated objects through active interactions, which enables sampling-based model-predictive control to plan trajectories achieving different goals without requiring demonstrations or reinforcement learning. It first predicts an interaction using an affordance network trained on self-supervised interaction data or videos of human manipulation. After executing the interactions on the real robot to move the object parts, we propose a novel modeling pipeline based on 3-D artificial intelligence generated content to build a digital twin of the object in simulation from multiple frames of observations. For dexterous hands, we utilize eigengrasp to reduce the action dimension, enabling more efficient trajectory searching. Experiments validate the framework’s effectiveness for precise manipulation using a suction gripper, a two-finger gripper, and two dexterous hands. The generalizability of the explicit world model also enables advanced manipulation strategies, such as manipulating with tools.
PaperID: 204,
Authors: Janak Panthi, Farshid Alambeigi, Mitch Pryor
Affiliations: Walker Department of Mechanical Engineering and Texas Robotics, The University of Texas at Austin, Austin, TX, USA
Abstract: Robots operating in unpredictable environments require versatile, hardware-agnostic frameworks capable of adapting to various tasks. While a recent screw-based affordance approach shows promise, it faces challenges in avoiding undesirable configurations, singularity navigation, and task success prediction. To address these limitations, we propose a novel framework that incorporates gripper orientation control and generates complete joint trajectories in real time for screw-based task affordance execution. Our method models the affordance and manipulator as a closed-chain mechanism, introducing an innovative approach to solving closed-chain inverse kinematics. It encapsulates task constraints and simplifies task definitions, while remaining hardware and robot agnostic, robust to errors, and invariant to the initial grasp. We validate our framework with simulations on a UR5 robot and real-world implementation on a Boston Dynamics Spot robot. Our experiments demonstrate rapid joint trajectory generation (0.0077–0.098 s) for various tasks, including a 420^\circ valve turn with consideration of the gripper orientation. Comparison with the state-of-the-art methods shows a 4x improvement in planning time, reduced joint movement, and achievement of greater task goals. Video demonstrations and the open-source code for this project are available online.
PaperID: 205,
Authors: Brian Acosta, Michael Posa
Affiliations: GRASP Laboratory, University of Pennsylvania, Philadelphia, PA, USA
Abstract: Traversing rough terrain requires dynamic bipeds to stabilize themselves through foot placement without stepping into unsafe areas. Planning these footsteps online is challenging given the nonconvexity of the safe terrain and imperfect perception and state estimation. This article addresses these challenges with a full-stack perception and control system for achieving underactuated walking on discontinuous terrain. First, we develop model-predictive footstep control, a single mixed-integer quadratic program, which assumes a convex polygon terrain decomposition to optimize over discrete foothold choice, footstep position, ankle torque, template dynamics, and footstep timing at over 100 Hz. We then propose a novel approach for generating convex polygon terrain decompositions online. Our perception stack decouples safe-terrain classification from fitting planar polygons, generating a temporally consistent terrain segmentation in real time using a single CPU thread. We demonstrate the performance of our perception and control stack through outdoor experiments with the underactuated biped Cassie, achieving state of the art perceptive bipedal walking on discontinuous terrain.
PaperID: 206,
Authors: Jinfei Hu, Zelong Chen, Yinjie Lin, Zheng Chen, Bin Yao, Xin Ma
Affiliations: T Stone Robotics Institute, Department of Mechanical and Automation Engineering, The Chinese University of Hong Kong, Hong Kong; State Key Laboratory of Fluid Power and Mechatronic Systems, Zhejiang University, Hangzhou, China; Hangzhou Hikvision Digital Technology Company Ltd., Hangzhou, China; School of Mechanical Engineering, Purdue University, West Lafayette, IN, USA; Shenzhen Key Laboratory of Intelligent Robotics and Flexible Manufacturing Systems, Institute for Robotics, Southern University of Science and Technology, Shenzhen, China
Abstract: Accurate rigid-body dynamics is crucial for serial industrial robot applications, such as force control and physical human–robot interaction. Despite decades of research, the precise identification of dynamic parameters—particularly low-magnitude inertia parameters—remains a challenge for serial industrial robots. Researchers usually focus on developing various parameter estimation methods, while optimizing exciting trajectories in similar ways, typically minimizing the condition number of the information matrix. However, such optimization usually fails to ensure sufficient excitation for each parameter, due to nonconvex coupling effects. To address this limitation, we propose a fully decoupled rigid-body dynamics identification (FDRDI) method in this article. This approach innovatively eliminates coupling effects by using novel symmetrical exciting trajectories based on reciprocating S-curve. This innovation enables the independent identification of dynamic parameters associated with joint friction, as well as the gravity and inertia of links and payloads. Comparative experiments show that FDRDI achieves superior identification accuracy, evidenced by reduced joint torque prediction errors and payload parameter estimation errors.
PaperID: 207,
Authors: Chao Zhao, Chunli Jiang, Lifan Luo, Shuai Yuan, Qifeng Chen, Hongyu Yu
Affiliations: School of Artificial Intelligence, Jilin University, Changchun, China; Hong Kong University of Science and Technology, Clear Water Bay, Hong Kong
Abstract: Robotic manipulation has made significant advancements, with systems demonstrating high precision and repeatability. However, this remarkable precision often fails to translate into efficient manipulation of thin deformable objects. Current robotic systems lack imprecise dexterity, the ability to perform dexterous manipulation through robust and adaptive behaviors that do not rely on precise control. This article explores the singulation and grasping of thin, deformable objects. Here, we propose a novel solution that incorporates passive compliance, touch, and proprioception into thin, deformable object manipulation. Our system employs a soft, underactuated hand that provides passive compliance, facilitating adaptive and gentle interactions to dexterously manipulate deformable objects without requiring precise control. The tactile and force/torque sensors equipped on the hand, along with a depth camera, gather sensory data required for manipulation via the proposed slip module. The manipulation policies are learned directly from raw sensory data via model-free reinforcement learning, bypassing explicit environmental and object modeling. We implement a hierarchical double-loop learning process to enhance learning efficiency by decoupling the action space. Our method was deployed on real-world robots and trained in a self-supervised manner. The resulting policy was tested on a variety of challenging tasks that were beyond the capabilities of prior studies, ranging from displaying suit fabric like a salesperson to turning pages of sheet music for violinists.
PaperID: 208,
Authors: Allen George Philip, Zhongqiang Ren, Sivakumar Rathinam, Howie Choset
Affiliations: Texas A&M University, College Station, TX, USA; Shanghai Jiao Tong University, Shanghai, China; Carnegie Mellon University, Pittsburgh, PA, USA
Abstract: We introduce a new bounding approach called Continuity (C^), which provides optimality guarantees for the moving-target traveling salesman problem (MT-TSP). Our approach relaxes the continuity constraints on the agent’s tour by partitioning the targets’ trajectories into smaller segments. This allows the agent to arrive at any point within a segment and depart from any point in the same segment when visiting each target. This formulation enables us to pose the bounding problem as a generalized traveling salesman problem on a graph, where the cost of traveling along an edge requires solving a new problem called the shortest feasible travel (SFT). We present various methods for computing bounds for the SFT problem, leading to several variants of C^. We first prove that the proposed algorithms provide valid lower bounds for the MT-TSP. In addition, we provide computational results to validate the performance of all C^ variants on instances with up to 15 targets. For the special case where targets move along straight lines, we compare our C^ variants with a mixed-integer second order conic program (SOCP)-based method, the current state-of-the-art solver for the MT-TSP. While the SOCP-based method performs well on instances with five and ten targets, C^ outperforms it on instances with 15 targets. For the general case, on average, our approaches find feasible solutions within approximately 4.5% of the lower bounds for the tested instances.
PaperID: 209,
Authors: Jun Huo, Jian Huang, Jie Zuo, Bo Yang, Zhongzheng Fu, Xi Li, Samer Mohammed
Affiliations: Key Laboratory of the Ministry of Education for Image Processing and Intelligent Control, Huazhong University of Science and Technology, Wuhan, China; School of Information Engineering, Wuhan University of Technology, Wuhan, China; State Key Laboratory of Intelligent Vehicle Safety Technology, Chongqing Changan Automobile Company Ltd., Chongqing, China; Univ Paris-Est Créteil, Vitry, France
Abstract: Supernumerary robotic limbs (SRLs) offer substantial potential in both the rehabilitation of hemiplegic patients and the enhancement of functional capabilities for healthy individuals. Designing a general-purpose SRL device is inherently challenging, particularly when developing a unified theoretical framework that meets the diverse functional requirements of both upper and lower limbs. In this article, we propose a multiobjective optimization (MOO) design theory that integrates grasping workspace similarity, walking workspace similarity, braced force for sit-to-stand (STS) movements, and overall mass and inertia. A geometric vector quantification method is developed using an ellipsoid to represent the workspace, aiming to reduce computational complexity and address quantification challenges. The ellipsoid envelope transforms workspace points into ellipsoid attributes, providing a parametric description of the workspace. Furthermore, the STS static braced force assesses the effectiveness of force transmission. The overall mass and inertia restricts excessive link length. To facilitate rapid and stable convergence of the model to high-dimensional irregular Pareto fronts, we introduce a multisubpopulation correction firefly algorithm. This algorithm incorporates a strategy involving attractive and repulsive domains to effectively handle the MOO task. The optimized solution is utilized to redesign the prototype for experimentation to meet specified requirements. Six healthy participants and two hemiplegia patients participated in real experiments. Compared to the preoptimization results, the average grasp success rate improved by 7.2%, while the muscle activity during walking and STS tasks decreased by an average of 12.7% and 25.1%, respectively. The proposed design theory offers an efficient option for the design of multifunctional SRL mechanisms.
PaperID: 210,
Authors: Juntao He, Baxi Chong, Jianfeng Lin, Zhaochen Xu, Hosain Bagheri, Esteban Flores, Daniel I. Goldman
Affiliations: Institute for Robotics and Intelligent Machines, Atlanta, GA, USA; School of Physic, Georgia Institute of Technology, Atlanta, GA, USA; School of Mechanical Engineering, Georgia Institute of Technology, Atlanta, GA, USA
Abstract: Achieving robust legged locomotion on complex terrains poses challenges due to the high uncertainty in robot–environment interactions. Recent advances in bipedal and quadrupedal robots demonstrate good mobility on rugged terrains but rely heavily on sensors for stability due to low static stability from a high center of mass and a narrow base of support (Ijspeert and Daley, 2023).We hypothesize that a multilegged robotic system can leverage morphological redundancy from additional legs to minimize sensing requirements when traversing challenging terrains. Studies suggest (Chong et al., 2023), (Chong et al., 2023) that a multilegged system with sufficient legs can reliably navigate noisy landscapes without sensing and control, albeit at a low speed of up to 0.1 body lengths per cycle (BLC). However, the feedback control framework to enhance speed of multilegged robots on challenging terrains remains underexplored due to diverse environmental interactions. Such complexity makes it difficult to identify the key parameters to control in these high-degree-of-freedom systems. Here, using laboratory and field experiments, we demonstrate that a vertical body undulation wave helps mitigate environmental disturbances that affect robot speed. These findings are supported by probabilistic models. Using such insights, we introduce a control framework, which monitors foot–ground contact patterns on rugose landscapes using binary foot–ground contact sensors to estimate terrain rugosity. The controller adjusts the vertical body wave based on the deviation of the limb’s averaged actual-to-ideal foot–ground contact ratio, achieving a significant enhancement of up to 0.235 BLC on rugose laboratory terrain. We observed a 50% to 60% increase in speed and a 30% to 50% reduction in speed variance compared to the open-loop controller. In addition, the controller operates in complex terrains outside the lab, including pine straw, robot-sized rocks, mud, and leaves.
PaperID: 211,
Authors: Christian Brommer, Alessandro Fornasier, Jan Steinbrener, Stephan Weiss
Affiliations: Control of Networked Systems Group, University of Klagenfurt, Klagenfurt am Wörthersee, Austria
Abstract: We present a method for the unattended gray-box identification of sensor models commonly used by localization algorithms in the field of robotics. The objective is to determine the most likely sensor model for a time series of unknown measurement data, given an extendable catalog of predefined sensor models. Sensor model definitions may require states for rigid-body calibrations and dedicated reference frames to replicate a measurement based on the robot’s localization state. A health metric is introduced, which verifies the outcome of the selection process in order to detect false positives and facilitate reliable decision-making. In the second stage, an initial guess for identified calibration states is generated, and the necessity of sensor world reference frames is evaluated. The identified sensor model with its parameter information is then used to parameterize and initialize a state estimation application, thus ensuring a more accurate and robust integration of new sensor elements. This method is helpful for inexperienced users who want to identify the source and type of a measurement, sensor calibrations, or sensor reference frames. It will also be important in the field of modular multiagent scenarios and modularized robotic platforms that are augmented by sensor modalities during runtime. Overall, this work aims to provide a simplified integration of sensor modalities to downstream applications and circumvent common pitfalls in the usage and development of localization approaches.
PaperID: 212,
Authors: Hongxin Huang, Qingqing Wang, Zhongtian Liu, Zhetian Ding, Fanghao Zhou, Zheng Chen, Tiefeng Li
Affiliations: State Key Laboratory of Ocean Sensing, Institute of Fundamental and Transdisciplinary Research, Zhejiang University, Hangzhou, China
Abstract: Twisted and coiled actuators (TCAs) are promising in soft robotics for their high energy density, light weight, and low voltage. However, current TCA-based soft robots face challenges of limited deformation and control precision, mainly due to the preloading requirement of TCAs and the lack of suitable intrinsic sensing capabilities. To address these issues, we designed a TCA with high load capacity, free stroke, and self-sensing capabilities, proposed flexible optical fiber-based posture and tactile sensing methods, and developed a multiloop feedback controller. Collectively, these enable millimeter-level tracking accuracy in a soft tentacle robot. The TCA, with an optimized manufacturing process, achieves a 30% free stroke without preloading, a 32% improvement in ultimate stress, and temperature self-sensing capabilities with a maximum error of less than 6% . Combining the TCAs with compliant macro-bend optical fibers and soft optical waveguides, we created a soft robotic tentacle with intrinsic posture and tactile sensing, and designed a multiinput–multioutput closed-loop and feedforward controller. Experiments demonstrate that the model-based feedforward significantly improves the control performance, reducing the rise time by 15.3% . The trajectory tracking error remains within the millimeter range, and the repetitive positioning error for hexagonal trajectories reaches submillimeter precision. The tactile sensor of the robot enables real-time perception of the object’s modulus and pressing states. These findings highlight the soft robotic tentacle’s potential for various applications, including underwater exploration, detection, and sampling.
PaperID: 213,
Authors: Amin Yazdanshenas, Reza Faieghi
Affiliations: Autonomous Vehicles Laboratory, Department of Aerospace Engineering, Toronto Metropolitan University, Toronto, ON, Canada
Abstract: This article presents a new adaptive sliding-mode control (SMC) framework for quadrotors that achieves robust and agile flight under tight computational constraints. The proposed controller addresses key limitations of prior SMC formulations, including, first, the slow convergence and almost-global stability of \mathrmSO(3)-based methods, second, the oversimplification of rotational dynamics in Euler-based controllers, third, the unwinding phenomenon in quaternion-based formulations, and fourth, the gain overgrowth problem in adaptive SMC schemes. Leveraging nonsmooth stability analysis, we provide rigorous global stability proofs for both the nonsmooth attitude sliding dynamics defined on \mathbb S^3 and the position sliding dynamics. Our controller is computationally efficient and runs reliably on a resource-constrained nano quadrotor, achieving 250 Hz and 500 Hz refresh rates for position and attitude control, respectively. In an extensive set of hardware experiments with over 130 flight trials, the proposed controller consistently outperforms three benchmark methods, demonstrating superior trajectory tracking accuracy and robustness with relatively low control effort. The controller enables aggressive maneuvers, such as dynamic throw launches, flip maneuvers, and accelerations exceeding 3 g, which is remarkable for a 32-gram nano quadrotor. These results highlight promising potential for real-world applications, particularly in scenarios requiring robust, high-performance flight control under significant external disturbances and tight computational constraints.
PaperID: 214,
Authors: Haicheng Liao, Zhenning Li, Kaiqun Zhu, Keqiang Li, Cheng-Zhong Xu
Affiliations: State Key Laboratory of Internet of Things for Smart City and the Department of Computer and Information Science, University of Macau, Macau, China; Department of Automotive Engineering, Tsinghua University, Beijing, China
Abstract: Trajectory prediction and planning remain key challenges for autonomous vehicles, particularly in complex and dynamic environments. Existing methods, typically based on static safety metrics like time-to-collision, fail to account for the evolving nature of risk in real-world traffic. This article proposes a novel safety-aware trajectory prediction and planning (SA-TP^2) model, which introduces an adaptive driver risk field to simulate human-like risk perception and decision-making. By dynamically modeling risk as a continuous variable, SA-TP^2 adjusts vehicle trajectories in real time, accounting for interactions with other agents, road conditions, and environmental uncertainties. The model integrates imitation learning, rule-based strategies, and physics-informed neural networks to ensure safe, efficient, and human-compatible behavior. A Linformer-based architecture and temporal hypergraph convolution network are introduced to optimize computational efficiency, enabling real-time operation in resource-constrained environments. Experimental results on benchmark datasets including next generation simulation (NGSIM), highway drone dataset (HighD), Macao connected autonomous driving (MoCAD), and NuScenes demonstrate that SA-TP^2 achieves the state-of-the-art performance in trajectory prediction. In addition, extensive closed-loop testing on the NuPlan and CommonRoad platforms further confirms that SA-TP^2 outperforms existing baselines, paving the way for safer navigation of autonomous driving systems.
PaperID: 215,
Authors: Nermin Covic, Bakir Lacevic, Dinko Osmankovic, Tarik Uzunovic
Affiliations: Faculty of Electrical Engineering, University of Sarajevo, Sarajevo, Bosnia and Herzegovina; Agile Robots SE, Munich, Germany
Abstract: In this article, we present the main features of the dynamic rapidly-exploring generalized bur tree (DRGBT) algorithm, a sampling-based planner for dynamic environments. We provide a detailed time analysis and appropriate scheduling to facilitate a real-time operation. To this end, an extensive analysis is conducted to identify the time-critical routines and their dependence on the number of obstacles. Furthermore, information about the distance to obstacles is used to compute a structure called dynamic expanded bubble of free configuration space, which is then utilized to establish sufficient conditions for a guaranteed safe motion of the robot while satisfying all kinematic constraints. An extensive comparative study is conducted to compare the proposed algorithm to competing state-of-the-art methods. Finally, an experimental study on a real robot is carried out covering a variety of scenarios including those with human presence. The results show the effectiveness and feasibility of real-time execution of the proposed motion planning algorithm within a typical sensor-based arrangement, using cheap hardware and sequential architecture, without the necessity for GPUs or heavy parallelization.
PaperID: 216,
Authors: Yubo Sheng, Yiwei Wang, Haoyuan Cheng, Huan Zhao, Han Ding
Affiliations: State Key Laboratory of Intelligent Manufacturing Equipment and Technology, Huazhong University of Science and Technology, Wuhan, China
Abstract: Harmonious human–robot collaboration requires the robot to behave like a human partner, which raises the critical question of what factors make the robot do so. This article proposes a series of policies based on empathetic and nonempathetic intent inference, proactive and reactive action planning, and ego and nonego action styles to examine, which modules enable robots to exhibit human-like behaviors. Two series of experiments are conducted with human subjects to test the performance of the proposed controllers. In Experiment 1, the participant must identify whether the collaborating partner is a human, similar to a turing test. The classification results empirically verify that the designed empathetic proactive policies enable the robot to exhibit human-like behaviors. Experiment 2 indicates that the proposed policy can be applied to complex collaborative tasks, and this result is consistent with the findings of Experiment 1. From empirical evidence from the experiments, we believe that empathy and proactive policies are essential elements to enable robots to perform human-like actions.
PaperID: 217,
Authors: Yuanshuai Ding, Yongmao Pei
Affiliations: State Key Laboratory for Turbulence and Complex Systems, College of Engineering, Peking University, Beijing, China
Abstract: The piezoelectric plate robot driven by traveling waves (TWs) boasts a compact structure and excellent load-bearing capacity. However, vibration coupling across the plate restricts its freedom of movement, along with issues, such as poor straight-line motion and a large turning radius. Inspired by octopus locomotion, we designed a three-degree-of-freedom (3-DoF) piezoelectric plate robot using a 0.7-mm metal plate. Through drive region division and dynamic modeling, our design achieves fully controllable motion in any planar direction, providing the highest DoF among TW-driven plate robots. Weighing just 14.2 g, it is lighter and easier to fabricate compared to other 3-DoF piezoelectric robots. The robot demonstrated 3-DoF movement capabilities, climbing (16.5^\circ), dragging (20 g), and carrying (180 g), and can bear over 5000 times its own body weight. A wireless drive prototype with closed-loop control reduced trajectory errors by more than 80% compared to open-loop control. Experiments involving high-curvature path movement, confined-space ‘‘search and rescue,’’ and light focusing highlight its potential in extreme environments and high-precision tasks.
PaperID: 218,
Authors: Ziyuan Jiao, Yida Niu, Zeyu Zhang, Yangyang Wu, Yao Su, Yixin Zhu, Hangxin Liu, Song-Chun Zhu
Affiliations: State Key Laboratory of General Artificial Intelligence, Beijing Institute for General Artificial Intelligence (BIGAI), Beijing, China; School of Psychological and Cognitive Sciences, Peking University, Beijing, China
Abstract: We present a sequential mobile manipulation planning framework that can solve long-horizon multistep mobile manipulation tasks with coordinated whole-body motion, even when interacting with articulated objects. By abstracting environmental structures as kinematic models and integrating them with the robot’s kinematics, we construct an augmented configuration space (A-Space) that unifies the previously separate task constraints for navigation and manipulation, while accounting for the joint reachability of the robot base, arm, and manipulated objects. This integration facilitates efficient planning within a tri-level framework: a task planner generates symbolic action sequences to model the evolution of A-Space, an optimization-based motion planner computes continuous trajectories within A-Space to achieve desired configurations for both the robot and scene elements, and an intermediate plan refinement stage selects action goals that ensure long-horizon feasibility. Our simulation studies first confirm that planning in A-Space achieves an 84.6% higher task success rate compared to baseline methods. Validation on real robotic systems demonstrates fluid mobile manipulation involving first, seven types of rigid and articulated objects across 17 distinct contexts, and second, long-horizon tasks of up to 14 sequential steps. Our results highlight the significance of modeling scene kinematics into planning entities, rather than encoding task-specific constraints, offering a scalable and generalizable approach to complex robotic manipulation.
PaperID: 219,
Authors: Yuan Gao, Victor Paredes, Yukai Gong, Zijian He, Ayonga Hereid, Yan Gu
Affiliations: College of Engineering, University of Massachusetts Lowell, Lowell, MA, USA; Department of Mechanical and Aerospace Engineering, The Ohio State University, Columbus, OH, USA; Robotics Department, University of Michigan, Ann Arbor, MI, USA; School of Mechanical Engineering, Purdue University, West Lafayette, IN, USA
Abstract: Locomotion on dynamic rigid surface (i.e., rigid surface accelerating in an inertial frame) presents complex challenges for controller design, which are essential to address for deploying humanoid robots in dynamic real-world environments such as moving trains, ships, and airplanes. This article introduces a real-time, provably stabilizing control approach for humanoid walking on periodically swaying rigid surface. The first key contribution is an analytical extension of the classical angular momentum-based linear inverted pendulum model from static to swaying grounds whose motion period may be different than the robot’s gait period. This extension results in a time-varying, nonhomogeneous robot model, which is fundamentally different from the existing pendulum models. We synthesize a discrete footstep control law for the model and derive a new set of sufficient stability conditions that verify the controller’s stabilizing effect. Finally, experiments conducted on a Digit humanoid robot, both in simulations and on hardware, demonstrate the framework’s effectiveness in addressing bipedal locomotion on swaying ground, even under uncertain surface motions and unknown external pushes.
PaperID: 220,
Authors: Matteo Scucchia, Davide Maltoni
Affiliations: University of Bologna, Department of Computer Science and Engineering, Bologna, Italy
Abstract: Most of today’s simultaneous localization and mapping (SLAM) approaches learn the map of the environment in the first stage (referred to as mapping) and subsequently use this static map for planning and navigation. This method is suboptimal in dynamic contexts because changes in the environment can result in poor performance of the localization components essential for loop closure detection and relocalization. To address the limitations of the mapping-navigation dualism, continual SLAM has been proposed, which focuses on methods that can continually update the knowledge of the environment and the corresponding map. However, continual SLAM poses challenges, particularly for real-time navigation of large maps, and many of the existing techniques are not yet mature for practical application. In this article, we present a continual learning approach aimed at accurate and efficient robot localization on large maps, advancing the goal of continual SLAM. Our approach incrementally trains a region prediction neural network to recognize familiar places and preselect a subset of map nodes for localization and map optimization. We integrate this method into RTAB-Map, a well-known graph-based SLAM system, and validate its practical applicability through assessments on several real-world SLAM datasets.
PaperID: 221,
Authors: Yinzhao Dong, Ji Ma, Liu Zhao, Wanyue Li, Peng Lu
Affiliations: Adaptive Robotic Controls Lab (ArcLab), Department of Mechanical Engineering, The University of Hong Kong, Hong Kong, SAR China
Abstract: Deep Reinforcement Learning (DRL) controllers for quadrupedal locomotion have demonstrated impressive performance on challenging terrains, allowing robots to execute complex skills such as climbing, running, and jumping. However, existing blind locomotion controllers often struggle to ensure safety and efficient traversal through risky gap terrains, which are typically highly complex, requiring robots to perceive terrain information and select appropriate footholds during locomotion accurately. Meanwhile, existing perception-based controllers still present several practical limitations, including a complex multisensor deployment system and expensive computing resource requirements. This article proposes a DRL controller named MAstering Risky Gap Terrains (MARG), which integrates terrain maps and proprioception to dynamically adjust the action and enhance the robot’s stability in these tasks. During the training phase, our controller accelerates policy optimization by selectively incorporating privileged information (e.g., center of mass, friction coefficients) that are available in simulation but unmeasurable directly in real-world deployments due to sensor limitations. We also designed three foot-related rewards to encourage the robot to explore safe footholds. More importantly, a terrain map generation model is proposed to reduce the drift existing in mapping and provide accurate terrain maps using only one LiDAR, providing a foundation for zero-shot transfer of the learned policy. The experimental results indicate that MARG maintains stability in various risky terrain tasks.
PaperID: 222,
Authors: Zhiping Wang, Zonggang Li, Bin Li, Guangqing Xia, Huifeng Kang
Affiliations: School of Mechatronic Engineering, Robotics Institute, Lanzhou Jiaotong University, Lanzhou, China; College of Mechatronic Engineering, Robotics Institute, Lanzhou Jiaotong University, Lanzhou, China; State Key Laboratory of Structural Analysis, Optimization, CAE Software for Industrial Equipment, Dalian University of Technology, Dalian, China; Hebei Key Laboratory of Trans-Media Aerial Underwater Vehicle, North China Institute of Aerospace Engineering, Hebei, China
Abstract: The robotic fish of BCF/MPF hybrid propulsion achieves efficient and stable swimming through the synergistic control of pectoral fins and body. However, the control problem has been less studied of fins-body coupling with multiple degrees of freedom. This study focuses on the development and synergistic control of robotic fish with fins-body. First, a robotic fish with pectoral fin and body co-propulsion was designed, and a gait controller of fin-body synergic was constructed by a central pattern generator. Specifically, the control parameters were simplified, and the synergic movement was realized of fin-body coupling with multiple degrees of freedom. Second, the dataset was obtained with computational fluid dynamics simulations and 6-D force sensors, and the offline hydrodynamic model was obtained by bidirectional long short-time memory networks identification, which is the relationship between the control parameters and force/torque of the robotic fish. The model parameters were updated online with experimental data. Finally, a control framework is constructed for offline–online model and event-triggered nonlinear model predictive control, which compensates for the driving force of the robotic fish, achieves tracking trajectory precisely, and reduces the computational cost.
PaperID: 223,
Authors: Xiaohui Zhang, Enrica Tricomi, Xunju Ma, Manuela Gomez-Correa, Alessandro Ciaramella, Francesco Missiroli, Luka Miskovic, Huimin Su, Lorenzo Masia
Affiliations: Institut für Technische Informatik (ZITI), Heidelberg University, Heidelberg, Germany; School of Mechatronical Engineering, Beijing Institute of Technology, Beijing, China; Medical Robotics and Biosignals Laboratory, Centro de Innovación y Desarrollo Tecnológico en Cómputo, Instituto Politécnico Nacional, Mexico City, Mexico; Institute of Mechanical Intelligence, Sant'Anna School of Advanced Studies, Pisa, Italia; Department of Automatics, Biocybernetics and Robotics, Jožef Stefan Institute, Ljubljana, Slovenia; Department of Computer Engineering, School of Computation, Information and Technology, Technical University of Munich, Munich, Germany
Abstract: Sitting, standing, and walking are fundamental activities crucial for maintaining independence in daily life. However, aging or lower limb injuries can impede these activities, posing obstacles to individuals' autonomy. In response to this challenge, we developed the LM-Ease (lower-limb movement ease), a compact and soft wearable robot designed to provide hip assistance. Its purpose is to aid users in carrying out essential daily activities such as sitting, standing, and walking. The LM-Ease features a fully actuated tendon-driven system that seamlessly transitions between assistance actuation profiles tailored for sitting, standing, and walking movements. This device provides the user with gravity support during stand-to-sit, and offers hip extension assistance pulling force during sit-to-stand and walking. Our preliminary results show that with the LM-Ease, healthy young adults (n = 8) had significantly lower muscle activation: average reduction of 15.6% during stand-to-sit and 17.8% during sit-to-stand. Furthermore, with LM-Ease, participants demonstrated a 12.7% reduction in metabolic cost during ground walking. These evidences suggest that the LM-Ease holds potential in reducing muscular activation and energy expenditure during these fundamental daily activities. It could serve as a valuable tool for individuals seeking assistance in enhancing lower limb mobility, thereby bolstering their independence and overall quality of life.
PaperID: 224,
Authors: Junwen Gu, Jian Wang, Zhijie Liu, Min Tan, Junzhi Yu, Zhengxing Wu
Affiliations: Key Laboratory of Cognition and Decision Intelligence for Complex Systems, Institute of Automation, Chinese Academy of Sciences, Beijing, China; School of Intelligence Science and Technology, the Institute of Artificial Intelligence, and Key Laboratory of Intelligent Bionic Unmanned Systems, Ministry of Education, University of Science and Technology Beijing, Beijing, China; State Key Laboratory for Turbulence and Complex Systems, Department of Advanced Manufacturing and Robotics, College of Engineering, Peking University, Beijing, China
Abstract: In nature, fish have evolved sophisticated muscular systems that enable them to dynamically regulate their body movements for efficient and agile swimming, which has inspired the development of compact and fast flexibility regulation mechanisms in robotic fish. While existing robotic fish have primarily relied on passive flexible mechanisms and tunable stiffness mechanisms, these approaches often lack the dynamic adjustment capabilities that are characteristic of living fish. This article proposes a novel biomimetic flexible fishtail capable of dynamically controlling its deformation through artificial muscles made from macrofiber composite. In detail, the fishtail is equipped with a servo motor as the sole driving joint, while the artificial muscles regulate the deformation to indirectly adjust stiffness. A dynamic model considering both flexibility and hydrodynamics is established, and a partial differential equation observer is particularly developed to estimate the tail's full states. Subsequently, a deformation control framework incorporating a deep reinforcement learning strategy is constructed and successfully deployed on an embedded platform via lightweight design. Simulation and experimental results validate the accuracy and effectiveness of the dynamic model, observer, and control strategy. Especially, the proposed fishtail demonstrates the ability to enhance propulsion in fishlike swimming modes across various frequencies, ranging from 15% to 203%. When assembled into an untethered robotic prototype, deformation control allows the prototype's swimming speed to vary, achieving up to 42% slower or 37% faster speeds compared to passive compliance. Its rapid adjustability and adaptability to different frequencies represent significant advancements not widely reported in previous studies. The obtained results will offer some significant insights for flexible robotic systems to enhance their agility and interactivity.
PaperID: 225,
Authors: Yi Xu, Weitao Zhang, Liang Peng, Qijie Zhou, Qi Li, Qing Shi
Affiliations: Key Laboratory of Biomimetic Robots and Systems, Beijing Institute of Technology, Ministry of Education, Beijing, China
Abstract: Locusts have various motion modes among which they continuously switch in terrestrial and aerial domains, hence achieving high environmental adaptability. Several robots have been developed to mimic the jump–gliding locomotion of locusts, but their mobility and transitional stability are limited because of structural and control limitations at a small scale. In this article, we develop a small-scale locust-inspired robot (LocustBot) that can not only jump and glide but also crawl. We propose a coordinately actuated mechanism that allows LocustBot to perform jump–gliding with few actuators. To achieve the stable and long-distance moving, a reinforcement-learning-based optimized control is used to generate then track the robot's position and orientation from take-off to landing. The jump–gliding distance of LocustBot reaches 5.39 m, revealing a high-energy utilization efficiency of the mobile strategy, which combines the spring-driven jumping with the propeller-driven gliding. Remarkably, without a high platform, the robot can still achieve a far moving range by continuous crawl–jump–gliding on horizontal planes and, thus, outperforms the state-of-art jump–gliding robots.
PaperID: 226,
Authors: Ahmad Bilal Asghar, Shreyas Sundaram, Stephen L. Smith
Affiliations: DEVCOM Army Research Laboratory, Adelphi, MD, USA; School of Electrical and Computer Engineering, Purdue University, West Lafayette, IN, USA; Department of Electrical and Computer Engineering, University of Waterloo, Waterloo, ON, Canada
Abstract: In this article, we study multirobot path planning for persistent monitoring tasks. We consider the case where robots have a limited battery capacity with a discharge time D. We represent the areas to be monitored as the vertices of a weighted graph. For each vertex, there is a constraint on the maximum allowable time between robot visits, called the latency. The objective is to find the minimum number of robots that can satisfy these latency constraints while also ensuring that the robots periodically charge at a recharging depot. The decision version of this problem is known to be PSPACE-complete. We present a O\left(\frac\log D\log \log D h \log \rho\right) approximation algorithm for the problem where \rho is the ratio of the maximum and the minimum latency constraints, and h reflects the ratio of distance of vertices from the depot to their latency constraints. We also present an orienteering-based heuristic to solve the problem and show empirically that it typically provides higher quality solutions than the approximation algorithm. We extend our results to provide an algorithm for the problem of minimizing the maximum weighted latency given a fixed number of robots. We evaluate our algorithms on large problem instances in a patrolling scenario and in a wildfire monitoring application. We also compare the algorithms with an existing solver on benchmark instances.
PaperID: 227,
Authors: Sajad Shahsavari, M. Hashem Haghbayan, Antonio Miele, Eero Immonen, Juha Plosila
Affiliations: Autonomous Systems Laboratory, Department of Computing, University of Turku, Turku, Finland; Dipartimento di Elettronica, Informazione e Bioingegneria, Politecnico di Milano, Milano, Italy; Turku University of Applied Sciences, Turku, Finland
Abstract: Energy management of mechanical and cyber parts in mobile robots consists of two processes operating concurrently at runtime. Both the two processes can significantly improve the robots' battery lifetime and further extend mission time. In each process, information on energy consumption of one of the two parts is captured and analyzed to manipulate various mechanical/computational actuators in a robot, such as motor speed and CPU voltage/frequency. In this article, we show that considering management of mechanical and computational segments separately does not necessarily result in an energy-optimal solution due to their co-dependence; as a consequence, a runtime co-management scheme is required. We propose a proactive energy optimization methodology in which dynamically trained internal models are utilized to predict the future energy consumption for the mechanical and computational parts of a mobile robot, and based on that, the optimal mechanical speed and CPU voltage/frequency are determined at runtime. The experimental results on a ground wheeled robot show up to 36.34% reduction in the overall energy consumption compared to the state-of-the-art methods.
PaperID: 228,
Authors: François Heremans, Jeanne Evrard, David Langlois, Renaud Ronsse
Affiliations: Institute of Mechanics, Material and Civil Engineering, Louvain Bionics, UCLouvain, Louvain-la-Neuve, Belgium; R&D, Össur, Grjótháls , Reykjavík, Iceland
Abstract: Powered ankle–foot prostheses offer the potential to emulate natural locomotion dynamics, thereby addressing the issues related to uneven gait and insufficient propulsion typically experienced by individuals with lower limb amputation wearing a passive prosthetic device. Despite significant progress, existing powered prostheses are often hindered by their substantial build height, bulky design, excessive weight, and noise level, limiting their widespread adoption. This work presents ELSA (Efficient and Lightweight Spring Ankle), a lightweight (1.15 kg) and compact (11 cm high) powered ankle–foot prosthesis fitting within the volume of a shoe and capable of providing a net positive mechanical energy over the gait cycle. This level of integration is achieved through an innovative arrangement of a spring and actuator mechanisms operating in synergy. This hybrid architecture offers users the choice to walk actively, with propulsive energy assistance; regeneratively, potentially allowing for energy harvesting to recharge the device battery; or completely turned off (passive). This prototype has been validated during benchtop experiments and through trials involving four amputated participants. These tests encompassed various scenarios, including treadmill walking and everyday ambulation tasks. In addition, a sensitivity analysis was conducted to assess how different control parameters impacted the provided mechanical energy and resulting gait performance.
PaperID: 229,
Authors: Kechun Xu, Zhongxiang Zhou, Jun Wu, Haojian Lu, Rong Xiong, Yue Wang
Affiliations: Zhejiang University, Hangzhou, China
Abstract: We focus on the task of unknown object rearrangement, where a robot is supposed to reconfigure the objects into a desired goal configuration specified by an RGB-D image. Recent works explore unknown object rearrangement systems by incorporating learning-based perception modules. However, they are sensitive to perception error, and pay less attention to task-level performance. In this article, we aim to develop an effective system for unknown object rearrangement amidst perception noise. We theoretically reveal that the noisy perception impacts grasp and place in a decoupled way, and show such a decoupled structure is valuable to improve task optimality. We propose grasp, see, and place (GSP), a dual-loop system with the decoupled structure as prior. For the inner loop, we learn a see policy for self-confident in-hand object matching. For the outer loop, we learn a grasp policy aware of object matching and grasp capability guided by task-level rewards. We leverage the foundation model CLIP for object matching, policy learning, and self-termination. A series of experiments indicate that GSP can conduct unknown object rearrangement with higher completion rates and fewer steps.
PaperID: 230,
Authors: Victor Klemm, Yvain de Viragh, David Rohr, Roland Siegwart, Marco Tognon
Affiliations: ASL, ETH Zürich, Zürich, Switzerland
Abstract: Recent years have seen a steady rise in the abilities of wheeled–legged balancing robots. Yet, their use is still severely restricted by the lack of efficient control algorithms for overcoming obstacles such as stairs. We take a considerable step toward closing this gap by presenting a fast trajectory optimizer for generating trajectories over a large class of challenging terrains. By limiting the underlying modeling to the planar, nonlinear rigid-body dynamics and subdividing the terrain into contact-phases, a tractable nonlinear programming problem is obtained. The model explicitly accounts for contact switches and impacts, traction limits, and actuation bounds. By introducing an arc-length-related parametrization, the trajectories are rendered inherently contact constraint-consistent. We apply our method to the specific case of the wheeled bipedal robot Ascento, for which we derive closed-form expressions of the dynamics equations, including the kinematic loops. To track the trajectories, we propose a simple LQR-based controller. The approach is validated in real-world experiments where we show the execution of trajectories for traversing steps, driving up ramps, jumping, standing up, and driving up entire stairways. To the best of our knowledge, enabling the latter by means of trajectory optimization (TO) is a novelty for wheeled–legged robots.
PaperID: 231,
Authors: Hyunkyu Park, Woojong Kim, Sangha Jeon, Youngjin Na, Jung Kim
Affiliations: Department of Mechanical Engineering, Korea Advanced Institute of Science and Technology, Daejeon, South Korea; Department of Mechanical Systems Engineering, Sookmyung Women's University, Seoul, South Korea
Abstract: Electrical impedance tomographic (EIT) tactile sensing holds great promise for whole-body coverage of contact-rich robotic systems, offering extensive flexibility in sensor geometry. However, low spatial resolution restricts its practical use, despite the existing deep-learning-based reconstruction methods. This study introduces EIT-GNN, a graph-structured data-driven EIT reconstruction framework that achieves super-resolution in large-area tactile perception on unbounded form factors of robots. EIT-GNN represents the arbitrary sensor shape into mesh connections, then employs a twofold architecture of transformer encoder and graph convolutional neural network to best manage such the geometrical prior knowledge, resulting in the accurate, generalized, and parameter-efficient reconstruction procedure. As a proof-of-concept, we demonstrate its application using large-area face-shaped sensor hardware, which represents one of the most complex geometries in human/humanoid anatomy. An extensive set of experiments, including simulation study, ablation analysis, single-touch indentation test, and latent feature analysis, confirm its superiority over alternative models. The beneficial features of the approach are demonstrated through its application in active tactile-servo control of humanoid head motion, paving the new way for integrating tactile sensors with intricate designs into robotic systems.
PaperID: 232,
Authors: Christopher K. Fourie, Nadia Figueroa, Julie A. Shah
Affiliations: Department of Aeronautics and Astronautics, Massachusetts Institute of Technology, Cambridge, MA, USA; Mechanical Engineering and Applied Mechanics, University of Pennsylvania, Philadelphia, PA, USA
Abstract: References were removed from the final submission that were part of the accepted paper. There were also two duplicative references.
PaperID: 233,
Authors: Lijun Han, Jinyu Zhang, Hesheng Wang
Affiliations: Department of Automation, Key Laboratory of System Control and Information Processing of Ministry of Education, State Key Laboratory of Avionics Integration and Aviation System-of-Systems Synthesis, Shanghai Jiao Tong University, Shanghai, China
Abstract: In this article, we propose a two-stage shared control framework for physical human–robot interaction (pHRI) that addresses the inconsistency of human–robot commands and consider the influence of environmental information. In the human–robot–environment system, based on the human intention measured by the interaction force, autonomy will actively initiate the replanning when the human control intention is strong, generating a feasible local desired trajectory of the robot. At the same time, we define an index called predicted safety index (PSI) to measure the safety of the system status. When the human has control intention but does not reach the threshold, we propose a shared controller based on cooperative-game theory and PSI. Specially, it is designed within the model predictive control framework, utilizing cooperative game theory to analyze human–robot interaction behavior and treating the Pareto optimal solution as the control input. We conduct comparative experiments to evaluate the assistive performance of the proposed shared control algorithm through a waypoint tracking task with naive human users. User study with objective and subjective measures demonstrate that the algorithm effectively reduces human effort while maintaining tracking accuracy, thus enhancing both performance and safety.
PaperID: 234,
Authors: Yanfei Cao, Mingxue Cai, Bonan Sun, Zhaoyang Qi, Junnan Xue, Yihang Jiang, Bo Hao, Jiaqi Zhu, Xurui Liu, Chaoyu Yang, Li Zhang
Affiliations: Department of Mechanical and Automation Engineering, The Chinese University of Hong Kong, Hong Kong SAR, China; Guangdong Provincial Key Laboratory of Robotics and Intelligent System and the CAS Key Laboratory of Human-Machine Intelligence-Synergic Systems, Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences, Shenzhen, China; Department of Mechanical and Automation Engineering, the Department of Surgery, CUHK T Stone Robotics Institute, Hong Kong SAR, China
Abstract: Magnetic continuum robots (MCRs) have become popular owing to their inherent advantages of easy miniaturization without requiring complicated transmission structures. The evolution of MCRs, from initial designs with one embedded magnet to current designs with specific magnetization profile configurations (MPCs), has significantly enhanced their dexterity. While much progress has been achieved, the quantitative index-based evaluation of deformability for different MPCs, which can assist in designing MPCs with enhanced robot deformability, has not been addressed before. Here, we use “deformability” to describe the capability for body deflection when an MCR forms different global shapes under an external magnetic field. Therefore, in this article, we propose methodologies to design and control an MCR composed of modular axially magnetized segments. To guide robot MPC design, for the first time, we introduce a quantitative index-based evaluation strategy to analyze and optimize robot deformability. In addition, a control framework with neural network-based controllers is developed to endow the robot with two control modes: the robot tip position and orientation (M_1) and the global shape (M_2). The excellent performance of the learnt controllers in terms of computation time and accuracy was validated via both simulation and experimental platforms. In the experimental results, the best closed-loop control performance metrics, indicated as the mean absolute errors, were 0.254 mm and 0.626^\circ for mode M_1 and 1.564 mm and 0.086^\circ for mode M_2.
PaperID: 235,
Authors: Yangjun Liu, Sheng Liu, Binghan Chen, Zhi-Xin Yang, Sheng Xu
Affiliations: State Key Laboratory of Internet of Things for Smart City, Centre for Artificial Intelligence and Robotics, Department of Electromechanical Engineering, University of Macau, Macau, China; Harbin Institute of Technology, Shenzhen, China; Guangdong Provincial Key Laboratory of Robotics and Intelligent System, Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences, Shenzhen, China
Abstract: Most prior robot learning methods focus on image-based observations, limiting their capability in 3-D robotic manipulation. Voxel representation naturally delivers rich spatial features but remains underutilized. Specifically, current voxel-based methods struggle with fine-grained tasks, since precise actions are not fully achievable. However, humans can accomplish these tasks well using vision and proprioception. Inspired by this, this article proposed a novel Fusion-Perception-to-Action Transformer (FP2AT) with cross-layer feature aggregation to handle fine-grained manipulation in 3-D space. In particular, a multiscale 3-D visual fusion attention mechanism is devised to draw attention to local regions of interest and maintain awareness of global scenes, thereby boosting the capabilities of visual perception and action planning. Meanwhile, a 3-D visual mutual attention mechanism is designed and it can also enhance spatial perception. Besides, we further explore the potential of FP2AT by developing its coarse-to-fine version, which progressively refines the action space for more precise predictions. In addition, a proprioceptive encoder is developed to mimic the perception of body movements and contact, elevating the effectiveness of the FP2AT. Furthermore, a new metric, the average number of key actions (ANKA), is introduced to evaluate efficiency and planning capability. In various simulated and real-robot examples, our methods significantly outperform state-of-the-art 3-D-vision-based methods in success rate and ANKA metrics.
PaperID: 236,
Authors: Bryan Habas, Bo Cheng
Affiliations: Department of Mechanical Engineering, Biological and Robotic Intelligent Fluid Locomotion Lab, The Pennsylvania State University, University Park, PA, USA
Abstract: Inverted landing is a routine behavior among a number of animal fliers. However, mastering this feat poses a considerable challenge for robotic fliers, especially to perform dynamic perching with rapid body rotations (or flips) and landing against gravity. Inverted landing in flies have suggested that optical flow senses are closely linked to the precise triggering and control of body flips that lead to a variety of successful landing behaviors. Building upon this knowledge, we aimed to replicate the flies' landing behaviors in small quadcopters by developing a control policy general to arbitrary ceiling-approach conditions. First, we employed reinforcement learning in simulation to optimize discrete sensory-motor pairs across a broad spectrum of ceiling-approach velocities and directions. Next, we converted the sensory-motor pairs to a two-stage control policy in a continuous optical flow space augmented by ceiling distance measurement. The control policy consists of a first-stage Flip-Trigger Policy, which employs a one-class support vector machine, and a second-stage Flip-Action Policy, implemented as a feed-forward neural network. To transfer the inverted-landing policy to physical systems, we utilized domain randomization and system identification techniques for a zero-shot sim-to-real transfer with emulated optical flow using external motion tracking. As a result, we successfully achieved a range of robust inverted-landing behaviors in small quadcopters, emulating those observed in flies.
PaperID: 237,
Authors: Shamak Dutta, Nils Wilde, Stephen L. Smith
Affiliations: Department of Electrical and Computer Engineering, University of Waterloo, Waterloo, ON, Canada; Department of Computer Science, Dalhousie University, Halifax, Canada
Abstract: We study informative path planning for active regression in Gaussian Processes (GP). Here, a resource constrained robot team collects measurements of an unknown function, assumed to be a sample from a GP, with the goal of minimizing the trace of the M-weighted expected squared estimation error covariance (where M is a positive semidefinite matrix) resulting from the GP posterior mean. While greedy heuristics are a popular solution in the case of length constrained paths, it remains a challenge to compute optimal solutions in the discrete setting subject to routing constraints. We show that this challenge is surprisingly easy to circumvent. Using the optimality of the posterior mean for a class of functions of the squared loss yields an exact formulation as a mixed integer program. We demonstrate that this approach finds optimal solutions in a variety of settings in seconds and when terminated early, it finds sub-optimal solutions of higher quality than existing heuristics.
PaperID: 238,
Authors: Alexander Antonio Oliva, Maarten Jongeneel, Alessandro Saccon
Affiliations: Department of Mechanical Engineering, Eindhoven University of Technology (TU/e), Eindhoven, The Netherlands
Abstract: Active suction cups are widely adopted in industrial and logistics automation. Despite that, validated dynamic models describing their 6D force/torque interaction with objects are rare. This work aims at filling this gap by showing that it is possible to employ a compact model for suction cups, providing good accuracy also for large deformations. Its potential use is for advanced manipulation, planning, and control. We model the interconnected object-suction cup system as a lumped 6D mass-spring-damper systems, employing a potential energy function on \text SE(3), parametrized by a 6× 6 stiffness matrix. By exploiting geometric symmetries of the suction cup, we reduce the parameter identification problem, from 6(6+1) / 2 = 21 to only \boldsymbol 5 independent parameters, greatly simplifying the parameter identification procedure, that is otherwise ill-conditioned. Experimental validation is provided and data is shared openly to further stimulate research. As an indication of the achievable pose prediction in steady state, for an object of about \boldsymbol 1.75 kg, we obtain a pose error in the order of \boldsymbol 5 mm and \boldsymbol 3 deg, with a gripper inclination of \boldsymbol 60 deg.
PaperID: 239,
Authors: Jeongseob Lee, Doyoon Kong, Hojun Cha, Jeongmin Lee, Dongseok Ryu, Hocheol Shin, Dongjun Lee
Affiliations: Department of Mechanical Engineering, IAMD and IOER, Seoul National University, Seoul, South Korea; Korea Atomic Energy Research Institute, Daejeon, South Korea
Abstract: We propose a novel high-force/high-precision interaction control framework of a dual-arm robot system on a flexible base, with one arm holding, or making contact with, a supporting surface, while the other arm can exert any arbitrary wrench in a certain polytope through a desired pose against environments or objects. Our proposed framework can achieve high-force/precision tasks by utilizing the supporting surface just as we humans do while taking into account various important constraints (e.g., system stability, joint angle/torque limits, friction-cone constraint, etc.) and the passive compliance of the flexible base. We first design the control as a combination of: 1) nominal control; 2) active stiffness control; and 3) feedback wrench control. We then sequentially perform optimizations of the nominal configuration (and its related wrenches) and the active stiffness control gain. We also design the proportional–integral type feedback wrench control to improve the robustness and precision of the control. The key theoretical enabler for our framework is a novel stiffness analysis of the dual-arm system with flexibility, which, when combined with certain constraints, provides some peculiar relations, that can effectively be used to significantly simplify the optimization problem-solving and to facilitate the feedback wrench control design by manifesting the compliance relation at the interaction port. The efficacy of the theory is then validated and demonstrated through simulations and experiments.
PaperID: 240,
Authors: Zhengyuan Xin, Shihao Zhong, Anping Wu, Zhiqiang Zheng, Qing Shi, Qiang Huang, Toshio Fukuda, Huaping Wang
Affiliations: Intelligent Robotics Institute, School of Mechatronical Engineering, Beijing Institute of Technology, Beijing, China; School of Medical Technology, Beijing Institute of Technology, Beijing, China; Department of Biomedical Engineering, City University of Hong Kong, Hong Kong; Beijing Advanced Innovation Center for Intelligent Robots and Systems, Beijing Institute of Technology, Beijing, China; Department of Micro-Nano Systems Engineering, Nagoya University, Nagoya, Japan; Key Laboratory of Biomimetic Robots and Systems, Beijing Institute of Technology, Ministry of Education, Beijing, China
Abstract: Soft millirobots are highly promising for biomedical applications due to their reconfigurability and multifunctionality within physiological environments. However, the diverse and narrow biological cavity environments pose significant adaptability challenges for these millirobots. Here, we present a dual-morphology, thin-film millirobot equipped with a magnetic drive head and a functional tail to facilitate multimodal motion and targeted cell delivery. The millirobot can reversibly switch between two distinct morphologies in response to environmental stimuli through the deformation of its hydrogel body. Utilizing these dual morphologies, the millirobot can perform robust multimodal fundamental motions controlled by magnetic fields. We encapsulate fundamental motions with specific programmable magnetic field parameters into motion primitives, allowing easy invocation and adjustment of motion modes on demand. A knowledge graph is established to map terrain features to motion units, enabling the identification of optimal motion modes based on typical terrain characteristics. Experimental results indicate that the millirobot can effectively switch its morphology and movement modes to navigate various terrains, including narrow and curved channels as small as 1 mm, 0.8 mm high stairs with a 15° incline, and even the complex environment of a swine intestinal lumen. Its functional tail can carry immune cells to target and kill cancer cells. This robot can transport drugs and cells while navigating complex terrains through multimodal motion, paving the way for targeted medical tasks in intricate human environments in the future.
PaperID: 241,
Authors: Mengchao Zhang, Devesh K. Jha, Arvind U. Raghunathan, Kris Hauser
Affiliations: Department of Mechanical Science and Engineering, University of Illinois at Urbana-Champaign, Urbana, IL, USA; Mitsubishi Electric Research Laboratories (MERL), Cambridge, MA, USA; School of Computing and Data Science, University of Illinois at Urbana-Champaign, Urbana, IL, USA
Abstract: Contact-implicit trajectory optimization (CITO) is an effective method to plan complex trajectories for various contact-rich systems including manipulation and locomotion. CITO formulates a mathematical program with complementarity constraints (MPCC) that enforces that contact forces must be zero when points are not in contact. However, MPCC solve times increase steeply with the number of allowable points of contact, which limits CITO's applicability to problems in which only a few, simple geometries are allowed us to make contact. This article introduces simultaneous trajectory optimization and contact selection (STOCS), as an extension of CITO that overcomes this limitation. The innovation of STOCS is to identify salient contact points and times inside the iterative trajectory optimization process. This effectively reduces the number of variables and constraints in each MPCC invocation. The STOCS framework, instantiated with key contact identification subroutines, renders the optimization of manipulation trajectories computationally tractable even for high-fidelity geometries consisting of tens of thousands of vertices.
PaperID: 242,
Authors: Maulik Bhatt, Yixuan Jia, Negar Mehr
Affiliations: Department of Mechanical Engineering, University of California at Berkeley, Berkeley, CA, USA; LIDS, Massachusetts Institute of Technology, Cambridge, MA, USA
Abstract: In interactive multiagent settings, decision-making and planning are challenging mainly due to the agents' interconnected objectives. Dynamic game theory offers a formal framework for analyzing such intricacies. Yet, solving constrained dynamic games and determining the interaction outcome in the form of generalized Nash equilibria (GNE) pose computational challenges due to the need for solving constrained coupled optimal control problems. In this article, we address this challenge by proposing to leverage the special structure of many real-world multiagent interactions. More specifically, our key idea is to leverage constrained dynamic potential games, which are games for which GNE can be found by solving a single constrained optimal control problem associated with minimizing the potential function. We argue that constrained dynamic potential games can effectively facilitate interactive decision-making in many multiagent interactions. We will identify structures in realistic multiagent interactive scenarios that can be transformed into weighted constrained potential dynamic games (WCPDGs). We will show that the GNE of the resulting WCPDG can be obtained by solving a single constrained optimal control problem. We will demonstrate the effectiveness of the proposed method through various simulation studies and show that we achieve significant improvements in solve time compared to state-of-the-art game solvers. We further provide experimental validation of our proposed method in a navigation setup involving two quadrotors carrying a rigid object while avoiding collisions with two humans.
PaperID: 243,
Authors: Vít Krátký, Robert Penicka, Jiri Horyna, Petr Stibinger, Tomás Báca, Matej Petrlík, Petr Stepan, Martin Saska
Affiliations: Department of Cybernetics, Faculty of Electrical Engineering, Czech Technical University in Prague, Prague, Czech Republic
Abstract: In this article, we introduce an algorithm designed to address the problem of time-optimal formation reshaping in three-dimensional environments while preventing collisions between agents. The utility of the proposed approach is particularly evident in mobile robotics, where agents benefit from being organized and navigated in formation for a variety of real-world applications requiring frequent alterations in formation shape for efficient navigation or task completion. Given the constrained operational time inherent to battery-powered mobile robots, the time needed to complete the formation reshaping process is crucial for their efficient operation, especially in case of multi-rotor uncrewed aerial vehicles (UAVs). The proposed collision-aware time-optimal formation reshaping algorithm (CAT-ORA) builds upon the Hungarian algorithm for the solution of the robot-to-goal assignment implementing the interagent collision avoidance through direct constraints on mutually exclusive robot-goal pairs combined with a trajectory generation approach minimizing the duration of the reshaping process. Theoretical validations confirm the optimality of CAT-ORA, with its efficacy further showcased through simulations, and a real-world outdoor experiment involving 19 UAVs. Thorough numerical analysis shows the potential of CAT-ORA to decrease the time required to perform complex formation reshaping tasks by up to 49%, and 12% on average compared to commonly used methods in randomly generated scenarios.
PaperID: 244,
Authors: Mengyue Li, Liang Li, Junjian Zhou, Lianqing Liu, Niandong Jiao
Affiliations: State Key Laboratory of Robotics, Shenyang Institute of Automation, Chinese Academy of Sciences, Shenyang, China
Abstract: Biohybrid microrobots with autonomous movement capabilities have broad application prospects in targeted delivery, attracting researchers to study their movement characteristics. However, its automatic control is still challenging, and exploring real-time detection of its environment for path planning to achieve stable closed-loop control is highly important for its practical application. Here, we applied deep learning for the detection of biohybrid microrobots and their targets and obstacles, followed by real-time path planning and trajectory tracking of biohybrid microrobots for targeted delivery. The proposed detection algorithm introduces attention and multiscale feature fusion mechanisms in YOLOv7 algorithm (AM-YOLOv7) with the aim of enhancing the precision of detecting small-scale targets when robots, obstacles and targets are displayed globally, and the detection capabilities are verified through simulations and experiments. The proposed planning algorithm introduces a turning penalty function and a path smoothing strategy into A algorithm (PS-A) to make the planned path short and smooth, which has been verified through simulation and experiments. The adaptive fuzzy PID method is used to track the robot's trajectory, and experiments and simulations show that the biohybrid microrobot can move according to the preset trajectory better. The final cell scene experimental results show that the biohybrid microrobot using this system can effectively avoid obstacle cells and be delivered to target cells. The system can detect biohybrid microrobots, obstacle cells and target cells, plan short and smooth trajectories, and track them accurately. The proposed method has certain generalizability and broad application prospects in targeted delivery.
PaperID: 245,
Authors: Lei Lei, Yu Zhou, Jianxing Zhang
Affiliations: Department of Systems Engineering, City University of Hong Kong, Hong Kong, SAR, China; College of Electrical Engineering and Automation, Fuzhou University, Fuzhou, China; Institute of Marine Mechatronics Equipment, School of Mechanical Science and Engineering, Huazhong University of Science and Technology, Wuhan, China
Abstract: Underwater robots are critical observation platforms for diverse ocean environments. However, existing robotic designs often lack long-range and deep-sea observation capabilities and overlook the effects of environmental uncertainties on robotic operations. This article presents a novel long-range underwater robot for extreme ocean environments, featuring a low-power dual-circuit buoyancy adjustment system, an efficient mass-based attitude adjustment system, flying wings, and an open sensor cabin. After that, an extended environment perception strategy with incremental updating is proposed to understand and predict full hydrological dynamics based on sparse observations. On this basis, a real-time dynamic modeling approach integrates multibody dynamics, perceived hydrological dynamics, and environment-robot interactions to provide accurate dynamics predictions and enhance motion efficiency. Extensive simulations and field experiments covering 600 km validated the reliability and autonomy of the robot in long-range ocean observations, highlighting the accuracy of the extended perception and real-time dynamics modeling methods.
PaperID: 246,
Authors: Mingxin Wei, Lanxiang Zheng, Ying Wu, Ruidong Mei, Hui Cheng
Affiliations: School of Artificial Intelligence, Sun Yat-Sen University, Zhuhai, China; School of Computer Science and Engineering, Sun Yat-Sen University, Guangzhou, China; School of Systems Science and Engineering, Sun Yat-Sen University, Guangzhou, China
Abstract: In agile quadrotor flight, accurately modeling the varying aerodynamic drag forces encountered at different speeds is critical. These drag forces significantly impact the performance and maneuverability of the quadrotor, especially during high-speed maneuvers. Traditional control models based on first principles struggle to capture these dynamics due to the complexity and variability of aerodynamic effects, which are challenging to model accurately. To address these challenges, this study proposes a meta-learning-based control strategy for accurately modeling quadrotor dynamics under varying speeds, treating each velocity condition as an independent learning task with a specifically trained neural network to ensure precise dynamic predictions. The meta-learning framework rapidly generates task-specific parameters adapted to speed variations by solving an optimization problem and employs an online incremental learning strategy to integrate real-time data for continuous model updates, enhancing system robustness. Regularization is introduced to prevent overfitting and improve generalizability. The integration of the meta-learned model into Model Predictive Contouring Control (MPCC) allows the system to achieve optimal control across different velocity levels, ensuring efficient and accurate flight control even during sharp turns and high-speed maneuvers. Extensive simulations and real-world experiments confirm that the proposed algorithm maintains a high level of control precision despite the nonlinear effects of rapid speed changes, complex flight trajectories and wind disturbances. The results highlight the advantages of combining meta-learning with adaptive control strategies, providing a robust framework for quadrotors operating in diverse and dynamic environments.
PaperID: 247,
Authors: Jiaqi Jiang, Xuyang Zhang, Daniel Fernandes Gomes, Thanh-Toan Do, Shan Luo
Affiliations: School of Aerospace Engineering, Beijing Institute of Technology, Beijing, China; Department of Engineering, King’s College London, London, U.K.; Department of Data Science and AI, Monash University, Clayton, Australia
Abstract: This article introduces RoTipBot, a novel robotic system for handling thin, flexible objects. Different from previous works that are limited to singulating them using suction cups or soft grippers, RoTipBot can count multiple layers and then grasp them simultaneously in a single grasp closure. Specifically, we first develop a vision-based tactile sensor named RoTip that can rotate and sense contact information around its tip. Equipped with two RoTip sensors, RoTipBot rolls and feeds multiple layers of thin, flexible objects into the centre between its fingers, enabling effective grasping. Moreover, we design a tactile-based grasping strategy that uses RoTip’s sensing ability to ensure both fingers maintain secure contact with the object while accurately counting the number of fed objects. Extensive experiments demonstrate the efficacy of the RoTip sensor and the RoTipBot approach. The results show that RoTipBot not only achieves a higher success rate but also grasps and counts multiple layers simultaneously—capabilities not possible with previous methods. Furthermore, RoTipBot operates up to three times faster than state-of-the-art methods. The success of RoTipBot paves the way for future research in object manipulation using mobilized tactile sensors.
PaperID: 248,
Authors: Lipu Zhou, Zhenzhong Wei, Xu Wang
Affiliations: Key Laboratory of Precision Opto-Mechatronics Technology, Ministry of Education, Beijing, China
Abstract: The perspective-n-point (PnP) problem, which estimates the camera pose through N 2-D/3-D point correspondences, has been extensively studied. Although minimizing the reprojection cost is regarded as the gold standard for solving the PnP problem, this cost lacks an analytic solution, leading previous works to focus on developing simpler costs. State-of-the-art PnP solutions are generally considered to be close to the gold-standard solution. However, this perception is based on limited experimental setups. Our extensive evaluations show that these solutions generally deviate from the gold-standard solution as the depth range of 3-D points increases. This article investigates two noise models of the PnP problem and provides a unified, accurate, and efficient solution. The main contributions of this article are threefold. First, we propose an efficient initialization method that compresses 2N constraints to three quadratic equations for rotation using principal component analysis. Second, we prove that our initialization algorithm provides a solution to the P3P problem, making it applicable to the full range N \geq 3 of the PnP problem. Third, we propose a novel iterative algorithm that approximates reprojection residuals using second-order polynomials and determines the optimal step size analytically, ensuring fast convergence. Extensive experiments on synthetic and real data demonstrate that our algorithm outperforms state-of-the-art methods in terms of accuracy and robustness, while achieving comparable efficiency.
PaperID: 249,
Authors: Jiexin Zhang, Tengyu Hou, Ye Ding, Bo Zhang, Honghai Liu
Affiliations: State Key Laboratory of Intelligent Manufacturing Equipment and Technology, Huazhong University of Science and Technology, Wuhan, China; State Key Laboratory of Mechanical System and Vibration, Shanghai Jiao Tong University, Shanghai, China; State Key Laboratory of Robotics and Systems, Harbin Institute of Technology Shenzhen, Shenzhen, China
Abstract: Increasing the viscosity of elastic joints can significantly improve the performance of elastic joint robots during physical human–robot interactions. However, current approaches for injecting viscous elements require an additional damper to be added in parallel with the elastic elements. In this article, we propose a new concept called virtual viscous element injection (VVI), which enables a robot to exhibit viscoelasticity without altering its mechanical structure. VVI relies only on motor-side dynamics reshaping and state feedback. Interestingly, the VVI method allows high-resolution joint torque measurements in elastic joint robots, unlike in physical viscoelastic joint robots, which measure joint torque using higher-order derivatives of the positions. Furthermore, the VVI method is proved to preserve the passivity of robot dynamics, which provides numerous possibilities for the applications of combined passivity-based controllers. Specifically, we first emphasize the impedance control method using VVI. The results demonstrate that the VVI-DF method, which combines the direct feedback method with VVI, addresses the issue of excessive acceleration feedback in the controller. This provides looser constraints for achieving a high-gain torque loop in impedance control. Moreover, this article also provides examples of the application of VVI combined with passivity-based position and torque controllers. Experiments and simulations demonstrate the effectiveness of the proposed methods. The proposed method can be extended to various robots, such as exoskeletons, and collaborative robots.
PaperID: 250,
Authors: Ali Noormohammadi-Asl, Stephen L. Smith, Kerstin Dautenhahn
Affiliations: Department of Electrical and Computer Engineering, Faculty of Engineering, University of Waterloo, Waterloo, ON, Canada
Abstract: Adaptive task planning is fundamental to ensuring effective and seamless human–robot collaboration. This article introduces a robot task planning framework that takes into account both human leading/following preferences and performance, specifically focusing on task allocation and scheduling in collaborative settings. We present a proactive task allocation approach with three primary objectives: 1) enhancing team performance; 2) incorporating human preferences; and 3) upholding a positive human perception of the robot and the collaborative experience. Through a user study, involving an autonomous mobile manipulator robot working alongside participants in a collaborative scenario, we confirm that the task planning framework successfully attains all three intended goals, thereby contributing to the advancement of adaptive task planning in human–robot collaboration. This article mainly focuses on the first two objectives, and we discuss the third objective, participants’ perception of the robot, tasks, and collaboration in a companion article.
PaperID: 251,
Authors: Sheng Hong, Chunran Zheng, Yishu Shen, Changze Li, Fu Zhang, Tong Qin, Shaojie Shen
Affiliations: Department of Electronic Computer Engineering, The Hong Kong University of Science and Technology, Hong Kong; Department of Mechanical Engineering, The University of Hong Kong, Hong Kong; Global Institute of Future Technology, Shanghai Jiao Tong University, Shanghai, China
Abstract: In recent years, 3-D Gaussian splatting (3D-GS) has emerged as a novel scene representation approach. However, existing vision-only 3D-GS methods often rely on hand-crafted heuristics for point-cloud densification and face challenges in handling occlusions and high graphics processing unit (GPU) memory and computation consumption. Light detection and ranging (LiDAR)-inertial-visual sensor configuration has demonstrated superior performance in precise localization and dense mapping by leveraging complementary sensing characteristics: rich texture information from cameras, precise geometric measurements from LiDAR, and high-frequency motion data from inertial measurement unit. Inspired by this, we propose a novel real-time Gaussian-based simultaneous localization and mapping system. Our map system comprises a global Gaussian map and a sliding window of Gaussians, along with an iterative error state Kalman filter (IESKF)-based real-time odometry utilizing Gaussian maps. The structure of the global Gaussian map consists of hash-indexed voxels organized in a recursive octree. This hierarchical structure effectively covers sparse spatial volumes while adapting to different levels of detail and scales in the environment. The Gaussian map is efficiently initialized through multisensor fusion and optimized with photometric gradients. Our system incrementally maintains a sliding window of Gaussians with minimal graphics memory usage, significantly reducing GPU computation and memory consumption by only optimizing the map within the sliding window, enabling real-time optimization. Moreover, we implement a tightly coupled multisensor fusion odometry with an IESKF, which leverages real-time updating and rendering of the Gaussian map to achieve competitive localization accuracy. Our system represents the first real-time Gaussian-based SLAM framework deployable on resource-constrained embedded systems (all implemented in C++/CUDA for efficiency), demonstrated on the NVIDIA Jetson Orin NX platform. The framework achieves real-time performance while maintaining robust multisensor fusion capabilities. All implementation algorithms, hardware designs, and CAD models and demo video of our GPU-accelerated system will be publicly available.
PaperID: 252,
Authors: Zhaoyuan Gu, Yuntian Zhao, Yipu Chen, Rongming Guo, Jennifer K. Leestma, Gregory S. Sawicki, Ye Zhao
Affiliations: Laboratory for Intelligent Decision and Autonomous Robots, Woodruff School of Mechanical Engineering, Georgia Institute of Technology, Atlanta, GA, USA; Physiology of Wearable Robotics Lab, Woodruff School of Mechanical Engineering, Georgia Institute of Technology, Atlanta, GA, USA
Abstract: This study introduces a robust planning framework that utilizes a model predictive control (MPC) approach, enhanced by incorporating signal temporal logic (STL) specifications. This marks the first-ever study to apply STL-guided trajectory optimization for bipedal locomotion, specifically designed to handle both translational and orientational perturbations. Existing recovery strategies often struggle with reasoning complex task logic and evaluating locomotion robustness systematically, making them susceptible to failures caused by inappropriate recovery strategies or lack of robustness. To address these issues, we design an analytical stability metric for bipedal locomotion and quantify this metric using STL specifications, which guide the generation of recovery trajectories to achieve maximum robustness degree. To enable safe and computational-efficient crossed-leg maneuver, we design data-driven self-leg-collision constraints that are 1000 times faster than the traditional inverse-kinematics-based approach. Our framework outperforms a state-of-the-art locomotion controller, a standard MPC without STL, and a linear-temporal-logic-based planner in a high-fidelity dynamic simulation, especially in scenarios involving crossed-leg maneuvers. In addition, the Cassie bipedal robot achieves robust performance under horizontal and orientational perturbations, such as those observed in ship motions. These environments are validated in simulations and deployed on hardware. Furthermore, our proposed method demonstrates versatility on stepping stones and terrain-agnostic features on inclined terrains.
PaperID: 253,
Authors: Leonard Klüpfel, Lukas Burkhard, Anne E. Reichert, Maximilian Durner, Rudolph Triebel
Affiliations: Institute of Robotics and Mechatronics, German Aerospace Center (DLR), Wessling, Germany
Abstract: We present prior knowledge robot keypoint detection (PK-ROKED), a learning-based pipeline for probabilistic robot pose estimation relative to a camera, addressing inaccuracies in forward kinematics, particularly in systems with elastic and lightweight modules. Our approach integrates a probabilistic 2-D keypoint detection mechanism that leverages prior knowledge derived from the robot’s imprecise kinematics. We further improve the detection accuracy and geometric understanding by incorporating segmentation of the robot arm. The method computes reliable uncertainty estimates, enabling a robust 2D–6D fusion for precise robot arm pose estimation from a single detected keypoint. PK-ROKED requires only synthetic training data, effectively exploits imperfect kinematics as valuable prior knowledge, and introduces a novel fusion framework for enhanced robot pose estimation. We validate our method on the Panda-Orb dataset, demonstrating competitive performance against state-of-the-art approaches. In addition, we evaluate on two other robotic systems in real-world scenarios and show its practicality by using the predictions to initialize a tracking algorithm. Code and pretrained models are available.
PaperID: 254,
Authors: Lun Luo, Si-Yuan Cao, Xiaorui Li, Jintao Xu, Rui Ai, Zhu Yu, Xieyuanli Chen
Affiliations: College of Intelligence Science and Technology, National University of Defense Technology, Hunan, China; Ningbo Innovation Center, Zhejiang University, Zhejiang, China; College of Instrument Science and Optoelectronics Engineering, Beihang University, Beijing, China; Haomo. AI Technology Company Ltd., Beijing, China; College of Information Science and Electronic Engineering, Zhejiang University, Zhejiang, China
Abstract: This article introduces BEVPlace++, a novel, fast, and robust light detection and ranging (LiDAR) global localization method for autonomous ground vehicles (AGV). It uses lightweight convolutional neural networks (CNNs) on bird’s eye view (BEV) image-like representations of LiDAR data to achieve accurate global localization through place recognition, followed by 3-degrees of freedom (DoF) pose estimation. Our detailed analyses reveal an interesting fact that CNNs are inherently effective at extracting distinctive features from LiDAR BEV images. Remarkably, keypoints of two BEV images with large translations can be effectively matched using CNN-extracted features. Building on this insight, we design a rotation equivariant module (REM) to obtain distinctive features while enhancing robustness to rotational changes. A rotation equivariant and invariant network (REIN) is then developed by cascading REM and a descriptor generator, NetVLAD, to sequentially generate rotation equivariant local features and rotation invariant global descriptors. The global descriptors are used first to achieve robust place recognition, and then local features are used for accurate pose estimation. Experimental results on seven public datasets and our AGV platform demonstrate that BEVPlace++, even when trained on a small dataset (3000 frames of KITTI) only with place labels, generalizes well to unseen environments, performs consistently across different days and years, and adapts to various types of LiDAR scanners. BEVPlace++ achieves state-of-the-art performance in multiple tasks, including place recognition, loop closure detection, and global localization. In addition, BEVPlace++ is lightweight, runs in real-time, and does not require accurate pose supervision, making it highly convenient for deployment.
PaperID: 255,
Authors: Ehsan Yousefi, Mo Chen, Inna Sharf
Affiliations: Department of Mechanical Engineering, McGill University, Montreal, QC, Canada; School of Computing Science, Simon Fraser University, Burnaby, BC, Canada
Abstract: This article presents a novel shared autonomy and baseline policy adapting framework for human–robot interactions in high-level context-aware robotic tasks. With a unique methodology that leverages hierarchies in decision-making as well as variational analysis of human policy, we propose a mathematical model of shared autonomy policy. The framework aims at interpretable high-level decision-making for efficient robot operation with human in the loop. We modeled the decision-making process using hierarchical Markov decision processes in an algorithm we called policy adapting, where the autonomous system policy is adapted, and hence shaped by incorporating design variables contextual to the robot, human, task, and pretraining. By integrating deep reinforcement learning within a multiagent hierarchical context, we present an end-to-end algorithm to train a baseline policy designed for shared autonomy. We showcase the effectiveness of our framework, and particularly the interplay between different design elements and human’s skill level, in a pilot study with a human user in a simulated sequence of high-level pick-and-place tasks. The proposed framework advances the state of the art in shared autonomy for robotic tasks, but can also be applied to other domains of autonomous operation.
PaperID: 256,
Authors: Guangming Wang, Yu Zheng, Yuxuan Wu, Yanfeng Guo, Zhe Liu, Yixiang Zhu, Wolfram Burgard, Hesheng Wang
Affiliations: Department of Engineering, University of Cambridge, Cambridge, U.K.; School of Automation and Intelligent Sensing, Shanghai Jiao Tong University, Shanghai, China; Electrical and Computer Engineering, University of California, Los Angeles, CA, USA; Computer Control and Automation, Nanyang Technological University, Singapore; Department of Computer Science and Artificial Intelligence, University of Technology Nuremberg, Nuremberg, Germany
Abstract: Robot localization using a built map is essential for a variety of tasks including accurate navigation and mobile manipulation. A popular approach to robot localization is based on image-to-point cloud registration, which combines illumination-invariant LiDAR-based mapping with economical image-based localization. However, the recent works for image-to-point cloud registration either divide the registration into separate modules or project the point cloud to the depth image to register the RGB and depth images. In this article, we present I2PNet, a novel end-to-end 2D-3D registration network, which directly registers the raw 3-D point cloud with the 2-D RGB image using differential modules with a united target. The 2D-3D cost volume module for differential 2D-3D association is proposed to bridge feature extraction and pose regression. The soft point-to-pixel correspondence is implicitly constructed on the intrinsic-independent normalized plane in the 2D-3D cost volume module. Moreover, we introduce an outlier mask prediction module to filter the outliers in the 2D-3D association before pose regression. Furthermore, we propose the coarse-to-fine 2D-3D registration architecture to increase localization accuracy. Extensive localization experiments are conducted on the KITTI, nuScenes, M2DGR, Argoverse, Waymo, and Lyft5 datasets. The results demonstrate that I2PNet outperforms the state-of-the-art by a large margin and has a higher efficiency than the previous works. Moreover, we extend the application of I2PNet to the camera-LiDAR online calibration and demonstrate that I2PNet outperforms recent approaches on the online calibration task.
PaperID: 257,
Authors: Jialei Shi, Korn Borvorntanajanya, Kaiwen Chen, Enrico Franco, Ferdinando Rodriguez y Baena
Affiliations: Hamlyn Centre for Robotic Surgery, Department of Mechanical Engineering, Imperial College London, London, U.K.; Department of Electrical and Electronic Engineering, Imperial College London, London, U.K.
Abstract: Colonoscopy is a medical procedure used to examine the inside of the colon for abnormalities, such as polyps or cancer. Traditionally, this is done by manually inserting a long, flexible tube called a colonoscope into the colon. However, this method can cause pain, discomfort, and even the risk of perforation. To address these shortcomings, advancements in technology are needed to develop safer, more intelligent colonoscopes. This article presents the design, control, and evaluation of a self-growing soft robotic colonoscope, leveraging the evertion principle. The device features a tube with an 18 mm diameter, constructed from stretchable fabric, which grows 1.6 m at the tip under pressurization. A pneumatically driven, elastomer-based manipulator enables omni-directional steering over 180° at the tip. An airtight base houses motors and spools that control the material and regulate growth speed. The robot operates in two modes: teleoperation via joysticks and autonomous navigation using sensor inputs, such as a tip-mounted camera. Thorough in-vitro experiments are conducted to assess the system’s functionality and performance. Results illustrate that the robot can achieve locomotion in confined spaces such as a colon phantom, while exerting contact forces averaging less than 0.3 N. Our soft robot shows potential for improving the safety and autonomy of colonoscopies, while reducing discomfort to patients.
PaperID: 258,
Authors: João Carvalho, An T. Le, Piotr Kicki, Dorothea Koert, Jan Peters
Affiliations: Intelligent Autonomous Systems Lab, Computer Science Department, Technical University of Darmstadt, Darmstadt, Germany; Poznan University of Technology, Poznan, Poland
Abstract: The performance of optimization-based robot motion planning algorithms is highly dependent on the initial solutions, commonly obtained by running a sampling-based planner to obtain a collision-free path. However, these methods can be slow in high-dimensional and complex scenes and produce nonsmooth solutions. Given previously solved path-planning problems, it is highly desirable to learn their distribution and use it as a prior for new similar problems. Several works propose utilizing this prior to bootstrap the motion planning problem, either by sampling initial solutions from it, or using its distribution in a maximum-a-posterior formulation for trajectory optimization. In this work, we introduce motion planning diffusion (MPD), an algorithm that learns trajectory distribution priors with diffusion models. These generative models have shown increasing success in encoding multimodal data and have desirable properties for gradient-based motion planning, such as cost guidance. Given a motion planning problem, we construct a cost function and sample from the posterior distribution using the learned prior combined with the cost function gradients during the denoising process. Instead of learning the prior on all trajectory waypoints, we propose learning a lower dimensional representation of a trajectory using linear motion primitives, particularly B-spline curves. This parametrization guarantees that the generated trajectory is smooth, can be interpolated at higher frequencies, and needs fewer parameters than a dense waypoint representation. We demonstrate the results of our method ranging from simple 2-D to more complex tasks using a 7-DOF robot arm manipulator. In addition to learning from simulated data, we also use human demonstrations on a real-world pick-and-place task. The experiment results show that diffusion models are strong priors for encoding multimodal trajectory distributions for optimization-based motion planning.
PaperID: 259,
Authors: Kithmi N. D. Widanage, Jingkang Xia, Rizuwana Parween, Hareesh Godaba, Nicolas Herzig, Romeo Glovnea, Deqing Huang, Yanan Li
Affiliations: Department of Engineering and Design, University of Sussex, Brighton, U.K.; School of Electrical Engineering, Southwest Jiaotong University, Chengdu, China; Department of Mechanical Engineering, University of Southampton, Southampton, U.K.
Abstract: Automation of abrasive machining operations has become a challenging aspect in the remanufacturing industry where it is required to conduct operations on a surface of which the exact dimensions are unknown. In such cases, skilled human workers have to step in to perform labor-intensive tasks with inconsistent quality. In existing research work, collaborative robots are used to partially automate such operations under human supervision. However, these methods do not perform learning and control simultaneously and are often affected by the interactions of the human operator. In this article, a novel learning and control scheme is proposed where the robot explores an unknown surface iteratively while achieving the desired contact control performance under supervision and occasional interference from the human operator. The unknown surface is divided into subregions, and the learning and control parameters are updated each time the robot visits each subregion. This method is independent of the path of the robot and, thus, is unaffected by the irregularities introduced by a human operator’s interactions. The proposed method is applied to force control, stiffness learning, and orientation adaptation cases. The validity of this method is shown via simulations as well as experiments conducted using a Kinova Gen3 7-degrees of freedom robot.
PaperID: 260,
Authors: Quan Khanh Luu, Dinh Quang Nguyen, Nhan Huu Nguyen, Nam Phuong Dam, Van Anh Ho
Affiliations: School of Materials Science, Japan Advanced Institute of Science and Technology, Nomi, Japan; VNU University of Engineering and Technology, Hanoi, Vietnam
Abstract: Soft-bodied robots with multimodal sensing capabilities hold promise for versatile and user-friendly robotics. However, seamlessly integrating multiple sensing functionalities into soft artificial skins remains a challenge due to compatibility issues between soft materials and conventional electronics. While vision-based tactile sensing has enabled simple and effective sensor designs for robotic touch, there has been limited exploration of this technique for intrinsic multimodal sensing in large-sized soft robot bodies. To address this gap, this article introduces a novel vision-based soft sensing technique, named ProTac, capable of operating either in tactile or proximity sensing modes. This vision-based sensing technology relies on a soft functional skin that can actively switch its optical properties between opaque and transparent states. Furthermore, this article develops efficient learning pipelines for proximity and tactile perceptions, as well as sensing strategies enabled through the timing activation of the two sensing modes. The effectiveness of the soft sensing technology is demonstrated through a soft ProTac link, which can be integrated into newly constructed or existing commercial robot arms. Results suggest that robots integrated with the ProTac link, along with rigorous control formulation can perform safe and purposeful control actions, which enhances human–robot interaction scenarios and facilitates motion control tasks that are challenging to achieve with conventional rigid links.
PaperID: 261,
Authors: Xusheng Luo, Changliu Liu
Affiliations: Robotics Institute, Carnegie Mellon University, Pittsburgh, PA, USA
Abstract: Research in robotic planning with temporal logic specifications, such as linear temporal logic (LTL), has relied on single formulas. However, as task complexity increases, LTL formulas become lengthy, making them difficult to interpret and generate, and straining the computational capacities of planners. To address this, we introduce a hierarchical structure for a widely used specification type—LTL on finite traces (LTL_f). The resulting language, termed H-LTL_f, is defined with both its syntax and semantics. We further prove that H-LTL_f is more expressive than its standard “flat” counterparts. Moreover, we conducted a user study that compared the standard LTL_f with our hierarchical version and found that users could more easily comprehend complex tasks using the hierarchical structure. We develop a search-based approach to synthesize plans for multirobot systems, achieving simultaneous task allocation and planning. This method approximates the search space by loosely interconnected subspaces, each corresponding to an LTL_f specification. The search primarily focuses on a single subspace, transitioning to another under conditions determined by the decomposition of automata. We develop multiple heuristics to significantly expedite the search. Our theoretical analysis, conducted under mild assumptions, addresses completeness and optimality. Compared to existing methods used in various simulators for service tasks, our approach improves planning times while maintaining comparable solution quality.
PaperID: 262,
Authors: Harry Holt, Roberto Armellin
Affiliations: Te Pūnaha Ātea—Space Institute, University of Auckland, Auckland, New Zealand
Abstract: Spacecraft autonomy is a major barrier to increasing the scope, ambition, and affordability of both Earth-based and deep-space missions. Reinforcement learning (RL) offers huge potential in solving this problem, however, their adoption is hampered by the lack of stability guarantees, search space size and the complexity of spacecraft optimal control problems. Control techniques, such as control Lyapunov functions and linear quadratic regulators can help the RL frameworks find the optimal solution. The combination of these controllers with RL is investigated in Clohessy-Wiltshire–Hill dynamics. Several different greedy control approaches, as well as a novel nongreedy formulation, are considered for time-optimal and fuel-optimal transfers. Comparisons with optimal control theory, particle swarm optimisation and RL-only simulations are presented, demonstrating the effectiveness of RL-enhanced control approaches.
PaperID: 263,
Authors: Matthew Cavorsi, Frederik Mallmann-Trenn, David Saldaña, Stephanie Gil
Affiliations: School of Engineering and Applied Sciences, Harvard University, Cambridge, MA, USA; Department of Informatics, King’s College, London, U.K.; Department of Computer Science and Engineering, Lehigh University, Bethlehem, PA, USA
Abstract: We are interested in the problem where robots traverse through an environment modeled by a graph of discrete sites, and an unknown subset of the multirobot team is malicious. Previous works require that each robot gathers information about the trustworthiness of all other robots, called trust observations, which can be time consuming in large networks. This article decreases the time required to estimate trustworthiness by building upon an algorithm that leverages the concept of “crowd vetting” and the opinion of trusted neighbors. This allows each robot to estimate trust in dynamic scenarios, where the team size, robot neighborhoods, and robot legitimacy can change. In particular, we employ an assumption that there exists quasi-dynamic time periods, where if a robot’s legitimacy remains fixed for a sufficient length of time, its trustworthiness can be characterized. In this setting, we develop a closed-form expression for the critical number of time-steps required for our algorithm to successfully identify the true legitimacy of each robot within a specified failure probability. We show that the number of time-steps required for robots to correctly estimate the trust of all other robots increases logarithmically with the number of robots when robots do not leverage neighboring opinions, called the direct protocol. Conversely, for most general graph topologies, the number of time-steps required remains constant as the number of robots increases when our proposed algorithm, called quasi-dynamic crowd vetting (DCV), is used, for a fixed ratio of legitimate to malicious robots. Finally, our theoretical results are successfully validated through simulated persistent surveillance tasks where robots maintain a desired distribution of robots over sites in the environment.
PaperID: 264,
Authors: Andrea Pergolini, Clara Beatriz Sanz-Morère, Chiara Livolsi, Matteo Fantozzi, Filippo Dell'Agnello, Tommaso Ciapetti, Alessandro Maselli, Andrea Baldoni, Emilio Trigili, Simona Crea, Nicola Vitiello
Affiliations: The BioRobotics Institute, Scuola Superiore Sant’Anna, Pisa, Italy; Institute of Recovery and Care of Scientific Character (IRCCS) Fondazione Don Carlo Gnocchi, Firenze, Italy; IRCCS Fondazione Don Carlo Gnocchi, Firenze, Italy
Abstract: Most individuals who experience a stroke exhibit several sensorimotor impairments that limit their independence in everyday activities. Hemiparetic gait is frequently characterized by reduced knee flexion in swing due to knee stiffness or muscle weakness and knee hyperextension or knee buckling in the stance phase. Recently, unilateral-powered orthoses have been designed to overcome the limitations of the passive knee–ankle–foot orthoses. This study presents a unilateral active knee orthosis exoskeleton, AKO-β, endowed with a series-elastic actuator and designed to assist the knee in flexion and extension movements. In this article, we describe the system mechatronic design and its characterization on the bench, the control system, and pilot experiments with three poststroke participants. The device has a weight of 1.78 kg on the user’s leg, with a lateral encumbrance of 76 mm. The pilot experiments aimed to verify the effects of the exoskeleton assistance in hemiparetic gait patterns. When walking with the device, participants on average increased the knee flexion on the paretic side by 18.70° (+44.9%) during swing and decreased knee hyperextension in stance by 4.50°, compared to walking without it. Overall, when walking with the exoskeleton, subjects showed an improved gait variable score of the paretic knee profile by 37.5% compared to walking without it. The temporal and spatial gait symmetry indices did not show clear changes, although an improvement in symmetry was observed in two of the three participants. These preliminary results suggest the potential benefits of the unilateral active knee orthosis exoskeleton to enhance and restore mobility in individuals with hemiparetic gait.
PaperID: 265,
Authors: Matthew Chignoli, Nicholas Adrian, Sangbae Kim, Patrick M. Wensing
Affiliations: Department of Mechanical Engineering, Massachusetts Institute of Technology (MIT), Cambridge, MA, USA; Department of Aerospace and Mechanical Engineering, University of Notre Dame, Notre Dame, IN, USA
Abstract: We revisit the concept of constraint embedding as a means for dealing with kinematic loop constraints during dynamics computations for rigid-body systems. Specifically, we consider the local loop constraints emerging from common actuation submechanisms in modern robotics systems (e.g., geared motors, differential drives, and four-bar mechanisms). As a complementary perspective to prior work on constraint embedding, we present an analysis that generalizes the traditional concepts of joint models and motion/force subspaces between individual rigid bodies to generalized joint models and motion/force subspaces between groups of rigid bodies subject to loop constraints. We then use these generalized concepts to derive the constraint-embedded recursive forward dynamics algorithm using multihandle articulated bodies. We demonstrate the broad applicability of the generalized joint concepts by showing how they also lead to the constraint-embedding-based recursive algorithm for inverse dynamics. Lastly, we benchmark our open-source implementation in C++ for the forward dynamics algorithm against state-of-the-art, sparsity-exploiting algorithms. Our alternative derivation is intended to make the constraint-embedding methodology more accessible to the broader robotics community, while the benchmarking study clarifies the relative strengths and limitations of constraint embedding versus sparsity-exploiting methods. Indeed, our benchmarking validates that constraint embedding outperforms the nonrecursive alternative in cases involving local kinematic loops.
PaperID: 266,
Authors: Ce Guo, Xieyuanli Chen, Zhiwen Zeng, Zirui Guo, Yihong Li, Haoran Xiao, Dewen Hu, Huimin Lu
Affiliations: College of Intelligence Science and Technology, National University of Defense Technology, Changsha, China
Abstract: Tactile and kinesthetic perceptions are crucial for human dexterous manipulation, enabling reliable grasping of objects via proprioceptive sensorimotor integration. For robotic hands, even though acquiring such tactile and kinesthetic feedback is feasible, establishing a direct mapping from this sensory feedback to motor actions remains challenging. In this article, we propose a novel glove-mediated tactile–kinematic perception–prediction framework for grasp skill transfer from human intuitive and natural operation to robotic execution based on imitation learning, and its effectiveness is validated through generalized grasping tasks, including those involving deformable objects. First, we integrate a data glove to capture tactile and kinesthetic data at the joint level. The glove is adaptable for both human and robotic hands, allowing data collection from natural human hand demonstrations across different scenarios. It ensures consistency in the raw data format, enabling evaluation of grasping for both human and robotic hands. Second, we establish a unified representation of multimodal inputs based on graph structures with polar coordinates. We explicitly integrate the morphological differences into the designed representation, enhancing the compatibility across different demonstrators and robotic hands. Furthermore, we introduce the tactile–kinesthetic spatio-temporal graph networks, which leverage multidimensional subgraph convolutions and attention-based long short-term memory (LSTM) layers to extract spatio-temporal features from graph inputs to predict node-based states for each hand joint. These predictions are then mapped to final commands through a force-position hybrid mapping. Comparative experiments and ablation studies demonstrate that our approach surpasses other methods in grasp success rate, finger coordination, contact force management, and both grasp and computational efficiency, achieving results most akin to human grasping. The robustness of our approach is also validated through multiple randomized experimental setups, and its generalization capability is tested across diverse objects and robotic hands.
PaperID: 267,
Authors: Hao Chen, Takuya Kiyokawa, Zhengtao Hu, Weiwei Wan, Kensuke Harada
Affiliations: Department of Systems Innovation, Graduate School of Engineering Science, Osaka University, Toyonaka, Japan
Abstract: Grasping unknown objects from a single view has remained a challenging topic in robotics due to the uncertainty of partial observation. Recent advances in large-scale models have led to benchmark solutions such as GraspNet-1Billion. However, such learning-based approaches still face a critical limitation in performance robustness for their sensitivity to sensing noise and environmental changes. To address this bottleneck in achieving highly generalized grasping, we abandon the traditional learning framework and introduce a new perspective: similarity matching, where similar known objects are utilized to guide the grasping of unknown target objects. We newly propose a method that robustly achieves unknown-object grasping from a single viewpoint through three key steps: 1) leverage the visual features of the observed object to perform similarity matching with an existing database containing various object models, identifying potential candidates with high similarity; 2) use the candidate models with pre-existing grasping knowledge to plan imitative grasps for the unknown target object; 3) optimize the grasp quality through a local fine-tuning process. To address the uncertainty caused by partial and noisy observation, we propose a multilevel similarity matching framework that integrates semantic, geometric, and dimensional features for comprehensive evaluation. Especially, we introduce a novel point cloud geometric descriptor, the clustered fast point feature histogram descriptor, which facilitates accurate similarity assessment between partial point clouds of observed objects and complete point clouds of database models. In addition, we incorporate the use of large language models, introduce the semioriented bounding box, and develop a novel point cloud registration approach based on plane detection to enhance matching accuracy under single-view conditions. Real-world experiments demonstrate that our proposed method significantly outperforms existing benchmarks in grasping a wide variety of unknown objects in both isolated and cluttered scenarios, showcasing exceptional robustness across varying object types and operating environments.
PaperID: 268,
Authors: Zhenliang Zheng, Ning Ding, Herbert Werner, Feng Ren, Yongyuan Xu, Wenchao Zhang, Xiaoli Hu, Jianguo Zhang, Tin Lun Lam
Affiliations: School of Science and Engineering, The Chinese University of Hong Kong (CUHK), Shenzhen, China; Shenzhen Institute of Artificial Intelligence and Robotics for Society (AIRS), Shenzhen, China; Institute of Control Systems, Hamburg University of Technology (TUHH), Hamburg, Germany
Abstract: This study introduces a novel climbing strategy, reconfigurable parallel-type cable-driven climbing designed for long-span, large-scale bridge stay cable robotic applications, which has the potential to revolutionize the stay cable inspection and maintenance practice. The proposed methodology features the development of a collaborative climbing robot squad (CCRobot-S), which builds upon the design principles of the previous CCRobot series. In this study, CCRobot-S implements a parallel-type cable-driven manipulation design, allowing for reconfigurable kinematic morphology by its movable anchor bases and realizing the capacity of crossing over the stay cables for its flying platform. The collaborative robot squad design liberates the dimensions and scales of the robot’s reachable workspace and moves the part of the robotic system that indeed needs to be moved, enhancing the working efficiency and climbing agility. This strategy also utilizes controllable adhesion instead of friction to interact with the bridge cable surface for the flying platform, realizing force multiplication for forceful manipulation. Toward bringing high efficiency and heavy-duty capacity, we propose the applicable climbing frameworks (zero-downtime climbing gait for cable inspection and spider-like climbing gait for cable maintenance) and the optimization frameworks (optimal anchor configuration for the movable anchor bases and optimal grasp arrangement for the flying gripper). This article includes the exploration of the design and climbing gaits of CCRobot-S, the formulation of the CCRobot-S model, a comprehensive analysis of its workspace, and its climbing strategy and optimization. Extensive experiments have assessed the proposed climbing strategy’s effectiveness and showcased CCRobot-S’ capabilities.
PaperID: 269,
Authors: Jonghyeok Kim, Minchang Sung, Youngjin Choi, Jonghoon Park, Wan Kyun Chung
Affiliations: Department of Mechanical Engineering, Pohang University of Science and Technology (POSTECH), Pohang, South Korea; Department of Electrical and Electronic Engineering, Hanyang University, Ansan, South Korea; Department of Robotics, Hanyang University, Ansan, South Korea; Neuromeka, Seoul, South Korea
Abstract: Impedance control is a widely adopted approach that ensures the compliant behavior of robot manipulators as they interact with their environment according to specifically designed dynamics. For tasks involving six degrees of freedom (DoF), it is crucial to appropriately manage the position and orientation of the end-effector by controlling dynamic behavior. However, describing orientational displacement and designing the corresponding rotational impedance can be challenging, especially when we use a minimal representation. The well-known minimal representation for orientation, the Euler angle, suffers from representation singularity. As a remedy, the quaternion or dual quaternion can be an alternative, but with nonminimal representations. This lack of minimal representation, which does not suffer from the representation singularity, often leads to handling the impedance design by directly defining the potential energy function in the matrix Lie group. This article proposes a framework for the six-DoF impedance control design that takes advantage of Lie group theory with minimal representation, known as the exponential coordinate. Since the exponential coordinate can be treated as the Euclidean variable within the injectivity radius, it allows for the formulation of the impedance control more systematically and familiarly. In our framework, a detour strategy is utilized; the impedance is designed in the Lie group SE(3), and the control is designed in the Lie algebra \mathfrak se(3), which is isomorphic to the vector space \mathbb R^6. The group structure of SE(3) can be maintained using the proposed conversion formula between the Lie group and the Lie algebra, called the differential of the exponential map and its time derivative, with a closed-form expression. Experiments with a 6-DoF robot manipulator verified that the proposed impedance control framework effectively reflects the SE(3) group structure and achieves the desired dynamic behavior as the functionality of the impedance control with minimal parameters.
PaperID: 270,
Authors: Sebastien Tiburzio, Tomás Coleman, Daniel Feliú-Talegon, Cosimo Della Santina
Affiliations: Department of Cognitive Robotics, Delft University of Technology, Delft, The Netherlands
Abstract: Model-based manipulation of deformable objects has traditionally dealt with objects while neglecting their dynamics, thus mostly focusing on very lightweight objects at steady state. At the same time, soft robotic research has made considerable strides toward general modeling and control, despite soft robots, and deformable objects being very similar from a mechanical standpoint. In this work, we leverage these recent results to develop a control-oriented, fully dynamic framework of slender deformable objects grasped at one end by a robotic manipulator. We introduce a dynamic model of this system using functional strain parameterizations and describe the manipulation challenge as a regulation control problem. This enables us to define a fully model-based control architecture, for which we can prove analytically closed-loop stability and provide sufficient conditions for steady state convergence to the desired state. The nature of this work is intended to be markedly experimental. We provide an extensive experimental validation of the proposed ideas, tasking a robot arm with controlling the distal end of six different cables, in a given planar position and orientation in space.
PaperID: 271,
Authors: Dongwhan Kim, Euncheol Im, Yujin Kim, Myotaeg Lim, Yisoo Lee
Affiliations: Korea Institute of Science and Technology, Seoul, South Korea; School of Electrical Engineering, Korea University, Seoul, South Korea
Abstract: This study presents a model predictive path integral (MPPI) method capable of conducting high-frequency real-time model predictive control (MPC) for robot manipulators. Real-time MPC-based manipulation holds significant potential for controlling an end-effector precisely and reactively while satisfying various constraints in dynamic environments. However, the optimization under a complex robot model and various constraints imposes a heavy computational burden, hindering the realization of high-frequency updates. To address this challenge, we propose a single-instance sampling-based MPPI algorithm and dynamic time horizon to significantly reduce the computational burden while enhancing control performance. The performance and efficacy of the proposed method are verified through experiments conducted on a 7-degree-of-freedom robotic arm, along with comparative simulations and analysis.
PaperID: 272,
Authors: Pol Mestres, Carlos Nieto-Granda, Jorge Cortés
Affiliations: Department of Mechanical and Civil Engineering, California Institute of Technology, Pasadena, CA, USA; DEVCOM U.S. Army Research Laboratory (ARL), Adelphi, MA, USA; Department of Mechanical and Aerospace Engineering, University of California San Diego, La Jolla, CA, USA
Abstract: This article considers the problem of designing motion planning algorithms for control-affine systems that generate collision-free paths from an initial to a final destination and can be executed using safe and dynamically feasible controllers. We introduce the compatible control lyapunov function control barrier function rapidly exploring random tree (C-CLF-CBF-RRT) algorithm, which produces paths with such properties and leverages rapidly exploring random trees (RRTs), control Lyapunov functions (CLFs), and control barrier functions (CBFs). For linear systems with polytopic and ellipsoidal constraints, C-CLF-CBF-RRT requires solving a quadratically constrained quadratic program at every iteration of the algorithm, which can be done efficiently. We prove the probabilistic completeness of C-CLF-CBF-RRT and showcase its performance in simulation and hardware experiments.
PaperID: 273,
Authors: Yongsheng Luo, Zhaokun Guo, Tao Liu, Kaixuan Li, Jinnong Liao, Lefan Guo, Yanhe Zhu, Gangfeng Liu, Jie Zhao
Affiliations: State Key Laboratory of Robotics and System, Harbin Institute of Technology, Harbin, China
Abstract: Existing polar robots are constrained by limited energy supply, making it difficult to carry out long-term scientific exploration missions, which highlights an urgent demand for energy conservation. An energy-efficient multimode motion polar robot is proposed to address this challenge. Both increasing external assistance and reducing the driving force are critical for lowering energy consumption. A foldable sail is designed to provide external assistance. When unfolded, the sail generates assistive force. When folded, it maintains stability in extreme polar climates. The sail shape is designed based on a symmetrically extended NACA0018 airfoil, and the influence of different sail parameters on performance is discussed. The transformable tracks realize switching between traction and sliding modes through the separation of the track and teeth chain, using the sliding mode to reduce driving force. The effect of teeth parameter variations on traction performance is analyzed. The system kinematics and dynamics are model, and stability conditions are determined. Based on this, an energy-saving motion control framework for multimode motion is proposed. Finally, experiments are conducted to evaluate the energy-saving contribution of each independent mode under different configurations. Comprehensive experiments in multimode motion demonstrate an overall energy-saving rate of approximately 24%, verifying the effectiveness of the energy-saving motion control strategy. With its energy-saving advantages, this robot shows strong potential for enabling long-term scientific exploration in polar regions.
PaperID: 274,
Authors: Cheonghwa Lee, Hyeongwon Kim, Midum Oh, Kisu Ok, Sung-Hoon Ahn
Affiliations: Department of the Electrical and Computer Engineering, Seoul National University, Seoul, South Korea; Department of Mechanical Engineering, Seoul National University, Seoul, South Korea; Department of Electrical and Computer Engineering, Seoul National University, Seoul, South Korea
Abstract: Global trend in robotics has shifted toward deploying humanoid robots and mobile manipulators in industrial settings to automate repetitive and structured tasks traditionally performed by human workers. However, most tools and equipment are designed for human hands, and current grippers or end-effectors are highly specialized, limiting their ability to fully replace human handling of simple tools and tasks. This study proposes a novel frictional and prismatic pin-array gripper developed for universal gripping and tool manipulation. A pin-array structure of the gripper mimics the behavior of soft grippers while incorporating rigid components, enabling adaptability to various shapes and sizes. Each pin features semiautomatic actuation through a compression spring, supporting the underactuated mechanism. Most existing studies on grippers focus on simple pick-and-place tasks, whereas the proposed gripper extends functionality to practical tool usage. Enabled by the pin-array structure, it provides increased contact surface and support points, ensuring stable gripping and enhanced manipulation performance. In the evaluation, the pin-array gripper achieved a payload capacity of 2400 g, significantly outperforming the conventional RG2-FT gripper and the frictional flat gripper, which reached maximum capacities of 800 and 400 g, respectively. It also exhibited higher grasping forces, measuring 1.17 times greater than the RG2-FT gripper and up to 23 times greater than the frictional flat gripper. For tool manipulation, the pin-array gripper exhibited significantly lower manipulation errors, with 21.67 and 6.59 times fewer errors than the RG2-FT and flat grippers, respectively, when handling the hammer, and 7.69 and 4.45 times fewer for the metal file. In addition, qualitative demonstrations in universal gripping, omnidirectional gripping, and tool usage further validated the gripper’s performance in mobile manipulator tasks.
PaperID: 275,
Authors: Min Jin Yang, Hyunjo Chung, Yoonjin Kim, Kyungseo Park, Jung Kim
Affiliations: Department of Mechanical Engineering, Korea Advanced Institute of Science and Technology, Daejeon, South Korea; Department of Robotics and Mechatronics Engineering, Daegu Gyeongbuk Institute of Science and Technology (DGIST), Daegu, South Korea
Abstract: Robotic systems start to coexist around humans but cannot physically interact as humans do due to the absence of tactile sensitivity across their bodies. Various studies have developed a scalable tactile sensor to grant a body-scale robotic skin, yet many faced drawbacks arising from the rapidly increasing number of sensing elements or a limited sensibility to a wide range of touches. This article proposes a body-scale robotic skin composed of multimodal sensing modules and a multilayered fabric, simultaneously utilizing superresolution and tomographic transducing mechanisms. These mechanisms employ fewer sensing elements across a large area and complement each other in perceiving a wide range of stimuli humans can sense. Their measurements are processed to encode spatiotemporal properties of touch, which are decoded by a trained convolutional neural network to classify the touch modality, while their computational costs are minimized for on-device computation. The robotic skin was demonstrated on a commercial robotic arm and interpreted human touches for tactile communication, suggesting its capability as a body-scale robotic skin for further physical interaction.
PaperID: 276,
Authors: Wanxin Chen, Bi Zhang, Xiaowei Tan, Yiwen Zhao, Lianqing Liu, Xingang Zhao
Affiliations: State Key Laboratory of Robotics, Shenyang Institute of Automation, Chinese Academy of Sciences, Shenyang, China
Abstract: While rehabilitation exoskeletons have been extensively studied, systematic design principles for effectively addressing heterogeneous bilateral locomotion in hemiplegia patients are poorly understood. In this article, a multijoint lower exoskeleton driven by series elastic actuators (SEAs) is developed, and the design philosophy of rehabilitation robots for hemiplegia patients is systematically explored. The exoskeleton has six powered joints for both lower limbs in a hip–knee–ankle configuration, and each joint incorporates a custom, lightweight SEA module. A unified interaction-oriented control framework is designed for exoskeleton-assisted walking, including gait generation, task scheduling, and advanced joint-level control. The closed-loop design provides methodical solutions to address hemiplegia rehabilitation needs and provides walking assistance for bilateral lower limbs. Moreover, a multitemplate gait generation approach is proposed to address the altered kinematics induced by exoskeleton-assisted walking and enhance the exoskeleton's adaptability to patient-specific kinematic variations in an iterative manner. Experiments are conducted with both healthy individuals and hemiplegia patients to verify the effectiveness of the exoskeleton system. The clinical outcomes demonstrate that the exoskeleton can achieve mechanical transparency, facilitate movement, and enable coordinated interjoint locomotion for bilateral gait assistance.
PaperID: 277,
Authors: Can Zhao, Jin Liu, Daolin Ma
Affiliations: School of Ocean and Civil Engineering, Shanghai Jiao Tong University, Shanghai, China; School of Mechanical Engineering, Shanghai Jiao Tong University, Shanghai, China
Abstract: Vision-based tactile sensors offer rich tactile information through high-resolution tactile images, enabling the reconstruction of dense contact force fields on the sensor surface. However, accurately reconstructing the 3-D contact force distribution remains a challenge. In this article, we propose the multilayer inverse finite-element method (iFEM2.0) as a robust and generalized approach to reconstruct dense contact force distribution. We systematically analyze various parameters within the iFEM2.0 framework, and determine the appropriate parameter combinations through simulation and in situ mechanical calibration. Our approach incorporates multilayer mesh constraints and ridge regularization to enhance robustness. Furthermore, as no off-the-shelf measurement equipment or criterion metrics exist for 3-D contact force distribution perception, we present a benchmark covering accuracy, fidelity, and noise resistance that can serve as a cornerstone for other future force distribution reconstruction methods. The proposed iFEM2.0 demonstrates good performance in both simulation- and experiment-based evaluations. Such dense 3-D contact force information is critical for enabling dexterous robotic manipulation that handles both rigid and soft materials.
PaperID: 278,
Authors: Sicong Pan, Hao Hu, Hui Wei, Nils Dengler, Tobias Zaenker, Murad Dawood, Maren Bennewitz
Affiliations: Laboratory of Algorithms for Cognitive Models, School of Computer Science, Fudan University, Shanghai, China; Humanoid Robots Lab, University of Bonn, Bonn, Germany
Abstract: Existing view planning systems either adopt an iterative paradigm using next-best views (NBV) or a one-shot pipeline relying on the set-covering view-planning (SCVP) network. However, neither of these methods can concurrently guarantee both high-quality and high-efficiency reconstruction of 3-D unknown objects. To tackle this challenge, we introduce a crucial hypothesis: with the availability of more information about the unknown object, the prediction quality of the SCVP network improves. There are two ways to provide extra information: first, leveraging perception data obtained from NBVs, and second, training on an expanded dataset of multiview inputs. In this work, we introduce a novel combined pipeline that incorporates a single NBV before activating the proposed multiview-activated (MA-)SCVP network. The MA-SCVP is trained on a multiview dataset generated by our long-tail sampling method, which addresses the issue of unbalanced multiview inputs and enhances the network performance. Extensive simulated experiments substantiate that our system demonstrates a significant surface coverage increase and a substantial 45% reduction in movement cost compared to state-of-the-art systems. Real-world experiments justify the capability of our system for generalization and deployment.
PaperID: 279,
Authors: David Rohr, Olov Andersson, Nicholas R. J. Lawrance, Thomas Stastny, Roland Siegwart
Affiliations: Autonomous Systems Lab, ETH Zurich, Zurich, Switzerland
Abstract: By combining rotary- and fixed-wing flight, hybrid uncrewed aerial vehicles (H-UAVs) can uniquely address missions combining long-range aerial transport and precise ground-relative tasks, such as the placement or retrieval of payloads. However, to leverage their full maneuverability, first, the fundamentally different operating modes of rotary- and fixed-wing vehicles need to be unified and second, the system be controlled precisely despite complex aerodynamic effects. This work presents a general and lightweight, cascaded control formulation for such versatile and accurate operation of H-UAVs. First, a novel guidance law unifies ground- and air-relative position control modes typical for the individual flight regimes. Second, we formulate a jerk-level feedback-linearization to accurately track the guidance outputs despite model errors and disturbances. In extensive real flight tests with a tiltwing H-UAV, we demonstrate the versatile allocation of (hybrid) flight states and the overall accuracy enabled by the control system. Position errors remain below 0.5 m (one quarter of the wingspan) in the full flight envelope, including accelerated maneuvers up to 10 ms2 and gusting wind reaching 12 m/s. Finally, the control system demonstrates exploiting hybrid flight for transport-related missions with a precise, in-flight pickup of a payload.
PaperID: 280,
Authors: Jian Hou, Xin Zhou, Adam Spiers
Affiliations: Manipulation and Touch Lab, Dept. Electrical and Electronic Engineering, Imperial College London, London, U.K.
Abstract: The adoption of tactile sensors in robotics is hindered by their high cost and fragility. We designed and validated a cost-effective and robust barometric tactile sensor array, whose material cost is below 80 USD. Unlike past work, we do not mold the rubber surface over the barometers but instead keep it as a separate element, leading to a design that is easy to fabricate and repair. Machine learning techniques are applied to enhance the sensor's localization precision, increasing the effective resolution from 6 mm (the distance between adjacent barometers) to 0.284 mm. To investigate the localization model's robustness, we utilized an E-TRoll robotic gripper to roll differently shaped prismatic objects across the sensing surface mounted on one finger. Under these uncontrolled settings, we achieved a satisfactory average real-time localization resolution of within 2.66 mm. Furthermore, we demonstrate a novel practical application: The E-TRoll mimics a one-DoF parallel gripper inferring a cube's orientation relative to the sensor. The range of orientations is split into four classes, which a trained CNN-LSTM model can predict with an 86.91% five-fold cross-validated accuracy.
PaperID: 281,
Authors: Sepehr Samavi, James R. Han, Florian Shkurti, Angela P. Schoellig
Affiliations: University of Toronto Robotics Institute, Vector Institute for Artificial Intelligence, Toronto, ON, Canada; Technical University of Munich, Munich, Germany
Abstract: Robots need to predict and react to human motions to navigate through a crowd without collisions. Many existing methods decouple prediction from planning, which does not account for the interaction between robot and human motions and can lead to the robot getting stuck. In this article, we propose safe and interactive crowd navigation (SICNav), a model predictive control (MPC) method that jointly solves for robot motion and predicted crowd motion in closed loop. We model each human in the crowd to be following an optimal reciprocal collision avoidance (ORCA) scheme and embed that model as a constraint in the robot's local planner, resulting in a bilevel nonlinear MPC optimization problem. We use a Karush–Kuhn–Tucker (KKT)-reformulation to cast the bilevel problem as a single level and use a nonlinear solver to optimize. Our MPC method can influence pedestrian motion while explicitly satisfying safety constraints in a single-robot multihuman environment. We analyze the performance of SICNav in two simulation environments and indoor experiments with a real robot to demonstrate safe robot motion that can influence the surrounding humans. We also validate the trajectory forecasting performance of ORCA on a human trajectory dataset.
PaperID: 282,
Authors: Minghao Lu, Xiyu Fan, Han Chen, Peng Lu
Affiliations: Adaptive Robotic Controls Lab (ArcLab), Department of Mechanical Engineering, The University of Hong Kong, Hong Kong; Huawei Technologies Company, Ltd., Xi'an, China
Abstract: Obstacle avoidance for uncrewed aerial vehicles (UAVs) in cluttered environments is significantly challenging. Existing obstacle avoidance for UAVs either focuses on fully static environments or static environments with only a few dynamic objects. In this article, we take the initiative to consider the obstacle avoidance of UAVs in dynamic cluttered environments in which dynamic objects are the dominant objects. This type of environment poses significant challenges to both perception and planning. Multiple dynamic objects possess various motions, making it extremely difficult to estimate and predict their motions using one motion model. The planning must be highly efficient to avoid cluttered dynamic objects. This article proposes fast and adaptive perception and planning for UAVs flying in complex dynamic cluttered environments. A novel and efficient point cloud segmentation strategy is proposed to distinguish static and dynamic objects. To address multiple dynamic objects with different motions, an adaptive estimation method with covariance adaptation is proposed to quickly and accurately predict their motions. Our proposed trajectory optimization algorithm is highly efficient, enabling it to avoid fast objects. Furthermore, an adaptive replanning method is proposed to address the case when the trajectory optimization cannot find a feasible solution, which is common for dynamic cluttered environments. Extensive validations in both simulation and real-world experiments demonstrate the effectiveness of our proposed system for highly dynamic and cluttered environments.
PaperID: 283,
Authors: AbdulAziz Y. AlKayas, Anup Teejo Mathew, Daniel Feliú-Talegon, Ping Deng, Thomas George Thuruthel, Federico Renda
Affiliations: Department Mechanical and Nuclear Engineering, Khalifa University, Abu Dhabi, UAE; Department of Computer Science, University College London, London, U.K.
Abstract: Soft robots offer remarkable adaptability and safety advantages over rigid robots, but modeling their complex, nonlinear dynamics remains challenging. Strain-based models have recently emerged as a promising candidate to describe such systems, however, they tend to be high-dimensional and time-consuming. This article presents a novel model order reduction approach for soft and hybrid robots by combining strain-based modeling with proper orthogonal decomposition (POD). The method identifies optimal coupled strain basis functions—or mechanical synergies—from simulation data, enabling the description of soft robot configurations with a minimal number of generalized coordinates. The reduced order model (ROM) achieves substantial dimensionality reduction in the configuration space while preserving accuracy. Rigorous testing demonstrates the interpolation and extrapolation capabilities of the ROM for soft manipulators under static and dynamic conditions. The approach is further validated on a snake-like hyper-redundant rigid manipulator and a closed-chain system with soft and rigid components, illustrating its broad applicability. Moreover, the approach is leveraged for shape estimation of a real six-actuator soft manipulator using only two position markers, showcasing its practical utility. Finally, the ROM's dynamic and static behavior is validated experimentally against a parallel hybrid soft-rigid system, highlighting its effectiveness in representing the high-order model and the real system. This POD-based ROM offers significant computational speed-ups, paving the way for real-time simulation and control of complex soft and hybrid robots.
PaperID: 284,
Authors: Songyuan Zhang, Oswin So, Kunal Garg, Chuchu Fan
Affiliations: Department of Aeronautics and Astronautics, Massachusetts Institute of Technology, Cambridge, MA, USA
Abstract: Distributed, scalable, and safe control of large-scale multiagent systems is a challenging problem. In this article, we design a distributed framework for safe multiagent control in large-scale environments with obstacles, where a large number of agents are required to maintain safety using only local information and reach their goal locations. We introduce a new class of certificates, termed graph control barrier function (GCBF), which are based on the well-established control barrier function theory for safety guarantees and utilize a graph structure for scalable and generalizable distributed control of MAS. We develop a novel theoretical framework to prove the safety of an arbitrary-sized MAS with a single GCBF. We propose a new training framework GCBF+ that uses graph neural networks to parameterize a candidate GCBF and a distributed control policy. The proposed framework is distributed and is capable of taking point clouds from LiDAR, instead of actual state information, for real-world robotic applications. We illustrate the efficacy of the proposed method through various hardware experiments on a swarm of drones with objectives ranging from exchanging positions to docking on a moving target without collision. In addition, we perform extensive numerical experiments, where the number and density of agents, as well as the number of obstacles, increase. Empirical results show that in complex environments with agents with nonlinear dynamics (e.g., Crazyflie drones), GCBF+ outperforms the hand-crafted CBF-based method with the best performance by up to 20% for relatively small-scale MAS with up to 256 agents, and leading reinforcement learning (RL) methods by up to 40% for MAS with 1024 agents. Furthermore, the proposed method does not compromise on the performance, in terms of goal reaching, for achieving high safety rates, which is a common tradeoff in RL-based methods. Project website: https://mit-realm.github.io/gcbfplus/
PaperID: 285,
Authors: Ying Zhang, Maoliang Yin, Wenfu Bi, Haibao Yan, Shaohan Bian, Cui-Hua Zhang, Changchun Hua
Affiliations: School of Electrical Engineering, Yanshan University, Qinhuangdao, China
Abstract: Service robots operating in unstructured environments must effectively recognize and segment unknown objects to enhance their functionality. Traditional supervised learning-based segmentation techniques require extensive annotated datasets, which are impractical for the diversity of objects encountered in real-world scenarios. Unseen object instance segmentation (UOIS) methods aim to address this by training models on synthetic data to generalize to novel objects, but they often suffer from the simulation-to-reality gap. This article proposes a novel approach (ZISVFM) for solving UOIS by leveraging the powerful zero-shot capability of the segment anything model (SAM) and explicit visual representations from a self-supervised vision transformer (ViT). The proposed framework operates in the following three stages: generating object-agnostic mask proposals from colorized depth images using SAM, refining these proposals using attention-based features from the self-supervised ViT to filter nonobject masks, and applying K-Medoids clustering to generate point prompts that guide SAM toward precise object segmentation. Experimental validation on two benchmark datasets and a self-collected dataset demonstrates the superior performance of ZISVFM in complex environments, including hierarchical settings such as cabinets, drawers, and handheld objects.
PaperID: 286,
Authors: Yeqing Yuan, Weichao Sun
Affiliations: Research Institute of Intelligent Control and Systems, Harbin Institute of Technology, Harbin, China
Abstract: Hierarchical control is widely employed for redundant robots to manage multiple simultaneous tasks with distinct priority levels. A novel hierarchical optimal control strategy was recently introduced to achieve performance-optimal tracking under static and strict priority constraints. However, in complex and dynamic environments, robots must possess the capability to switch hierarchical behaviors online to adapt to varying operational scenarios. Existing continuous priority-switching methods often sacrifice hierarchical control performance and fail to asymptotically track the hierarchical optimal trajectory. In this article, a continuously shaping prioritized Jacobian algorithm is proposed and integrated into a newly developed continuous hierarchical optimal control framework with priority transitions. This approach not only ensures optimal control performance but also facilitates continuous priority switching. The continuity and accuracy of the proposed algorithm, as well as the bounded stability of the closed-loop system state variables, are thoroughly analyzed in this work. The effectiveness of the proposed method is validated through simulations and experiments on the Franka Emika Panda robot.
PaperID: 287,
Authors: Daniele Cattaneo, Abhinav Valada
Affiliations: Department of Computer Science, University of Freiburg, Freiburg im Breisgau, Germany
Abstract: Light detection and rangings (LiDARs) are widely used for mapping and localization in dynamic environments. However, their high cost limits their widespread adoption. On the other hand, monocular localization in LiDAR maps using inexpensive cameras is a cost-effective alternative for large-scale deployment. Nevertheless, most existing approaches struggle to generalize to new sensor setups and environments, requiring retraining or fine-tuning. In this article, we present CMRNext, a novel approach for camera-LiDAR matching that is independent of sensor-specific parameters, generalizable, and can be used in the wild for monocular localization in LiDAR maps and camera-LiDAR extrinsic calibration. CMRNext exploits recent advances in deep neural networks for matching cross-modal data and standard geometric techniques for robust pose estimation. We reformulate the point-pixel matching problem as an optical flow estimation problem and solve the perspective-n-point problem based on the resulting correspondences to find the relative pose between the camera and the LiDAR point cloud. We extensively evaluate CMRNext on six different robotic platforms, including three publicly available datasets and three in-house robots. Our experimental evaluations demonstrate that CMRNext outperforms existing approaches on both tasks and effectively generalizes to previously unseen environments and sensor setups in a zero-shot manner.
PaperID: 288,
Authors: Ruxiang Jiang, Lanhui Fu, Yanan Li, Hareesh Godaba
Affiliations: Department of Engineering and Informatics, University of Sussex, Brighton, U.K.; School of Electronics and Information Engineering, Wuyi University, Jiangmen, China; Department of Mechanical Engineering, University of Southampton, Southampton, U.K.
Abstract: The ability to precisely perceive external physical interactions would enable robots to interact effectively with the environment and humans. While vision-based tactile sensing has improved robotic grippers, it is challenging to realize high resolution vision-based tactile sensing in robot arms due to presence of curved surfaces, difficulty in uniform illumination, and large distance of sensing area from the cameras. In this article, we propose a novel piezoluminescent skin that transduces external applied pressures into changes in light intensity on the other side for viewing by a camera for pressure estimation. By engineering elastomer layers with specific optical properties and integrating a flexible electroluminescent panel as a light source, we develop a compact tactile sensing layer that resolves the layout issues in curved surfaces. We achieved multipoint pressure estimation over an expansive area of 502 cm2 with high spatial resolution, a two-point discrimination distance of 3 mm horizontally and 5 mm vertically which is comparable to that of human fingers as well as a high localization accuracy (RMSE of 1.92 mm). These promising attributes make this tactile sensing technique suitable for use in robot arms and other applications requiring high resolution tactile information over a large area.
PaperID: 289,
Authors: Yaolei Shen, Antonio Franchi, Chiara Gabellieri
Affiliations: Robotics and Mechatronics Department, Electrical Engineering, Mathematics, and Computer Science (EEMCS) Faculty, University of Twente, Enschede, AE, The Netherlands
Abstract: In this work, we present a model-based optimal boundary control design for an aerial robotic system composed of a quadrotor carrying a flexible cable. The whole system is modeled by partial differential equations combined with boundary conditions described by ordinary differential equations. The proper orthogonal decomposition (POD) method is adopted to project the original infinite-dimensional system on a finite low-dimensional space spanned by orthogonal basis functions. Based on such a reduced-order model, nonlinear model predictive control is implemented online to realize both position and shape trajectory tracking of the flexible cable in an optimal predictive fashion. The proposed POD-based reduced modeling and optimal control paradigms are verified in simulation using an accurate high-dimensional finite difference method-based model and experimentally using a real quadrotor and a cable. The results show the viability of the POD-based predictive control approach (allowing to close the control loop on the full system state) and its superior performance compared to an optimally tuned proportional–integral–derivative (PID) controller (allowing to close the control loop on the quadrotor state only).
PaperID: 290,
Authors: Iman Askari, Ali Vaziri, Xuemin Tu, Shen Zeng, Huazhen Fang
Affiliations: Department of Mechanical Engineering, University of Kansas, Lawrence, KS, USA; Department of Mathematics, University of Kansas, Lawrence, KS, USA; Department of Electrical and Systems Engineering, Washington University in St. Louis, St. Louis, MO, USA
Abstract: Model predictive control (MPC) has proven useful in enabling safe and optimal motion planning for autonomous vehicles. In this article, we investigate how to achieve MPC-based motion planning when a neural state-space model represents the vehicle dynamics. As the neural state-space model will lead to highly complex, nonlinear, and nonconvex optimization landscapes, mainstream gradient-based MPC methods will struggle to provide viable solutions due to heavy computational load. In a departure, we propose the idea of model predictive inferential control (MPIC), which seeks to infer the best control decisions from the control objectives and constraints. Following this idea, we convert the MPC problem for motion planning into a Bayesian state estimation problem. Then, we develop a new implicit particle filtering/smoothing approach to perform the estimation. This approach is implemented as banks of unscented Kalman filters/smoothers and offers high sampling efficiency, fast computation, and estimation accuracy. We evaluate the MPIC approach through a simulation study of autonomous driving in different scenarios, along with an exhaustive comparison with gradient-based MPC. The simulation results show that the MPIC approach has considerable computational efficiency despite complex neural network architectures and the capability to solve large-scale MPC problems for neural state-space models.
PaperID: 291,
Authors: Qianhao Wang, Zhepei Wang, Mingyang Wang, Jialin Ji, Zhichao Han, Tianyue Wu, Rui Jin, Yuman Gao, Chao Xu, Fei Gao
Affiliations: State Key Laboratory of Industrial Control Technology, Institute of Cyber-Systems and Control, Zhejiang University, Hangzhou, China
Abstract: Convex polytopes have compact representations and exhibit convexity, which makes them suitable for abstracting obstacle-free spaces from various environments. Existing generation methods struggle with balancing high-quality output and efficiency. Moreover, another crucial requirement for convex polytopes to accurately contain certain seed point sets, such as a robot or a front-end path, is proposed in various tasks, which we refer to as manageability. In this article, we propose fast iterative regional inflation (FIRI) to generate high-quality convex polytope while ensuring efficiency and manageability simultaneously. FIRI consists of two iteratively executed submodules: restrictive inflation (RsI) and maximum volume inscribed ellipsoid (MVIE) computation. By explicitly incorporating constraints that include the seed point set, RsI guarantees manageability. Meanwhile, iterative MVIE optimization ensures high-quality result through monotonic volume bound improvement. In terms of efficiency, we design methods tailored to the low-dimensional and multiconstrained nature of both modules, resulting in orders of magnitude improvement compared to generic solvers. Notably, in 2-D MVIE, we present the first linear complexity analytical algorithm for maximum area inscribed ellipse, further enhancing the performance in 2-D cases. Extensive benchmarks conducted against state-of-the-art methods validate the superior performance of FIRI in terms of quality, manageability, and efficiency. Furthermore, various real-world applications showcase the generality and practicality of FIRI. The high-performance code of FIRI will be open-sourced.
PaperID: 292,
Authors: Hongliang Guo, Qi Kang, Wei-Yun Yau, Chee-Meng Chew, Daniela Rus
Affiliations: Institute for Infocomm Research (IR), Agency for Science, Technology and Research (A*STAR), Singapore; National University of Singapore (NUS), Singapore; Computer Science, and Artificial Intelligence Laboratory (CSAIL), Massachusetts Institute of Technology (MIT), Cambridge, MA, USA
Abstract: This article investigates the resilient multirobot efficient search problem (R-MuRES), which aims at coordinating multiple robots to detect a “nonadversarial” moving target with the minimal expected time. One unique characteristic of R-MuRES among others is the possibility of individual robot's malfunction and withdrawal from the team during task execution, which results in a variable number of searchers in the deployment phase and entails that the possibility of team member failures must be considered during the planning stage, particularly in the training phase. We propose a resilient value function factorization (R-FAC) paradigm, which constructs the central value function from individual ones in a resilient manner, taking into account individual robots' failures, and ensures that the constructed central value function has the minimal mean squared temporal difference error across various team compositions. R-FAC stipulates that the individual global maximum principle is satisfied for whichever team configuration and thus any functioning robot contributes positively to the remaining team, as long as it executes the greedy policy with respect to the factorized individual value function. Subsequently, we introduce the variational value decomposition network (V2DN) as one of the instantiated R-FAC algorithms. V2DN employs the \log-sum-\exp mechanism to construct the central value function from individual ones, enabling it to take a varying number of robots' individual value functions as inputs. Then, we explain why, specifically for the multirobot search task, the \log-sum-\exp mechanism is superior to the brute-force summation operation used in the canonical value decomposition network (VDN), and compare V2DN with state-of-the-art MuRES solutions as well as the vanilla VDN algorithm in two canonical MuRES testing environments and show that it achieves the best resiliency score when one or several individual robots quit the team during task execution. Furthermore, we validate V2DN with a real multirobot system in a self-constructed indoor environment as the proof of concept.
PaperID: 293,
Authors: Enrica Tricomi, Giuseppe Piccolo, Federica Russo, Xiaohui Zhang, Francesco Missiroli, Sandro Ferrari, Letizia Gionfrida, Fanny Ficuciello, Michele Xiloyannis, Lorenzo Masia
Affiliations: Institut für Technische Informatik (ZITI), Heidelberg University, Heidelberg, Germany; ICAROS Center, University of Naples Federico II, Naples, Italy; Department of Informatics, Faculty of Natural Mathematics and Engineering Sciences, King's College London, London, U.K.; Akina AG, Zürich, Switzerland; Department of Computer Engineering, School of Computation, Information and Technology, Technical University of Munich (TUM), Munich, Germany
Abstract: Human beings adapt their motor patterns in response to their surroundings, utilizing sensory modalities such as visual inputs. This context-informed adaptive motor behavior has increased interest in integrating computer vision (CV) algorithms into robotic assistive technologies, marking a shift toward context aware control. However, such integration has rarely been achieved so far, with current methods mostly relying on data-driven approaches. In this study, we introduce a novel control framework for a soft hip exosuit, employing instead a physics-informed CV method grounded on geometric modeling of the captured scene for assistance tuning during stairs and level walking. This approach promises to provide a viable solution that is more computationally efficient and does not depend on training examples. Evaluating the controller with six subjects on a path comprising level walking and stairs, we achieved an overall detection accuracy of 93.0\pm 1.1%. CV-based assistance provided significantly greater metabolic benefits compared to non-vision-based assistance, with larger energy reductions relative to being unassisted during stair ascent (-18.9 \pm 4.1% versus -5.2 \pm 4.1%) and descent (-10.1 \pm 3.6% versus -4.7 \pm 4.8%). Such a result is a consequence of the adaptive nature of the device, enabled by the context aware controller that allowed for more effective walking support, i.e., the assistive torque showed a significant increase while ascending stairs (+33.9\pm 8.8%) and decrease while descending stairs (-17.4\pm 6.0%) compared to a condition without assistance modulation enabled by vision. These results highlight the potential of the approach, promoting effective real-time embedded applications in assistive robotics.
PaperID: 294,
Authors: Yidi Zhang, Fulin Tang, Yihong Wu
Affiliations: School of Artificial Intelligence, University of Chinese Academy of Sciences, Beijing, China
Abstract: A compact and consistent map of surroundings is critical for intelligent robots to understand their situations and realize robust navigation. Most existing techniques rely on infinite planes, which are sensitive to pose drift and may lead to confusing maps. Toward high-level perception in indoor environments, we propose CornerVINS, an innovative RGB-D inertial localization and layout mapping method leveraging hierarchical geometric features, i.e., points, planes, and box corners. Specifically, points are enhanced by fusing depth information, and planes are modeled as bounded patches using convex hulls to increase their discriminability. More importantly, box corners, lying at the intersection of three orthogonal planes, are parameterized with a 6-D vector and integrated into the extended Kalman filter for the first time. We introduce a hierarchical mechanism to effectively extract and associate planes and corners, which are considered as layout components of scenes and serve as long-term landmarks to correct camera poses. Extensive experiments prove that the proposed box corners bring significant improvements, enabling accurate localization and consistent layout mapping at low computational cost. Overall, the proposed CornerVINS outperforms state-of-the-art systems in both accuracy and efficiency.
PaperID: 295,
Authors: Qian Gao, Guanglin Ji, Minyi Sun, Yin Xiao, Huaiyuan Rao, Zhenglong Sun
Affiliations: School of Science and Engineering, The Chinese University of Hong Kong, Shenzhen, China; School of Electrical and Computer Engineering, Georgia Institute of Technology, Atlanta, GA, USA
Abstract: The accurate position transmission of tendon-sheath mechanisms (TSMs) is challenging but of significance to the flexible robot for minimally invasive surgery. The challenges are mainly attributed to the following: first, the tendon-elongation and its caused hysteresis that depend on the route configuration of the TSM and could result in misaligned position transmission; second, realistic surgical scenarios requiring the TSM with arbitrary and even time-varying route configurations; and third the absence of distal sensory feedback due to strict spatial constraints. Existing works are always devoted to tackling the first challenge yet evade the second and third. Here, a route-related tendon-elongation model is formulated to resolve the first challenge, and in response to the second, a route-sensing optical fiber is used. Obeying the third challenge, a feedforward hysteresis compensator is then developed to align the distal position of the tendon with the desired position. Our final contribution gives an application-oriented remedy for the foregoing methodologies. Applying our compensator on the challenging position transmission tasks subject to second and third challenges, the positional accuracy can be still maintained at around 97.50%; guided by the provided remedy, the surgical end-effector achieves submillimeter tip position accuracy. Extensive tests demonstrate that the pending concerns yet of great practical importance in existing related works are well resolved.
PaperID: 296,
Authors: Florian Voigt, Abdeldjallil Naceri, Sami Haddadin
Affiliations: Munich Institute of Robotics and Machine Intelligence, Chair of Robotics and Systems Intelligence, Technical University of Munich, München, Germany; Mohamed Bin Zayed University of Artificial Intelligence, Abu Dhabi, UAE
Abstract: In this work, we advance robotic grasping by incorporating wrist compliance in a unified hand–arm system inspired by human limb coordination. This integration improves grasping reliability and robustness through impedance and force learning in robotic arms. The compliant wrist system effectively compensates for uncertainties in object position and orientation. Employing a combined impedance-force control approach, we address diverse grasping and manipulation tasks in simulation. Successfully transferring the learned policy to a service humanoid mobile robot enables the seamless execution of grasping and opening tasks for various doors and handles without additional learning, using both fully actuated and underactuated robotic hands. Remarkably, our robust strategies yielded only one failure in 30 trials for the underactuated hand, even with up to 8 cm translation normal to the handle and 33^\circ rotation errors, and no failures for the fully actuated one with up to 12 cm translation and 30^\circ rotation. This significantly outperforms state-of-the-art end-to-end reinforcement learning approaches. Furthermore, we successfully tested and validated our approach across various constrained everyday tasks in different environments. Our proposed framework represents an advancement in the learning and execution of power grasping with compliant manipulation, achieving practically relevant performance.
PaperID: 297,
Authors: Jason J. Choi, Fernando Castañeda, Wonsuhk Jung, Bike Zhang, Claire J. Tomlin, Koushil Sreenath
Affiliations: University of California, Berkeley, CA, USA; Georgia Institute of Technology, Atlanta, GA, USA
Abstract: As the use of autonomous robots expands in tasks that are complex and challenging to model, the demand for robust data-driven control methods that can certify safety and stability in uncertain conditions is increasing. However, the practical implementation of these methods often faces scalability issues due to the growing amount of data points with system complexity and a significant reliance on high-quality training data. In response to these challenges, this study presents a scalable data-driven controller that efficiently identifies and infers from the most informative data points for implementing data-driven safety filters. Our approach is grounded in the integration of a model-based certificate function-based method and Gaussian Process regression, reinforced by a novel online data selection algorithm that reduces time complexity from quadratic to linear relative to dataset size. Empirical evidence, gathered from successful real-world cart–pole swing-up experiments and simulated locomotion of a five-link bipedal robot, demonstrates the efficacy of our approach. Our findings reveal that our efficient online data selection algorithm, which strategically selects key data points, enhances the practicality and efficiency of data-driven certifying filters in complex robotic systems, significantly mitigating scalability concerns inherent in nonparametric learning-based control methods.
PaperID: 298,
Authors: Haiming Gao, Qibo Qiu, Hongyan Liu, Dingkun Liang, Chaoqun Wang, Xuebo Zhang
Affiliations: ZJU-Hangzhou Global Scientific and Technological Innovation Center, Zhejiang University, Hangzhou, China; State Key Lab of CAD&CG, Zhejiang University, Hangzhou, China; Beijing Fuyouhua Intelligent Technology Company Ltd, Beijing, China; College of Information Engineering, Zhejiang University of Technology, Hangzhou, China; School of Control Science and Engineering, Shandong University, Jinan, China; Institute of Robotics and Automatic Information System (IRAIS), Tianjin Key Laboratory of Intelligent Robotics (TJKLIR), Nankai University, Tianjin, China
Abstract: This article presents an effective and reliable pose tracking solution, termed ERPoT, for mobile robots operating in large-scale outdoor and challenging indoor environments, underpinned by an innovative prior polygon map. Especially, to overcome the challenge that arises as the map size grows with the expansion of the environment, the novel form of a prior map composed of multiple polygons is proposed. Benefiting from the use of polygons to concisely and accurately depict environmental occupancy, the prior polygon map achieves long-term reliable pose tracking while ensuring a compact form. More importantly, pose tracking is carried out under pure LiDAR mode, and the dense 3-D point cloud is transformed into a sparse 2-D scan through ground removal and obstacle selection. On this basis, a novel cost function for pose estimation through point-polygon matching is introduced, encompassing two distinct constraint forms: point-to-vertex and point-to-edge. In this study, our primary focus lies on two crucial aspects: lightweight and compact prior map construction, as well as effective and reliable robot pose tracking. Both aspects serve as the foundational pillars for future navigation across diverse mobile platforms equipped with different LiDAR sensors in varied environments. Comparative experiments based on the publicly available datasets and our self-recorded datasets are conducted, and evaluation results show the superior performance of ERPoT on reliability, prior map size, pose estimation error, and runtime over the other six approaches. The corresponding code can be accessed online.
PaperID: 299,
Authors: Yuxiang Li, Kun Chen, Yifei Wang, Weifan Zhang, Jiancheng Wang, Haoyao Chen, Yunhui Liu
Affiliations: School of Intelligence Science and Engineering, Harbin Institute of Technology (Shenzhen), Shenzhen, China; Department of Mechanical and Automation Engineering, The Chinese University of Hong Kong, Hong Kong, China
Abstract: Autonomous ground mobile robots rely on their configuration characteristics to prevent tip-overs and collisions, ensuring safe navigation in complex environments. However, complex configurations with specially designed links and joints produce a higher dimensional workspace and bring significant challenges for path planning, especially in large-scale rough terrains. To address this, we propose a real-time multilevel terrain-aware path planning framework that integrates different levels of terrain awareness into the global and local layers. An implicit map representation is introduced at the global layer to enable efficient terrain analysis and path planning, while an iterative geometric evaluation is designed at the local layer to estimate configuration stability and improve path smoothness. By sharing the global layer information with the local layer, the framework enhances path planning efficiency and adaptability in complex environments. Its modular design supports diverse robot configurations and pathfinding algorithms, enabling effective autonomous navigation in large-scale 3-D terrains with online or offline maps. Simulations and real-world experiments demonstrated that our approach outperforms state of the art across diverse environments, including uneven terrains, multilayered structures, and complex debris fields. The results highlighted that our approach provides faster and safer path planning, more accurate and robust configuration-stability estimation, and higher success rates in traversing complex 3-D environments.
PaperID: 300,
Authors: Jonathan B. Michaux, Seth Isaacson, Challen Enninful Adu, Adam Li, Rahul Kashyap Swayampakula, Parker Ewen, Sean Rice, Katherine A. Skinner, Ram Vasudevan
Affiliations: Department of Robotics, University of Michigan, Ann Arbor, MI, USA
Abstract: Neural radiance fields and Gaussian splatting have recently transformed computer vision by enabling photorealistic representations of complex scenes. However, they have seen limited applications in real-world robotics tasks, such as trajectory optimization. This is due to the difficulty in reasoning about collisions in radiance models and the computational complexity associated with operating in dense models. This article addresses these challenges by proposing SPLANNING, a risk-aware trajectory optimizer operating in a Gaussian Splatting model. This article first derives a method to rigorously upper bound the probability of collision between a robot and a radiance field. Then, this article introduces a normalized reformulation of Gaussian splatting that enables efficient computation of this collision bound. Finally, this article presents a method to optimize trajectories that avoid collisions in a Gaussian splat. Experiments show that SPLANNING outperforms state-of-the-art methods in generating collision-free trajectories in cluttered environments. The proposed system is also tested on a real-world robot manipulator.
PaperID: 301,
Authors: Xianlong Mai, Jian Yang, Lei Li, Bin Zi, Shiwu Zhang, Xinglong Gong, Weihua Li, Guolin Yun, Shuaishuai Sun
Affiliations: CAS Key Laboratory of Mechanical Behavior and Design of Materials, Institute of Humanoid Robots, School of Engineering Sciences, University of Science and Technology of China, Hefei, China; School of Electrical Engineering and Automation, Anhui University, Hefei, China; School of Mechanical Engineering, Hefei University of Technology, Hefei, China; CAS Key Laboratory of Mechanical Behavior and Design of Materials, Department of Modern Mechanics, University of Science and Technology of China, Hefei, China; School of Mechanical, Materials, Mechatronic, and Biomedical Engineering, University of Wollongong, Wollongong, NSW, Australia
Abstract: Robotic hand exoskeletons hold immense potential for enhancing human hand functionality, addressing the hand’s strength limitations and fatigue during physically-demanding tasks. However, most existing hand exoskeletons are motorized, being weak in generating high supporting force for gripping augmentation. We present a nonmotorized hand exoskeleton based on magnetorheological (MR) actuators to provide high gripping support and elevate grip endurance. Meanwhile, it ingeniously harnesses human energy for actuation and energy storage, enhancing grip strength without external power. The MR actuator demonstrates a peak holding force of 1046 N with merely 5 W power input, boasting a force-to-power ratio one-order-of-magnitude higher than conventional approaches, and 97.7% energy reduction for same holding force compared to other approaches. Participants wearing the hand exoskeletons experience a 41.8% enhancement in grip strength without external power and reduced hand muscle fatigue during prolonged physical labor. In rescuing scenarios such as postearthquake rescue, debris clearance, and casualty evacuation, our exoskeleton effectively supports gripping and improves working efficiency.
PaperID: 302,
Authors: Huiseok Moon, Oussama Bey, Abderrahmane Boubezoul, Latifa Oukhellou, Samer Mohammed
Affiliations: LISSI, Université Paris-Est Créteil, Vitry-sur-Seine, France; SATIE, Université Gustave Eiffel, Gif-sur-Yvette, France; COSYS-GRETTIA, Université Gustave Eiffel, Marne-la-Vallée, France
Abstract: The implementation of real-time gait mode detection is paramount for providing tailored support to individuals utilizing actuated ankle-foot orthoses (AAFOs), enhancing their walking and mobility. However, existing systems often rely on multiple sensors and struggle with accurate and prompt detection of gait transitions, especially in varied environments. This study develops a novel real-time gait mode detection system that accurately identifies five daily living gait modes including level walking, ramp ascent and descent, and stair ascent and descent using only two foot-mounted inertial measurement units. A long short-term memory based algorithm, trained on data from ten healthy subjects, extracts six kinematic features to predict gait modes. The proposed method integrates this detection system with a taskoriented control strategy to adapt AAFO control according to identified gait modes. Real-time experiments with three healthy participants demonstrated robust gait mode detection, achieving an average accuracy of 98 \pm 1% across the five modes, even under assistive torque. In trials mimicking abnormal gait, the system maintained an accuracy of 93 \pm 3%. Additionally, transition delays were analyzed, showing detection can occur between transitions of the leading and trailing foot. The control strategy reduced dorsiflexor and plantar-flexor muscle activation, measured by electromyography, and improved swing phase tracking performance. Detection robustness was further evaluated by walking with obstacles and changes in environmental dimensions.
PaperID: 303,
Authors: Charles L. Clark, Biyun Xie
Affiliations: Electrical and Computer Engineering Department, University of Kentucky, Lexington, KY, USA
Abstract: The recently developed approach to motion planning in graphs of convex sets (GCS) provides an efficient framework for computing shortest-distance collision-free paths using convex optimization. This new motion planner is notably more computationally efficient than popular sampling-based motion planners, but it does not support nonconvex cost functions. This article develops a novel motion planning algorithm, graph of convex sets with general costs (GCSGC), to solve this problem. A given nonconvex cost function is accurately approximated by a multiple-layer ReLU neural network and the configuration space is decomposed into a set of linear-cost regions using the hidden layers of the neural network. These linear-cost regions are intersected with a set of collision-free regions, and the resulting collision-free linear-cost regions are intersected to form the vertices and edges of the motion planner’s underlying graph structure. The edge costs have a closed-form solution within each collision-free linear-cost region, but it is nonconvex, so the McCormick relaxation is applied to convexify the edge costs. Finally, a graph preprocessing technique is developed to compute a representative graph structure that acts as a heuristic for the edge costs of the underlying GCS and then simplify the underlying graph structure by removing cycles and high-cost paths, which can significantly improve the efficiency of the planner and quality of the produced trajectories. The proposed motion planner is first validated in a 2-D configuration space with comparisons between different sized neural networks with and without preprocessing, comparisons between optimal trajectories from GCSGC with shortest-distance trajectories, and comparisons between GCSGC and GCS-Sequential linear programming (SLP). The GCSGC planner is further validated in a complex 7-D configuration space by comparing to state-of-the-art multiquery (PRM, GCS-SLP) and single-query (TrajOpt, BIT, AIT, RRT) planners. The results show that the proposed motion planner is very competitive in terms of computational efficiency, trajectory cost, and memory footprint. Two physical experiments further validate the effectiveness of the proposed motion planner in real-world motion planning applications.
PaperID: 304,
Authors: Hui Zhao, Fuqiang Gu, Jianga Shang, Xianlei Long, Jiarui Dou, Chao Chen, Huayan Pu, Jun Luo
Affiliations: School of Geography and Information Engineering, China University of Geosciences, Wuhan, China; College of Computer Science, Chongqing University, Chongqing, China; School of Computer Science, China University of Geosciences, Wuhan, China; State Key Laboratory of Mechanical Transmissions, Chongqing University, Chongqing, China
Abstract: Visual simultaneous localization and mapping (SLAM) is crucial to many applications such as self-driving vehicles and robot tasks. However, it is still challenging for existing visual SLAM approaches to achieve good performance in low-texture or illumination-changing scenes. In recent years, some researchers have turned to edge-based SLAM approaches to deal with the challenging scenes, which are more robust than feature-based and direct SLAM methods. Nevertheless, existing edge-based methods are computationally expensive and inferior than other visual SLAM systems in terms of accuracy. In this study, we propose EdgeSLAM, a novel RGB-D edge-based SLAM approach to deal with challenging scenarios that is efficient, accurate, and robust. EdgeSLAM is built on two innovative modules: efficient edge selection and adaptive robust motion estimation. The edge selection module can efficiently select a small set of edge pixels, which significantly improves the computational efficiency without sacrificing the accuracy. The motion estimation module improves the system’s accuracy and robustness by adaptively handling outliers in motion estimation. Extensive experiments were conducted on technical university of munich (TUM) RGBD, imperial college london (ICL)-National University of Ireland Maynooth (NUIM), and ETH zurich 3D reconstruction (ETH3D) datasets, and experimental results show that EdgeSLAM significantly outperforms five state-of-the-art methods in terms of efficiency, accuracy, and robustness, which achieves 29.17% accuracy improvements with a high processing speed of up to 120 frames/s and a high positioning success rate of 97.06%.
PaperID: 305,
Authors: Rodrigue de Schaetzen, Alexander Botros, Ninghan Zhong, Kevin Murrant, Robert Gash, Stephen L. Smith
Affiliations: Department of Computer Science and Operations Research, Université de Montréal, Montréal, QC, Canada; Integrus Solutions, Sheung Wan, Hong Kong; Institute for Robotics and Intelligent Machines, Georgia Institute of Technology, Atlanta, GA, USA; National Research Council Canada, St. John’s, NL, Canada; Department of Electrical and Computer Engineering, University of Waterloo, Waterloo, ON, Canada
Abstract: Ice conditions often require ships to reduce speed and deviate from their main course to avoid damage to the ship. In addition, broken ice fields are becoming the dominant ice conditions encountered in the Arctic, where the effects of collisions with ice are highly dependent on where contact occurs and on the particular features of the ice floes. In this article, we present AUTO-IceNav, a framework for the autonomous navigation of ships operating in ice floe fields. Trajectories are computed in a receding-horizon manner, where we frequently replan given updated ice field data. During a planning step, we assume a nominal speed that is safe with respect to the current ice conditions, and compute a reference path. We formulate a novel cost function that minimizes the kinetic energy loss of the ship from ship-ice collisions and incorporate this cost as part of our lattice-based path planner. The solution computed by the lattice planning stage is then used as an initial guess in our proposed optimization-based improvement step, producing a locally optimal path. Extensive experiments were conducted both in simulation and in a physical testbed to validate our approach.
PaperID: 306,
Authors: Simon Boche, Jae-Hyung Jung, Sebastián Barbas Laina, Stefan Leutenegger
Affiliations: Mobile Robotics Lab, School of Computation, Information and Technology (CIT), Technical University of Munich (TUM), Munich, Germany; Mobile Robotics Lab, ETH Zurich, Zurich, Switzerland
Abstract: To empower mobile robots with usable maps as well as highest state estimation accuracy and robustness, we present OKVIS2-X: a state-of-the-art multisensor simultaneous localization and mapping (SLAM) system building dense volumetric occupancy maps, while scalable to large environments and operating in realtime. Our unified SLAM framework seamlessly integrates different sensor modalities: visual, inertial, measured or learned depth, LiDAR, and Global Navigation Satellite System (GNSS) measurements. Unlike most state-of-the-art SLAM systems, we advocate using dense volumetric map representations when leveraging depth or range-sensing capabilities. We employ an efficient submapping strategy that allows our system to scale to large environments, showcased in sequences of up to 9 km. OKVIS2-X enhances its accuracy and robustness by tightly-coupling the estimator and submaps through map alignment factors. Our system provides globally consistent maps, directly usable for autonomous navigation. To further improve the accuracy of OKVIS2-X, we also incorporate the option of performing online calibration of camera extrinsics. Our system achieves the highest trajectory accuracy in EuRoC against state-of-the-art alternatives, outperforms all competitors in the Hilti22 VI-only benchmark, while also proving competitive in the LiDAR version, and showcases state of the art accuracy in the diverse and large-scale sequences from the VBR dataset.
PaperID: 307,
Authors: Jiyu Cheng, Junhui Fan, Xiaolei Li, Paul L. Rosin, Yibin Li, Wei Zhang
Affiliations: School of Control Science and Engineering, Shandong University, Shandong, China; School of Computer Science and Informatics, Cardiff University, Cardiff, U.K.
Abstract: Despite significant advancements in multirobot technologies, efficiently and collaboratively exploring an unknown environment remains a major challenge. In this article, we propose AIM-Mapping, an Asymmetric InforMation enhanced Mapping framework based on deep reinforcement learning. The framework fully leverages the privileged information to help construct the environmental representation as well as the supervised signal in an asymmetric actor–critic training framework. Specifically, privileged information is used to evaluate exploration performance through an asymmetric feature representation module and a mutual information evaluation module. The decision-making network employs the trained feature encoder to extract structural information of the environment and integrates it with a topological map constructed based on geometric distance. By leveraging this topological map representation, we apply topological graph matching to assign corresponding boundary points to each robot as long-term goal points. We conduct experiments in both iGibson simulation environments and real-world scenarios. The results demonstrate that the proposed method achieves significant performance improvements compared to existing approaches.
PaperID: 308,
Authors: Ang Liu, Xianrui Zhang, Haozhi Huang, Fengqi Xiao, Zhuang Zhang, Guangming Cui, Baijin Mao, Yining Xu, Juntian Qu
Affiliations: Shenzhen International Graduate School, Tsinghua University, Shenzhen, China; Institute of AI and Robotics, Fudan University, Shanghai, China
Abstract: This study presents an intelligent bionic amphibious turtle robot (IBATR) featuring a three-degree-of-freedom bionic flipper mechanism, designed to achieve high maneuverability, agility, and adaptive locomotion in dynamic aquatic–terrestrial environments. Specifically, mechanical testing and hydrodynamic analysis validate the robot’s operational capabilities in granular media and aquatic settings. Subsequently, Bayesian optimization generates energy-efficient gait parameters, enabling flexible motion under low-power constraints. To further bridge perception and action, a terrain classification framework is implemented by fusing visual data from an onboard camera and tactile feedback from pressure sensors, enhancing environmental adaptability. This framework utilizes a dual-stream convolutional neural network, achieving 99.17% classification accuracy across four terrestrial substrates and one aquatic condition. Experimental results demonstrate that terrain-aware gait adaptation improves energy efficiency by 19.1% and movement speed by 9.2% compared to static gait configurations. Field tests under wave disturbances further confirm the robot’s capability for seamless land–water transitions. Collectively, this work advances biomimetic robotics by unifying perception-driven control, terrain-optimized actuation, and lightweight structural design, offering novel methodologies for resilient operations in complex amphibious environments.
PaperID: 309,
Authors: Jialei Shi, Hanyu Jin, Sara-Adela Abad, Wenlong Gaozhang, Ge Shi, Helge A. Wurdemann
Affiliations: Department of Mechanical Engineering, University College London, London, U.K.; Department of Mechanical Engineering, Carnegie Mellon University, Pittsburgh, PA, USA
Abstract: Elastomer-based soft manipulators with fibre-reinforced chambers, represent a prevalent design paradigm in soft robotics. These robots incorporate multiple actuation chambers, enabling elongation and bending motions. However, the inherent compliance of materials and the pressurized chambers inevitably introduce significant nonlinearity to these robots. Moreover, design of such robots often relies on a trial-and-error approach. Consequently, a comprehensive robot prototyping framework is of paramount importance. To achieve this, we present a static modeling, design and evaluation framework for soft robots with densely reinforced chambers (i.e., the angle between the reinforcement fibre and the axial direction of soft robots is \text90^\circ). We first propose a static analytical modeling framework to achieve both the forward kinematics and the tip force generation modeling. This modeling framework accommodates the effects of pressurized chambers and (non)linear material behaviors. Furthermore, our design and evaluation framework incorporates an open-accessible simulation toolbox with a user-friendly graphical interface, along with a physical evaluation platform. The entire framework is validated by eight kinds of manipulators with varying diameters and lengths. Meanwhile, the nonlinearity introduced by geometrical deformation resulting from the elongation, the pressurized actuation chambers (i.e., the chamber stiffening effect), and material hyperelasticity are investigated. Results also enable informed decision-making on design specifications prior to robot fabrication.
PaperID: 310,
Authors: Dianye Huang, Nassir Navab, Zhongliang Jiang
Affiliations: Chair for Computer Aided Medical Procedures and Augmented Reality, Technical University of Munich, Garching bei München, Germany
Abstract: Integrating generative models with action chunking has shown significant promise in imitation learning for robotic manipulation. However, the existing diffusion-based paradigm often struggles to capture strong temporal dependencies across multiple steps, particularly when incorporating proprioceptive input. This limitation can lead to task failures, where the policy overfits to proprioceptive cues at the expense of capturing the visually derived features of the task. To overcome this challenge, we propose the deep Koopman-boosted dual-branch diffusion policy (D3P) algorithm. D3P introduces a dual-branch architecture to decouple the roles of different sensory modality combinations. The visual branch encodes the visual observations to indicate task progression, while the fused branch integrates both visual and proprioceptive inputs for precise manipulation. Within this architecture, when the robot fails to accomplish intermediate goals, such as grasping a drawer handle, the policy can dynamically switch to execute action chunks generated by the visual branch, allowing recovery to previously observed states and facilitating retrial of the task. To further enhance visual representation learning, we incorporate a deep Koopman operator module that captures structured temporal dynamics from visual inputs. During inference, we use the test-time loss of the generative model as a confidence signal to guide the aggregation of the temporally overlapping predicted action chunks, thereby enhancing the reliability of policy execution. In simulation experiments across six RLBench tabletop tasks, D3P outperforms the state-of-the-art diffusion policy by an average of 14.6%. On three real-world robotic manipulation tasks, it achieves a 15.0% improvement.
PaperID: 311,
Authors: Baskin Senbaslar, Gaurav S. Sukhatme
Affiliations: NVIDIA, Santa Clara, CA, USA; Department of Computer Science, University of Southern California, Los Angeles, CA, USA
Abstract: Collision-free navigation in cluttered environments with static and dynamic obstacles is essential for many multirobot tasks. Dynamic obstacles may also be interactive, i.e., their behavior varies based on the behavior of other entities. We propose a novel representation for interactive behavior of dynamic obstacles and a decentralized real-time multirobot trajectory planning algorithm allowing interrobot collision avoidance as well as static and dynamic obstacle avoidance. Our planner simulates the behavior of dynamic obstacles, accounting for interactivity. We account for the perception inaccuracy of static and prediction inaccuracy of dynamic obstacles. We handle asynchronous planning between teammates and message delays, drops, and reorderings. We evaluate our algorithm in simulations using 25400 random cases and compare it against three state-of-the-art baselines using 2100 random cases. Our algorithm achieves up to 1.68× success rate using as low as 0.28× time in single-robot, and up to 2.15× success rate using as low as 0.36× time in multirobot cases compared to the best baseline. We implement our planner on real quadrotors to show its real-world applicability.
PaperID: 312,
Authors: Dandan Zhang, Wen Fan, Jialin Lin, Haoran Li, Qingzheng Cong, Weiru Liu, Nathan F. Lepora, Shan Luo
Affiliations: Imperial-X Initiative, and Department of Bioengineering, Imperial College London, London, U.K.; School of Engineering Mathematics and Technology, Bristol Robotics Laboratory, University of Bristol, Bristol, U.K.; Department of Engineering, King's College London, London, U.K.
Abstract: In this article, we present the design and benchmark of an innovative sensor, ViTacTip, which fulfills the demand for advanced multimodal sensing in a compact design. A notable feature of ViTacTip is its transparent skin, which incorporates a “see-through-skin” mechanism. This mechanism aims at capturing detailed object features upon contact, significantly improving both vision-based and proximity perception capabilities. In parallel, the biomimetic tips embedded in the sensor's skin are designed to amplify contact details, thus substantially augmenting tactile and derived force perception abilities. To demonstrate the multimodal capabilities of ViTacTip, we developed a multitask learning model that enables simultaneous recognition of hardness, material, and textures. To assess the functionality and validate the versatility of ViTacTip, we conducted extensive benchmarking experiments, including object recognition, contact point detection, pose regression, and grating identification. To facilitate seamless switching between various sensing modalities, we employed a generative adversarial network (GAN)-based approach. This method enhances the applicability of the ViTacTip sensor across diverse environments by enabling cross-modality interpretation.
PaperID: 313,
Authors: Dante Archangeli, Brendon Ortolano, Rosemarie C. Murray, Lukas Gabert, Tommaso Lenzi
Affiliations: Department of Mechanical Engineering and the Robotics Center, University of Utah, Salt Lake City, UT, USA
Abstract: Wearable robots and powered exoskeletons may improve ambulation for millions of individuals with poor mobility. Powered exoskeletons primarily assist in the sagittal plane to improve walking efficiency and speed. However, individuals with poor mobility often have limited mediolateral balance, which requires torque generation in the frontal plane. Existing hip exoskeletons that assist in both the sagittal and frontal planes are too heavy and bulky for use in the real world. Here we present the kinematic model, mechatronic design, and benchtop and human testing of a powered hip exoskeleton with a unique parallel kinematic actuator. The exoskeleton is lightweight (5.3 kg), has a slim profile, and can generate 30 N·m and 20 N·m of torque during gait in the sagittal and frontal planes. The exoskeleton torque density is 5.7 N·m/kg—53% higher than previously possible with series kinematic design. Testing with five healthy subjects indicate that frontal plane torques applied during stance or swing can alter step width, while sagittal plane torque can assist with hip flexion and extension. A device with these characteristics may improve both gait economy and balance in the real world.
PaperID: 314,
Authors: Tanguy Navez, Etienne Ménager, Paul Chaillou, Olivier Goury, Alexandre Kruszewski, Christian Duriez
Affiliations: CNRS, Centrale Lille, UMR CRIStAL, Univ. Lille, Inria, France; Inria, Département d'informatique de l'ENS, Ecole normale supérieur, CNRS, PSL Research University, Paris, France
Abstract: The finite element method (FEM) is a powerful modeling tool for predicting soft robots' behavior, but its computation time can limit practical applications. In this article, a learning-based approach based on condensation of the FEM model is detailed. The proposed method handles several kinds of actuators and contacts with the environment. We demonstrate that this compact model can be learned as a unified model across several designs and remains very efficient in terms of modeling since we can deduce the direct and inverse kinematics of the robot. Building upon the intuition introduced in (Ménager et al., 2023), the learned model is presented as a general framework for modeling, controlling, and designing soft manipulators. First, the method's adaptability and versatility are illustrated through optimization-based control problems involving positioning and manipulation tasks with mechanical contact-based coupling. Second, the low-memory consumption and the high prediction speed of the learned condensed model are leveraged for real-time embedding control without relying on costly online FEM simulation. Finally, the ability of the learned condensed FEM model to capture soft robot design variations and its differentiability are leveraged in calibration and design optimization applications.
PaperID: 315,
Authors: Hirokazu Ishida, Naoki Hiraoka, Kei Okada, Masayuki Inaba
Affiliations: University of Tokyo, Tokyo, Japan
Abstract: Library-based methods are known to be very effective for fast motion planning by adapting an experience retrieved from a precomputed library. This article presents CoverLib, a principled approach for constructing and utilizing such a library. CoverLib iteratively adds an experience-classifier-pair to the library, where each classifier corresponds to an adaptable region of the experience within the problem space. This iterative process is an active procedure, as it selects the next experience based on its ability to effectively cover the uncovered region. During the query phase, these classifiers are utilized to select an experience that is expected to be adaptable for a given problem. Experimental results demonstrate that CoverLib effectively mitigates the tradeoff between plannability and speed observed in global (e.g., sampling-based) and local (e.g., optimization-based) methods. As a result, it achieves both fast planning and high success rates over the problem domain. Moreover, due to its adaptation-algorithm-agnostic nature, CoverLib seamlessly integrates with various adaptation methods, including nonlinear programming-based and sampling-based algorithms.
PaperID: 316,
Authors: Yu Cao, Mengshi Zhang, Jian Huang, Samer Mohammed
Affiliations: Hubei Key Laboratory of Brain-inspired Intelligent Systems, School of Artificial Intelligence and Automation, Huazhong University of Science and Technology, Wuhan, China; Wuhan United Imaging Healthcare Surgical Technology Company Ltd., Wuhan, China; Laboratoire Images, Signaux et Systèmes Intelligents, University of Paris-Est Cretéil, Cretéil, France
Abstract: Active suspended backpacks represent a promising solution to mitigate the impact of inertial forces on individuals engaged in load carriage. However, identifying effective control objectives aimed at enhancing human carrying capacity remains a significant challenge. In this study, we introduce a novel approach by integrating a limb-like structure-type (LLS) bioinspired vibration isolator, modeled using Lagrangian mechanics, into an active load-transfer suspended backpack to primarily alleviate human shoulder pressure, thereby constructing a human–robot interaction control framework for the system. Drawing from a double-mass coupled oscillator model, this approach formulates a vertical dynamics model for the human-backpack system, systematically exploring the principles of both static load transfer and dynamic load reduction on the human shoulder. Subsequently, a series-elastic-actuator-based controller with prescribed performance is proposed to simultaneously achieve trajectory tracking and ensure load motion within the limited range. Theoretically, we validate the input–output stability of the LLS model and guarantee the ultimate uniform boundedness of the closed-loop system. Simulation and experimental trials conducted across different terrain scenarios validate the effectiveness of the proposed method, highlighting reductions of 18.68% in metabolic rate during level ground walking, 9.58% in a staircase scenario, and 12.35% in a complex terrain, involving uphill, downstairs, and flat ground walking.
PaperID: 317,
Authors: Federico Thomas, Jaume Franch
Affiliations: Institut de Robòtica i Informàtica Industrial (CSIC-UPC), ETSEIB, Barcelona, Spain
Abstract: The kinematics, dynamics, and control of a unicycle moving without slipping on a plane has been extensively studied in the literature of nonholonomic mechanical systems. However, since planar motion can be seen as a limiting case of the motion on a sphere, we focus our analysis on the more general spherical case. This article introduces a novel approach to path planning for a unicycle rolling on a sphere while satisfying the nonslipping constraint. Our method is based on a simple yet effective idea: first, we model the system as a linear time-varying dynamic system. Then, leveraging the fact that certain such systems can be integrated under specific algebraic conditions, we derive a closed-form expression for the control variables. This formulation includes three free parameters, which can be tuned to generate a path connecting any two configurations of the unicycle. Notably, our approach requires no prior knowledge of nonholonomic system analysis, making it accessible to a broader audience.
PaperID: 318,
Authors: Kaidi Wang, Ganghua Lai, Yushu Yu, Jianrui Du, Jiali Sun, Bin Xu, Antonio Franchi, Fuchun Sun
Affiliations: School of Mechatronical Engineering, Beijing Institute of Technology, Beijing, China; School of Mechanical Engineering, Beijing Institute of Technology, Beijing, China; Robotics and Mechatronics Lab, Faculty of Electrical Engineering, Mathematics and Computer Science, University of Twente, Enschede, The Netherlands; Department of Computer Science and Technology, Tsinghua University, Beijing, China
Abstract: Connecting multiple aerial vehicles to a rigid central platform through passive spherical joints holds the potential to construct a fully actuated aerial platform. The integration of multiple vehicles enhances efficiency in tasks like mapping and object reconnaissance. This article proposes a control and state estimation framework for the integrated aerial platform (IAP), enabling it to perform versatile tasks like object reconnaissance and physical interactive tasks with only onboard sensors. In the framework, the 6-D motion control serves as the low-level controller, while the high-level controller comprises a 6-D admittance filter and a perception-aware attitude correction module. The 6-D admittance filter, serving as the interaction controller, is adaptable for aerial interaction tasks. The perception-aware attitude correction algorithm is carefully designed by adopting a geometric model predictive controller (MPC). This algorithm, incorporating both offline and online calculations, proves to be well-suited for the intricate dynamics of an IAP. A 6-D direct wrench controller is also developed for the IAP. Notably, both the interaction controller and the direct wrench controller operate without reliance on force/torque sensors. Instead, a wrench observer algorithm is devised, considering external disturbances. In addition, based on the kinematics constraints of the multiple aerials in the platform, a fusion algorithm for multiple visual-inertial odometry and kinematics constraints is developed, providing more accurate localization. A prototype of the IAP is constructed, and its capabilities are demonstrated through experiments including perception-aware object reconnaissance, aerial mapping, aerial peg-in-hole task, and 6-D contact wrench generation. All experiments are conducted exclusively with onboard sensors. These tasks exemplify the merits of the proposed IAP and validate the effectiveness of the proposed control framework and fusion algorithm.
PaperID: 319,
Authors: Yingyu Wang, Liang Zhao, Shoudong Huang
Affiliations: Robotics Institute, University of Technology Sydney, Sydney, Australia
Abstract: Joint optimization of poses and features has been extensively studied and demonstrated to yield more accurate results in feature-based SLAM problems. However, research on jointly optimizing poses and non-feature-based maps remains limited. Occupancy maps are widely used non-feature-based environment representations because they effectively classify spaces into obstacles, free, and uknown regions, providing robots with spatial information for various tasks. In this article, we propose Occupancy-SLAM, a novel optimization-based SLAM method enabling the joint optimization of robot trajectory and the occupancy map through a parameterized map representation. The key novelty lies in optimizing both robot poses and occupancy values at different cell vertices simultaneously, a significant departure from existing methods, where the robot poses need to be optimized first before the map can be estimated. In our formulation, the state variables in optimization include both robot poses and occupancy values at cell vertices in the map. Moreover, a multi-resolution optimization framework utilizing occupancy maps with varying resolutions in different stages is introduced. A variation of GaussNewton method is proposed to solve the optimization problem at different stages. The proposed algorithm efficiently converges with initialization from odometry inputs. Furthermore, we propose an occupancy submap joining method within Occupancy-SLAM framework to handle large-scale problems effectively. Evaluations using simulations and practical 2D datasets demonstrate that the proposed approach can robustly obtain more accurate results than state-of-the-art techniques, with comparable computational time. Preliminary 3D results further confirm the potential of the proposed method in practical 3D applications, achieving more accurate results than existing methods.
PaperID: 320,
Authors: Qinbo Sun, Weimin Qi, Huihuan Qian
Affiliations: Shenzhen Institute of Artificial Intelligence and Robotics for Society, The Chinese University of Hong Kong, Shenzhen, China
Abstract: Sailboats are purely wind-driven, and thus, have great potential for long-term voyaging. For robotic sailboats, the constraints on the energy of the control boards, sensors, communication modules, and actuators are crucial to the sustainability of automation. Reducing the control frequency of actuators is crucial for energy conservation. This study proposes an energy-efficient long-short term (EeLsT) approach for sustainable sailing. In EeLsT, long-term and short-term observers are designed to adaptively take control decisions for time-varying environmental influences (e.g., waves and currents). Our approach can be generally applied as an energy management module in sailing robots. It explicitly leverages the sailing motion characteristics and the dynamic model of the robot considering marine disturbances. We have designed an experimental enhanced simulation platform to evaluate motion performance and energy consumption. Both baseline approach and the scheme incorporating EeLsT method (refered to as EeLsT approach in the subsequent sections) have been conducted. In simulation, the EeLsT approach saves 31.8% energy. In the real marine environment, experiments are conducted with OceanVoy, a catamaran sailing robot. The results show that 27.4% of the energy is saved during stable sailing. In long-term sailing, compared to the standby mode when the motors are not working, the average power of the full automation mode has increased by no more than 1 W, i.e., 4% relatively.
PaperID: 321,
Authors: Kanwal Naveed, Wajahat Hussain, Irfan Hussain, Donghwan Lee, Muhammad Latif Anjum
Affiliations: Robotics and Machine Intelligence Lab, School of Electrical Engineering and Computer Science, National University of Sciences and Technology, Islamabad, Pakistan; Khalifa University Center for Autonomous Robotic Systems, Khalifa University, Abu Dhabi, UAE; Reinforcement Learning Research Lab, School of Electrical Engineering, Korea Advanced Institute of Science and Technology, Daejeon, South Korea
Abstract: Large-scale evaluation of state-of-the-art visual simultaneous localization and mapping (SLAM) has shown that its tracking performance degrades considerably if the camera view is not adjusted to avoid the low-texture areas. Deep reinforcement learning (RL)-based approaches have been proposed to improve the robustness of visual tracking in such unsupervised settings. Our extensive analysis reveals the fundamental limitations of RL-based active view planning, especially in transition scenarios (entering/exiting the room, texture-less walls, and lobbies). In challenging transition scenarios, the agent generally remains unable to cross the transition during training, limiting its ability to learn the maneuver. We propose human-supervised RL training (imitation learning) and achieve significantly improved performance after ~50 h of supervised training. To reduce longer human supervision requirements, we also explore fine-tuning our network with an online learning policy. Here, we use limited human-supervised training (~20 h), and fine-tune the network with unsupervised training (~45 h), obtaining encouraging results. We also release our multimodel, human supervised training dataset. The dataset contains challenging and diverse transition scenarios and can aid the development of imitation learning policies for consistent visual tracking. We also release our implementation.
PaperID: 322,
Authors: Guangming Cui, Haozhi Huang, Xianrui Zhang, Yueyue Liu, Qigao Fan, Yining Xu, Ang Liu, Baijin Mao, Tian Qiu, Juntian Qu
Affiliations: Shenzhen International Graduate School, Tsinghua University, Shenzhen, China; School of Internet of Things Engineering, Jiangnan University, Wuxi, China; Division of Smart Technologies for Tumor Therapy, German Cancer Research Center (DKFZ), Dresden, Germany
Abstract: Programmable manipulation of fluid-based soft robots has recently attracted considerable attention. Achieving parallel control of large-scale ferrofluid droplet robots (FDRs) is still one of the major challenges that remain unsolved. In this article, we develop a distributed magnetic field control platform to generate a series of localized magnetic fields that enable the simultaneous control of many FDRs, allowing teams of FDRs to collaborate in parallel for multifunctional manipulation tasks. Based on the mathematical model using the finite element method, we first evaluate the distribution properties of the local magnetic fields as well as the gradients generated by individual electromagnets. Meanwhile, the locomotion and deformation behavior of the FDR is also characterized to verify the actuation performance of the developed system. Subsequently, a vision-based closed-loop feedback control strategy is then presented, which aims to achieve path tracking of multiple robot formations. Thermal analysis shows that the system’s low output power enables reliable and sustained long-term operation. Finally, the developed system is tested through extensive physical experiments with different numbers of FDRs. The results demonstrate the potential of the designed setup in manipulating dozens of FDRs for digital display, message encoding, and microfluidic logistics. To the best of authors’ knowledge, this is the first attempt that allows independent control of such scale droplet robots (up to 72) for cooperative applications.
PaperID: 323,
Authors: Jonathan B. Michaux, Patrick D. Holmes, Bohao Zhang, Che Chen, Baiyue Wang, Shrey Sahgal, Tiancheng Zhang, Sidhartha Dey, Shreyas Kousik, Ram Vasudevan
Affiliations: Robotics Institute, University of Michigan, Ann Arbor, MI, USA; Agility Robotics, Albany, OR, USA; Mechanical Engineering, Georgia Institute of Technology, Atlanta, GA, USA
Abstract: Ensuring safe, real-time motion planning in arbitrary environments requires a robotic manipulator to avoid collisions, obey joint limits, and account for uncertainties in the mass and inertia of objects and the robot itself. This article proposes autonomous robust manipulation via optimization with uncertainty-aware reachability (ARMOUR), a provably-safe, receding-horizon trajectory planner and tracking controller framework for robotic manipulators to address these challenges. ARMOUR first constructs a robust controller that tracks desired trajectories with bounded error despite uncertain dynamics. ARMOUR then uses a novel recursive Newton–Euler method to compute all inputs required to track any trajectory within a continuum of desired trajectories. Finally, ARMOUR overapproximates the swept volume of the manipulator; this enables one to formulate an optimization problem that can be solved in real time to synthesize provably-safe motions. This article compares ARMOUR to state of the art methods on a set of challenging manipulation examples in simulation and demonstrates its ability to ensure safety on real hardware in the presence of model uncertainty without sacrificing performance.
PaperID: 324,
Authors: Hao Cheng, Feitian Zhang
Affiliations: Robotics and Control Laboratory, School of Advanced Manufacturing and Robotics, and the State Key Laboratory of Turbulence and Complex Systems, Peking University, Beijing, China
Abstract: Robotic blimps, as lighter-than-air aerial platforms, offer extended operational duration and enhanced safety in human–robot interactions due to their buoyant lift. However, achieving robust flight performance under environmental airflow disturbances remains a critical challenge, thereby limiting their broader deployment. Inspired by avian flight mechanics, particularly the ability of birds to perch and stabilize in turbulent wind conditions, this article introduces RGBlimp-Q—a robotic gliding blimp equipped with a bird-inspired continuum arm featuring a novel moving mass actuation mechanism. This continuum arm enables flexible attitude regulation through internal mass redistribution, significantly enhancing the system’s resilience to external disturbances. In addition, it facilitates aerial manipulation by employing end-effector claws that interact with the environment in a manner analogous to avian perching behavior. This article presents the design, modeling, and prototyping of RGBlimp-Q, supported by comprehensive experimental evaluation and comparative analysis. To the best of the authors’ knowledge, this represents the first interdisciplinary integration of continuum mechanisms into a lighter-than-air robotic platform, where the continuum arm simultaneously functions as both an actuation and manipulation module. This design establishes a novel paradigm for robotic blimps, expanding their applicability to complex and dynamic environments.
PaperID: 325,
Authors: Fengkang Ying, Hanwen Zhang, Haozhe Wang, Huishi Huang, Marcelo H. Ang
Affiliations: Integrative Sciences and Engineering Programme, NUS Graduate School, National University of Singapore, Singapore; Department of Mechanical Engineering, College of Design and Engineering, National University of Singapore, Singapore
Abstract: Collision-free motion planning for redundant robot manipulators in complex environments is yet to be explored. Although recent advancements at the intersection of deep reinforcement learning (DRL) and robotics have highlighted its potential to handle versatile robotic tasks, current DRL-based collision-free motion planners for manipulators are highly costly, hindering their deployment and application. This is due to an overreliance on the minimum distance between the manipulator and obstacles, inadequate exploration and decision making by DRL, and inefficient data acquisition and utilization. In this article, we propose URPlanner, a universal paradigm for collision-free robotic motion planning based on DRL. URPlanner offers several advantages over existing approaches: it is platform agnostic, cost-effective in both training and deployment, and applicable to arbitrary manipulators without solving inverse kinematics. To achieve this, we first develop a parameterized task space and a universal obstacle avoidance reward that is independent of minimum distance. Second, we introduce an augmented policy exploration and evaluation algorithm that can be applied to various DRL algorithms to enhance their performance. Third, we propose an expert data diffusion strategy for efficient policy learning, which can produce a large-scale trajectory dataset from only a few expert demonstrations. Finally, the superiority of the proposed methods is comprehensively verified through experiments.
PaperID: 326,
Authors: Shuyu Wang, Dongling Liu, Changzeng Fu, Xiaoming Yuan, Peng Shan, Victor C. M. Leung
Affiliations: College of Information Science and Engineering, Northeastern University, Shenyang, China; Computer Science and Software Engineering, Shenzhen University, Shenzhen, China
Abstract: Predicting the causal flow by fusing multimodal perception is fundamental for constructing the bodily awareness of soft robots. However, forming such a predictive model while fusing the multimodal sensory data of soft robots remains challenging and less explored. In this study, we leverage the free energy principle within a Bayesian probabilistic deep learning framework to merge visual, pressure, and flex sensing signals. Our proposed multimodal association mechanism enhances the fusion process, establishing a robust computational methodology. We train the model using a newly collected dataset that captures the grasping dynamics of a soft gripper equipped with multimodal perception capabilities. By incorporating the current state and image differences, the forward model can predict the soft gripper’s physical interaction and movement in the image flow, which amounts to imagining future motion events. Moreover, we showcase effective predictions across modalities as well as for grasping outcomes. Notably, our enhanced variational autoencoder approach can pave the way for unprecedented possibilities of bodily awareness in soft robotics.
PaperID: 327,
Authors: Zhaoran Yin, Chao Zhou, Xiaocun Liao, Xiaofei Wang, Zhuoliang Zhang, Long Cheng, Junfeng Fan, Jian Wang
Affiliations: Laboratory of Cognition and Decision Intelligence for Complex Systems, Institute of Automation, Chinese Academy of Sciences, Beijing, China; Department of Automation, Tsinghua University, Beijing, China; State Key Laboratory of Multimodal Artificial Intelligence Systems, Institute of Automation, Chinese Academy of Sciences, Beijing, China
Abstract: Recent advances in underwater robotics highlight the potential of fish-like robots for efficient propulsion. However, their motion performance still lags behind real fish due to an incomplete understanding of fluid dynamics. This article hypothesizes that chordwise-oriented vortices dominate the flow evolution of low-aspect-ratio flapping foils. Based on this, a quasi 3-D hydrodynamic model, the Sliding Strip Discrete Vortex Method (SSDVM), is developed. SSDVM tracks chordwise-oriented vortex evolution along the chord, enabling hydrodynamic force calculations for heave and pitch motions across various parameters. The model closely agrees with Computational Fluid Dynamics simulations while reducing computational cost. When integrated with a dynamic model, SSDVM accurately predicts robotic fish motion, with simulated speeds closely matching experimental results and achieving a Mean Absolute Percentage Error of 6.54% . SSDVM offers an analytical tool for robotic fish hydrodynamics, balancing accuracy and efficiency, with potential applications in optimizing bioinspired underwater propulsion.
PaperID: 328,
Authors: Jianshu Zhou, Wei Chen, Junda Huang, Boyuan Liang, Yunhui Liu, Masayoshi Tomizuka
Affiliations: Department of Mechanical Engineering, University of California, Berkeley, CA, USA; Department of Mechanical and Automation Engineering, The Chinese University of Hong Kong, Hong KongSAR, China
Abstract: Robotic systems operating in unstructured environments require the ability to switch between compliant and rigid states to perform diverse tasks, such as adaptive grasping, high-force manipulation, shape holding, and navigation in constrained spaces, among others. However, many existing variable stiffness solutions rely on complex actuation schemes, continuous input power, or monolithic designs, limiting their modularity and scalability. This article presents the programmable locking cell (PLC)—a modular, tendon-driven unit that achieves discrete stiffness modulation through mechanically interlocked joints actuated by cable tension. Each unit transitions between compliant and firm states via structural engagement, and the assembled system exhibits high stiffness variation—up to 950% per unit—without susceptibility to damage under high payload in the firm state. Multiple PLC units can be assembled into reconfigurable robotic structures with spatially programmable stiffness. We validate the design through two functional prototypes: first, a variable-stiffness gripper capable of adaptive grasping, firm holding, and in-hand manipulation, and second, a pipe-traversing robot composed of serial PLC units that achieves shape adaptability and stiffness control in confined environments. These results demonstrate the PLC as a scalable, structure-centric mechanism for programmable stiffness and motion, enabling robotic systems with reconfigurable morphology and task-adaptive interaction.
PaperID: 329,
Authors: Alessandro Mancinelli, Bart D. W. Remes, Guido C. H. E. de Croon, Ewoud J. J. Smeur
Affiliations: Faculty of Aerospace Engineering, Delft University Of Technology, Delft, Zuid Holland, The Netherlands
Abstract: Hybrid overactuated tilt rotor uncrewed aerial vehicles (TRUAVs) are a category of versatile UAVs known for their exceptional wind resistance capabilities. However, their extensive operational range, combined with thrust vectoring capabilities, presents complex control challenges due to nonaffine dynamics and the necessity to coordinate lift and thrust for controlling accelerations at varying airspeeds. Traditionally, these vehicles rely on switched logic controllers with two or more intermediate states to control transitions. In this study, we introduce an innovative, unified incremental nonlinear controller designed to seamlessly control an overactuated dual-axis tilting rotor quad-plane throughout its entire flight envelope. Our controller is based on an incremental nonlinear control allocation algorithm to simultaneously generate pitch and roll commands, along with physical actuator commands. The control allocation problem is solved using a sequential quadratic programming (SQP) iterative optimization algorithm making it well-suited for the nonlinear actuator effectiveness typical of thrust vectoring vehicles. The controller's design integrates desired roll and pitch angle inputs. These desired attitude angles are managed by the controller and then conveyed to the vehicle during slow airspeed phases, when the vehicle maintains its 6-degrees of freedom (6-DOF). As the airspeed increases, the controller seamlessly shifts its focus to generating attitude commands for lift production, consequently smoothly disregarding the desired roll and pitch angles. Furthermore, our controller integrates an angle of attack (AoA) protection logic to mitigate wing stalling risks during transitions. It also features a yaw rate reference model to enable coordinated turns and minimize side-slip. The effectiveness of our proposed control technique has been confirmed through comprehensive flight tests. These tests demonstrated the successful transition from hovering flight to forward flight, the attainment of vertical and lateral accelerations, and the ability to revert to hovering.
PaperID: 330,
Authors: Theodora Kastritsi, Theofanis Prapavesis Semetzidis, Zoe Doulgeri
Affiliations: Human-Robot Interfaces and Interaction Laboratory, Istituto Italiano di Tecnologia, Genoa, Italy; Department of Electrical and Computer Engineering, Aristotle University of Thessaloniki, Thessaloniki, Greece
Abstract: The primary issue in bilateral teleportation setups is the existence of communication delays, which can destabilize the system. We are addressing this challenge in the case of a bilateral leader–follower surgical setup, where the surgeon uses a haptic device as the leader robot to manipulate the surgical instrument held by a general-purpose manipulator, the follower robot. The follower robot is equipped with an elongated tool that through a small incision passes inside the patient's body, where sensitive structures may exist. These structures may include organs, arteries, or veins that require protection during surgery. To address this challenge, we propose a bilateral control framework that is proven to maintain passivity, ensure bounded tracking errors between the leader and follower robots, and impose remote center of motion and spatial constraints related with the sensitive structures, all in the presence of constant and variable communication delays. Experimental results in a virtual intraoperative environment, using a point cloud of a kidney and its surrounding vessels, demonstrate the effectiveness of our control scheme under various communication delay scenarios.
PaperID: 331,
Authors: Giulia Pagnanelli, Lucia Zinelli, Nathan F. Lepora, Manuel G. Catalano, Antonio Bicchi, Matteo Bianchi
Affiliations: Centro di Ricerca “Enrico Piaggio”, Universita‘ di Pisa, Pisa, Italy; Department of Engineering Mathematics, Bristol Robotics Laboratory, University of Bristol, Bristol, U.K.
Abstract: Endowing robots with advanced tactile abilities based on biomimicry involves designing human-like tactile sensors, computational models, and motor control policies to enhance contact information retrieval. Here, we consider compliance discrimination with a soft biomimetic tactile optical sensor (TacTip). In previous work, we proposed a vision-based approach derived from a computational model of human tactile perception to discriminate object compliance with the TacTip, based on contact area spread computation over the indenting force. In this work, we first increased the robustness of our vision-based method with a more precise estimation of the initial contact area condition, which enables correct compliance estimation also when the probing direction is other than normal to the specimen surface. Then, we integrated within our validated framework the mechanisms of internal muscular regulation (co-contraction) that humans adopt during object compliance probing, to maximize the information uptake. To this aim, we used human co-contraction patterns extracted during object softness probing to control a Variable Stiffness Actuator (that emulates the agonistic-antagonistic behavior of human muscles), which is used to actuate the indenter system endowed with the TacTip for object compliance exploration. We found that our model-based approach for compliance discrimination, fed with more precisely estimated initial conditions, significantly improves with the human-inspired impedance regulation, with respect to the usage of a rigid actuator.
PaperID: 332,
Authors: Ruiqian Wang, Chuang Zhang, Wenjun Tan, Yiwei Zhang, Lianchao Yang, Wenyuan Chen, Feifei Wang, Jiandong Tian, Lianqing Liu
Affiliations: State Key Laboratory of Robotics, Shenyang Institute of Automation, Chinese Academy of Sciences, Shenyang, China; Department of Electrical and Electronic Engineering, University of Hong Kong, Hong Kong
Abstract: Fish can adaptively adjust their body kinematics and swimming modes by sensing to realize optimal propulsion. However, most soft robotic fish have an unchangeable swimming mode through simple structure design, making them difficult to adapt to dynamic and complex fluid environments. Here, inspired by the multiple muscle synergy and lateral line sensing function of fish, we developed a soft robotic fish with multiple actuating units and embedded sensing elements. By collaboratively controlling the amplitude and phase of excitation from the multiple flexible actuating units, the soft robotic fish can successfully realize various swimming modes very similar to those of natural fish. Additionally, the embedded flexible sensing elements enable the robotic fish to sense the swimming state and the surrounding fluid environment in real time. The multiple actuation and embedded sensing allow the soft robotic fish to adaptively switch to an optimal swimming mode in a certain fluid environment. The multimode swimming and perception capabilities proposed in this work not only make soft robotic fish more intelligent and adaptable to complex fluid environments, but also contribute to the future implementation of autonomous control capabilities for robotic fish.
PaperID: 333,
Authors: Junling Fu, Giorgia Maimone, Elisa Iovene, Jianzhuang Zhao, Alberto Redaelli, Giancarlo Ferrigno, Elena De Momi
Affiliations: Department of Electronics, Information, and Bioengineering, Politecnico di Milano, Milan, Italy
Abstract: This work presents a compliant and passive shared control framework for teleoperated robot-assisted tasks. Inspired by the human operator's capability of continuously regulating the arm impedance to perform contact-rich tasks, a novel control schema, exploiting the variable impedance control framework for force tracking is proposed. Moreover, bilateral teleoperation and shared control strategies are implemented to alleviate the human operator's workload. Furthermore, a global energy tank-based approach is integrated to enforce the system's passivity. The proposed framework is first evaluated to assess the force-tracking capability when the robot autonomously performs contact-rich tasks, e.g., in an ultrasound scanning scenario. Then, a validation experiment is conducted utilizing the proposed shared control framework. Finally, the system's usability is investigated with 12 users. The experiment results in system assessment revealed a maximum median error of 0.25 N across all the force-tracking experiment setups, i.e., constant and time-varying ones. Then, the validation experiment demonstrated significant improvements regarding the force tracking tasks compared to conventional control methods, and the system passivity was preserved during the task execution. Finally, the usability experiment shows that the human operator workload is significantly reduced by 54.6 % compared to the other two control modalities. The proposed framework holds significant potential for the execution of remote robot-assisted medical procedures, such as palpation and ultrasound scanning, particularly in addressing deformation challenges while ensuring safety, compliance, and system passivity.
PaperID: 334,
Authors: Charles L. Clark, Biyun Xie
Affiliations: Electrical and Computer Engineering Department, University of Kentucky, Lexington, KY, USA
Abstract: The focus of this research is to develop a learning-based method that computes self-motion manifolds (SMMs) efficiently and accurately to enable real-time global fault-tolerant motion planning. The proposed method first develops a learnable, closed-form representation of SMMs based on Fourier series. A cellular automaton is then applied to cluster workspace locations having the same number of SMMs and group SMMs with similar shape by homotopy classes, such that the SMMs of each homotopy class can be accurately learned by a neural network. To approximate the SMMs of an arbitrary workspace location, a neural network is first trained to predict the set of homotopy classes belonging to this workspace location. For each set of homotopy classes, another neural network is trained to approximate the Fourier series coefficients of the SMMs, and the joint configurations along the SMMs can be retrieved using the inverse Fourier transform. The proposed method is validated on planar 3R positioning, spatial 4R positioning, and spatial 7R positioning and orienting robots, using 10 000 randomly sampled workspace locations each. The results show that the proposed method can approximate SMMs with high accuracy and is much faster than the traditionally used nullspace projection method, a sampling-based method, and a grid-based method. The performance of the proposed method in real-time fault-tolerant motion planning applications is also demonstrated using the simulation of the spatial 7R robot and physical experiments on a planar 3R robot. Due to the computational efficiency of the proposed method, both robots are able to quickly plan trajectories which maximize the likelihood of task completion after the failure of one arbitrary joint.
PaperID: 335,
Authors: Yusheng Wang, Yonghoon Ji, Hiroshi Tsuchiya, Jun Ota, Hajime Asama, Atsushi Yamashita
Affiliations: Research into Artifacts, Center for Engineering, The University of Tokyo, Bunkyo, Japan; Graduate School of Advanced Science and Technology, Japan Advanced Institute of Science and Technology, Ishikawa, Japan; Research Institute, Wakachiku Construction Company Ltd., Chiba, Japan; Tokyo College, The University of Tokyo, Bunkyo, Japan; Department of Human and Engineered Environmental Studies, Graduate School of Frontier Sciences, The University of Tokyo, Bunkyo, Japan
Abstract: We present a novel acoustic camera simulator that generates realistic sonar images by incorporating recursive ray tracing and sonar artifact modeling and provides various ground truth labels, enabling benchmarking and learning purposes. The 2-D forward-looking sonar, also known as the acoustic camera, produces high-quality 2-D images. Conducting real-world underwater experiments is challenging, making realistic sonar image simulation a necessary alternative. However, existing simulators often lack sufficient realism or are limited to specific scenes and phenomena. As a result, training on simulations and testing on real sonar images (i.e., sim-to-real) remain open problems for deep learning-based applications. Our work introduces a novel sonar simulator with a customized rendering engine. We use recursive ray tracing to model multipath reflections in arbitrary scenes and propose physics-based shading for intensity computation. We propose a resampling method for antialiasing and model significant artifacts, such as rolling shutter distortions and crosstalk noise. The simulator provides various ground truths for benchmarking and deep learning applications. We tested several tasks by training on synthetic images and demonstrated that the models also work on real images. We developed a Blender add-on for an enhanced user interface and will make the simulator open-source to advance future research.
PaperID: 336,
Authors: Fei Suo, Xiaolong Hui, Peixin Hua, Xuejian Bai, Jin Ma, Min Tan, Yu Wang
Affiliations: School of Artificial Intelligence, University of Chinese Academy of Sciences, Beijing, China; State Key Laboratory of Multimodal Artificial Intelligence Systems, Institute of Automation, Chinese Academy of Sciences, Beijing, China; School of Electrical Engineering, Liaoning University of Technology, Jinzhou, China
Abstract: The complex underwater environment presents numerous challenges for the design of soft grippers, which often suffer from limited load capacity, poor stability, low portability, and imprecise control. This article proposes a novel rigid-soft hybrid gripper specifically designed for underwater use. The gripper's finger is constructed from silicone, reinforced with a multilink rigid exoskeleton on the outside, and actuated by tendons. This design provides three key advantages: compliance (capable of handling fragile objects such as a piece of tofu), heavy lifting (demonstrated by lifting an 80-kg barbell with three fingers), and precise, stable operation (the hybrid gripper maintains its shape despite water flow disturbances). In addition, the gripper is compact and lightweight, with the driving system powered by just four 23-g servo motors, making it easy to mount on various underwater robots. To enable precise control, both specialized kinematic and mechanics models were developed, allowing accurate predictions of the relationships among tendon displacement, exoskeleton deformation, soft material deformation, and tendon tension. This study thoroughly considers the challenges of underwater environments, offering new insights for advancing the field of underwater soft grasping.
PaperID: 337,
Authors: Xuguang Dong, Yixin Wang, Jingyi Zhou, Xin An, Yinglei Zhu, Fugui Xie, Xin-Jun Liu, Huichan Zhao
Affiliations: Department of Mechanical Engineering, State Key Laboratory of Tribology in Advanced Equipment, Beijing Key Laboratory of Transformative High-end Manufacturing Equipment and Technology, Tsinghua University, Beijing, China
Abstract: The development of high-performance bionic legged robots can benefit from the continued advancements in various actuation methods, such as artificial muscles. This work presents a musculoskeletal bionic leg driven by fluidic elastomer actuators (FEAs), showcasing their potential as artificial muscles for legged robots. Our approach integrates three key innovations: First, we established a mechanics model using thin plate theory to optimize the bellows shell structure of the FEAs, achieving high force output while maintaining inherent compliance. Second, we developed a lightweight embedded optoelectronic sensing system that enables closed-loop control without significantly increasing mass. Third, we designed a two-joint leg in the sagittal plane that utilizes a bionic configuration incorporating both monoarticular and biarticular FEAs. The leg demonstrated robust performance across various tasks including extreme positional movements, load-bearing squats supporting up to 2.45 times its body weight, vertical jumping with 147 mm ground clearance, and stable walking. Notably, our embedded sensing system successfully detected ground contact states without additional foot sensors, enabling reliable gait control while minimizing complexity and weight. The experimental results validate both the mechanical capabilities of the optimized FEAs and their controllability through embedded sensing, laying a foundation for developing full legged robots with muscle-like actuation.
PaperID: 338,
Authors: Zixi Chen, Qinghua Guan, Josie Hughes, Arianna Menciassi, Cesare Stefanini
Affiliations: Department of Excellence in Robotics and AI, Biorobotics Institute, Scuola Superiore Sant’Anna, Pisa, Italy; CREATE Lab, EPFL, Lausanne, Switzerland
Abstract: Modular soft robot arms (MSRAs) are composed of multiple modules connected in a sequence, and they can bend at different angles in various directions. This capability allows MSRAs to perform more intricate tasks than single-module robots. However, the modular structure also induces challenges in accurate planning and control. Nonlinearity and hysteresis complicate the physical model, while the modular structure and increased degrees of freedom further lead to cumulative errors along the sequence. To address these challenges, we propose a versatile configuration space planning and control strategy for MSRAs, named state to configuration to action. Our approach formulates an optimization problem, state to configuration planning, which integrates various loss functions and a forward model based on biLSTM to generate configuration trajectories based on target states. A configuration controller configuration to action control based on biLSTM is implemented to follow the planned configuration trajectories, leveraging only inaccurate internal sensing feedback. We validate our strategy using a cable-driven MSRA, demonstrating its ability to perform diverse offline tasks such as position and orientation control and obstacle avoidance. Furthermore, our strategy endows MSRA with online interaction capability with targets and obstacles. Future work focuses on addressing MSRA challenges, such as more accurate physical models.
PaperID: 339,
Authors: Inseung Kang, Dean D. Molinaro, Dongho Park, Dawit Lee, Pratik Kunapuli, Kinsey R. Herrin, Aaron J. Young
Affiliations: Department of Mechanical Engineering, Carnegie Mellon University, Pittsburgh, PA, USA; Robotics and AI (RAI) Institute, Cambridge, MA, USA; School of Mechanical Engineering, Georgia Institute of Technology, Atlanta, GA, USA; Department of Bioengineering, Stanford University, Stanford, CA, USA; Department of Computer and Information Science, University of Pennsylvania, Philadelphia, PA, USA
Abstract: Robotic exoskeletons can transform mobility for individuals with lower limb disabilities. However, their widespread adoption is limited by controller degradation caused by varying gait dynamics across different users and environments. Here, we propose an online adaptation framework that leverages real-time data streams to continuously update the user state estimator model. This approach allows the exoskeleton to learn the user-specific gait patterns, effectively customizing the model for each new user. In addition, we demonstrate a sensor signal transformation technique that enables model transfer across different exoskeleton hardware (from a research-grade exoskeleton to a commercial device). With less than one minute of adaptation, our framework improved gait phase estimation, which directly affects assistance timing, by 40.9% for able-bodied subjects and 65.9% for stroke survivors (p < 0.05), and reduced torque profile error by 32.7% compared to the baseline model (p < 0.05). Furthermore, in a pilot test, we applied our adaptation framework with human-in-the-loop optimization for control tuning. In a single stroke survivor, this approach led to a 21.8% increase in walking speed and a 6.5% reduction in metabolic cost compared to walking without exoskeleton. While preliminary, these results suggest the potential for personalized exoskeleton assistance in clinical populations.
PaperID: 340,
Authors: Nathan Goulet, Beshah Ayalew
Affiliations: Buildings and Transportation Science Division, Oak Ridge National Laboratory, Oak Ridge, TN, USA; Applied Dynamics and Control Group, Clemson University—International Center for Automotive Research, Greenville, SC, USA
Abstract: Teams of automated battery-powered electric vehicles have the potential to execute complex mission tasks in off-road environments for agriculture, military, and other applications. Limited onboard energy reserves hinder their adoption in large-scale resource-constrained environments, where recharging is a necessity. It may be infeasible to install a network of static charging stations in off-road environments. For this reason, dedicated mobile host vehicles with charging capabilities are proposed as a means to increase range and capabilities of the multivehicle team. Here, we consider an ad hoc planning framework, where results from a high-confidence trajectory planner are leveraged to plan charging rendezvous between a host and other worker vehicles in a receding horizon fashion to provide high confidence that energy reserves will not be prematurely exhausted. The core problem is posed so as to minimize the impact of recharging on the mission in terms of task delays, overall energy utilization, and costs of fast charging. Through extensive Monte Carlo simulations of an off-road mission, we show a decrease in task delays without substantial increases in energy needs by updating the charging rendezvous plan during the mission. However, if updates are made too often, model mismatch may cause unnecessary cycling and mission failure.
PaperID: 341,
Authors: Mingjie Dong, Hanwei Ruan, Zeyu Wang, Chenyang Sun, Shiping Zuo, Yi-Feng Chen, Jianfeng Li, Mingming Zhang
Affiliations: Beijing Key Laboratory of Advanced Manufacturing Technology, College of Mechanical & Energy Engineering, Beijing University of Technology, Beijing, P.R. China; Department of Biomedical Engineering, Southern University of Science and Technology, Shenzhen, China
Abstract: Robot-assisted ankle rehabilitation training imitating physician’s professional techniques is highly important for promoting personalized training and improving clinical outcomes. In this work, we propose a two-level kernelized movement primitives (2-level-KMP) imitation learning algorithm under the kernelized movement primitives (KMP) framework, which reproduces physician’s experience and optimizes the imitation trajectory during rehabilitation, to realize physician-level performance in robot-assisted ankle rehabilitation training. First, a KMP process combined with a Bayesian optimizer is used to imitate the rehabilitation trajectory. Second, the other KMP process is used to smooth the imitation trajectory further. Then the two KMP processes combined with patient-in-the-loop optimization (PILO) realize temporal rehabilitation adaptation. Finally, the 2-level-KMP algorithm is reproduced on a parallel ankle rehabilitation robot (PARR), which enables the patient’s passive rehabilitation training to be empirical and adaptive. Ten ankle dysfunction patients were involved in clinical experiments, with the results showing that the proposed algorithm can accurately reproduce physician’s trajectories and modulate trajectories based on patient’s feedback. After ten rehabilitation exercises, the number of modulation points calculated from patient’s torque feedback decreases by 85.19% on average compared with the beginning stage. A comparison between the 2-level KMP algorithm and existing algorithms shows that the 2-level-KMP algorithm can better ensure smoothness and retain the shape of the trajectory during trajectory modulation, ensuring the safety of ankle rehabilitation and retaining the experience of the physician.
PaperID: 342,
Authors: Xianda Wu, Ming Xu, Zhihao Zhou, Wenjie Lou, Teng Zhang, Yalei Zhou, Jingeng Mai, Qining Wang
Affiliations: School of Advanced Manufacturing and Robotics, Peking University, Beijing, China; Institute for Artificial Intelligence, Peking University, Beijing, China
Abstract: Evolutionary pressures have pushed humans to become efficient walkers, but inefficient divers. People consume more energy to travel the same distance underwater than on land. In diverse overground locomotion, emerging exoskeletons have reduced the metabolic cost of humans. Can we also improve the energy economy in underwater locomotion via exoskeletons? Here, we propose an underwater exoskeleton to assist scuba diving using flutter kick, by applying assistive knee extension torque during the strike phase of the diving kick cycle. When divers wore the powered exoskeleton, the average net air cost across six experienced divers was reduced by 22.7 \pm 10.0%, and the peak quadriceps activation was decreased by 20.9 \pm 7.5%, compared with normal diving without the exoskeleton. The average gastrocnemius activation also decreased by 20.6 \pm 5.3%, suggesting that the divers sufficiently utilized the exoskeleton assistance. These results indicate that applying exoskeleton assistance is conducive to improving the endurance of human underwater diving and enhancing our ability to explore the underwater world. Our study extends the application boundary of wearable robots, and provides a reference for the design and assessment of future underwater assistive devices, with the potential to strengthen the connection between humans and the ocean.
PaperID: 343,
Authors: Jianshu Zhou, Junda Huang, Qi Dou, Pieter Abbeel, Yunhui Liu
Affiliations: Department of Mechanical and Automation Engineering, The Chinese University of Hong Kong, Hong Kong; Department of Computer Science and Engineering, The Chinese University of Hong Kong, Hong Kong; University of California at Berkeley, Berkeley, CA, USA
Abstract: Human beings possess a remarkable skill for fine in-hand manipulation, utilizing both intrafinger interactions (in-finger) and finger–environment interactions across a wide range of daily tasks. These tasks range from skilled activities like screwing light bulbs, picking and sorting pills, and in-hand rotation, to more complex tasks such as opening plastic bags, cluttered bin picking, and counting cards. Despite its prevalence in human activities, replicating these fine motor skills in robotics remains a substantial challenge. This study tackles the challenge of fine in-hand manipulation by introducing the dexterous and compliant (DexCo) hand system. The DexCo hand mimics human dexterity, replicating the intricate interaction between the thumb, index, and middle fingers, with a contractable palm. The key to maneuverable fine in-hand manipulation lies in its innovative soft hydraulic actuation, which strikes a balance between control complexity, dexterity, compliance, and motion accuracy within a compact structure, enhancing the overall performance of the system. The model of soft hydraulic actuation, based on hydrostatic force analysis, reveals the compliance of hand joints, which is also further extended to a dedicated robot operating system (ROS) package for DexCo hand simulation, considering both motion and stiffness aspects. Dedicated velocity and position teleoperation controllers are designed for implementing real physical manipulation tasks. The benchmark results show that the fingertip achieves a maximum repeatable finger strength of 34.4 N, a grasp cycle time of less than 2.04 s, and a maximum repeatability accuracy of 0.03 mm. Experimental results demonstrate the DexCo hand successfully performs complex fine in-hand manipulation tasks, providing a promising solution for advancing robotic manipulation capabilities toward the human level.
PaperID: 344,
Authors: Pau Vial, Joan Solà, Narcís Palomeras, Marc Carreras
Affiliations: Institut VICOROB, Universitat de Girona, Girona, Catalunya, Spain; Institut de Robòtica i Informàtica Industrial (IRII), CSIC-Universitat Politècnica de Catalunya, Barcelona, Catalunya, Spain
Abstract: Robot localization is a fundamental task in achieving true autonomy. Recently, many graph-based navigators have been proposed that combine an inertial measurement unit (IMU) with an exteroceptive sensor applying IMU preintegration to synchronize both sensors. IMUs are affected by biases that also have to be estimated. To increase the navigator robustness when faults appear on the perception system, IMU preintegration can be complemented with linear velocity measurements obtained from visual odometry, leg odometry, or a Doppler Velocity Log (DVL), depending on the robotic application. Moreover, higher grade IMUs are sensitive to the Earth rotation rate, which must be compensated in the preintegrated measurements. In this article, we propose a general purpose preintegration methodology formulated on a compact Lie group to set motion constraints on graph simultaneous localization and mapping problems considering the Earth rotation effect. We introduce the SE_N(3) group to jointly preintegrate IMU data and linear velocity measurements to preserve all the existing correlation within the preintegrated quantity. Field experiments using an autonomous underwater vehicle equipped with a DVL and a navigational grade IMU are provided and results are benchmarked against a commercial filter-based inertial navigation system to prove the effectiveness of our methodology.
PaperID: 345,
Authors: Shifeng Huang, Fan Li, Xing Zhou, Molong Duan
Affiliations: Department of Computer Science, University of Exeter, Exeter, U.K.; Foshan Institute of Intelligent Equipment Technology, Foshan, China; Hong Kong University of Science and Technology, Hong Kong
Abstract: Generating optimal excitation trajectories is crucial for ensuring that the observation matrix is well conditioned in robot dynamic identification. This task is a typical optimization problem involving explicit physical constraints defined by initial conditions (zero initial joint velocity and acceleration) and physical limits (joint position, velocity, and acceleration within specified bounds). Physical constraints complicate problem-solving, necessitating the use of heuristic or gradient-based iteration methods. Despite extensive study of this problem over many years, the success rate of finding feasible solutions that do not violate physical constraints within a limited number of iteration steps is lower than desired, and two major challenges remain: 1) a low success rate; and 2) high time consumption, which adversely affect practical applications. This article presents an analytical approach to address these physical constraints for excitation optimization. Feasible solutions are ensured through a deterministic calculation of the Fourier series-based parameterization rather than relying on iterative searches. Specifically, initial conditions are met by assigning offsets directly, while scaling and central-translation operations ensure adherence to physical limits. Our approach achieves a 100% success rate in generating physically executable excitation trajectories. Extensive experiments indicate that our approach has improved optimization efficiency by an order of magnitude compared to available methods, while delivering excellent excitation performance. For practitioners, our method renders excitation optimization a viable approach for time-critical payload identification tasks.
PaperID: 346,
Authors: Dejan Milojevic, Gioele Zardini, Miriam Elser, Andrea Censi, Emilio Frazzoli
Affiliations: Institute for Dynamic Systems and Control, ETH Zürich, Zürich, Switzerland; Laboratory for Information and Decision Systems, Massachusetts Institute of Technology, Cambridge, MA, USA; Chemical Energy Carriers and Vehicle Systems Laboratory, Empa—Swiss Federal Laboratories for Materials Science and Technology, Dübendorf, Switzerland
Abstract: This article discusses the integration challenges and strategies for designing mobile robots, by focusing on the task-driven, optimal selection of hardware and software to balance safety, efficiency, and minimal usage of resources such as costs, energy, computational requirements, and weight. We emphasize the interplay between perception and motion planning in decision-making by introducing the concept of occupancy queries to quantify the perception requirements for sampling-based motion planners. Sensor and algorithm performance are evaluated using false negative rate and false positive rate across various factors such as geometric relationships, object properties, sensor resolution, and environmental conditions. By integrating perception requirements with perception performance, an integer linear programming approach is proposed for efficient sensor and algorithm selection and placement. This forms the basis for a co-design optimization that includes the robot body, motion planner, perception pipeline, and computing unit. We refer to this framework for solving the co-design problem of mobile robots as CODEI, short for co-design of embodied intelligence. A case study on developing an autonomous vehicle for urban scenarios provides actionable information for designers, and shows that complex tasks escalate resource demands, with task performance affecting choices of the autonomy stack. The study demonstrates that resource prioritization influences sensor choice: cameras are preferred for cost-effective and lightweight designs, while lidar sensors are chosen for better energy and computational efficiency.
PaperID: 347,
Authors: Yu Zheng
Affiliations: Research Institute, UBTECH Robotics Inc., Shenzhen, China
Abstract: In this article, we present an efficient unified algorithm for the minimum Euclidean distance between two collections of compact convex sets, each of which can be a collection of convex primitives, such as ellipsoids, capsules, and cylinders, or a collection of triangles (i.e., triangle mesh) or a collection of points (i.e., point cloud) as special cases. The Euclidean distance between two compact convex sets is defined to be the smallest translation to bring them into intersection if they are separated or to separate them if they intersect, which can be computed by the well-known Gilbert–Johnson–Keerthi and expanding polytope algorithms, respectively. While existing algorithms are aimed at computing the minimum Euclidean distance for a specific type of collections, algorithms for mixed situations always remain vacant. We discover that the smallest translation direction between any two compact convex sets determines the planes to bound and separate some other sets in two collections and can help quickly identify sets that do not have the minimum distance. In this way, the minimum distance between two collections can be efficiently computed, hundreds to thousands of times faster than the brute-force search. The computational efficiency of the proposed algorithm is verified with a number of numerical experiments in various scenarios.
PaperID: 348,
Authors: Shaopeng Liu, Chao Huang, Hailong Huang
Affiliations: Department of Industrial and Systems Engineering, The Hong Kong Polytechnic University, Hong Kong, SAR, China; Department of Aeronautical and Aviation Engineering, The Hong Kong Polytechnic University, Hong Kong, SAR, China
Abstract: Given the limitations of current methods in terms of accuracy and efficiency for robot scene recognition (SR) in domestic environments, this article proposes an active scene recognition (ASR) approach that allows the robot to recognize scenes correctly using less images, even when the robot’s position and observation direction are uncertain. ASR includes a behavior cloning-based action classification model, which can adjust the robot view actively to capture beneficial images for SR. To address the lack of essential expert data for training the action model, we introduce an expert data generation method that avoids time-consuming and inefficient manual data collection. In addition, we present a multiview SR method to handle the multiple images resulting from view changes. This method includes an SR model that scores each image and a revision and prediction method to mitigate the compounding error introduced by behavior cloning as well as output the finial recognition result. We conducted numerous comparative experiments and an ablation study in various domestic environments using a publicly simulated platform to validate our ASR method. The experimental results demonstrate that our proposed approach outperforms state-of-the-art methods in terms of both accuracy and efficiency for SR. Furthermore, our method, trained in simulated environments, demonstrates excellent generalization capabilities, allowing it to be directly transferred to the real world without the need for fine-tuning. When deployed on a TurtleBot 4 robot, it achieves precise and efficient SR in diverse real-world environments.
PaperID: 349,
Authors: Lelai Zhou, Zhengmao Li, Yibin Li, Shaoping Bai
Affiliations: Center for Robotics, School of Control Science and Engineering, Shandong University, Jinan, China; Department of Materials and Production, Aalborg University, Aalborg, Denmark
Abstract: Real-time motion planning in dynamic environments presents a significant challenge for robotic manipulators. This article introduces an innovative parallel model predictive path integral (MPPI) algorithm enabling the robot to navigate swiftly and safely in such environments. Unlike the conventional MPPI methods that rely on a single sequence of Gaussian means for trajectory sampling, the proposed parallel MPPI (PMPPI) concurrently runs multiple planners with different strategies and adaptively integrates planned paths based on the current state, leveraging the advantages of different strategies and greatly improving the MPPI’s exploration capability. Moreover, a gradient-velocity modulated signed distance field (SDF) cost function that dynamically adjusts costs based on the robot’s velocity and the SDF gradient is defined, thereby promoting safer and purposeful motion planning. In the implementation, techniques like utilizing inverse kinematics solver for path guidance and sparse reward to expedite reaching time are integrated into the MPPI cost function design. Comparative evaluations against the traditional MPPI architecture and standard SDF cost designs demonstrate the superiority of the new method. Real-world experiments, including human–robot interaction, obstacle-crossing, and grasping tasks, validate the robustness and universality of our methodology, with average and maximum end effector speeds of 0.523 m/s and 1.225 m/s respectively.
PaperID: 350,
Authors: Mingyue Wang, Siyuan An, Zhenhuan Sun, Jiaqi Li, Yang Wang, Song Liu
Affiliations: School of Information Science and Technology, ShanghaiTech University, Shanghai, China
Abstract: The noncontact acoustic manipulation of particles, biosamples, droplets, and air bubbles has emerged as a promising technology in the fields of biology, chemistry, medicine, etc. The noncontact nature offers significant advantages in terms of biocompatibility, contamination free, and material versatility. However, current noncontact acoustic manipulation techniques still lack adequate selectivity, robustness, and precision controllability in complex environments. To this end, in this article, we propose an automated noncontact manipulation system that leverages a high-density ultrasonic phased transducer array in combination with a microscope to further optimize and enhance the controllability and flexibility of noncontact particle manipulation. This work presents several notable contributions. First, we successfully realized selective particle manipulation, allowing instantaneous interaction with users to perform user-designated and objective-oriented manipulation tasks. Second, we integrated a closed-loop control strategy into the system that effectively mitigates misalignment errors induced by the trapping stiffness heterogeneity of acoustic trap and enables automated precision position control of particles in complex environments (in 30-mm-wide workspace, positioning precision is 1/40 of the wavelength). Third, we proposed a reconfigurable acoustic trap design method, named pseudovortex trap, featuring real-time computing and trapping particles larger than the wavelength. The system setup, the calibration specifics, the acoustic trap design methodology, and the corresponding visual servo control scheme (in terms of selective trapping, precision positioning, and dynamic trajectory planning) are given in detail in the article. Meanwhile, the trapping stiffness and the manipulation stability are also analyzed in this work. Experimental results well demonstrated the effectiveness of the proposed system.
PaperID: 351,
Authors: Zeyu Wang, Wenchuan Jia, Yi Sun, Tianxu Bao, Zihan Ding, Qi Chen
Affiliations: Shanghai Key Laboratory of Intelligent Manufacturing and Robotics, School of Mechatronic Engineering and Automation, Shanghai University, Shanghai, China
Abstract: Kinematic performance of a quadruped robot is determined by the mechanical structure. This article presents a novel leg structure for legged robots that integrates a differential mechanism into the conventional design. This approach enables all actuators to be positioned within the robot's torso at fixed locations, significantly reducing the leg's inertia. Furthermore, the new structure introduces a parallel transmission system that balances motion and torque distribution among the joint actuators, effectively reducing torque peaks and enhancing the drive capability during dynamic motions. A family of configurations of differential leg structures is constructed, and their mapping to the classic serial leg structure is dissected in kinematic and mathematic. Simulations of various single-leg models are conducted to validate the performance of the new configuration under typical gait conditions. Subsequently, a leg prototype is designed, manufactured, and tested in experiments involving tasks, such as trajectory tracking, weighted squats, and squat jumps. The development of a prototype quadruped robot featuring this novel leg structure is also presented.
PaperID: 352,
Authors: Xiangyang Wang, Chunjie Chen, Jianquan Sun, Sida Du, Yue Ma, Xinyu Wu
Affiliations: Guangdong Provincial Key Lab of Robotics and Intelligent System, Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences, Shenzhen, China
Abstract: Assisting underwater movements improves divers' efficiency and reduces the risk of decompression sickness from physical activity. Although exoskeletons have been developed for numerous land-based scenarios, their application in underwater diving remains unexplored. This article proposes a soft underwater lower-limb exosuit designed to assist three aquatic movements: flutter kick, breaststroke kick, and underwater walk. We presented the mechanical design of the exosuit that is capable of assisting bidirectional leg movements in full kicking/gait cycle, while ensuring natural leg mobility without impeding normal leg function. A cascade force integral controller is also designed to resolve issues related to uncontrollable states and stiffness variations within the system. To verify the assistive performance of the system, experiments were conducted with nine participants to assess how the proposed exosuit aids in reducing metabolic cost across various motion patterns and frequencies. The findings indicate that the underwater exosuit effectively reduces the air consumption rate by 29.77\pm 7.68% during flutter kick, 25.70\pm 5.99% during breaststroke kick, and 18.35\pm 4.53% during underwater walk.
PaperID: 353,
Authors: Bike Zhu, Jun He, Zhicheng Yuan, Feng Gao
Affiliations: State Key Laboratory of Mechanical System and Vibration, School of Mechanical Engineering, Shanghai Jiao Tong University, Shanghai, China
Abstract: Wheel-legged planetary rovers possess superb locomotion capabilities. This article combines an offline predefined motion planning library with online path planning, integrating energy consumption and probabilistic aspects of the robotic system. The primary focus is on addressing the planning challenges in dense environments, where the distance between any adjacent obstacles is smaller than the width of the prototype. Therefore, it is necessary to consider the interaction between the prototype and the environment. First, the generalized function set theory and the configuration topology theory are utilized to mathematically describe the motions of multilimbed systems. Based on the representation, an offline planning library is established. Second, the Markov-decision-process-based path planning method is extended by incorporating the platform's geometry and locomotion capabilities. The concept of “limb-travel relevant nodes” is introduced. To address the numerous iteration problems, the informed value iteration algorithm is proposed. Third, a multilayered map is evaluated to further enhance computational efficiency. Finally, the proposed algorithm is implemented on the terrain adaptive wheel-legged rover. Experimental results demonstrate that the proposed algorithm is capable of finding the optimal path with high computational efficiency, and it exhibits excellent adaptability on nonuniform maps.
PaperID: 354,
Authors: Wenbin Zhu, Jing Yuan, Xuebo Zhang, Fei Chen
Affiliations: College of Artificial Intelligence, Nankai University, Tianjin, China
Abstract: Existing object-level simultaneous localization and mapping (SLAM) methods often overlook the correspondence between semantic information and geometric features, resulting in a significant gap between them within SLAM frameworks. To tackle this issue, this article proposes, a semantic-geometric tight-coupling monocular visual object SLAM system, (TiMoSLAM), which considers a rigorous correspondence between semantics and geometry across all steps of SLAM. Initially, a general semantic relation graph (SRG) is developed to consistently represent semantic information alongside geometric features. Detailed analyzes on complete constraints of the geometric feature combinations on estimation of 3-D cuboid model are performed. Subsequently, a compound hypothesis tree is proposed to incrementally construct the object-specific SRG and concurrently estimate the 3-D cuboid model of an object, ensuing semantic-geometric consistency in object representation and estimation. Special attention is given to the matching errors between geometric features and objects during the optimization of camera poses and object parameters. The effectiveness of this method is validated on various datasets, as well as in real-world environments.
PaperID: 355,
Authors: Enhao Zheng, Xiaodong Liu, Chenfeng Xu, Zhihao Zhou, Qining Wang
Affiliations: State Key Laboratory of Multimodal Artificial Intelligence Systems, Institute of Automation, Chinese Academy of Sciences, Beijing, China; Institute for Artificial Intelligence, Peking University, Beijing, China
Abstract: Representing human arm dynamic intent is essential for effective human–robot interaction. Accurately and robustly decoding these intentions through mathematical modeling of neuromuscular processes poses significant challenges. This study introduces an electrical impedance tomography (EIT)-driven musculoskeletal model, which integrates an EIT sensing system with methods for muscle identification, parameter estimation, and musculoskeletal system modeling. Unlike existing muscle-signal techniques, EIT captures muscle activities from the anatomical cross-sectional plane, providing both activation dynamics and morphological features. We validated our method through multiDoF wrist kinematics estimation under varying contraction intensities, arm endpoint stiffness estimation, and robotic variable admittance control. Our approach achieves accuracy comparable to state-of-the-art methods while requiring fewer training samples and a more compact sensing system. The model incorporates physiological constraints, minimizing decoding errors, and ensuring interaction safety. This method enables reliable intent decoding with practical training demands. Future work will enhance the EIT system for complex tasks.
PaperID: 356,
Authors: Myeong-Ju Kim, Daegyu Lim, Gyeongjae Park, Kwanwoo Lee, Jaeheung Park
Affiliations: Department of Intelligence and Information, Seoul National University, Seoul, South Korea
Abstract: The robust balancing capability of humanoids is essential for mobility in real environments. Many studies focus on implementing human-inspired ankle, hip, and stepping strategies to achieve human-level balance. In this article, a robust balance control framework for humanoids is proposed. First, a model predictive control (MPC) framework is proposed for capture point (CP) tracking control, enabling the integration of ankle, hip, and stepping strategies within a single framework. In addition, a variable weighting method is introduced that adjusts the weighting parameters of the centroidal angular momentum damping control. Second, a hierarchical structure of the MPC and a stepping controller was proposed, allowing for the step time optimization. The robust balancing performance of the proposed method is validated through simulations and real robot experiments. Furthermore, a superior balancing performance is demonstrated compared to a state-of-the-art quadratic programming-based CP controller that employs the ankle, hip, and stepping strategies.
PaperID: 357,
Authors: Zhenyuan Zhang, Yu Zhang, Darong Huang, Xin Fang, Mu Zhou, Ying Zhang
Affiliations: School of Traffic and Transportation, Chongqing Jiaotong University, Chongqing, China; School of Information Science and Engineering, Chongqing Jiaotong University, Chongqing, China; School of Artificial Intelligence, Anhui University, Hefei, China; School of Mechanical and Electrical Engineering, Southwest Petroleum University, Chengdu, China; School of Communications and Information Engineering, Chongqing University of Posts and Telecommunications, Chongqing, China; School of Electrical and Computer Engineering, Georgia Institute of Technology, Atlanta, GA, USA
Abstract: Extended object detection and tracking (EODT) is becoming a promising alternative for autonomous perception, which provides not only common motion states but also accurate spatial extent information, such as shape and size estimations. However, due to uncoordinated radar transmissions in zero-trust autonomous driving scenarios, radar-based EODT systems suffer from mutual radio frequency (RF) interference launched by attackers, leading to ghost targets and increased noise. On this account, a novel joint anti-interference detection and tracking system for weak extended targets is presented in this article. In contrast to pioneering works that treat object detection and tracking as two separate steps, the proposed method handles them jointly by integrating a continuous detection process into tracking, improving the detectability of weak targets. More specifically, to accommodate the time-varying number and extended size of radar reflections, an adaptive spatial distribution model representing the deformable extents is incorporated to capture the contour evolution over time. The key insight is that by accumulating the reflected power, all backscattered points are regarded as one entity to match the real target so that the intractable data association problem can be circumvented in the proposed method. Unlike the prominent random matrix model-based approaches that split motion and extent states into independent parts, this study explores the interdependencies between the states and updates them simultaneously. In addition, the proposed system has been deployed on a low-cost automotive radar platform. Experimental results confirm that the proposed approach can achieve accurate and resilient EODT against RF interference attacks, especially in occlusion, dynamic motion switching, and complex multiple extended target tracking scenarios.
PaperID: 358,
Authors: Dongjiao He, Haotian Li, Jie Yin
Affiliations: Department of Mechanical Engineering, University of Hong Kong, Hong Kong, SAR, China; Department of Electronic Engineering, Shanghai Jiao Tong University, Shanghai, China
Abstract: This article introduces a method for tightly fusing sensors with diverse characteristics to maximize their complementary properties, thereby surpassing the performance of individual components. Specifically, we propose a tightly coupled light detection and ranging (LiDAR)-inertial-global navigation satellite system (GNSS) odometry (LIGO) system, which synthesizes the advantages of LiDAR, inertial measurement unit (IMU), and GNSS. Integrating LiDAR with IMU demonstrates remarkable precision and robustness in high-dynamics and high-speed motions. However, LiDAR-Inertial systems encounter limitations in feature-scarce environments or during large-scale movements. GNSS integration overcomes these challenges by providing global and absolute measurements. LIGO employs an innovative hierarchical fusion approach with both front-end and back-end components to achieve synergistic performance. The front-end of LIGO utilizes a tightly coupled, extended Kalman filter (EKF)-based LiDAR-Inertial system for high-bandwidth localization and real-time mapping within a local-world frame. The back-end tightly integrates the filtered LiDAR-Inertial factors from the front-end with GNSS observations in an extensive factor graph, being more robust to outliers and noises in GNSS observations and producing optimized globally referenced state estimates. These optimized back-end results are then fed back to the front-end through the EKF to ensure a drift-free trajectory, particularly in degenerate and large-scale scenarios. Real-world experiments validate the effectiveness of LIGO, especially when applied to aerial vehicles with outlier-prone GNSS data, demonstrating its resilience to signal losses and data quality fluctuations. LIGO outperforms comparable systems, offering enhanced accuracy and reliability across varying conditions.
PaperID: 359,
Authors: Zhongjin Ju, Ke Wei, Yundou Xu
Affiliations: Parallel Robot and Mechatronic System Laboratory of Hebei Province, Yanshan University, Qinhuangdao, China
Abstract: This study introduces a novel quadruped robot, the TerraAdapt, furnished with an innovative deformable wheel–foot integrated structure. This unique design grants the robot the flexibility to alternate between wheeled and footed modes of locomotion, making it efficient in traversing diverse terrains, from smooth indoor floors to challenging outdoor landscapes laden with obstacles. The study delineates an in-depth design and analysis of the deformable wheel and its integrated wheel–foot structure using screw theory. We engineer a 2 R: Rotational, P: Prismatic (RRR-RP) wheel–foot mode-switching mechanism by modifying a 2RRR spatial six-bar mechanism with an additional RP branch. This mechanism aids in seamless transitioning between different movement modes. Moreover, a 2RRR parallel structure is employed to construct the footed mode structure.To substantiate the viability and efficacy of the proposed design, we carry out extensive motion simulations and construct an experimental prototype for field testing. The field trials reveal the robot's adeptness in adapting to varied terrains, highlighting the possible advantages of incorporating the proposed deformable wheel into micro mobile robot designs.
PaperID: 360,
Authors: Jun Ueda, Hyukbin Kwon
Affiliations: George W. Woodruff School of Mechanical Engineering, Georgia Institute of Technology, Atlanta, GA, USA
Abstract: With the increasing integration of cyber-physical systems (CPS) into critical applications, ensuring their resilience against cyberattacks is paramount. A particularly concerning threat is the vulnerability of CPS to deceptive attacks that degrade system performance while remaining undetected. This article investigates perfectly undetectable false data injection attacks (FDIAs) targeting the trajectory tracking control of a nonholonomic mobile robot. The proposed attack method utilizes affine transformations of intercepted signals, exploiting weaknesses inherent in the partially linear dynamic properties and symmetry of the nonlinear plant. The feasibility and potential impact of these attacks are validated through experiments using a Turtlebot 3 platform, highlighting the urgent need for sophisticated detection mechanisms and resilient control strategies to safeguard CPS against such threats. Furthermore, a novel approach for detection of these attacks called the state monitoring signature function (SMSF) is introduced. An example SMSF, a carefully designed function resilient to FDIA, is shown to be able to detect the presence of an FDIA through signatures based on system states.
PaperID: 361,
Authors: Alp Sahin, Subhrajit Bhattacharya
Affiliations: Department of Mechanical Engineering and Mechanics, Lehigh University, Bethlehem, PA, USA
Abstract: Many robotics applications benefit from being able to compute multiple geodesic paths in a given configuration space. Existing paradigm is to use topological path planning, which can compute optimal paths in distinct topological classes. However, these methods usually require nontrivial geometric constructions, which are prohibitively expensive in 3-D, and are unable to distinguish between distinct topologically equivalent geodesics that are created due to high-cost/curvature regions or prismatic obstacles in 3-D. In this article, we propose an approach to compute k geodesic paths using the concept of a novel neighborhood-augmented graph, on which graph search algorithms can compute multiple optimal paths that are topo-geometrically distinct. Our approach does not require complex geometric constructions, and the resulting paths are not restricted to distinct topological classes, making the algorithm suitable for problems where finding and distinguishing between geodesic paths are of interest. We demonstrate the application of our algorithm to planning shortest traversible paths for a tethered robot in 3-D with cable-length constraint.
PaperID: 362,
Authors: Marta Lagomarsino, Marta Lorenzini, Elena De Momi, Arash Ajoudani
Affiliations: Human-Robot Interfaces and Interaction Laboratory, Istituto Italiano di Tecnologia, Genoa, Italy; Department of Electronics, Information and Bioengineering, Politecnico di Milano, Milan, Italy
Abstract: Despite impressive advancements of industrial collaborative robots, their potential remains largely untapped due to the difficulty in balancing human safety and comfort with fast production constraints. To help address this challenge, we present PRO-MIND, a novel human-in-the-loop framework that exploits valuable data about the human coworker to optimize robot trajectories. By estimating human attention and mental effort, our method dynamically adjusts safety zones and enables on-the-fly alterations of the robot path to enhance human comfort and optimal stopping conditions. Moreover, we formulate a multiobjective optimization to adapt the robot's trajectory execution time and smoothness based on the current human psychophysical stress, estimated from heart rate variability and frantic movements. These adaptations exploit the properties of B-spline curves to preserve continuity and smoothness, which are crucial factors in improving motion predictability and comfort. Evaluation in two realistic case studies showcases the framework's ability to restrain the operators' workload and stress and to ensure their safety while enhancing human–robot productivity. Further strengths of PRO-MIND include its adaptability to each individual's specific needs and sensitivity to variations in attention, mental effort, and stress during task execution.
PaperID: 363,
Authors: Yu Zheng
Affiliations: Research Institute, UBTECH Robotics Inc., Shenzhen, China
Abstract: This article presents a new efficient way to quantitatively measure separation and penetration between a collection of convex primitives (including ellipsoids, capsules, cylinders, convex polyhedra, and triangles) and a point cloud or a triangle mesh. First, the minimum scaling factor of a convex primitive with respect to its centroid to contact a point or a triangle is proposed as a new distance metrics, which can be greater than, equal to, or less than one, implying that the point or the triangle is separated from, just contacts, or penetrates into the convex primitive. It can be computed mostly in closed form or occasionally with a 1-D gradient descent search, which is much faster than computing the Euclidean distance. Furthermore, an efficient algorithm is proposed to compute the smallest minimum scaling factor of convex primitives in a collection to a point cloud or a triangle mesh. It is based on the discovery that computing the minimum scaling factor of a convex primitive to a point or a triangle yields a plane separating more points or triangles from this or other convex primitives. Then, the overall smallest scaling factor can be found by checking only a few pairs of primitives and points or triangles, being significantly faster than the exhaustive search. In various numerical examples and comparison with the existing algorithms, the proposed metrics and algorithm show superior or comparable efficiency.