Proceedings of the XMO Industrial Seminar 2026: Excellence in Manufacturing and Operations

Keywords

Object 6D pose estimation; Point cloud registration; Industrial automation

Tracks

TESTING AND VALIDATION STRATEGIES

DOI

10.5703/1288284318682

Abstract

Precise estimation of a workpiece’s position and orientation is indispensable for advanced robotic automation in a smart manufacturing setting, where robotic manufacturing tasks such as welding, spraying, and assembly must be performed on complex three-dimensional components with higher adaptability and guaranteed repeatability. In response to such requirements, this paper presents an integrated vision-based pipeline for object 6D pose estimation using an RGB-D camera mounted on a collaborative robot arm (FANUC CRX-10iA/L). The proposed system applies the Segment Anything Model 2 (SAM2) to isolate the target object from a single RGB frame via a single point prompt, generating a segmentation mask that filters the corresponding depth data. The masked depth image is converted into a target point cloud, which is registered against a pre-captured reference point cloud through a coarse-to-fine registration pipeline comprising FPFH-RANSAC initialization and ICP-based refinement. The resulting homogeneous transformation matrix encodes the full displacement of the target object relative to its reference pose. Experimental results confirm robust registration and pose estimation, achieving a mean fine ICP fitness of 0.971 ± 0.047 and a mean inlier RMSE of 3.17e-3 ± 4.18e-4 m across ten trials on two distinct objects. The proposed pipeline is object-agnostic and does not require any task-specific training, making it well-suited for flexible industrial environments.

Share

COinS
 

RGB-D Object Pose Estimation via SAM2 Segmentation and Coarse-to-Fine Point Cloud Registration

Precise estimation of a workpiece’s position and orientation is indispensable for advanced robotic automation in a smart manufacturing setting, where robotic manufacturing tasks such as welding, spraying, and assembly must be performed on complex three-dimensional components with higher adaptability and guaranteed repeatability. In response to such requirements, this paper presents an integrated vision-based pipeline for object 6D pose estimation using an RGB-D camera mounted on a collaborative robot arm (FANUC CRX-10iA/L). The proposed system applies the Segment Anything Model 2 (SAM2) to isolate the target object from a single RGB frame via a single point prompt, generating a segmentation mask that filters the corresponding depth data. The masked depth image is converted into a target point cloud, which is registered against a pre-captured reference point cloud through a coarse-to-fine registration pipeline comprising FPFH-RANSAC initialization and ICP-based refinement. The resulting homogeneous transformation matrix encodes the full displacement of the target object relative to its reference pose. Experimental results confirm robust registration and pose estimation, achieving a mean fine ICP fitness of 0.971 ± 0.047 and a mean inlier RMSE of 3.17e-3 ± 4.18e-4 m across ten trials on two distinct objects. The proposed pipeline is object-agnostic and does not require any task-specific training, making it well-suited for flexible industrial environments.