Proceedings of the XMO Industrial Seminar 2026: Excellence in Manufacturing and Operations

Keywords

Robotics, Space, Machine Learning, AI, Computer Vision, Perception, Fiducial Marker, Manipulation, Manufacturing, Assembly

Tracks

INTEGRATION AND SYSTEMS

DOI

10.5703/1288284318664

Abstract

Visual 6D object pose estimation is a critical capability in modern robotics and computer vision. This is most commonly achieved by exploiting an object's identifiable features and known geometry. While classical computer vision algorithms  perform reliably under controlled conditions, they are sensitive to common real-world degradations such as occlusion and adverse illumination. This paper presents a framework for robust and accurate pose estimation that combines deep learning and classical algorithms. Our approach features a U-Net-based segmentation model and a MobileNetV3-based keypoint regression network, both trained entirely on synthetic data with extensive domain randomization. An iterative PnP refinement further enhances accuracy and reliability. We focus on fiducial markers as a natural instantiation of this framework, given their ubiquitous use across robotics and computer vision and their purpose-designed high-contrast patterns with known geometry. We evaluate the method across four challenging conditions — truncation, underexposure, glare, and shadow— and compare it against classical detection. Our method achieves a higher detection rate (99%) than classical approaches (64%) and demonstrates strong resilience to partial visibility and environmental degradation, while maintaining pose accuracy.

Share

COinS
 

Robust and Accurate Pose Estimation for In-Space Robotic Manufacturing under Challenging Lighting Conditions

Visual 6D object pose estimation is a critical capability in modern robotics and computer vision. This is most commonly achieved by exploiting an object's identifiable features and known geometry. While classical computer vision algorithms  perform reliably under controlled conditions, they are sensitive to common real-world degradations such as occlusion and adverse illumination. This paper presents a framework for robust and accurate pose estimation that combines deep learning and classical algorithms. Our approach features a U-Net-based segmentation model and a MobileNetV3-based keypoint regression network, both trained entirely on synthetic data with extensive domain randomization. An iterative PnP refinement further enhances accuracy and reliability. We focus on fiducial markers as a natural instantiation of this framework, given their ubiquitous use across robotics and computer vision and their purpose-designed high-contrast patterns with known geometry. We evaluate the method across four challenging conditions — truncation, underexposure, glare, and shadow— and compare it against classical detection. Our method achieves a higher detection rate (99%) than classical approaches (64%) and demonstrates strong resilience to partial visibility and environmental degradation, while maintaining pose accuracy.