Proceedings of the XMO Industrial Seminar 2026: Excellence in Manufacturing and Operations
Keywords
Robotics, Space, Machine Learning, AI, Computer Vision, Perception, Fiducial Marker, Manipulation, Manufacturing, Assembly
Tracks
INTEGRATION AND SYSTEMS
DOI
10.5703/1288284318664
Abstract
Visual 6D object pose estimation is a critical capability in modern robotics and computer vision. This is most commonly achieved by exploiting an object's identifiable features and known geometry. While classical computer vision algorithms perform reliably under controlled conditions, they are sensitive to common real-world degradations such as occlusion and adverse illumination. This paper presents a framework for robust and accurate pose estimation that combines deep learning and classical algorithms. Our approach features a U-Net-based segmentation model and a MobileNetV3-based keypoint regression network, both trained entirely on synthetic data with extensive domain randomization. An iterative PnP refinement further enhances accuracy and reliability. We focus on fiducial markers as a natural instantiation of this framework, given their ubiquitous use across robotics and computer vision and their purpose-designed high-contrast patterns with known geometry. We evaluate the method across four challenging conditions — truncation, underexposure, glare, and shadow— and compare it against classical detection. Our method achieves a higher detection rate (99%) than classical approaches (64%) and demonstrates strong resilience to partial visibility and environmental degradation, while maintaining pose accuracy.
Robust and Accurate Pose Estimation for In-Space Robotic Manufacturing under Challenging Lighting Conditions
Visual 6D object pose estimation is a critical capability in modern robotics and computer vision. This is most commonly achieved by exploiting an object's identifiable features and known geometry. While classical computer vision algorithms perform reliably under controlled conditions, they are sensitive to common real-world degradations such as occlusion and adverse illumination. This paper presents a framework for robust and accurate pose estimation that combines deep learning and classical algorithms. Our approach features a U-Net-based segmentation model and a MobileNetV3-based keypoint regression network, both trained entirely on synthetic data with extensive domain randomization. An iterative PnP refinement further enhances accuracy and reliability. We focus on fiducial markers as a natural instantiation of this framework, given their ubiquitous use across robotics and computer vision and their purpose-designed high-contrast patterns with known geometry. We evaluate the method across four challenging conditions — truncation, underexposure, glare, and shadow— and compare it against classical detection. Our method achieves a higher detection rate (99%) than classical approaches (64%) and demonstrates strong resilience to partial visibility and environmental degradation, while maintaining pose accuracy.