Enhanced YOLOV5-6D for robust 6D pose estimation using attention guided and transformer-based feature fusion

Authors

  • Mohd Hairi Mohd Zaman Universiti Kebangsaan Malaysia Author
  • Yuanwei Chen Universiti Kebangsaan Malaysia, Guangdong Technology College Author
  • Mohd Faisal Ibrahim Universiti Kebangsaan Malaysia Author

DOI:

https://doi.org/10.52152/D11518

Keywords:

6D Pose Estimation, YOLOv5-6D, PANet, CBAM, Transformer, PnP

Abstract

Six degree-of-freedom (6D) pose estimation plays a crucial role in robotic grasping, autonomous driving, and augmented reality. Its primary objective is to estimate the three-dimensional (3D) translation and rotation of target objects. Traditional approaches, such as template matching and feature point matching, exhibit limited robustness in complex environments with occlusions and lighting variations. The You Only Look Once (YOLO) algorithm series, known for its efficiency in object detection, has been extended to 6D pose estimation, as seen in YOLO6D and YOLOv5-6D. However, these methods still suffer from insufficient feature modeling, inadequate multi-scale feature fusion, and difficulties in capturing long-range dependencies. To address these issues, this study proposes an optimized YOLOv5-6D model. First, path aggregation network (PANet) is introduced to replace the original bi-directional feature pyramid network (BiFPN) structure, enhancing multi-scale feature fusion and improving detection accuracy. Second, convolutional block attention module (CBAM) mechanisms are incorporated into the feature layers to enhance the model’s ability to focus on key regions, thereby improving feature extraction. Additionally, a transformer module is integrated into the upper layers of the backbone to enhance global context modeling and optimize two-dimensional (2D) keypoint detection. Finally, the perspective-n-point (PnP) algorithm is utilized for 6D pose computation, ensuring more accurate and robust pose estimation. Experimental results on the LINEMOD dataset demonstrate that the proposed method improves the ADD metric from 71.63% to 74.22%, achieving a 2.59% performance gain while maintaining real-time efficiency. This enhancement significantly improves the accuracy and robustness of 6D pose estimation.

Published

2026-05-28

Issue

Section

Articles

Categories