Abstract
In modern industrial manufacturing, augmented reality (AR) assembly guidance systems offer substantial productivity enhancements by overlaying contextual spatial instructions directly onto physical workpieces. However, traditional interaction modalities, such as handheld controllers or static graphical menus, disrupt natural bimanual workflows and induce cognitive friction. This article presents a robust, real-time hand tracking and dynamic gesture recognition framework engineered specifically for intuitive six-degree-of-freedom (6-DoF) object manipulation in AR assembly tasks. Our architecture couples a lightweight, dual-stage 3D hand pose estimation network with a spatial-temporal graph convolutional network (ST-GCN) optimized for edge deployment on optical see-through head-mounted displays (OST-HMDs). By integrating kinematic joint constraints and a physics-based virtual coupling mechanism, our system delivers sub-millimeter precision during micro-positioning while maintaining high frame rates. Experimental evaluations conducted on a high-fidelity industrial pump assembly benchmark demonstrate an end-to-end tracking latency of 17.8 ms, a mean per-joint position error of 6.2 mm, and a gesture recognition accuracy of 98.4%. Furthermore, user studies reveal a 31.5% reduction in overall task completion time, a 42.1% decrease in assembly errors, and significantly reduced subjective cognitive workload compared to conventional input methods, validating the framework's viability for next-generation smart manufacturing environments.