HUBEI AGRICULTURAL SCIENCES ›› 2026, Vol. 65 ›› Issue (9): 200-208.doi: 10.14088/j.cnki.issn0439-8114.2026.09.032

• Information Engineering • Previous Articles     Next Articles

Monocular depth estimation for cotton fields fused with attention distillation

LIU Taoa, LI Ying-keb   

  1. a. College of Computer and Information Engineering; b. College of Mathematics and Physics, Xinjiang Agricultural University, Urumqi 830052, China
  • Received:2026-05-22 Online:2026-09-25 Published:2026-09-17

Abstract: To address the balance between accuracy and computation for self-supervised monocular depth estimation (MDE) and meet edge device deployment requirements, a lightweight self-supervised MDE framework fusing knowledge distillation and attention was proposed. Based on MobileViT, the framework used the inverted residual mobile block (iRMB) of the efficient model (EMO) as the backbone to unify CNN local features and ViT global modeling. An edge-aware consistency distillation loss was designed to strengthen depth alignment in key regions, and a bidirectional selective channel attention (BSCA) module was embedded to alleviate insufficient spatial and channel feature coupling in lightweight models. The results showed that on KITTI, the model with 5.8 M parameters achieved Abs Rel 0.103 and δ1 0.893, outperforming MViTDepth with 6.5% fewer parameters and 0.2% higher accuracy, at only 3.3 G FLOPs. It also showed good generalization on Make3D and ran at 2.4 ms per frame on NVIDIA Jetson. The framework effectively balanced accuracy and efficiency, providing reliable depth perception support for visual perception and intelligent agricultural machinery operations in cotton fields.

Key words: monocular depth estimation, self-supervised learning, knowledge distillation, lightweight model, cotton field

CLC Number: