湖北农业科学 ›› 2026, Vol. 65 ›› Issue (9): 200-208.doi: 10.14088/j.cnki.issn0439-8114.2026.09.032

• 信息工程 • 上一篇    下一篇

融合注意力蒸馏的棉花田间单目深度估计

刘韬a, 李盈科b   

  1. 新疆农业大学,a.计算机与信息工程学院; b.数理学院,乌鲁木齐 830052
  • 收稿日期:2026-05-22 出版日期:2026-09-25 发布日期:2026-09-17
  • 通讯作者: 李盈科(1979-),男,甘肃庆阳人,副教授,博士,主要从事生物数学及智慧农业研究,(电话)15809910835(电子信箱)xjaulyk@163.com。
  • 作者简介:刘韬(2001-),男,湖北利川人,在读硕士研究生,研究方向为计算机视觉与智能算法,(电话)13997757266(电子信箱)1760927149@qq.com
  • 基金资助:
    国家自然科学基金项目(12561093)

Monocular depth estimation for cotton fields fused with attention distillation

LIU Taoa, LI Ying-keb   

  1. a. College of Computer and Information Engineering; b. College of Mathematics and Physics, Xinjiang Agricultural University, Urumqi 830052, China
  • Received:2026-05-22 Published:2026-09-25 Online:2026-09-17

摘要: 为解决自监督单目深度估计(MDE)精度与计算量难以平衡的问题,满足边缘设备部署需求,提出一种融合知识蒸馏与轻量化注意力的自监督MDE框架。该框架以MobileViT 为基础,采用高效轻量级模型(EMO)的倒残差移动模块(iRMB)构建骨干网络,统一CNN局部特征提取与ViT全局建模能力;设计边缘感知一致性蒸馏损失强化关键区域深度对齐,并嵌入双向选择性通道注意力(BSCA)模块缓解轻量化模型空间与通道特征耦合不足的问题。结果显示,在KITTI数据集上,模型以5.8 M参数实现绝对相对误差(Abs Rel)0.103、精度δ1=0.893,较MViTDepth参数量减少6.5%、精度提升0.2%,FLOPs仅3.3 G;在Make3D数据集验证良好泛化性,在NVIDIA Jetson设备上单帧推理时间2.4 ms。该框架有效兼顾精度与效率,可为棉花田间视觉感知与农机智能作业提供可靠的深度感知支撑。

关键词: 单目深度估计, 自监督学习, 知识蒸馏, 轻量化模型, 棉花田间

Abstract: To address the balance between accuracy and computation for self-supervised monocular depth estimation (MDE) and meet edge device deployment requirements, a lightweight self-supervised MDE framework fusing knowledge distillation and attention was proposed. Based on MobileViT, the framework used the inverted residual mobile block (iRMB) of the efficient model (EMO) as the backbone to unify CNN local features and ViT global modeling. An edge-aware consistency distillation loss was designed to strengthen depth alignment in key regions, and a bidirectional selective channel attention (BSCA) module was embedded to alleviate insufficient spatial and channel feature coupling in lightweight models. The results showed that on KITTI, the model with 5.8 M parameters achieved Abs Rel 0.103 and δ1 0.893, outperforming MViTDepth with 6.5% fewer parameters and 0.2% higher accuracy, at only 3.3 G FLOPs. It also showed good generalization on Make3D and ran at 2.4 ms per frame on NVIDIA Jetson. The framework effectively balanced accuracy and efficiency, providing reliable depth perception support for visual perception and intelligent agricultural machinery operations in cotton fields.

Key words: monocular depth estimation, self-supervised learning, knowledge distillation, lightweight model, cotton field

中图分类号: