湖北农业科学 ›› 2026, Vol. 65 ›› Issue (8): 180-186.doi: 10.14088/j.cnki.issn0439-8114.2026.08.026

• 信息工程 • 上一篇    下一篇

基于主动学习的棉花病虫害图像数据集标注方法

刘凯1a, 肖国锋1a, 迪力夏提·多力昆1a,1b,1c, 赵新苗1a,1b,1c, 徐金1a,1b,1c, 汤丽斯2   

  1. 1.新疆农业大学,a.计算机与信息工程学院; b.新疆农业信息化工程技术研究中心; c.智能农业教育部工程研究中心,乌鲁木齐 830052;
    2.新疆维吾尔自治区农业科学院,乌鲁木齐 830091
  • 收稿日期:2026-06-18 发布日期:2026-09-02
  • 通讯作者: 迪力夏提·多力昆(1993-),男(维吾尔族),新疆莎车人,讲师,硕士,主要从事人工智能、农业信息化研究,(电话)15199099067(电子信箱)dlxt.dlk@xjau.edu.cn。
  • 作者简介:刘 凯(2003-),男,吉林白山人,在读本科生,专业方向为机器视觉,(电话)15104395720(电子信箱)d1539995352@163.com
  • 基金资助:
    国家科技部专项资金项目(2022ZD0115805); 新疆维吾尔自治区重大科技专项(2022A02011-4); 新疆农业信息化工程中心开放课题(XJAIEC2026K004); 新疆维吾尔自治区大学生创新训练计划项目(S202510758037)

An image dataset annotation method for cotton pests and diseases based on active learning

LIU Kai1a, XIAO Guo-feng1a, Dilixiati Duolikun1a, 1b, 1c, ZHAO Xin-miao1a, 1b, 1c, XU Jin1a, 1b, 1c, TANG Li-si2   

  1. 1. a. College of Computer and Information Engineering; b. Xinjiang Agricultural Information Engineering Technology Research Center; c. Engineering Research Center of Intelligent Agriculture, Ministry of Education, Xinjiang Agricultural University, Urumqi 830052, China;
    2. Xinjiang Academy of Agricultural Sciences, Urumqi 830091, China
  • Received:2026-06-18 Online:2026-09-02

摘要: 为了降低大规模图像数据集构建的人工标注成本,并解决传统主动学习方法易忽略长尾场景导致模型性能受限的问题,构建了一种三阶段渐进式图像数据集智能标注方法。首先,在冷启动阶段引入DINOv2进行特征预提取,并提出自适应少数簇增强算法(AMCB);其次,在主动学习循环阶段设计了两阶段混合采样算法,通过MiniBatch K-means进行多样性粗筛,并融合预测不确定性与局部特征密度进行精选,以高效驱动模型迭代;最后,在模型性能达标后,利用基于置信度分层的筛选机制对剩余数据进行全量自动标注。在棉花病虫害图像数据集上的试验结果表明,与随机采样算法相比,AMCB+两阶段混合采样算法能够更均衡地覆盖稀有类别,在达到同等检测精度的前提下明显降低了人工标注量。

关键词: 主动学习, 棉花病虫害, 图像数据集, 标注方法

Abstract: To reduce the manual annotation cost of large-scale image dataset construction and to solve the problem that traditional active learning methods tended to ignore long-tail scenarios, which limited model performance, a three-stage progressive intelligent annotation method for image datasets was constructed. First, in the cold start stage, DINOv2 was introduced for feature pre-extraction, and an adaptive minority cluster boosting algorithm (AMCB) was proposed; then, in the active learning cycle stage, a two-stage hybrid sampling algorithm was designed, which performed diversity coarse screening through MiniBatch K-means and combined prediction uncertainty with local feature density for fine selection, so as to efficiently drive model iteration; finally, after the model performance reached the standard, a confidence-based stratified filtering mechanism was used to perform full automatic annotation on the remaining data. The experimental results on the cotton pests and diseases image dataset showed that compared with the random sampling algorithm, the AMCB + two-stage hybrid sampling algorithm could cover rare classes more evenly and significantly reduced the manual annotation workload while achieving the same detection accuracy.

Key words: active learning, cotton pests and diseases, image dataset, annotation method

中图分类号: