高级检索

水下声学目标识别技术综述:从机器学习到深度学习

Review of Underwater Acoustic Target Recognition Technology: from Machine Learning to Deep Learning

  • 摘要: 为系统梳理水下声学目标识别领域中机器学习与深度学习方法的发展脉络、关键技术及其应用瓶颈,明确当前研究热点与未来发展方向,首先围绕水下声学信号的数据预处理与特征提取过程,综述了具有物理意义的特征、时频特征、听觉感知特征及多维特征融合方法;在此基础上,剖析了基于支持向量机、隐马尔可夫等模型的传统机器学习基线方法,并归纳了以卷积神经网络、循环神经网络、注意力机制及Transformer为首的深度学习架构;同时,讨论了迁移学习与数据增强在小样本条件下的应用。现有研究表明:水下声学目标识别已由浅层统计学习逐步发展为以深层神经网络特征表征为主的数据驱动范式,其中“时频图+深度网络”是当前常用的技术路线;多特征融合、迁移学习和数据增强能够在一定程度上缓解低信噪比、样本不足和类别不平衡等问题。与此同时,模型性能对数据集划分方式高度敏感,在更严格的录音级划分条件下,识别结果更能反映模型的真实泛化能力,仅采用准确率评价模型性能也存在一定局限。未来水下声学目标识别研究应进一步面向复杂海洋环境与工程部署需求,重点突破标准化基准构建、评价指标完善、跨域泛化、模型可解释性以及轻量化边缘部署等关键问题,推动物理先验与数据驱动模型的深度融合。

     

    Abstract: This paper aims to systematically review the development, key techniques, and major challenges of machine learning and deep learning methods for underwater acoustic target recognition, and to clarify current research hotspots and future directions. It first reviews the pipeline of underwater acoustic signal processing, including data preprocessing, feature extraction, and related representation methods, covering physically meaningful features, time-frequency features, auditory perceptual features, and multi-feature fusion strategies. On this basis, traditional machine learning baseline methods based on models such as Support Vector Machines (SVM) and Hidden Markov Models (HMM) are analyzed, and deep learning architectures led by Convolutional Neural Networks (CNNs), Recurrent Neural Networks (RNNs), attention mechanisms, and Transformers are summarized. Meanwhile, the application of transfer learning and data augmentation under few-shot conditions is discussed. Existing studies indicate that underwater acoustic target recognition (UATR) has gradually evolved from shallow statistical learning toward a data-driven paradigm centered on deep neural feature representation, with the combination of time-frequency representations and deep networks becoming a commonly adopted technical route. Multi-feature fusion, transfer learning, and data augmentation can effectively alleviate the problems of low signal-to-noise ratio, limited samples, and class imbalance. Meanwhile, model performance is highly sensitive to dataset partition protocols, and recording-level evaluation provides a more realistic measure of generalization. In addition, accuracy alone is insufficient for evaluating imbalanced multi-class UATR tasks. Future research should focus on standardized benchmarks, more comprehensive evaluation metrics, cross-domain generalization, model interpretability, and lightweight deployment for edge platforms, while promoting a deeper integration of physical priors with deep learning models.

     

/

返回文章
返回