Cetacean Call Recognition and Classification Based on Multimodal MAE Data Augmentation Network
-
摘要: 被动声学监测中的叫声识别与分类是海洋动物保护与种群调查的重要手段。针对鲸豚类动物叫声识别与分类中存在的数据稀缺与类间不平衡问题, 数据增强方法具有重要的实用价值与研究意义。然而海洋动物叫声拥有丰富的声学信息, 仅依赖于频域特征提取缺乏对音频结构与语义的建模能力, 难以有效捕捉叫声的深层特征。为此, 文中提出了一种基于多模态掩码自编码器(MAE-MF)的数据增强网络, 突破单模态信息局限, 以梅尔频谱图为主模态, 融合时序特征与帧级统计指标构成多模态输入, 并引入语义标签作为条件引导重建; 同时结合门控融合与交叉注意力机制,强化多模态信息交互与特征表达能力。基于Watkins数据集开展实验, 结果表明: 相较于主流算法, 文中方法的频谱图重建效果更佳;实验测得6类鲸豚物种平均识别准确率达97.6%, 较基础MAE方法提升6.72%。该方案可有效改善样本类别失衡问题, 提升复杂声学特征与微弱叫声的识别能力, 为鲸豚保护相关工作提供可靠的技术支撑。Abstract: Passive acoustic monitoring-based call recognition and classification are essential means for marine animal conservation and population surveys. To address the issues of data scarcity and inter-class imbalance in call recognition and classification, data augmentation methods hold significant practical value and research importance. However, marine animal calls contain rich acoustic information, and relying solely on frequency-domain feature extraction lacks the capability to model audio structure and semantics, making it difficult to effectively capture the deep features of calls. To this end, a data augmentation network based on a multi-modal masked autoencoder with multi-modal fusion(MAE-MF) is proposed in this paper, which breaks through the limitations of single-modal information. The network employed Mel-spectrograms as the primary modality, integrated temporal features and frame-level statistical metrics to form multimodal inputs, and incorporated semantic labels as conditional guidance for reconstruction. Meanwhile, gated fusion and cross-attention mechanisms were combined to enhance multimodal information interaction and feature representation capability. Experiments conducted on the Watkins dataset show that compared with mainstream algorithms, the proposed method achieves better spectrogram reconstruction performance. The average recognition accuracy for six cetacean species reaches 97.6%, which is 6.72% higher than that of the baseline MAE method. This scheme can effectively alleviate the class imbalance problem, enhance the recognition capability for complex acoustic features and weak calls, and provide reliable technical support for cetacean conservation efforts.
-
表 1 Watkins数据集物种分布与样本量
Table 1. Species distribution and sample size of Watkins dataset
序号 物种 样本量 1 虎鲸 2 134 2 座头鲸 1 204 3 抹香鲸 1 922 4 长鳍领航鲸 1 616 5 小须鲸 607 6 宽吻海豚 965 表 2 数据增强系数
Table 2. Data augmentation coefficients
序号 物种 样本量 增强系数 增强后样本量 1 虎鲸 2 134 1.2 2 560 2 座头鲸 1 204 2.1 2 528 3 抹香鲸 1 922 1.3 2 498 4 长鳍领航鲸 1 616 1.6 2 586 5 小须鲸 607 4.0 2 428 6 宽吻海豚 965 2.6 2 509 表 3 模型训练超参数设置
Table 3. Hyperparameter setting for model training
阶段 优化器 初始
学习率批量
大小训练
轮数权重衰
减系数训练 AdamW 0.000 2 128 100 $ 5\times {10}^{-2} $ 微调 AdamW 0.000 1 64 50 $ 1\times {10}^{-2} $ 表 4 不同模型频谱重建质量评估对比
Table 4. Evaluation comparison of spectrogram reconstruction quality among different models
模型 MSE PSNR/dB SSIM DCGAN 0.30×10−2 25.23 0.82 MAE-Res2Net 0.20×10−2 26.99 0.88 MAE 0.18×10−2 27.45 0.90 MAE-MF 0.14×10−2 28.66 0.92 表 5 不同物种频谱重建质量评估结果
Table 5. Evaluation results of spectrogram reconstruction quality of different species
物种 MSE PSNR/dB SSIM 虎鲸 0.12×10−2 29.21 0.94 座头鲸 0.10×10−2 30.00 0.95 抹香鲸 0.14×10−2 28.53 0.92 长鳍领航鲸 0.15×10−2 28.24 0.90 小须鲸 0.16×10−2 27.96 0.89 宽吻海豚 0.19×10−2 27.21 0.86 -
[1] Montgomery J C, Radford C A. Marine bioacoustics[J]. Current Biology, 2017, 27(11): 502-507. doi: 10.1016/j.cub.2017.01.041 [2] Verfuss U K, Gillespie D, Gordon J, et al. Comparing methods suitable for monitoring marine mammals in low visibility conditions during seismic surveys[J]. Marine Pollution Bulletin, 2018, 126: 1-18. doi: 10.1016/j.marpolbul.2017.10.034 [3] Tyack P L. Implications for marine mammals of large-scale changes in the marine acoustic environment[J]. Journal of Mammalogy, 2008, 89(3): 549-558. doi: 10.1644/07-MAMM-S-307R.1 [4] Qiao Z, Liu S, Wang D, et al. Bio-inspired underwater acoustic communication through PCHIP-based whistle generation and improved CSS modulation[J]. Applied Acoustics, 2025, 235: 110673. doi: 10.1016/j.apacoust.2025.110673 [5] Li L, Qiao G, Liu S, et al. Automated classification of Tursiops aduncus whistles based on a depth-wise separable convolutional neural network and data augmentation[J]. The Journal of the Acoustical Society of America, 2021, 150(5): 3861-3873. doi: 10.1121/10.0007291 [6] Abayomi-Alli O O, Damaševičius R, Qazi A, et al. Data augmentation and deep learning methods in sound classification: A systematic review[J]. Electronics, 2022, 11(22): 3795. doi: 10.3390/electronics11223795 [7] Park D S, Chan W, Zhang Y, et al. SpecAugment: A simple data augmentation method for automatic speech recognition[PP/OL]. arXiv(2019-04-18)[2026-04-16]. https://arxiv.org/abs/1904.08779. [8] Zhang H, Cisse M, Dauphin Y N, et al. Mixup: beyond empirical risk minimization[PP/OL]. arXiv(2017-10-25)[2026-04-16]. https://arxiv.org/abs/1710.09412. [9] Kopets E, Shpilevaya T, Vasilchenko O, et al. Generating synthetic sperm whale voice data using StyleGAN2-ADA[J]. Big Data and Cognitive Computing, 2024, 8(4): 40. doi: 10.3390/bdcc8040040 [10] Li P, Roch M A, Klinck H, et al. Learning stage-wise GANs for whistle extraction in time-frequency spectrograms[J]. IEEE Transactions on Multimedia, 2023, 25: 9302-9314. doi: 10.1109/TMM.2023.3251109 [11] Li D, Liao J, Jiang H, et al. A classification method of marine mammal calls based on two-channel fusion network[J]. Applied Intelligence, 2024, 54(4): 3017-3039. doi: 10.1121/1.429434 [12] Wahlberg M, Jensen F H, Aguilar S N, et al. Source parameters of echolocation clicks from wild bottlenose dolphins(tursiops aduncus and tursiops truncatus)[J]. The Journal of the Acoustical Society of America, 2011, 130(4): 2263-2274. doi: 10.1121/1.3624822 [13] Li X, Dong C, Dong G, et al. Marine mammal call classification using a multi-scale two-channel fusion network (MT-resformer)[J]. Journal of Marine Science and Engineering, 2025, 13(5): 944. doi: 10.3390/jmse13050944 [14] Vester H, Hammerschmidt K, Timme M, et al. Bag-of-calls analysis reveals group-specific vocal repertoire in long-finned pilot whales[PP/OL]. arXiv(2014-10-17) [2026-04-16]. https://arxiv.org/abs/1410.4711. [15] Constantinescu C, Brad R. An overview of sound features in time and frequency domain[J]. International Journal of Advanced Statistics and IT&C for Economics and Life Sciences, 2023, 13(1): 45-58. [16] Peeters G. A large set of audio features for sound description(similarity and classification) in the CUIDADO project[R]. Paris: IRCAM, 2004: 1-25. [17] Baumann-Pickering S, McDonald M A, Simonis A E, et al. Species-specific beaked whale echolocation signals[J]. The Journal of the Acoustical Society of America, 2013, 134(3): 2293-2301. doi: 10.1121/1.4817832 [18] He K, Chen X, Xie S, et al. Masked autoencoders are scalable vision learners[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022: 16000-16009. [19] 刘凇佐, 王蕴聪, 青昕, 等. 鲸豚动物吸附式声学行为记录器综述[J]. 水下无人系统学报, 2023, 31(1): 152-166.Liu S Z, Wang Y C, Qing X. et al, Review of attached acoustic behavior recorders for cetaceans[J]. Journal of Unmanned Undersea Systems, 2023, 31(1): 152-166. [20] Radford A, Metz L, Chintala S. Unsupervised representation learning with deep convolutional generative adversarial networks[PP/OL]. arXiv(2015-11-19)[2026-04-16]. https://arxiv.org/abs/1511.06434. [21] Gao S H, Cheng M M, Zhao K, et al. Res2net: a new multi-scale backbone architecture[J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2019, 43(2): 652-662. -

下载: