Development of a Hybrid ConvNeXt-BiLSTM Model Based on Segment-Level Mel-Spectrograms for AI-Generated Song Detection
DOI:
https://doi.org/10.47709/brilliance.v6i3.9372Keywords:
AI-generated music, audio classification, bidirectional long short-term memory, ConvNeXt, mel-spectrogramAbstract
AI-generated music increasingly resembles human-produced songs, creating a need for reliable source detection in audio classification. This study aims to develop and evaluate a segment-level hybrid model that combines ConvNeXt with bidirectional long short-term memory (BiLSTM) to capture local spectral patterns and temporal relationships across a song. Each song is represented by 24 five-second segments. Every segment is converted into a 128 x 128 mel-spectrogram, processed by ConvNeXt to obtain 768 features, and arranged as a sequence for BiLSTM modeling. The base dataset contains 22,844 songs divided into 16,522 training, 877 validation, and 5,445 test songs. Under the same data split and training protocol, the proposed ConvNeXt-BiLSTM model achieved accuracy of 0.8408, balanced accuracy of 0.8450, sensitivity of 0.7002, and an F1-score of 0.8190. The ConvNeXt-only baseline achieved 0.8228, 0.8276, 0.6613, and 0.7934 for the same metrics. Fine-tuning used 3,200 additional balanced samples, while final evaluation used a separate balanced test set of 200 songs. With a calibrated threshold of 0.484, the fine-tuned model achieved an F1-score of 0.8556 and balanced accuracy of 0.8700, correctly identifying 97 non-AI songs and 77 AI-generated songs. These results show that temporal modeling across mel-spectrogram segments improves detection, although evaluation using additional music generators, codecs, recording qualities, and adaptive audio contexts remains necessary.References
Agostinelli, A., Denk, T. I., Borsos, Z., Engel, J., Verzetti, M., Caillon, A., … Frank, C. (2023). MusicLM: Generating Music From Text. arXiv. https://doi.org/10.48550/arXiv.2301.11325
Almutairi, Z., & Elgibreen, H. (2022). A Review of Modern Audio Deepfake Detection Methods: Challenges and Future Directions. Algorithms, 15(5), 155. https://doi.org/10.3390/a15050155
Ashraf, M., Abid, F., Din, I. U., Rasheed, J., Yesiltepe, M., Yeo, S. F., & Ersoy, M. T. (2023). A Hybrid CNN and RNN Variant Model for Music Classification. Applied Sciences, 13(3), 1476. https://doi.org/10.3390/app13031476
Chen, K., Du, X., Zhu, B., Ma, Z., Berg-Kirkpatrick, T., & Dubnov, S. (2022). HTS-AT: A Hierarchical Token-Semantic Audio Transformer for Sound Classification and Detection. ICASSP 2022 - 2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 646–650. IEEE. https://doi.org/10.1109/ICASSP43922.2022.9746312
Comanducci, L., Bestagini, P., & Tubaro, S. (2025). FakeMusicCaps: A Dataset for Detection and Attribution of Synthetic Music Generated via Text-to-Music Models. Journal of Imaging, 11(7), 242. https://doi.org/10.3390/jimaging11070242
Copet, J., Kreuk, F., Gat, I., Remez, T., Kant, D., Synnaeve, G., … Defossez, A. (2023). Simple and Controllable Music Generation. arXiv. https://doi.org/10.48550/arXiv.2306.05284
Evans, Z., Parker, J. D., Carr, C. J., Zukowski, Z., Taylor, J., & Pons, J. (2024). Stable Audio Open. arXiv. https://doi.org/10.48550/arXiv.2407.14358
Ghosh, K., Bellinger, C., Corizzo, R., Branco, P., Krawczyk, B., & Japkowicz, N. (2024). The Class Imbalance Problem in Deep Learning. Machine Learning, 113(7), 4845–4901. https://doi.org/10.1007/s10994-022-06268-8
Gong, Y., Chung, Y.-A., & Glass, J. (2021a). AST: Audio Spectrogram Transformer. Interspeech 2021, 571–575. ISCA. https://doi.org/10.21437/Interspeech.2021-698
Gong, Y., Chung, Y.-A., & Glass, J. (2021b). PSLA: Improving Audio Tagging With Pretraining, Sampling, Labeling, and Aggregation. IEEE/ACM Transactions on Audio, Speech, and Language Processing, 29, 3292–3306. https://doi.org/10.1109/TASLP.2021.3120633
Jung, J., Heo, H.-S., Tak, H., Shim, H., Chung, J. S., Lee, B.-J., … Evans, N. (2022). AASIST: Audio Anti-Spoofing Using Integrated Spectro-Temporal Graph Attention Networks. ICASSP 2022 - 2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 6367–6371. IEEE. https://doi.org/10.1109/ICASSP43922.2022.9747766
Kapoor, S., & Narayanan, A. (2023). Leakage and the reproducibility crisis in machine-learning-based science. Patterns, 4(9), 100804. https://doi.org/10.1016/j.patter.2023.100804
Khanjani, Z., Watson, G., & Janeja, V. P. (2023). Audio deepfakes: A survey. Frontiers in Big Data, 5, 1001063. https://doi.org/10.3389/fdata.2022.1001063
Koutini, K., Schluter, J., Eghbal-zadeh, H., & Widmer, G. (2022). Efficient Training of Audio Transformers with Patchout. Interspeech 2022, 2753–2757. ISCA. https://doi.org/10.21437/Interspeech.2022-227
Liu, H., Chen, Z., Yuan, Y., Mei, X., Liu, X., Mandic, D., … Plumbley, M. D. (2023). AudioLDM: Text-to-Audio Generation with Latent Diffusion Models. arXiv. https://doi.org/10.48550/arXiv.2301.12503
Liu, H., Yuan, Y., Liu, X., Mei, X., Kong, Q., Tian, Q., … Plumbley, M. D. (2024). AudioLDM 2: Learning Holistic Audio Generation With Self-Supervised Pretraining. IEEE/ACM Transactions on Audio, Speech, and Language Processing, 32, 2871–2883. https://doi.org/10.1109/TASLP.2024.3399607
Liu, Z., Mao, H., Wu, C.-Y., Feichtenhofer, C., Darrell, T., & Xie, S. (2022). A ConvNet for the 2020s. 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 11966–11976. IEEE. https://doi.org/10.1109/CVPR52688.2022.01167
Muruganandham, P., Thangasamy, G. R., Jayaraman, S., & Dharmarajan, R. (2025). LSTM Autoencoder Based Parallel Architecture for Deepfake Audio Detection with Dynamic Residual Encoding and Feature Fusion. Scientific Reports, 15(1). https://doi.org/10.1038/s41598-025-08198-6
Pargent, F., Schoedel, R., & Stachl, C. (2023). Best Practices in Supervised Machine Learning: A Tutorial for Psychologists. Advances in Methods and Practices in Psychological Science, 6(3). https://doi.org/10.1177/25152459231162559
Pfob, A., Lu, S.-C., & Sidey-Gibbons, C. (2022). Machine learning in medicine: a practical introduction to techniques for data pre-processing, hyperparameter tuning, and model comparison. BMC Medical Research Methodology, 22(1), 282. https://doi.org/10.1186/s12874-022-01758-8
Rahman, M. A., Hakim, Z. I. A., Sarker, N. H., Paul, B., & Fattah, S. A. (2024). SONICS: Synthetic Or Not – Identifying Counterfeit Songs. arXiv. https://doi.org/10.48550/arXiv.2408.14080
Rainio, O., Teuho, J., & Klen, R. (2024). Evaluation metrics and statistical tests for machine learning. Scientific Reports, 14(1), 6086. https://doi.org/10.1038/s41598-024-56706-x
Schroer, C., Kruse, F., & Gomez, J. M. (2021). A systematic literature review on applying CRISP-DM process model. Procedia Computer Science, 181, 526–534. https://doi.org/10.1016/j.procs.2021.01.199
Tak, H., Kamble, M., Patino, J., Todisco, M., & Evans, N. (2022). RawBoost: A Raw Data Boosting and Augmentation Method Applied to Automatic Speaker Verification Anti-Spoofing. ICASSP 2022 - 2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 6382–6386. IEEE. https://doi.org/10.1109/ICASSP43922.2022.9746213
Tak, H., Patino, J., Todisco, M., Nautsch, A., Evans, N., & Larcher, A. (2021). End-to-End anti-spoofing with RawNet2. ICASSP 2021 - 2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 6369–6373. IEEE. https://doi.org/10.1109/ICASSP39728.2021.9414234
Tarawneh, A. S., Hassanat, A. B., Altarawneh, G. A., & Almuhaimeed, A. (2022). Stop Oversampling for Class Imbalance Learning: A Review. IEEE Access, 10, 47643–47660. https://doi.org/10.1109/ACCESS.2022.3169512
Vila, L. C., Sturm, B. L. T., Casini, L., & Dalmazzo, D. (2025). The AI Music Arms Race: On the Detection of AI-Generated Music. Transactions of the International Society for Music Information Retrieval, 8(1), 179–194. https://doi.org/10.5334/tismir.254
Wani, T. M., Qadri, S. A. A., Comminiello, D., & Amerini, I. (2024). Detecting Audio Deepfakes: Integrating CNN and BiLSTM with Multi-Feature Concatenation. Proceedings of the 2024 ACM Workshop on Information Hiding and Multimedia Security, 271–276. ACM. https://doi.org/10.1145/3658664.3659647
Xie, Y., Zhou, J., Lu, X., Jiang, Z., Yang, Y., Cheng, H., & Ye, L. (2023). FSD: An Initial Chinese Dataset for Fake Song Detection. arXiv. https://doi.org/10.48550/arXiv.2309.02232
Yamagishi, J., Wang, X., Todisco, M., Sahidullah, M., Patino, J., Nautsch, A., … Delgado, H. (2021). ASVspoof 2021: accelerating progress in spoofed and deepfake speech detection. 2021 Edition of the Automatic Speaker Verification and Spoofing Countermeasures Challenge, 47–54. ISCA. https://doi.org/10.21437/ASVSPOOF.2021-8
Yang, Y.-Y., Hira, M., Ni, Z., Astafurov, A., Chen, C., Puhrsch, C., … Quenneville-Belair, V. (2022). Torchaudio: Building Blocks for Audio and Speech Processing. ICASSP 2022 - 2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 6982–6986. IEEE. https://doi.org/10.1109/ICASSP43922.2022.9747236
Yi, J., Fu, R., Tao, J., Nie, S., Ma, H., Wang, C., … Li, H. (2022). ADD 2022: the first Audio Deep Synthesis Detection Challenge. ICASSP 2022 - 2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 9216–9220. IEEE. https://doi.org/10.1109/ICASSP43922.2022.9746939
Zaman, K., Sah, M., Direkoglu, C., & Unoki, M. (2023). A Survey of Audio Classification Using Deep Learning. IEEE Access, 11, 106620–106649. https://doi.org/10.1109/ACCESS.2023.3318015
Zang, Y., Zhang, Y., Heydari, M., & Duan, Z. (2024). SingFake: Singing Voice Deepfake Detection. ICASSP 2024 - 2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 12156–12160. IEEE. https://doi.org/10.1109/ICASSP48485.2024.10448184
Zhang, C.-B., Jiang, P.-T., Hou, Q., Wei, Y., Han, Q., Li, Z., & Cheng, M.-M. (2021). Delving Deep Into Label Smoothing. IEEE Transactions on Image Processing, 30, 5984–5996. https://doi.org/10.1109/TIP.2021.3089942
Zhen, C., & Changhui, L. (2021). Music Audio Sentiment Classification Based on CNN-BiLSTM and Attention Model. 2021 4th International Conference on Robotics, Control and Automation Engineering (RCAE), 156–160. IEEE. https://doi.org/10.1109/RCAE53607.2021.9638811
Zhuang, Z., Liu, M., Cutkosky, A., & Orabona, F. (2022). Understanding AdamW through Proximal Methods and Scale-Freeness. arXiv. https://doi.org/10.48550/arXiv.2202.00089
Downloads
Published
How to Cite
Issue
Section
License
Copyright (c) 2026 Dimas Aditya Saputra, Nugroho Dwi Saputro, Ramadhan Renaldy

This work is licensed under a Creative Commons Attribution-NonCommercial 4.0 International License.















