Development of a Hybrid ConvNeXt-BiLSTM Model Based on Segment-Level Mel-Spectrograms for AI-Generated Song Detection

Authors

  • Dimas Aditya Saputra Universitas PGRI Semarang, Indonesia
  • Nugroho Dwi Saputro Universitas PGRI Semarang, Indonesia
  • Ramadhan Renaldy Universitas PGRI Semarang, Indonesia

DOI:

https://doi.org/10.47709/brilliance.v6i3.9372

Keywords:

AI-generated music, audio classification, bidirectional long short-term memory, ConvNeXt, mel-spectrogram

Abstract

AI-generated music increasingly resembles human-produced songs, creating a need for reliable source detection in audio classification. This study aims to develop and evaluate a segment-level hybrid model that combines ConvNeXt with bidirectional long short-term memory (BiLSTM) to capture local spectral patterns and temporal relationships across a song. Each song is represented by 24 five-second segments. Every segment is converted into a 128 x 128 mel-spectrogram, processed by ConvNeXt to obtain 768 features, and arranged as a sequence for BiLSTM modeling. The base dataset contains 22,844 songs divided into 16,522 training, 877 validation, and 5,445 test songs. Under the same data split and training protocol, the proposed ConvNeXt-BiLSTM model achieved accuracy of 0.8408, balanced accuracy of 0.8450, sensitivity of 0.7002, and an F1-score of 0.8190. The ConvNeXt-only baseline achieved 0.8228, 0.8276, 0.6613, and 0.7934 for the same metrics. Fine-tuning used 3,200 additional balanced samples, while final evaluation used a separate balanced test set of 200 songs. With a calibrated threshold of 0.484, the fine-tuned model achieved an F1-score of 0.8556 and balanced accuracy of 0.8700, correctly identifying 97 non-AI songs and 77 AI-generated songs. These results show that temporal modeling across mel-spectrogram segments improves detection, although evaluation using additional music generators, codecs, recording qualities, and adaptive audio contexts remains necessary.

References

Agostinelli, A., Denk, T. I., Borsos, Z., Engel, J., Verzetti, M., Caillon, A., … Frank, C. (2023). MusicLM: Generating Music From Text. arXiv. https://doi.org/10.48550/arXiv.2301.11325

Almutairi, Z., & Elgibreen, H. (2022). A Review of Modern Audio Deepfake Detection Methods: Challenges and Future Directions. Algorithms, 15(5), 155. https://doi.org/10.3390/a15050155

Ashraf, M., Abid, F., Din, I. U., Rasheed, J., Yesiltepe, M., Yeo, S. F., & Ersoy, M. T. (2023). A Hybrid CNN and RNN Variant Model for Music Classification. Applied Sciences, 13(3), 1476. https://doi.org/10.3390/app13031476

Chen, K., Du, X., Zhu, B., Ma, Z., Berg-Kirkpatrick, T., & Dubnov, S. (2022). HTS-AT: A Hierarchical Token-Semantic Audio Transformer for Sound Classification and Detection. ICASSP 2022 - 2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 646–650. IEEE. https://doi.org/10.1109/ICASSP43922.2022.9746312

Comanducci, L., Bestagini, P., & Tubaro, S. (2025). FakeMusicCaps: A Dataset for Detection and Attribution of Synthetic Music Generated via Text-to-Music Models. Journal of Imaging, 11(7), 242. https://doi.org/10.3390/jimaging11070242

Copet, J., Kreuk, F., Gat, I., Remez, T., Kant, D., Synnaeve, G., … Defossez, A. (2023). Simple and Controllable Music Generation. arXiv. https://doi.org/10.48550/arXiv.2306.05284

Evans, Z., Parker, J. D., Carr, C. J., Zukowski, Z., Taylor, J., & Pons, J. (2024). Stable Audio Open. arXiv. https://doi.org/10.48550/arXiv.2407.14358

Ghosh, K., Bellinger, C., Corizzo, R., Branco, P., Krawczyk, B., & Japkowicz, N. (2024). The Class Imbalance Problem in Deep Learning. Machine Learning, 113(7), 4845–4901. https://doi.org/10.1007/s10994-022-06268-8

Gong, Y., Chung, Y.-A., & Glass, J. (2021a). AST: Audio Spectrogram Transformer. Interspeech 2021, 571–575. ISCA. https://doi.org/10.21437/Interspeech.2021-698

Gong, Y., Chung, Y.-A., & Glass, J. (2021b). PSLA: Improving Audio Tagging With Pretraining, Sampling, Labeling, and Aggregation. IEEE/ACM Transactions on Audio, Speech, and Language Processing, 29, 3292–3306. https://doi.org/10.1109/TASLP.2021.3120633

Jung, J., Heo, H.-S., Tak, H., Shim, H., Chung, J. S., Lee, B.-J., … Evans, N. (2022). AASIST: Audio Anti-Spoofing Using Integrated Spectro-Temporal Graph Attention Networks. ICASSP 2022 - 2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 6367–6371. IEEE. https://doi.org/10.1109/ICASSP43922.2022.9747766

Kapoor, S., & Narayanan, A. (2023). Leakage and the reproducibility crisis in machine-learning-based science. Patterns, 4(9), 100804. https://doi.org/10.1016/j.patter.2023.100804

Khanjani, Z., Watson, G., & Janeja, V. P. (2023). Audio deepfakes: A survey. Frontiers in Big Data, 5, 1001063. https://doi.org/10.3389/fdata.2022.1001063

Koutini, K., Schluter, J., Eghbal-zadeh, H., & Widmer, G. (2022). Efficient Training of Audio Transformers with Patchout. Interspeech 2022, 2753–2757. ISCA. https://doi.org/10.21437/Interspeech.2022-227

Liu, H., Chen, Z., Yuan, Y., Mei, X., Liu, X., Mandic, D., … Plumbley, M. D. (2023). AudioLDM: Text-to-Audio Generation with Latent Diffusion Models. arXiv. https://doi.org/10.48550/arXiv.2301.12503

Liu, H., Yuan, Y., Liu, X., Mei, X., Kong, Q., Tian, Q., … Plumbley, M. D. (2024). AudioLDM 2: Learning Holistic Audio Generation With Self-Supervised Pretraining. IEEE/ACM Transactions on Audio, Speech, and Language Processing, 32, 2871–2883. https://doi.org/10.1109/TASLP.2024.3399607

Liu, Z., Mao, H., Wu, C.-Y., Feichtenhofer, C., Darrell, T., & Xie, S. (2022). A ConvNet for the 2020s. 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 11966–11976. IEEE. https://doi.org/10.1109/CVPR52688.2022.01167

Muruganandham, P., Thangasamy, G. R., Jayaraman, S., & Dharmarajan, R. (2025). LSTM Autoencoder Based Parallel Architecture for Deepfake Audio Detection with Dynamic Residual Encoding and Feature Fusion. Scientific Reports, 15(1). https://doi.org/10.1038/s41598-025-08198-6

Pargent, F., Schoedel, R., & Stachl, C. (2023). Best Practices in Supervised Machine Learning: A Tutorial for Psychologists. Advances in Methods and Practices in Psychological Science, 6(3). https://doi.org/10.1177/25152459231162559

Pfob, A., Lu, S.-C., & Sidey-Gibbons, C. (2022). Machine learning in medicine: a practical introduction to techniques for data pre-processing, hyperparameter tuning, and model comparison. BMC Medical Research Methodology, 22(1), 282. https://doi.org/10.1186/s12874-022-01758-8

Rahman, M. A., Hakim, Z. I. A., Sarker, N. H., Paul, B., & Fattah, S. A. (2024). SONICS: Synthetic Or Not – Identifying Counterfeit Songs. arXiv. https://doi.org/10.48550/arXiv.2408.14080

Rainio, O., Teuho, J., & Klen, R. (2024). Evaluation metrics and statistical tests for machine learning. Scientific Reports, 14(1), 6086. https://doi.org/10.1038/s41598-024-56706-x

Schroer, C., Kruse, F., & Gomez, J. M. (2021). A systematic literature review on applying CRISP-DM process model. Procedia Computer Science, 181, 526–534. https://doi.org/10.1016/j.procs.2021.01.199

Tak, H., Kamble, M., Patino, J., Todisco, M., & Evans, N. (2022). RawBoost: A Raw Data Boosting and Augmentation Method Applied to Automatic Speaker Verification Anti-Spoofing. ICASSP 2022 - 2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 6382–6386. IEEE. https://doi.org/10.1109/ICASSP43922.2022.9746213

Tak, H., Patino, J., Todisco, M., Nautsch, A., Evans, N., & Larcher, A. (2021). End-to-End anti-spoofing with RawNet2. ICASSP 2021 - 2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 6369–6373. IEEE. https://doi.org/10.1109/ICASSP39728.2021.9414234

Tarawneh, A. S., Hassanat, A. B., Altarawneh, G. A., & Almuhaimeed, A. (2022). Stop Oversampling for Class Imbalance Learning: A Review. IEEE Access, 10, 47643–47660. https://doi.org/10.1109/ACCESS.2022.3169512

Vila, L. C., Sturm, B. L. T., Casini, L., & Dalmazzo, D. (2025). The AI Music Arms Race: On the Detection of AI-Generated Music. Transactions of the International Society for Music Information Retrieval, 8(1), 179–194. https://doi.org/10.5334/tismir.254

Wani, T. M., Qadri, S. A. A., Comminiello, D., & Amerini, I. (2024). Detecting Audio Deepfakes: Integrating CNN and BiLSTM with Multi-Feature Concatenation. Proceedings of the 2024 ACM Workshop on Information Hiding and Multimedia Security, 271–276. ACM. https://doi.org/10.1145/3658664.3659647

Xie, Y., Zhou, J., Lu, X., Jiang, Z., Yang, Y., Cheng, H., & Ye, L. (2023). FSD: An Initial Chinese Dataset for Fake Song Detection. arXiv. https://doi.org/10.48550/arXiv.2309.02232

Yamagishi, J., Wang, X., Todisco, M., Sahidullah, M., Patino, J., Nautsch, A., … Delgado, H. (2021). ASVspoof 2021: accelerating progress in spoofed and deepfake speech detection. 2021 Edition of the Automatic Speaker Verification and Spoofing Countermeasures Challenge, 47–54. ISCA. https://doi.org/10.21437/ASVSPOOF.2021-8

Yang, Y.-Y., Hira, M., Ni, Z., Astafurov, A., Chen, C., Puhrsch, C., … Quenneville-Belair, V. (2022). Torchaudio: Building Blocks for Audio and Speech Processing. ICASSP 2022 - 2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 6982–6986. IEEE. https://doi.org/10.1109/ICASSP43922.2022.9747236

Yi, J., Fu, R., Tao, J., Nie, S., Ma, H., Wang, C., … Li, H. (2022). ADD 2022: the first Audio Deep Synthesis Detection Challenge. ICASSP 2022 - 2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 9216–9220. IEEE. https://doi.org/10.1109/ICASSP43922.2022.9746939

Zaman, K., Sah, M., Direkoglu, C., & Unoki, M. (2023). A Survey of Audio Classification Using Deep Learning. IEEE Access, 11, 106620–106649. https://doi.org/10.1109/ACCESS.2023.3318015

Zang, Y., Zhang, Y., Heydari, M., & Duan, Z. (2024). SingFake: Singing Voice Deepfake Detection. ICASSP 2024 - 2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 12156–12160. IEEE. https://doi.org/10.1109/ICASSP48485.2024.10448184

Zhang, C.-B., Jiang, P.-T., Hou, Q., Wei, Y., Han, Q., Li, Z., & Cheng, M.-M. (2021). Delving Deep Into Label Smoothing. IEEE Transactions on Image Processing, 30, 5984–5996. https://doi.org/10.1109/TIP.2021.3089942

Zhen, C., & Changhui, L. (2021). Music Audio Sentiment Classification Based on CNN-BiLSTM and Attention Model. 2021 4th International Conference on Robotics, Control and Automation Engineering (RCAE), 156–160. IEEE. https://doi.org/10.1109/RCAE53607.2021.9638811

Zhuang, Z., Liu, M., Cutkosky, A., & Orabona, F. (2022). Understanding AdamW through Proximal Methods and Scale-Freeness. arXiv. https://doi.org/10.48550/arXiv.2202.00089

Downloads

Published

2026-08-07

How to Cite

Saputra, D. A., Saputro, N. D., & Renaldy, R. (2026). Development of a Hybrid ConvNeXt-BiLSTM Model Based on Segment-Level Mel-Spectrograms for AI-Generated Song Detection. Brilliance: Research of Artificial Intelligence, 6(3), 473–484. https://doi.org/10.47709/brilliance.v6i3.9372

Similar Articles

1 2 3 4 5 6 7 8 9 10 > >> 

You may also start an advanced similarity search for this article.