Islamic Sound Recognition Using MFCC and SVM: Case Study on Takbir and Sholawat
DOI:
https://doi.org/10.47709/cnahpc.v7i3.6523Keywords:
Takbir, Sholawat, MFCC, Support Vector Machine, speech recognition, Sound ClassificationAbstract
This study aims to develop an identification model for Islamic religious sounds, specifically Takbir and Sholawat pronunciations, using audio signal processing and machine learning techniques. With the increasing need for intelligent systems capable of recognizing speech patterns in religious contexts, the implementation of reliable audio classification methods becomes essential. This research utilizes Mel-Frequency Cepstral Coefficients (MFCC) to extract relevant spectral features from audio samples, representing the unique characteristics of Takbir and Sholawat utterances. The dataset consists of 300 audio recordings, evenly distributed between the two classes. Each audio file is preprocessed and converted into a fixed-length MFCC feature vector, which is then labeled accordingly. The feature vectors are split into training and testing sets using an 70:30 ratio. A Support Vector Machine (SVM) classifier is trained using the training data to recognize the distinction between Takbir and Sholawat patterns based on their acoustic signatures. Performance evaluation is carried out using accuracy, precision, recall, and F1-score metrics. The dataset used consists of 300 audio recordings with a division of 200 takbir recordings and 100 sholawat recordings. The MFCC feature extraction process uses 13 coefficients with optimized parameters to capture discriminative spectral characteristics. As a baseline, a Support Vector Machine (SVM) implementation with Radial Basis Function (RBF) kernel was performed for performance comparison.
Downloads
References
Helmiyah, S., Riadi, I., Umar, Rusydi., Hanif, A., Yudhana, A., & Fadlil, A. (2020). Identifikasi Emosi Manusia Berdasarkan Ucapan Menggunakan Metode Ekstraksi Ciri LPC Dan Metode Euclidean Distance. Jurnal Teknologi Informasi dan Ilmu Komputer (JTIIK), 7(6), 1177-1186. Doi: 10.25126/jtiik.202072693.
Aswari, P. and Diana, N. E. (2016) ‘Identifikasi Emosi Berdasarkan Action Unit Menggunakan Metode Bézier Curve’, Sinergi. Mercu Buana University, 20(1), pp. 74–84. Doi: 10.22441/sinergi.2016.1.010.
Chamoli, A., Semwal, A. and Saikia, N. (2017) ‘Detection Of Emotion In Analysis Of Speech Using Linear Predictive Coding Techniques (L.P.C)’, in Proceedings of the International Conference on Inventive Systems and Control, ICISC 2017, pp. 1–4. Doi: 10.1109/ICISC.2017.8068642.
Chaudhari, P. R. and Alex, J. S. R. (2016) ‘Selection of Features for Emotion Recognition from Speech’, Indian Journal of Science and Technology, 9(39), pp. 1–5. Doi: 10.17485/ijst/2016/v9i39/95585.
Deza, M. M. and Deza, E. (2009). ‘Encyclopedia of distances’, in Encyclopedia of distances. Springer, pp. 1–583.
Dewi, I. A., Zulkarnain, A. and Lestari, A. A. (2018) ‘Identifikasi Suara Tangisan Bayi menggunakan Metode LPC dan Euclidean Distance’, ELKOMIKA: Jurnal Teknik Energi Elektrik, Teknik Telekomunikasi, & Teknik Elektronika, 6(1), p. 153. doi: 10.26760/elkomika.v6i1.153.
Gumelar, A. B. et al. (2019) ‘Human Voice Emotion Identification Using Prosodic and Spectral Feature Extraction Based on Deep Neural Networks’, in 2019 IEEE 7th International Conference on Serious Games and Applications for Health (SeGAH). IEEE, pp. 1–8.
Wibowo, A., Santosa, P. I., & Nugroho, L. E. (2020). Speech emotion recognition using MFCC and SVM on Indonesian dataset. International Journal of Advanced Computer Science and Applications (IJACSA), 11(5), 346–351. https://doi.org/10.14569/IJACSA.2020.0110544
Patil, S. S., & Jain, R. (2020). Emotion recognition using MFCC and machine learning techniques. Procedia Computer Science, 167, 2255–2263. https://doi.org/10.1016/j.procs.2020.03.276
S. Vaishnav and S. Mitra, “Speech Emotion Recognition: A Review,” Int. Res. J.Eng. Technol. IRJET, vol. 3, no. 04, pp. 313–316, 2016.
S. Lalitha, A. Madhavan, B. Bhushan, and S. Saketh, “Speech Emotion Recognition,” 2014 Int. Conf. Adv. Electron. Comput. Commun. ICAECC 2014, vol.7, 2015, doi: 10.1109/ICAECC.2014.7002390.
A. B. Gumelar, “Human Voice Emotion Identification Using Prosodic and Spectral Feature Extraction Based on Deep Neural Networks,” 2019 IEEE 7th Int. Conf. Serious Games Appl. Health SeGAH, pp. 1–8, 2019.
M. D. Pell and S. A. Kotz, “On the time course of vocal emotion recognition,” PLoS One, vol. 6, no. 11, p. e27256, 2011.
I. Idrisa, M. S. H. Salamb, and M. S. Sunarc, “Speech Emotion Classification Using SVM and MLP on Prosodic and Voice Quality Features,” J. Teknol., vol. 78, 2015.
G. Liu, W. He, and B. Jin, “Feature Fusion of Speech Emotion Recognition Based Deep Learning,” in 2018 International Conference on Network Infrastructure and Digital Content (IC-NIDC), 2018, pp. 193–197, doi:
10.1109/ICNIDC.2018.8525706.
Y. Sun, G. Wen, and J. Wang, “Weighted spectral features based on local Hu moments for speech emotion recognition,” Biomed. Signal Process. Control, vol. 18, pp. 80–90, 2015, doi: 10.1016/j.bspc.2014.10.008.
S. S. Swaminathan and J. Thangaiyan, “Emotion Speech Recognition using MFCC and Residual Phase in Artificial Neural Network,” Int. J. Eng. Res. Sci. Technol., vol. 4 No.3, no. August, pp. 106–113, 2015.
Irmawan, H. Hikmarika, D. W. Sari, and M. C. Tammimi, “Pengenalan Kata dengan Metode Linear Predictive Coding dan Jaringan Syaraf Tiruan Pada Mobile Robot,” in Conference on Information Technology and Electrical Engineering, 2014, no. October 2014, pp. 139–144.
K. V. Krishna Kishore and P. Krishna Satish, “Emotion Recognition in Speech using MFCC and Wavelet Features,” Proc. 2013 3rd IEEE Int. Adv. Comput. Conf. IACC 2013, pp. 842–847, 2013, doi: 10.1109/IAdCC.2013.6514336.
Downloads
Published
How to Cite
Issue
Section
License
Copyright (c) 2025 Indah Purnama Sari, Abdul Fadlil, Tole Sutikno

This work is licensed under a Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International License.

