Analysis Of Tokopedia Product Clustering Using The K-Means And K-Medoids Algorithms
DOI:
https://doi.org/10.47709/brilliance.v5i2.6992Keywords:
Clustering, K-Means, K-Medoids, Silhouette Score, Davies-Bouldin IndexAbstract
The Indonesian e-commerce market has experienced extraordinary growth, driven by increasing internet penetration and smartphone adoption, which necessitates advanced data analysis for competitive advantage. Clustering is a crucial data mining technique used to group products based on similar characteristics, providing in-depth insights into product performance. Previous studies often focused on single performance metrics, overlooking the nuances of combining multiple variables. This study aims to address this gap by implementing and comparing the K-Means and K-Medoids clustering algorithms on Tokopedia product data using a combination of numerical attributes: Price, Customer Rating, Number Sold, and Total Review. The methodology involved data preprocessing, Min-Max Scaling for normalization, and using the Elbow Method to determine the optimal number of clusters, which was found to be K=2. The clustering quality was rigorously evaluated using the Davies-Bouldin Index (DBI) and Silhouette Score. The results demonstrate that K-Means exhibits superior performance, achieving a lower DBI of 0.5717 and a higher Silhouette Score of 0.6012, compared to K-Medoids (DBI: 0.5870; Silhouette Score: 0.5857). Furthermore, K-Means proved significantly more efficient computationally, with an execution time of 0.0947 seconds versus 0.1622 seconds for K-Medoids. The main conclusion is that K-Means is more effective in creating compact and clearly separated clusters. This research contributes a valuable analytical framework for e-commerce managers to comprehensively understand product profiles, guiding more effective marketing and recommendation strategies.
References
Ainur Rahman, & Suroyo, H. (2021). Analisis Data Produk Elektronik Di E-Commerce Dengan Metode Algoritma K-Means Menggunakan Python. Journal of Advances in Information and Industrial Technology, 3(2), 11–18. https://doi.org/10.52435/jaiit.v3i2.158
Arbelaitz, O., Gurrutxaga, I., Muguerza, J., Pérez, J. M., & Perona, I. (2013). An extensive comparative study of cluster validity indices. Pattern Recognition, 46(1), 243–256. https://doi.org/10.1016/j.patcog.2012.07.021
Arora, P., Deepali, & Varshney, S. (2016). Analysis of K-Means and K-Medoids Algorithm for Big Data. Physics Procedia, 78(December 2015), 507–512. https://doi.org/10.1016/j.procs.2016.02.095
Gymnastiar, S., & Bahtiar, A. (2024). Penerapan Algorima K-Means Clustering Untuk Mengelompokan Data Kejadian Kekeringan Di Kabupaten Cirebon. JATI (Jurnal Mahasiswa Teknik Informatika), 8(2), 2325–2331. https://doi.org/10.36040/jati.v8i2.8948
Harjono, S. W., Utami, N. W., & Putri, I. G. A. P. D. (2023). Klasterisasi Tingkat Penjualan pada Startup Panak.id dengan Algoritma K-Means. Jurnal Ilmiah Teknologi Informasi Asia, 17(1), 55–66. https://doi.org/10.32815/jitika.v17i1.888
Hermawati, A., Jumini, S., Astuti, M., Ismail, F., & Rahim, R. (2020). Unsupervised Data Mining with K-Medoids Method in Mapping Areas of Student and Teacher Ratio in Indonesia. TEM Journal, 9(4), 1614–1618. https://doi.org/10.18421/TEM94-37
Idrus, A., Tarihoran, N., Supriatna, U., Tohir, A., Suwarni, S., & Rahim, R. (2022). Distance Analysis Measuring for Clustering using K-Means and Davies Bouldin Index Algorithm. TEM Journal, 11(4), 1871–1876. https://doi.org/10.18421/TEM114-55
Kulkanjanapiban, P., & Silwattananusarn, T. (2025). A Performance-Driven Exploration of Combining Topic Modeling and Machine Learning for Online Learning Data Analysis. TEM Journal, 14(1), 511–527. https://doi.org/10.18421/TEM141-46
Li, S., Lim, C. Y., & Ang, S. L. (2024). An Analysis of Technostress Factors Among Teachers in Hunan, China Through Statistical Methods and K-means Clustering. TEM Journal, 13(4), 3231–3240. https://doi.org/10.18421/TEM134-57
Luká?, J., Kudlová, Z., Kop?áková, J., & Gallo, P. (2025). Impact of Socio-Economic Factors on Digital Literacy and Security. TEM Journal, 14(1), 925–932. https://doi.org/10.18421/TEM141-81
Purba, D. S., Dwi Permatasari, P., Tanjung, N., Rahayu, P., Fitriani, R., Wulandari, S., Universitas, ), Negeri, I., Utara, S., Muslim, U., & Al Washliyah, N. (2025). Analisis Perkembangan Ekonomi Digital Dalam Meningkatkan Pertumbuhan Ekonomi Di Indonesia. Jurnal Masharif Al-Syariah: Jurnal Ekonomi Dan Perbankan Syariah, 10(1), 126–139.
Rhomadhona, H., Kusrini, W., Aprianti, W., & Permadi, J. (2025). Implementation of K-Means Clustering for Social Assistance Recipients with Silhouette Score Evaluation. Brilliance: Research of Artificial Intelligence, 5(1), 136–143. https://doi.org/10.47709/brilliance.v5i1.5900
Riyahi, M., & Martín, A. G. (2025). Optimizing capacity expansion modeling with a novel hierarchical clustering and systematic elbow method: A case study on power and storage units in Spain. Energy, 323(March). https://doi.org/10.1016/j.energy.2025.135788
Saputri, F. W., & Arianto, D. B. (2023). Perbandingan Performa Algoritma K-Means, K- Medoids, Dan Dbscan Dalam Penggerombolan Provinsi Di Indonesia Berdasarkan Indikator Kesejahteraan Masyarakat. Jurnal Teknologi Informasi, 17(2), 138–151.
Sari, V. K., & Nasution, M. I. P. (2024). Dampak E-commerce Terhadap Perkembangan Digital. Jurnal Akademik Ekonomi Dan Manajemen, 1(4), 18–24.
Syahkur, M. R., & Hartama, D. (2024). Evaluasi Jumlah Cluster pada Algoritma K-Means ++ Menggunakan Silhouette dan Elbow dengan Validasi Nilai DBI dalam Mengelompokkan Gizi Balita. Jurnal Sains Dan Teknologi, 13(3), 487–496.
Umagapi, I. T., & Umaternate, B. (2023). Uji Kinerja K-Means Clustering Menggunakan Davies-Bouldin Index Pada Pengelompokan Data Prestasi Siswa. Prosiding Sisfotek, 303–308.
Downloads
Published
How to Cite
Issue
Section
License
Copyright (c) 2025 Raihan Malik, Pradita Eko Prasetyo Utomo, Benedika Ferdian Hutabarat

This work is licensed under a Creative Commons Attribution-NonCommercial 4.0 International License.















