Analysis Of Tokopedia Product Clustering Using The K-Means And K-Medoids Algorithms

Authors

  • Raihan Malik Universitas Jambi, Indonesia
  • Pradita Eko Prasetyo Utomo Universitas Jambi, Indonesia
  • Benedika Ferdian Hutabarat Universitas Jambi, Indonesia

DOI:

https://doi.org/10.47709/brilliance.v5i2.6992

Keywords:

Clustering, K-Means, K-Medoids, Silhouette Score, Davies-Bouldin Index

Abstract

The Indonesian e-commerce market has experienced extraordinary growth, driven by increasing internet penetration and smartphone adoption, which necessitates advanced data analysis for competitive advantage. Clustering is a crucial data mining technique used to group products based on similar characteristics, providing in-depth insights into product performance. Previous studies often focused on single performance metrics, overlooking the nuances of combining multiple variables. This study aims to address this gap by implementing and comparing the K-Means and K-Medoids clustering algorithms on Tokopedia product data using a combination of numerical attributes: Price, Customer Rating, Number Sold, and Total Review. The methodology involved data preprocessing, Min-Max Scaling for normalization, and using the Elbow Method to determine the optimal number of clusters, which was found to be K=2. The clustering quality was rigorously evaluated using the Davies-Bouldin Index (DBI) and Silhouette Score. The results demonstrate that K-Means exhibits superior performance, achieving a lower DBI of 0.5717 and a higher Silhouette Score of 0.6012, compared to K-Medoids (DBI: 0.5870; Silhouette Score: 0.5857). Furthermore, K-Means proved significantly more efficient computationally, with an execution time of 0.0947 seconds versus 0.1622 seconds for K-Medoids. The main conclusion is that K-Means is more effective in creating compact and clearly separated clusters. This research contributes a valuable analytical framework for e-commerce managers to comprehensively understand product profiles, guiding more effective marketing and recommendation strategies.

References

Ainur Rahman, & Suroyo, H. (2021). Analisis Data Produk Elektronik Di E-Commerce Dengan Metode Algoritma K-Means Menggunakan Python. Journal of Advances in Information and Industrial Technology, 3(2), 11–18. https://doi.org/10.52435/jaiit.v3i2.158

Arbelaitz, O., Gurrutxaga, I., Muguerza, J., Pérez, J. M., & Perona, I. (2013). An extensive comparative study of cluster validity indices. Pattern Recognition, 46(1), 243–256. https://doi.org/10.1016/j.patcog.2012.07.021

Arora, P., Deepali, & Varshney, S. (2016). Analysis of K-Means and K-Medoids Algorithm for Big Data. Physics Procedia, 78(December 2015), 507–512. https://doi.org/10.1016/j.procs.2016.02.095

Gymnastiar, S., & Bahtiar, A. (2024). Penerapan Algorima K-Means Clustering Untuk Mengelompokan Data Kejadian Kekeringan Di Kabupaten Cirebon. JATI (Jurnal Mahasiswa Teknik Informatika), 8(2), 2325–2331. https://doi.org/10.36040/jati.v8i2.8948

Harjono, S. W., Utami, N. W., & Putri, I. G. A. P. D. (2023). Klasterisasi Tingkat Penjualan pada Startup Panak.id dengan Algoritma K-Means. Jurnal Ilmiah Teknologi Informasi Asia, 17(1), 55–66. https://doi.org/10.32815/jitika.v17i1.888

Hermawati, A., Jumini, S., Astuti, M., Ismail, F., & Rahim, R. (2020). Unsupervised Data Mining with K-Medoids Method in Mapping Areas of Student and Teacher Ratio in Indonesia. TEM Journal, 9(4), 1614–1618. https://doi.org/10.18421/TEM94-37

Idrus, A., Tarihoran, N., Supriatna, U., Tohir, A., Suwarni, S., & Rahim, R. (2022). Distance Analysis Measuring for Clustering using K-Means and Davies Bouldin Index Algorithm. TEM Journal, 11(4), 1871–1876. https://doi.org/10.18421/TEM114-55

Kulkanjanapiban, P., & Silwattananusarn, T. (2025). A Performance-Driven Exploration of Combining Topic Modeling and Machine Learning for Online Learning Data Analysis. TEM Journal, 14(1), 511–527. https://doi.org/10.18421/TEM141-46

Li, S., Lim, C. Y., & Ang, S. L. (2024). An Analysis of Technostress Factors Among Teachers in Hunan, China Through Statistical Methods and K-means Clustering. TEM Journal, 13(4), 3231–3240. https://doi.org/10.18421/TEM134-57

Luká?, J., Kudlová, Z., Kop?áková, J., & Gallo, P. (2025). Impact of Socio-Economic Factors on Digital Literacy and Security. TEM Journal, 14(1), 925–932. https://doi.org/10.18421/TEM141-81

Purba, D. S., Dwi Permatasari, P., Tanjung, N., Rahayu, P., Fitriani, R., Wulandari, S., Universitas, ), Negeri, I., Utara, S., Muslim, U., & Al Washliyah, N. (2025). Analisis Perkembangan Ekonomi Digital Dalam Meningkatkan Pertumbuhan Ekonomi Di Indonesia. Jurnal Masharif Al-Syariah: Jurnal Ekonomi Dan Perbankan Syariah, 10(1), 126–139.

Rhomadhona, H., Kusrini, W., Aprianti, W., & Permadi, J. (2025). Implementation of K-Means Clustering for Social Assistance Recipients with Silhouette Score Evaluation. Brilliance: Research of Artificial Intelligence, 5(1), 136–143. https://doi.org/10.47709/brilliance.v5i1.5900

Riyahi, M., & Martín, A. G. (2025). Optimizing capacity expansion modeling with a novel hierarchical clustering and systematic elbow method: A case study on power and storage units in Spain. Energy, 323(March). https://doi.org/10.1016/j.energy.2025.135788

Saputri, F. W., & Arianto, D. B. (2023). Perbandingan Performa Algoritma K-Means, K- Medoids, Dan Dbscan Dalam Penggerombolan Provinsi Di Indonesia Berdasarkan Indikator Kesejahteraan Masyarakat. Jurnal Teknologi Informasi, 17(2), 138–151.

Sari, V. K., & Nasution, M. I. P. (2024). Dampak E-commerce Terhadap Perkembangan Digital. Jurnal Akademik Ekonomi Dan Manajemen, 1(4), 18–24.

Syahkur, M. R., & Hartama, D. (2024). Evaluasi Jumlah Cluster pada Algoritma K-Means ++ Menggunakan Silhouette dan Elbow dengan Validasi Nilai DBI dalam Mengelompokkan Gizi Balita. Jurnal Sains Dan Teknologi, 13(3), 487–496.

Umagapi, I. T., & Umaternate, B. (2023). Uji Kinerja K-Means Clustering Menggunakan Davies-Bouldin Index Pada Pengelompokan Data Prestasi Siswa. Prosiding Sisfotek, 303–308.

Downloads

Published

2025-10-10

How to Cite

Malik, R., Utomo, P. E. P., & Hutabarat, B. F. (2025). Analysis Of Tokopedia Product Clustering Using The K-Means And K-Medoids Algorithms. Brilliance: Research of Artificial Intelligence, 5(2), 896–903. https://doi.org/10.47709/brilliance.v5i2.6992

Most read articles by the same author(s)

Similar Articles

1 2 3 4 5 6 7 8 9 10 > >> 

You may also start an advanced similarity search for this article.