Principal Component Optimization for Random Forest-Based Classification of Palm Oil Leaf Disease Images
DOI:
https://doi.org/10.47709/brilliance.v6i3.7552Keywords:
principal component analysis, random forest, dimensionality reduction, trade-offAbstract
Principal Component Analysis (PCA) is a widely used dimensionality reduction technique for mitigating high-dimensional feature spaces, while Random Forest is a robust ensemble classifier that can naturally handle many input variables. However, the effect of the number of PCA components on the predictive performance and generalization ability of Random Forest models is still not well quantified, especially in terms of its trade-off between information preservation and noise reduction. This study investigates how varying the number of PCA components from 2 to 10 influences the performance of a Random Forest classifier on a multiclass dataset. The experimental design employs k-fold cross-validation and multiple values of the number of trees (n_estimators), and evaluates models using Accuracy, Precision, Recall, F1-score, and training time. The results exhibit an inverted U-shaped relationship, where 6–7 PCA components yield the highest and most stable performance, with average Accuracy around 0.96 and F1-score around 0.97, while very low (2–3) and high (?8) numbers of components lead to underfitting and structural overfitting, respectively. These findings suggest that PCA-based dimensionality reduction should be tuned with respect to discriminative performance rather than solely maximizing explained variance, and that a moderate number of components can best exploit the synergy between PCA and Random Forest.
References
Breiman, L. (2001). Random forests. Machine Learning, 45(1), 5–32.
Cai, H., Li, Y., Qian, Y., Li, B., & Fu, Y. (2016). Dimensionality reduction: PCA for network intrusion detection. International Journal of Computer Applications, 177(5), 6–13.
Fan, J., Wang, W., Xie, Y., & Yang, Z. (2018). A shrinkage principle component analysis based random forest with the nearest neighbor method for hyperspectral image classification. Applied Soft Computing, 65, 513–524.
Hamdani, H., Septiarini, A., Sunyoto, A., Suyanto, S., & Utaminingrum, F. (2021). Detection of oil palm leaf disease based on color histogram and supervised classifier. Optik, 245, 167753. https://doi.org/10.1016/j.ijleo.2021.167753
Khalid, S., Khalil, T., & Nasreen, S. (2014). A survey of feature selection and feature extraction techniques in machine learning. 2014 Science and Information Conference, 372–378.
Kurita, T. (2021). Principal component analysis (PCA). In K. Ikeuchi (Ed.), Computer vision: A reference guide (pp. 1013–1016). Springer. https://doi.org/10.1007/978-3-030-63416-2_649
Liaw, A., & Wiener, M. (2002). Classification and regression by randomForest. R News, 2(3), 18–22.
Lu, Y., Chen, M., & Zhang, X. (2024). A comparative study of PCA and LDA for dimensionality reduction in a 4 way classification problem. International Journal of Data Science and Analytics, 10(1), 45–59.
Ningsih, R. (2023). Optimasi metode Random Forest menggunakan principal component analysis pada data kualitas air. Jurnal Sains Data Indonesia, 4(2), 77–88.
Rachmawati, S., Pramudita, D., & Kurniawan, A. (2022). Dimensional reduction of QSAR features using a combination of PCA and Random Forest for toxicity prediction. Jurnal Penelitian Pendidikan IPA, 8(4), 2150–2158.
Rahmanto, O., Julianto, V., & Arrahimi, A. R. (2024). Machine learning to detect palm oil diseases based on leaf extraction features and principal component analysis (PCA). KLIK: Kumpulan Jurnal Ilmu Komputer, 11(1), 26–36. https://doi.org/10.20527/klik.v11i1.659
Saha, R., Ghosh, S., & Banerjee, S. (2023). A comparative analysis of supervised and unsupervised dimensionality reduction techniques for high dimensional biomedical data. BMC Medical Informatics and Decision Making, 23, 41.
Septiarini, A., Hamdani, H., Junirianto, E., Thayf, M. S. S., Triyono, G., & Henderi. (2022). Oil palm leaf disease detection on natural background using convolutional neural networks. 2022 IEEE International Conference on Communication, Networks and Satellite (COMNETSAT), 388–392. https://doi.org/10.1109/COMNETSAT56033.2022.9994555
Sitorus, A. (2023). Analisis akurasi Random Forest menggunakan principal component analysis pada data berdimensi tinggi. Indonesian Journal of Computer Science, 8(3), 155–166.
Sugiarti, L. (2023). Random Forest with PCA for breast cancer classification. International Journal of Advanced Computer Science, 13(2), 90–99.
Surya Wijaya, S., Fauziah, F., & Putri, A. D. (2023). Analysis of the comparison between linear regression, Random Forest, and logistic regression methods in predicting crude palm oil (CPO) price. Brilliance: Research of Artificial Intelligence, 3(2), 812–819.
Suryadi, S., Syahputra, D., Astrianda, N., Syahputra, R. A., & Suhendra, R. (2024). Leveraging machine learning for sentiment analysis in hotel applications: A comparative study of support vector machine and Random Forest algorithms. Brilliance: Research of Artificial Intelligence, 4(2), 930–938.
Yuan, Q., Shen, H., Li, T., & Zhang, L. (2022). PCA based random forest for hyperspectral image classification. Knowledge Based Systems, 241, 108264.
Zheng, J., Fu, H., Li, W., Wu, W., Yu, L., Yuan, S., & Kanniah, K. D. (2021). Growing status observation for oil palm trees using unmanned aerial vehicle (UAV) images. ISPRS Journal of Photogrammetry and Remote Sensing, 173, 95–121. https://doi.org/10.1016/j.isprsjprs.2021.01.008
Downloads
Published
How to Cite
Issue
Section
License
Copyright (c) 2026 Oky Rahmanto, Veri Julianto, Ahmad Rusadi Arrahimi

This work is licensed under a Creative Commons Attribution-NonCommercial 4.0 International License.















