Robust MRCD-PCA in Machine Learning for Breast Cancer Classification
DOI:
https://doi.org/10.35877/454RI.asci4755Keywords:
Artificial neural network, K-means, Principal component analysis, Robust principal component analysis, Support vector machineAbstract
Breast cancer classification plays a critical role in early diagnosis and clinical decision-making. However, medical datasets are often high-dimensional and contaminated with outliers, which can degrade classification performance. While Principal Component Analysis (PCA) is commonly used for dimensionality reduction, its sensitivity to outliers limits its effectiveness in medical data analysis. To address this limitation, this study proposes a robust PCA (RPCA) approach based on the Minimum Regularized Determinant (MRCD) estimator for dimensionality reduction and named the proposed method as Robust MRCD-PCA (RMPCA). The classification using proposed RMPCA is evaluated using Support Vector Machine (SVM), Artificial Neural Network (ANN), and K-means classifiers, and compared against baseline and PCA-based models. A total of nine classification models is examined using a breast cancer dataset, with performance assessed using accuracy, precision, recall, and F1-score. Experimental results demonstrate that RMPCA achieves more consistent and reliable classification performance, particularly when combined with supervised classifiers such as SVM, outperforming both baseline and PCA-based approaches. These findings highlight the importance of robust dimensionality reduction as an effective preprocessing strategy for improving machine learning-based breast cancer classification.
Downloads
References
Adiwijaya, Wisesty, U. N., Lisnawati, E., Aditsania, A., & Kusumo, D. S. (2018). Dimensionality reduction using Principal Component Analysis for cancer detection based on microarray data classification. Journal of Computer Science, 14(11), 1521–1530. https://doi.org/10.3844/jcssp.2018.1521.1530
Aggarwal, C. C. (2020). Machine Learning for Data Mining (2nd Editio). Springer.
Alessandrini, M., Biagetti, G., Crippa, P., Falaschetti, L., Luzzi, S., & Turchetti, C. (2022). Robust-PCA and LSTM Recurrent Neural Network. 1–18.
Almansour, N. A., Syed, H. F., Khayat, N. R., Altheeb, R. K., Juri, R. E., Alhiyafi, J., Alrashed, S., & Olatunji, S. O. (2019). Neural network and support vector machine for the prediction of chronic kidney disease: A comparative study. Computers in Biology and Medicine, 109(October 2018), 101–111. https://doi.org/10.1016/j.compbiomed.2019.04.017
An, Q., Rahman, S., Zhou, J., & Kang, J. J. (2023). A Comprehensive Review on Machine Learning in Healthcare Industry: Classification, Restrictions, Opportunities and Challenges. Sensors, 23(9). https://doi.org/10.3390/s23094178
Azkal, M., Aziz, A., Ismail, S., & Allias, N. (2022). Deep Learning in Face Recognition for Attendance System?: An Exploratory Study. 7(2), 74–81. https://doi.org/10.24191/jcrinn.v7i2.288
Boudt, K., Rousseeuw, P. J., Vanduffel, S., & Verdonck, T. (2020). The minimum regularized covariance determinant estimator. Statistics and Computing, 30(1), 113–128. https://doi.org/10.1007/s11222-019-09869-x
Chen, X., Zhang, B., Wang, T., Bonni, A., & Zhao, G. (2020). Robust principal component analysis for accurate outlier sample detection in RNA-Seq data. BMC Bioinformatics, 21(1), 1–20. https://doi.org/10.1186/s12859-020-03608-0
Chhajer, P., Shah, M., & Kshirsagar, A. (2022). The applications of artificial neural networks , support vector machines , and long – short term memory for stock market prediction. Decision Analytics Journal, 2(November 2021), 100015. https://doi.org/10.1016/j.dajour.2021.100015
Dhomse Kanchan, B., & Mahale Kishor, M. (2017). Study of machine learning algorithms for special disease prediction using principal of component analysis. Proceedings - International Conference on Global Trends in Signal Processing, Information Computing and Communication, ICGTSPICC 2016, (February), 5–10. https://doi.org/10.1109/ICGTSPICC.2016.7955260
Filzmoser, P., & Todorov, V. (2013). Robust tools for the imperfect world. Information Sciences, 245, 4–20. https://doi.org/10.1016/j.ins.2012.10.017
Gad, A. F. (2018). Practical Computer Vision Applications Using Deep Learning with CNNs Applications Using Deep Learning with CNNs. Springer.
He, L., Yang, Y., & Zhang, B. (2023). Robust PCA for high-dimensional data based on characteristic transformation. Australian and New Zealand Journal of Statistics, 65(2), 127–151. https://doi.org/10.1111/anzs.12385
Hubert, M., Rousseeuw, P. J., & Branden, K. Vanden. (2005). ROBPCA: A new approach to robust principal component analysis. Technometrics, 47(1), 64–79. https://doi.org/10.1198/004017004000000563
Johnson, R. A., & Wichern, D. W. (2002). Applied multivariate statistical analysis (Fifth Edit). New Jersey: Hall.
Jollife, I. T., & Cadima, J. (2016). Principal component analysis: A review and recent developments. Philosophical Transactions of the Royal Society A: Mathematical, Physical and Engineering Sciences, 374(2065). https://doi.org/10.1098/rsta.2015.0202
Kartini, D., Ahmad Yani, J. K., Selatan, K., Amin Badali, R., Turianto Nugrahadi, D., Indriani, F., & Wahyu Saputro, S. (2025). Dimensionality Reduction Using Principal Component Analysis and Feature Selection Using Genetic Algorithm with Support Vector Machine for Microarray Data Classification. Indonesian Journal of Electronics, Electromedical Engineering, and Medical Informatics, 7(1), 154–166.
Khan, I. K., Daud, H. B., Zainuddin, N. B., Sokkalingam, R., Abdussamad, Museeb, A., & Inayat, A. (2024). Addressing limitations of the K-means clustering algorithm: outliers, non-spherical data, and optimal cluster selection. AIMS Mathematics, 9(9), 25070–25097. https://doi.org/10.3934/math.20241222
Koca, Y. B., & Aktepe, E. (2024). Effect of dimension reduction with PCA and machine learning algorithms on diabetes diagnosis performance. Turkish Journal of Engineering, 8(3), 447–456. https://doi.org/10.31127/tuje.1413087
Lakshminarayanan, S. K., & Mccrae, J. (2019). A Comparative Study of SVM and LSTM Deep Learning Algorithms for Stock Market Prediction.
Li, S., Song, W., Fang, L., Chen, Y., Ghamisi, P., & Benediktsson, J. A. (2019). Deep Learning for Hyperspectral Image Classification: An Overview. IEEE Transactions on Geoscience and Remote Sensing, 57(9).
Mienye, I. D. (2024). A Comprehensive Review of Deep Learning?: Architectures , Recent Advances , and Applications.
Rahman, M. A., Khan, M. S. H., Watanobe, Y., Prioty, J. T., Annita, T. T., Rahman, S., Hossain, M. S., Aitijjo, S. A., Taskin, R. I., Dhrubo, V., Hanip, A., & Bhuiyan, T. (2025). Advancements in Breast Cancer Detection: A Review of Global Trends, Risk Factors, Imaging Modalities, Machine Learning, and Deep Learning Approaches. In BioMedInformatics (Vol. 5, Number 3). https://doi.org/10.3390/biomedinformatics5030046
Rencher, A. C. (2002). Methods of Multivariate Analysis. https://doi.org/10.2307/2669873
Shahare, P. D., & Giri, R. N. (2015). Comparative Analysis of Artificial Neural Network and Support Vector Machine Classification for Breast Cancer Detection. 2114–2119.
Shrifan, N. H. M. M., Akbar, M. F., Ashidi, N., & Isa, M. (2022). An adaptive outlier removal aided k-means clustering algorithm. Journal of King Saud University - Computer and Information Sciences, 34(8), 6365–6376. https://doi.org/10.1016/j.jksuci.2021.07.003
Ulku, I., Akagündüz, E., & Ulku, I. (2022). A Survey on Deep Learning-based Architectures for Semantic Segmentation on 2D Images A Survey on Deep Learning-based Architectures for Semantic Segmentation on 2D Images. Applied Artificial Intelligence, 00(00), 1–45. https://doi.org/10.1080/08839514.2022.2032924
Yaqoob, A., Verma, N. K., Mir, M. A., Tejani, G. G., Eisa, N. H. B., Mamoun Hussien Osman, H., & Shah, M. A. (2025). SGA-Driven feature selection and random forest classification for enhanced breast cancer diagnosis: A comparative study. Scientific Reports, 15(1), 1–23. https://doi.org/10.1038/s41598-025-95786-1
Zahariah, S., & Midi, H. (2023). Minimum regularized covariance determinant and principal component analysis-based method for the identification of high leverage points in high dimensional sparse data. Journal of Applied Statistics, 50(13), 2817–2835. https://doi.org/10.1080/02664763.2022.2093842
Zhu, A., Hua, Z., Shi, Y., & Tang, Y. (2021). An Improved K-Means Algorithm Based on Evidence Distance.
Published
How to Cite
Issue
Section
License
Copyright (c) 2026 Sharifah Sakinah Syed Abd Mutalib, Maharani Abu Bakar, Danang Adi Pratama, Muhamad Safiih Lola (Author)

This work is licensed under a Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International License.


