Enhanced Breast Cancer Classification Using Ensemble Learning and Data Resampling Techniques: A Machine Learning Approach

Authors
  • Mohammed K. BASHIR

    Author

  • Sanusi A. DARMA

    Author

  • Usman MAHMUD

    Author

  • Mansir ABUBAKAR

    Author

Keywords:
Breast cancer; Machine learning; XGBoost; SMOTE–Tomek; Random Forest; Classification; Diagnosis.
Abstract

Early and reliable breast cancer classification is essential for reducing unnecessary biopsies and ensuring timely diagnosis, particularly in low-resource settings where late detection drives high mortality. This study developed and compared an ensemble machine-learning framework against standard baseline classifiers for binary breast cancer classification using a publicly available clinical breast cancer dataset of 213 records (120 benign and 93 malignant) obtained from the Kaggle repository. To avoid information leakage, all preprocessing—median/mode imputation, Interquartile Range (IQR) outlier capping, label encoding, and z-score standardisation—and a Synthetic Minority Oversampling Technique combined with Tomek Links (SMOTE–Tomek) were fitted on the training partition only, which was balanced to 90 benign and 90 malignant cases, while an 80:20 stratified split reserved 43 cases for testing. Random Forest feature-importance analysis identified tumour size, involved lymph nodes, and metastasis as the most influential predictors, consistent with clinical knowledge. An Extreme Gradient Boosting (XGBoost) classifier achieved an accuracy of 0.81, recall of 0.74, F1-score of 0.78, and a ROC–AUC of 0.91 on the held-out test set. On this small dataset, XGBoost performed comparably to, but did not outperform, a regularised logistic regression baseline (accuracy 0.91, ROC–AUC 0.95), and McNemar’s test found no statistically significant difference between the models; the wide bootstrap confidence intervals reflect the limited test-set size. The results indicate that, for small structured datasets, a carefully validated and interpretable pipeline—rather than model complexity alone—is the key requirement, and that external validation on larger, multi-institutional datasets is needed before any clinical use.

References
Cover Image
Published
21-08-2026
Section
Articles
License

Copyright (c) 2026 Mohammed K. BASHIR, Sanusi A. DARMA, Usman MAHMUD, Mansir ABUBAKAR (Author)

Creative Commons License

This work is licensed under a Creative Commons Attribution-NonCommercial 4.0 International License.

How to Cite

[1]
M. K. BASHIR, S. A. DARMA, U. MAHMUD, and M. ABUBAKAR, “Enhanced Breast Cancer Classification Using Ensemble Learning and Data Resampling Techniques: A Machine Learning Approach”, FJET, vol. 2, no. 2, pp. 292–302, Aug. 2026, doi: 10.33003/nx2w6h79.

Similar Articles

1-10 of 65

You may also start an advanced similarity search for this article.