Replacement-based soft frequency-regulated resampling (RSFR): a probabilistic adaptive oversampling method for imbalanced classification
JOURNAL OF SUPERCOMPUTING, cilt.82, sa.14, 2026 (SCI-Expanded, Scopus)
- Yayın Türü: Makale / Tam Makale
- Cilt numarası: 82 Sayı: 14
- Basım Tarihi: 2026
- Doi Numarası: 10.1007/s11227-026-08839-1
- Dergi Adı: JOURNAL OF SUPERCOMPUTING
- Derginin Tarandığı İndeksler: Science Citation Index Expanded (SCI-EXPANDED), Scopus, Aerospace Database, Applied Science & Technology Source, Compendex, INSPEC, zbMATH, Academic Search Ultimate (EBSCO), Biomedical Reference Collection: Corporate Edition (EBSCO), Engineering Source (EBSCO), Materials Science & Engineering Collection (ProQuest), Technology Collection (ProQuest)
- Karadeniz Teknik Üniversitesi Adresli: Evet
Özet
This study introduces a probabilistic replacement-based resampling algorithm, called Replacement-based Soft Frequency-regulated Resampling (RSFR), designed to address class imbalance in supervised learning. RSFR adaptively regulates the selection probabilities of minority instances through a soft frequency-regulated replacement mechanism controlled by a penalty smoothness parameter. Unlike interpolation-based oversampling methods, RSFR does not generate synthetic feature vectors; instead, it operates on existing minority samples and penalizes instances that have been repeatedly selected, thereby reducing excessive concentration on a small subset of minority observations. The proposed method was evaluated on 45 benchmark datasets with varying imbalance ratios and dimensional characteristics using multiple classifiers, repeated stratified cross-validation, and six performance metrics. Additional analyses were conducted to examine parameter sensitivity and performance under severe class imbalance. Comparative experiments with established resampling methods showed that RSFR achieved stable and competitive classification performance, particularly for balance-sensitive metrics such as F1, G-Mean, and Matthews correlation coefficient. Statistical analyses indicated that RSFR performs comparably to strong oversampling baselines such as ROS and SMOTE, while providing adaptive control over minority-instance reuse. The experimental design, involving multiple datasets, classifiers, resampling strategies, cross-validation folds, repetitions, and statistical comparisons, constitutes a computationally intensive learning workflow with many independent training and evaluation tasks. This structure provides a direct motivation for parallel and distributed execution, particularly in large-scale benchmarking scenarios and scalable machine learning environments.