Academic Journal
DENGESİZ VERİ SETLERİ İÇİN İKİ AŞAMALI DENGELEME STRATEJİSİ: ADASYN İLE ÖRNEKLEM ARTIRMA, SVM TABANLI ÖRNEKLEM AZALTMA.
| Title: | DENGESİZ VERİ SETLERİ İÇİN İKİ AŞAMALI DENGELEME STRATEJİSİ: ADASYN İLE ÖRNEKLEM ARTIRMA, SVM TABANLI ÖRNEKLEM AZALTMA. (Turkish) |
|---|---|
| Alternate Title: | A Two-Stage Balancing Strategy for Imbalanced Datasets: Oversampling with ADASYN and Undersampling Based on SVM. (English) |
| Authors: | YILMAZ EROĞLU, Duygu |
| Source: | Uludag University Journal of the Faculty of Engineering (UUJFE); 2025, Vol. 30 Issue 3, p825-843, 19p |
| Subject Terms: | Machine learning, Support vector machines, Data quality, Data augmentation, Data reduction |
| Abstract (English): | This study tackles the frequently encountered problem of imbalanced datasets in machine learning, focusing on cases where minority-class examples are overshadowed by the majority class. Such imbalance significantly undermines model performance in diverse fields, including healthcare, fraud detection, and IoT-based industrial processes. To address this issue, we combine the ADASYN method— enriching the minority class with synthetic samples—with the removal of the "most distant" 10% of majority-class instances identified via an SVM-based distance measure. The proposed approach is tested with SVM, RF, XGBoost, and KNN classifiers on ten different datasets. Among these is the "Textile" dataset, which includes both quality control data and IoT sensor measurements and was collected from a real world production environment. Notably, this dataset includes rare yet critical events such as yarn breakage, which standard methods fail to detect effectively due to pronounced class imbalance. Our approach achieves considerable enhancements in the G-Mean metric, thereby improving the detection of minority cases and securing the highest G-Mean values on five out of ten datasets. [ABSTRACT FROM AUTHOR] |
| Abstract (Turkish): | Bu çalışma, makine öğrenmesi alanında sıkça karşılaşılan dengesiz veri sorununu ele alarak, azınlık sınıf örneklerinin çoğunluk sınıf tarafından gölgede bırakıldığı durumlara odaklanmaktadır. Böyle bir dengesizlik, sağlık hizmetlerinden finansal sahtekârlık tespitine ve IoT tabanlı endüstriyel süreçlere kadar pek çok alanda model performansını ciddi biçimde zayıflatır. Sorunu gidermek için, azınlık sınıfını sentetik örneklerle zenginleştiren ADASYN yöntemi, SVM tabanlı uzaklık ölçümüyle belirlenen "en uzak" %10’luk çoğunluk örneklerinin çıkarılmasıyla birleştirilmiştir. Önerilen yaklaşım, SVM, RF, XGBoost ve KNN sınıflandırıcılarıyla on farklı veri seti üzerinde test edilmiştir. Bunlar arasında hem kalite kontrol verilerini hem de IoT sensör ölçümlerini içeren, gerçek üretim ortamından elde edilmiş 'Tekstil' veri seti de yer almaktadır. Özellikle iplik kopması gibi nadir ancak üretim açısından kritik olayları barındıran bu veri seti, yoğun dengesizlik nedeniyle standart yöntemlerde düşük başarı sergilemektedir. G-Ortalamalar metriğinde önemli iyileşmeler sunan yöntem, azınlık sınıfın daha başarılı tespitine katkıda bulunmuş ve on veri setinden beşinde en yüksek G-Ortalamalar değerini elde etmiştir. [ABSTRACT FROM AUTHOR] |
| Copyright of Uludag University Journal of the Faculty of Engineering (UUJFE) is the property of Uludag Universitesi, Muhendislik Fakultesi and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.) | |
| Database: | Complementary Index |
| FullText | Text: Availability: 0 |
|---|---|
| Header | DbId: edb DbLabel: Complementary Index An: 190595731 RelevancyScore: 1023 AccessLevel: 6 PubType: Academic Journal PubTypeId: academicJournal PreciseRelevancyScore: 1023.08752441406 |
| IllustrationInfo | |
| Items | – Name: Title Label: Title Group: Ti Data: DENGESİZ VERİ SETLERİ İÇİN İKİ AŞAMALI DENGELEME STRATEJİSİ: ADASYN İLE ÖRNEKLEM ARTIRMA, SVM TABANLI ÖRNEKLEM AZALTMA. (Turkish) – Name: TitleAlt Label: Alternate Title Group: TiAlt Data: A Two-Stage Balancing Strategy for Imbalanced Datasets: Oversampling with ADASYN and Undersampling Based on SVM. (English) – Name: Author Label: Authors Group: Au Data: <searchLink fieldCode="AR" term="%22YILMAZ+EROĞLU%2C+Duygu%22">YILMAZ EROĞLU, Duygu</searchLink> – Name: TitleSource Label: Source Group: Src Data: Uludag University Journal of the Faculty of Engineering (UUJFE); 2025, Vol. 30 Issue 3, p825-843, 19p – Name: Subject Label: Subject Terms Group: Su Data: <searchLink fieldCode="DE" term="%22Machine+learning%22">Machine learning</searchLink><br /><searchLink fieldCode="DE" term="%22Support+vector+machines%22">Support vector machines</searchLink><br /><searchLink fieldCode="DE" term="%22Data+quality%22">Data quality</searchLink><br /><searchLink fieldCode="DE" term="%22Data+augmentation%22">Data augmentation</searchLink><br /><searchLink fieldCode="DE" term="%22Data+reduction%22">Data reduction</searchLink> – Name: AbstractNonEng Label: Abstract (English) Group: Ab Data: This study tackles the frequently encountered problem of imbalanced datasets in machine learning, focusing on cases where minority-class examples are overshadowed by the majority class. Such imbalance significantly undermines model performance in diverse fields, including healthcare, fraud detection, and IoT-based industrial processes. To address this issue, we combine the ADASYN method— enriching the minority class with synthetic samples—with the removal of the "most distant" 10% of majority-class instances identified via an SVM-based distance measure. The proposed approach is tested with SVM, RF, XGBoost, and KNN classifiers on ten different datasets. Among these is the "Textile" dataset, which includes both quality control data and IoT sensor measurements and was collected from a real world production environment. Notably, this dataset includes rare yet critical events such as yarn breakage, which standard methods fail to detect effectively due to pronounced class imbalance. Our approach achieves considerable enhancements in the G-Mean metric, thereby improving the detection of minority cases and securing the highest G-Mean values on five out of ten datasets. [ABSTRACT FROM AUTHOR] – Name: AbstractNonEng Label: Abstract (Turkish) Group: Ab Data: Bu çalışma, makine öğrenmesi alanında sıkça karşılaşılan dengesiz veri sorununu ele alarak, azınlık sınıf örneklerinin çoğunluk sınıf tarafından gölgede bırakıldığı durumlara odaklanmaktadır. Böyle bir dengesizlik, sağlık hizmetlerinden finansal sahtekârlık tespitine ve IoT tabanlı endüstriyel süreçlere kadar pek çok alanda model performansını ciddi biçimde zayıflatır. Sorunu gidermek için, azınlık sınıfını sentetik örneklerle zenginleştiren ADASYN yöntemi, SVM tabanlı uzaklık ölçümüyle belirlenen "en uzak" %10’luk çoğunluk örneklerinin çıkarılmasıyla birleştirilmiştir. Önerilen yaklaşım, SVM, RF, XGBoost ve KNN sınıflandırıcılarıyla on farklı veri seti üzerinde test edilmiştir. Bunlar arasında hem kalite kontrol verilerini hem de IoT sensör ölçümlerini içeren, gerçek üretim ortamından elde edilmiş 'Tekstil' veri seti de yer almaktadır. Özellikle iplik kopması gibi nadir ancak üretim açısından kritik olayları barındıran bu veri seti, yoğun dengesizlik nedeniyle standart yöntemlerde düşük başarı sergilemektedir. G-Ortalamalar metriğinde önemli iyileşmeler sunan yöntem, azınlık sınıfın daha başarılı tespitine katkıda bulunmuş ve on veri setinden beşinde en yüksek G-Ortalamalar değerini elde etmiştir. [ABSTRACT FROM AUTHOR] – Name: Abstract Label: Group: Ab Data: <i>Copyright of Uludag University Journal of the Faculty of Engineering (UUJFE) is the property of Uludag Universitesi, Muhendislik Fakultesi and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract.</i> (Copyright applies to all Abstracts.) |
| PLink | https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=edb&AN=190595731 |
| RecordInfo | BibRecord: BibEntity: Identifiers: – Type: doi Value: 10.17482/uumfd.1722270 Languages: – Code: tur Text: Turkish PhysicalDescription: Pagination: PageCount: 19 StartPage: 825 Subjects: – SubjectFull: Machine learning Type: general – SubjectFull: Support vector machines Type: general – SubjectFull: Data quality Type: general – SubjectFull: Data augmentation Type: general – SubjectFull: Data reduction Type: general Titles: – TitleFull: DENGESİZ VERİ SETLERİ İÇİN İKİ AŞAMALI DENGELEME STRATEJİSİ: ADASYN İLE ÖRNEKLEM ARTIRMA, SVM TABANLI ÖRNEKLEM AZALTMA. Type: main BibRelationships: HasContributorRelationships: – PersonEntity: Name: NameFull: YILMAZ EROĞLU, Duygu IsPartOfRelationships: – BibEntity: Dates: – D: 01 M: 09 Text: 2025 Type: published Y: 2025 Identifiers: – Type: issn-print Value: 21484147 Numbering: – Type: volume Value: 30 – Type: issue Value: 3 Titles: – TitleFull: Uludag University Journal of the Faculty of Engineering (UUJFE) Type: main |
| ResultId | 1 |