Academic Journal

DENGESİZ VERİ SETLERİ İÇİN İKİ AŞAMALI DENGELEME STRATEJİSİ: ADASYN İLE ÖRNEKLEM ARTIRMA, SVM TABANLI ÖRNEKLEM AZALTMA.

Bibliographic Details
Title: DENGESİZ VERİ SETLERİ İÇİN İKİ AŞAMALI DENGELEME STRATEJİSİ: ADASYN İLE ÖRNEKLEM ARTIRMA, SVM TABANLI ÖRNEKLEM AZALTMA. (Turkish)
Alternate Title: A Two-Stage Balancing Strategy for Imbalanced Datasets: Oversampling with ADASYN and Undersampling Based on SVM. (English)
Authors: YILMAZ EROĞLU, Duygu
Source: Uludag University Journal of the Faculty of Engineering (UUJFE); 2025, Vol. 30 Issue 3, p825-843, 19p
Subject Terms: Machine learning, Support vector machines, Data quality, Data augmentation, Data reduction
Abstract (English): This study tackles the frequently encountered problem of imbalanced datasets in machine learning, focusing on cases where minority-class examples are overshadowed by the majority class. Such imbalance significantly undermines model performance in diverse fields, including healthcare, fraud detection, and IoT-based industrial processes. To address this issue, we combine the ADASYN method— enriching the minority class with synthetic samples—with the removal of the "most distant" 10% of majority-class instances identified via an SVM-based distance measure. The proposed approach is tested with SVM, RF, XGBoost, and KNN classifiers on ten different datasets. Among these is the "Textile" dataset, which includes both quality control data and IoT sensor measurements and was collected from a real world production environment. Notably, this dataset includes rare yet critical events such as yarn breakage, which standard methods fail to detect effectively due to pronounced class imbalance. Our approach achieves considerable enhancements in the G-Mean metric, thereby improving the detection of minority cases and securing the highest G-Mean values on five out of ten datasets. [ABSTRACT FROM AUTHOR]
Abstract (Turkish): Bu çalışma, makine öğrenmesi alanında sıkça karşılaşılan dengesiz veri sorununu ele alarak, azınlık sınıf örneklerinin çoğunluk sınıf tarafından gölgede bırakıldığı durumlara odaklanmaktadır. Böyle bir dengesizlik, sağlık hizmetlerinden finansal sahtekârlık tespitine ve IoT tabanlı endüstriyel süreçlere kadar pek çok alanda model performansını ciddi biçimde zayıflatır. Sorunu gidermek için, azınlık sınıfını sentetik örneklerle zenginleştiren ADASYN yöntemi, SVM tabanlı uzaklık ölçümüyle belirlenen "en uzak" %10’luk çoğunluk örneklerinin çıkarılmasıyla birleştirilmiştir. Önerilen yaklaşım, SVM, RF, XGBoost ve KNN sınıflandırıcılarıyla on farklı veri seti üzerinde test edilmiştir. Bunlar arasında hem kalite kontrol verilerini hem de IoT sensör ölçümlerini içeren, gerçek üretim ortamından elde edilmiş 'Tekstil' veri seti de yer almaktadır. Özellikle iplik kopması gibi nadir ancak üretim açısından kritik olayları barındıran bu veri seti, yoğun dengesizlik nedeniyle standart yöntemlerde düşük başarı sergilemektedir. G-Ortalamalar metriğinde önemli iyileşmeler sunan yöntem, azınlık sınıfın daha başarılı tespitine katkıda bulunmuş ve on veri setinden beşinde en yüksek G-Ortalamalar değerini elde etmiştir. [ABSTRACT FROM AUTHOR]
Copyright of Uludag University Journal of the Faculty of Engineering (UUJFE) is the property of Uludag Universitesi, Muhendislik Fakultesi and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.)
Database: Complementary Index
FullText Text:
  Availability: 0
Header DbId: edb
DbLabel: Complementary Index
An: 190595731
RelevancyScore: 1023
AccessLevel: 6
PubType: Academic Journal
PubTypeId: academicJournal
PreciseRelevancyScore: 1023.08752441406
IllustrationInfo
Items – Name: Title
  Label: Title
  Group: Ti
  Data: DENGESİZ VERİ SETLERİ İÇİN İKİ AŞAMALI DENGELEME STRATEJİSİ: ADASYN İLE ÖRNEKLEM ARTIRMA, SVM TABANLI ÖRNEKLEM AZALTMA. (Turkish)
– Name: TitleAlt
  Label: Alternate Title
  Group: TiAlt
  Data: A Two-Stage Balancing Strategy for Imbalanced Datasets: Oversampling with ADASYN and Undersampling Based on SVM. (English)
– Name: Author
  Label: Authors
  Group: Au
  Data: <searchLink fieldCode="AR" term="%22YILMAZ+EROĞLU%2C+Duygu%22">YILMAZ EROĞLU, Duygu</searchLink>
– Name: TitleSource
  Label: Source
  Group: Src
  Data: Uludag University Journal of the Faculty of Engineering (UUJFE); 2025, Vol. 30 Issue 3, p825-843, 19p
– Name: Subject
  Label: Subject Terms
  Group: Su
  Data: <searchLink fieldCode="DE" term="%22Machine+learning%22">Machine learning</searchLink><br /><searchLink fieldCode="DE" term="%22Support+vector+machines%22">Support vector machines</searchLink><br /><searchLink fieldCode="DE" term="%22Data+quality%22">Data quality</searchLink><br /><searchLink fieldCode="DE" term="%22Data+augmentation%22">Data augmentation</searchLink><br /><searchLink fieldCode="DE" term="%22Data+reduction%22">Data reduction</searchLink>
– Name: AbstractNonEng
  Label: Abstract (English)
  Group: Ab
  Data: This study tackles the frequently encountered problem of imbalanced datasets in machine learning, focusing on cases where minority-class examples are overshadowed by the majority class. Such imbalance significantly undermines model performance in diverse fields, including healthcare, fraud detection, and IoT-based industrial processes. To address this issue, we combine the ADASYN method— enriching the minority class with synthetic samples—with the removal of the "most distant" 10% of majority-class instances identified via an SVM-based distance measure. The proposed approach is tested with SVM, RF, XGBoost, and KNN classifiers on ten different datasets. Among these is the "Textile" dataset, which includes both quality control data and IoT sensor measurements and was collected from a real world production environment. Notably, this dataset includes rare yet critical events such as yarn breakage, which standard methods fail to detect effectively due to pronounced class imbalance. Our approach achieves considerable enhancements in the G-Mean metric, thereby improving the detection of minority cases and securing the highest G-Mean values on five out of ten datasets. [ABSTRACT FROM AUTHOR]
– Name: AbstractNonEng
  Label: Abstract (Turkish)
  Group: Ab
  Data: Bu çalışma, makine öğrenmesi alanında sıkça karşılaşılan dengesiz veri sorununu ele alarak, azınlık sınıf örneklerinin çoğunluk sınıf tarafından gölgede bırakıldığı durumlara odaklanmaktadır. Böyle bir dengesizlik, sağlık hizmetlerinden finansal sahtekârlık tespitine ve IoT tabanlı endüstriyel süreçlere kadar pek çok alanda model performansını ciddi biçimde zayıflatır. Sorunu gidermek için, azınlık sınıfını sentetik örneklerle zenginleştiren ADASYN yöntemi, SVM tabanlı uzaklık ölçümüyle belirlenen "en uzak" %10’luk çoğunluk örneklerinin çıkarılmasıyla birleştirilmiştir. Önerilen yaklaşım, SVM, RF, XGBoost ve KNN sınıflandırıcılarıyla on farklı veri seti üzerinde test edilmiştir. Bunlar arasında hem kalite kontrol verilerini hem de IoT sensör ölçümlerini içeren, gerçek üretim ortamından elde edilmiş 'Tekstil' veri seti de yer almaktadır. Özellikle iplik kopması gibi nadir ancak üretim açısından kritik olayları barındıran bu veri seti, yoğun dengesizlik nedeniyle standart yöntemlerde düşük başarı sergilemektedir. G-Ortalamalar metriğinde önemli iyileşmeler sunan yöntem, azınlık sınıfın daha başarılı tespitine katkıda bulunmuş ve on veri setinden beşinde en yüksek G-Ortalamalar değerini elde etmiştir. [ABSTRACT FROM AUTHOR]
– Name: Abstract
  Label:
  Group: Ab
  Data: <i>Copyright of Uludag University Journal of the Faculty of Engineering (UUJFE) is the property of Uludag Universitesi, Muhendislik Fakultesi and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract.</i> (Copyright applies to all Abstracts.)
PLink https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=edb&AN=190595731
RecordInfo BibRecord:
  BibEntity:
    Identifiers:
      – Type: doi
        Value: 10.17482/uumfd.1722270
    Languages:
      – Code: tur
        Text: Turkish
    PhysicalDescription:
      Pagination:
        PageCount: 19
        StartPage: 825
    Subjects:
      – SubjectFull: Machine learning
        Type: general
      – SubjectFull: Support vector machines
        Type: general
      – SubjectFull: Data quality
        Type: general
      – SubjectFull: Data augmentation
        Type: general
      – SubjectFull: Data reduction
        Type: general
    Titles:
      – TitleFull: DENGESİZ VERİ SETLERİ İÇİN İKİ AŞAMALI DENGELEME STRATEJİSİ: ADASYN İLE ÖRNEKLEM ARTIRMA, SVM TABANLI ÖRNEKLEM AZALTMA.
        Type: main
  BibRelationships:
    HasContributorRelationships:
      – PersonEntity:
          Name:
            NameFull: YILMAZ EROĞLU, Duygu
    IsPartOfRelationships:
      – BibEntity:
          Dates:
            – D: 01
              M: 09
              Text: 2025
              Type: published
              Y: 2025
          Identifiers:
            – Type: issn-print
              Value: 21484147
          Numbering:
            – Type: volume
              Value: 30
            – Type: issue
              Value: 3
          Titles:
            – TitleFull: Uludag University Journal of the Faculty of Engineering (UUJFE)
              Type: main
ResultId 1