Academic Journal

SHAP as a Data Reduction Technique for Highly Imbalanced Big Data.

Λεπτομέρειες βιβλιογραφικής εγγραφής
Τίτλος: SHAP as a Data Reduction Technique for Highly Imbalanced Big Data.
Συγγραφείς: Hancock III, John T., Bauder, Richard A., Khoshgoftaar, Taghi M.
Πηγή: International Journal on Artificial Intelligence Tools; Jun-Aug2025, Vol. 34 Issue 4/5, p1-27, 27p
Θεματικοί όροι: Data reduction, Feature selection, Anomaly detection (Computer security), Outlier detection, Machine learning, Data quality, Big data
Περίληψη: Fraud detection through the classification of highly imbalanced Big Data is an exciting area of Machine Learning research. On the one hand, in certain fraud detection application domains, the use of One-Class classifiers is an overlooked opportunity. On the other hand, for researchers faced with the task of building Machine Learning models for identifying fraud, when only legitimate transaction data is available, One-Class Classifiers are indispensable. We investigate the efficacy of SHapley Additive exPlanations (SHAP) as a feature selection technique for One-Class classification tasks. In this study we utilize authentic data from the Credit Card fraud and Medicare insurance fraud application domains. Our contribution is to show that researchers can use SHAP in conjunction with One-Class Classifiers to do feature selection on highly imbalanced datasets, and then build models, with the selected features, that yield performance similar to, or better than, models built using all features. Our results in Big Medicare data fraud detection show that an over 90% data reduction through feature selection can nevertheless coincide with the best performance in terms of Area under the Precision Recall Curve. [ABSTRACT FROM AUTHOR]
Copyright of International Journal on Artificial Intelligence Tools is the property of World Scientific Publishing Company and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.)
Βάση Δεδομένων: Complementary Index
FullText Text:
  Availability: 0
Header DbId: edb
DbLabel: Complementary Index
An: 188764529
RelevancyScore: 1007
AccessLevel: 6
PubType: Academic Journal
PubTypeId: academicJournal
PreciseRelevancyScore: 1007.33093261719
IllustrationInfo
Items – Name: Title
  Label: Title
  Group: Ti
  Data: SHAP as a Data Reduction Technique for Highly Imbalanced Big Data.
– Name: Author
  Label: Authors
  Group: Au
  Data: <searchLink fieldCode="AR" term="%22Hancock+III%2C+John+T%2E%22">Hancock III, John T.</searchLink><br /><searchLink fieldCode="AR" term="%22Bauder%2C+Richard+A%2E%22">Bauder, Richard A.</searchLink><br /><searchLink fieldCode="AR" term="%22Khoshgoftaar%2C+Taghi+M%2E%22">Khoshgoftaar, Taghi M.</searchLink>
– Name: TitleSource
  Label: Source
  Group: Src
  Data: International Journal on Artificial Intelligence Tools; Jun-Aug2025, Vol. 34 Issue 4/5, p1-27, 27p
– Name: Subject
  Label: Subject Terms
  Group: Su
  Data: <searchLink fieldCode="DE" term="%22Data+reduction%22">Data reduction</searchLink><br /><searchLink fieldCode="DE" term="%22Feature+selection%22">Feature selection</searchLink><br /><searchLink fieldCode="DE" term="%22Anomaly+detection+%28Computer+security%29%22">Anomaly detection (Computer security)</searchLink><br /><searchLink fieldCode="DE" term="%22Outlier+detection%22">Outlier detection</searchLink><br /><searchLink fieldCode="DE" term="%22Machine+learning%22">Machine learning</searchLink><br /><searchLink fieldCode="DE" term="%22Data+quality%22">Data quality</searchLink><br /><searchLink fieldCode="DE" term="%22Big+data%22">Big data</searchLink>
– Name: Abstract
  Label: Abstract
  Group: Ab
  Data: Fraud detection through the classification of highly imbalanced Big Data is an exciting area of Machine Learning research. On the one hand, in certain fraud detection application domains, the use of One-Class classifiers is an overlooked opportunity. On the other hand, for researchers faced with the task of building Machine Learning models for identifying fraud, when only legitimate transaction data is available, One-Class Classifiers are indispensable. We investigate the efficacy of SHapley Additive exPlanations (SHAP) as a feature selection technique for One-Class classification tasks. In this study we utilize authentic data from the Credit Card fraud and Medicare insurance fraud application domains. Our contribution is to show that researchers can use SHAP in conjunction with One-Class Classifiers to do feature selection on highly imbalanced datasets, and then build models, with the selected features, that yield performance similar to, or better than, models built using all features. Our results in Big Medicare data fraud detection show that an over 90% data reduction through feature selection can nevertheless coincide with the best performance in terms of Area under the Precision Recall Curve. [ABSTRACT FROM AUTHOR]
– Name: Abstract
  Label:
  Group: Ab
  Data: <i>Copyright of International Journal on Artificial Intelligence Tools is the property of World Scientific Publishing Company and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract.</i> (Copyright applies to all Abstracts.)
PLink https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=edb&AN=188764529
RecordInfo BibRecord:
  BibEntity:
    Identifiers:
      – Type: doi
        Value: 10.1142/S0218213025400019
    Languages:
      – Code: eng
        Text: English
    PhysicalDescription:
      Pagination:
        PageCount: 27
        StartPage: 1
    Subjects:
      – SubjectFull: Data reduction
        Type: general
      – SubjectFull: Feature selection
        Type: general
      – SubjectFull: Anomaly detection (Computer security)
        Type: general
      – SubjectFull: Outlier detection
        Type: general
      – SubjectFull: Machine learning
        Type: general
      – SubjectFull: Data quality
        Type: general
      – SubjectFull: Big data
        Type: general
    Titles:
      – TitleFull: SHAP as a Data Reduction Technique for Highly Imbalanced Big Data.
        Type: main
  BibRelationships:
    HasContributorRelationships:
      – PersonEntity:
          Name:
            NameFull: Hancock III, John T.
      – PersonEntity:
          Name:
            NameFull: Bauder, Richard A.
      – PersonEntity:
          Name:
            NameFull: Khoshgoftaar, Taghi M.
    IsPartOfRelationships:
      – BibEntity:
          Dates:
            – D: 01
              M: 06
              Text: Jun-Aug2025
              Type: published
              Y: 2025
          Identifiers:
            – Type: issn-print
              Value: 02182130
          Numbering:
            – Type: volume
              Value: 34
            – Type: issue
              Value: 4/5
          Titles:
            – TitleFull: International Journal on Artificial Intelligence Tools
              Type: main
ResultId 1