Academic Journal
SHAP as a Data Reduction Technique for Highly Imbalanced Big Data.
| Τίτλος: | SHAP as a Data Reduction Technique for Highly Imbalanced Big Data. |
|---|---|
| Συγγραφείς: | Hancock III, John T., Bauder, Richard A., Khoshgoftaar, Taghi M. |
| Πηγή: | International Journal on Artificial Intelligence Tools; Jun-Aug2025, Vol. 34 Issue 4/5, p1-27, 27p |
| Θεματικοί όροι: | Data reduction, Feature selection, Anomaly detection (Computer security), Outlier detection, Machine learning, Data quality, Big data |
| Περίληψη: | Fraud detection through the classification of highly imbalanced Big Data is an exciting area of Machine Learning research. On the one hand, in certain fraud detection application domains, the use of One-Class classifiers is an overlooked opportunity. On the other hand, for researchers faced with the task of building Machine Learning models for identifying fraud, when only legitimate transaction data is available, One-Class Classifiers are indispensable. We investigate the efficacy of SHapley Additive exPlanations (SHAP) as a feature selection technique for One-Class classification tasks. In this study we utilize authentic data from the Credit Card fraud and Medicare insurance fraud application domains. Our contribution is to show that researchers can use SHAP in conjunction with One-Class Classifiers to do feature selection on highly imbalanced datasets, and then build models, with the selected features, that yield performance similar to, or better than, models built using all features. Our results in Big Medicare data fraud detection show that an over 90% data reduction through feature selection can nevertheless coincide with the best performance in terms of Area under the Precision Recall Curve. [ABSTRACT FROM AUTHOR] |
| Copyright of International Journal on Artificial Intelligence Tools is the property of World Scientific Publishing Company and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.) | |
| Βάση Δεδομένων: | Complementary Index |
| FullText | Text: Availability: 0 |
|---|---|
| Header | DbId: edb DbLabel: Complementary Index An: 188764529 RelevancyScore: 1007 AccessLevel: 6 PubType: Academic Journal PubTypeId: academicJournal PreciseRelevancyScore: 1007.33093261719 |
| IllustrationInfo | |
| Items | – Name: Title Label: Title Group: Ti Data: SHAP as a Data Reduction Technique for Highly Imbalanced Big Data. – Name: Author Label: Authors Group: Au Data: <searchLink fieldCode="AR" term="%22Hancock+III%2C+John+T%2E%22">Hancock III, John T.</searchLink><br /><searchLink fieldCode="AR" term="%22Bauder%2C+Richard+A%2E%22">Bauder, Richard A.</searchLink><br /><searchLink fieldCode="AR" term="%22Khoshgoftaar%2C+Taghi+M%2E%22">Khoshgoftaar, Taghi M.</searchLink> – Name: TitleSource Label: Source Group: Src Data: International Journal on Artificial Intelligence Tools; Jun-Aug2025, Vol. 34 Issue 4/5, p1-27, 27p – Name: Subject Label: Subject Terms Group: Su Data: <searchLink fieldCode="DE" term="%22Data+reduction%22">Data reduction</searchLink><br /><searchLink fieldCode="DE" term="%22Feature+selection%22">Feature selection</searchLink><br /><searchLink fieldCode="DE" term="%22Anomaly+detection+%28Computer+security%29%22">Anomaly detection (Computer security)</searchLink><br /><searchLink fieldCode="DE" term="%22Outlier+detection%22">Outlier detection</searchLink><br /><searchLink fieldCode="DE" term="%22Machine+learning%22">Machine learning</searchLink><br /><searchLink fieldCode="DE" term="%22Data+quality%22">Data quality</searchLink><br /><searchLink fieldCode="DE" term="%22Big+data%22">Big data</searchLink> – Name: Abstract Label: Abstract Group: Ab Data: Fraud detection through the classification of highly imbalanced Big Data is an exciting area of Machine Learning research. On the one hand, in certain fraud detection application domains, the use of One-Class classifiers is an overlooked opportunity. On the other hand, for researchers faced with the task of building Machine Learning models for identifying fraud, when only legitimate transaction data is available, One-Class Classifiers are indispensable. We investigate the efficacy of SHapley Additive exPlanations (SHAP) as a feature selection technique for One-Class classification tasks. In this study we utilize authentic data from the Credit Card fraud and Medicare insurance fraud application domains. Our contribution is to show that researchers can use SHAP in conjunction with One-Class Classifiers to do feature selection on highly imbalanced datasets, and then build models, with the selected features, that yield performance similar to, or better than, models built using all features. Our results in Big Medicare data fraud detection show that an over 90% data reduction through feature selection can nevertheless coincide with the best performance in terms of Area under the Precision Recall Curve. [ABSTRACT FROM AUTHOR] – Name: Abstract Label: Group: Ab Data: <i>Copyright of International Journal on Artificial Intelligence Tools is the property of World Scientific Publishing Company and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract.</i> (Copyright applies to all Abstracts.) |
| PLink | https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=edb&AN=188764529 |
| RecordInfo | BibRecord: BibEntity: Identifiers: – Type: doi Value: 10.1142/S0218213025400019 Languages: – Code: eng Text: English PhysicalDescription: Pagination: PageCount: 27 StartPage: 1 Subjects: – SubjectFull: Data reduction Type: general – SubjectFull: Feature selection Type: general – SubjectFull: Anomaly detection (Computer security) Type: general – SubjectFull: Outlier detection Type: general – SubjectFull: Machine learning Type: general – SubjectFull: Data quality Type: general – SubjectFull: Big data Type: general Titles: – TitleFull: SHAP as a Data Reduction Technique for Highly Imbalanced Big Data. Type: main BibRelationships: HasContributorRelationships: – PersonEntity: Name: NameFull: Hancock III, John T. – PersonEntity: Name: NameFull: Bauder, Richard A. – PersonEntity: Name: NameFull: Khoshgoftaar, Taghi M. IsPartOfRelationships: – BibEntity: Dates: – D: 01 M: 06 Text: Jun-Aug2025 Type: published Y: 2025 Identifiers: – Type: issn-print Value: 02182130 Numbering: – Type: volume Value: 34 – Type: issue Value: 4/5 Titles: – TitleFull: International Journal on Artificial Intelligence Tools Type: main |
| ResultId | 1 |