Academic Journal

Data cleaning and spam filtering methods in the TAWOS database.

Λεπτομέρειες βιβλιογραφικής εγγραφής
Τίτλος: Data cleaning and spam filtering methods in the TAWOS database.
Συγγραφείς: Szabó, Márk1 szabo.mark@uni-eszterhazy.hu, Kovács, Ádám1,2 kovacs2.adam@uni-eszterhazy.hu, Kusper, Gábor1 kusper.gabor@uni-eszterhazy.hu
Πηγή: Annales Mathematicae et Informaticae. 2026, Vol. 63, p132-148. 17p.
Θεματικοί όροι: *Spam filtering (Email), *Defect tracking (Computer software development), *Classification, *Relational databases, *Machine learning, *Data scrubbing, *Software engineering
Περίληψη: The TAWOS dataset is a large relational database of issue-tracking data mined from open-source projects. While it is valuable for empirical software engineering, the issue texts also contain spam, placeholders, and off-topic noise that can distort downstream analytics and fine-tuning tasks. We present a four-stage filtering pipeline for cleaning TAWOS and focus on a validity scoring step that assigns each issue a 0-100 score from its title and description. We compare a deterministic rule-based classifier, OwnMetrics, against four local small language models (Llama-3.1-8B, Mistral-7B, Phi-3.5-mini, and Gemma-3-4B) prompted for JSON scores. Evaluation first uses a nearbalanced labeled benchmark of 947 GitHub issues (481 spam/noise and 466 legitimate issues), collected from moderator-locked spam issues and legitimate issues from popular repositories. On this GitHub-based threshold-selection benchmark, OwnMetrics obtains the highest accuracy, 91.0%, at threshold 75, while the best LLM configurations reach 89.2% (Gemma-3-4B) and 89.1% (Mistral-7B). A separate manually labeled in-domain validation on 300 Jira/TAWOS issues (39 spam/noise and 261 non-spam) provides an in-domain sanity check consistent with the benchmark results, with OwnMetrics giving the strongest spam-class F1 among the selected configurations. On a 10,000-item sample, OwnMetrics yields zero parsing failures and an average execution time of 2.28 ms per item, whereas the LLMs require seconds per item and Meta-Llama-3.1-8B exhibits failure rates above 54%. Applying the full filtering pipeline to the processed TAWOS issue corpus flags 16.2% of records for filtering. The results show that a lightweight domain-tailored scorer can be both accurate and robust for large-scale issue cleaning. [ABSTRACT FROM AUTHOR]
Βάση Δεδομένων: Academic Search Index
Περιγραφή
ISSN:17875021
DOI:10.33039/ami.2026.06.004