Λεπτομέρειες βιβλιογραφικής εγγραφής
| Τίτλος: |
Data cleaning and spam filtering methods in the TAWOS database. |
| Συγγραφείς: |
Szabó, Márk1 szabo.mark@uni-eszterhazy.hu, Kovács, Ádám1,2 kovacs2.adam@uni-eszterhazy.hu, Kusper, Gábor1 kusper.gabor@uni-eszterhazy.hu |
| Πηγή: |
Annales Mathematicae et Informaticae. 2026, Vol. 63, p132-148. 17p. |
| Θεματικοί όροι: |
*Spam filtering (Email), *Defect tracking (Computer software development), *Classification, *Relational databases, *Machine learning, *Data scrubbing, *Software engineering |
| Περίληψη: |
The TAWOS dataset is a large relational database of issue-tracking data mined from open-source projects. While it is valuable for empirical software engineering, the issue texts also contain spam, placeholders, and off-topic noise that can distort downstream analytics and fine-tuning tasks. We present a four-stage filtering pipeline for cleaning TAWOS and focus on a validity scoring step that assigns each issue a 0-100 score from its title and description. We compare a deterministic rule-based classifier, OwnMetrics, against four local small language models (Llama-3.1-8B, Mistral-7B, Phi-3.5-mini, and Gemma-3-4B) prompted for JSON scores. Evaluation first uses a nearbalanced labeled benchmark of 947 GitHub issues (481 spam/noise and 466 legitimate issues), collected from moderator-locked spam issues and legitimate issues from popular repositories. On this GitHub-based threshold-selection benchmark, OwnMetrics obtains the highest accuracy, 91.0%, at threshold 75, while the best LLM configurations reach 89.2% (Gemma-3-4B) and 89.1% (Mistral-7B). A separate manually labeled in-domain validation on 300 Jira/TAWOS issues (39 spam/noise and 261 non-spam) provides an in-domain sanity check consistent with the benchmark results, with OwnMetrics giving the strongest spam-class F1 among the selected configurations. On a 10,000-item sample, OwnMetrics yields zero parsing failures and an average execution time of 2.28 ms per item, whereas the LLMs require seconds per item and Meta-Llama-3.1-8B exhibits failure rates above 54%. Applying the full filtering pipeline to the processed TAWOS issue corpus flags 16.2% of records for filtering. The results show that a lightweight domain-tailored scorer can be both accurate and robust for large-scale issue cleaning. [ABSTRACT FROM AUTHOR] |
| Βάση Δεδομένων: |
Academic Search Index |