N-gram Feature Selection for Authorship Identification

Automatic authorship identification offers a valuable tool for supporting crime investigation and security. It can be seen as a multi-class, single-label text categorization task. Automatic authorship identification depends on selecting stylisticfeatures that would capture an authors writing style i...

Πλήρης περιγραφή

Αποθηκεύτηκε σε:
Λεπτομέρειες βιβλιογραφικής εγγραφής
Κύριος συγγραφέας: Χουβαρδάς, Ιωάννης - Γεώργιος
Άλλοι συγγραφείς: Σταματάτος, Ευστάθιος
Γλώσσα:English
Δημοσίευση: 2015
Θέματα:
Διαθέσιμο Online:https://vsmart.lib.aegean.gr/webopac/FullBB.csp?WebAction=ShowFullBB&EncodedRequest=*AAmk*3D*D3w*10*C9*89*84*5D*5DJ*C9*197&Profile=Default&OpacLanguage=gre&NumberToRetrieve=50&StartValue=1&WebPageNr=1&SearchTerm1=2006%20.1.44709&SearchT1=&Index1=Keywordsbib&SearchMethod=Find_1&ItemNr=1
http://hdl.handle.net/11610/12497
Ετικέτες: Προσθήκη ετικέτας
Δεν υπάρχουν, Καταχωρήστε ετικέτα πρώτοι!
_version_ 1828460093385998336
author Χουβαρδάς, Ιωάννης - Γεώργιος
author2 Σταματάτος, Ευστάθιος
author_facet Σταματάτος, Ευστάθιος
Χουβαρδάς, Ιωάννης - Γεώργιος
author_sort Χουβαρδάς, Ιωάννης - Γεώργιος
collection DSpace
description Automatic authorship identification offers a valuable tool for supporting crime investigation and security. It can be seen as a multi-class, single-label text categorization task. Automatic authorship identification depends on selecting stylisticfeatures that would capture an authors writing style independent of the content or genre of text. Character n-grams have been used successfully to represent text for stylistic purposes in literature.They seem to be able to capture nuances in lexical, syntactical, and structural level. To date character n-grams of fixed length have been used for authorship identification. In this thesis: we introduce a new approach for selecting variable length n-grams inspired by previous work for selecting variable-length word sequences. We propose the use of variable-length n-grams to represent the stylistic information of the documents to be classified. We explore the significance of digits as stylistic features for distinguishing between authors and show that an increase in performance can be achieved using simple text pre-processing. Using a subset of the new Reuters corpus, consisting of texts on thesame topic by 50 different authors, we show that the proposed feature selection method is at least as effective as information gain for selecting the most significant n-grams, although the feature sets produced by the two methods have few common members.
id oai:hellanicus.lib.aegean.gr:11610-12497
institution Hellanicus
language English
publishDate 2015
record_format dspace
spelling oai:hellanicus.lib.aegean.gr:11610-124972022-02-09T11:13:15Z N-gram Feature Selection for Authorship Identification Χουβαρδάς, Ιωάννης - Γεώργιος Σταματάτος, Ευστάθιος Επιλογή Χαρακτηριστικών Αναγνώριση συγγραφέα Feature Selection Authorship identification Integrated software Automatic authorship identification offers a valuable tool for supporting crime investigation and security. It can be seen as a multi-class, single-label text categorization task. Automatic authorship identification depends on selecting stylisticfeatures that would capture an authors writing style independent of the content or genre of text. Character n-grams have been used successfully to represent text for stylistic purposes in literature.They seem to be able to capture nuances in lexical, syntactical, and structural level. To date character n-grams of fixed length have been used for authorship identification. In this thesis: we introduce a new approach for selecting variable length n-grams inspired by previous work for selecting variable-length word sequences. We propose the use of variable-length n-grams to represent the stylistic information of the documents to be classified. We explore the significance of digits as stylistic features for distinguishing between authors and show that an increase in performance can be achieved using simple text pre-processing. Using a subset of the new Reuters corpus, consisting of texts on thesame topic by 50 different authors, we show that the proposed feature selection method is at least as effective as information gain for selecting the most significant n-grams, although the feature sets produced by the two methods have few common members. 2015-11-18T10:39:45Z 2015-11-18T10:39:45Z 2006 https://vsmart.lib.aegean.gr/webopac/FullBB.csp?WebAction=ShowFullBB&EncodedRequest=*AAmk*3D*D3w*10*C9*89*84*5D*5DJ*C9*197&Profile=Default&OpacLanguage=gre&NumberToRetrieve=50&StartValue=1&WebPageNr=1&SearchTerm1=2006%20.1.44709&SearchT1=&Index1=Keywordsbib&SearchMethod=Find_1&ItemNr=1 http://hdl.handle.net/11610/12497 en Σάμος
spellingShingle Επιλογή Χαρακτηριστικών
Αναγνώριση συγγραφέα
Feature Selection
Authorship identification
Integrated software
Χουβαρδάς, Ιωάννης - Γεώργιος
N-gram Feature Selection for Authorship Identification
title N-gram Feature Selection for Authorship Identification
title_full N-gram Feature Selection for Authorship Identification
title_fullStr N-gram Feature Selection for Authorship Identification
title_full_unstemmed N-gram Feature Selection for Authorship Identification
title_short N-gram Feature Selection for Authorship Identification
title_sort n gram feature selection for authorship identification
topic Επιλογή Χαρακτηριστικών
Αναγνώριση συγγραφέα
Feature Selection
Authorship identification
Integrated software
url https://vsmart.lib.aegean.gr/webopac/FullBB.csp?WebAction=ShowFullBB&EncodedRequest=*AAmk*3D*D3w*10*C9*89*84*5D*5DJ*C9*197&Profile=Default&OpacLanguage=gre&NumberToRetrieve=50&StartValue=1&WebPageNr=1&SearchTerm1=2006%20.1.44709&SearchT1=&Index1=Keywordsbib&SearchMethod=Find_1&ItemNr=1
http://hdl.handle.net/11610/12497
work_keys_str_mv AT choubardasiōannēsgeōrgios ngramfeatureselectionforauthorshipidentification