Dissertation/ Thesis

Feature engineering for author profiling and identification: on the relevance of syntax and discourse

Bibliographic Details
Title: Feature engineering for author profiling and identification: on the relevance of syntax and discourse
Authors: Soler Company, Juan
Contributors: University/Department: Universitat Pompeu Fabra. Departament de Tecnologies de la Informació i les Comunicacions
Thesis Advisors: Wanner, Leo
Source: TDX (Tesis Doctorals en Xarxa)
Publisher Information: Universitat Pompeu Fabra, 2017.
Publication Year: 2017
Physical Description: 188 p.
Subject Terms: Author profiling, Author identification, Text classification, Stylometry, Gender identification, Machine learning, Natural language processing, Syntax, Discourse, Feature engineering, Perfilament d'autors, Identificació d'autors, Classificació de text, Identificació de gènere, Estilometría, Aprenentatge automàtic, Processat del llenguatge, Sintaxis, Discurs
Description: Author profiling and identification are two areas of data-driven computational linguistics that have gained a lot of relevance due to their potential applications in, e.g., forensic linguistic studies, marketing analysis, and historic/literary authorship verification. Author profiling aims to identify demographic traits of the authors, while author identification aims to identify the authors themselves by searching for distinctive linguistic patterns that distinguish them. The majority of approaches in the related work tends to focus on the content of the texts. We argue that focusing on structure rather than content can be more effective. The main focus of the thesis is thus on feature engineering, the development, evaluation and application of the feature set in the context of machine learning techniques to author profiling and identification. We prove the profiling potential of syntactic and iscourse features, which achieve state-of-the-art performance in many different scenarios, especially when combined with other features.
Description (Translated): El perfilament i la identificació d’autors són camps de la lingüística computacional que han guanyat rellevància als últims anys gràcies a les seves potencials aplicacions al camp de la lingüística forense o a la verificació d’autoria de textos històrics. El perfilament d’autors té com a objectiu identificar trets demogràfics dels autors; la identificació d’autors tracta d’identificar l’autor del text. Per fer-ho, es busquen automàticament patrons lingüístics per diferenciar entre autors/trets demogràfics. La majoria de treballs anteriors, es centren en el contingut dels texts. Nosaltres argumentem que analitzar l’estructura del text pot ser una alternativa més efectiva. El focus d’aquesta tesi està per tant, al feature engineering: la extracció avaluació i utilització d’un conjunt de característiques lingüístiques amb algoritmes d’aprenentatge automàtic per a perfilar/identificar autors. Demostrem que les característiques sintàctiques i discursives són rellevants i que combinades amb altres, obtenen resultats a l’altura de l’estat de l’art.
Programa de doctorat en Tecnologies de la Informació i les Comunicacions
Document Type: Dissertation/Thesis
File Description: application/pdf
Language: English
Access URL: http://hdl.handle.net/10803/404984
Rights: ADVERTIMENT. L'accés als continguts d'aquesta tesi doctoral i la seva utilització ha de respectar els drets de la persona autora. Pot ser utilitzada per a consulta o estudi personal, així com en activitats o materials d'investigació i docència en els termes establerts a l'art. 32 del Text Refós de la Llei de Propietat Intel·lectual (RDL 1/1996). Per altres utilitzacions es requereix l'autorització prèvia i expressa de la persona autora. En qualsevol cas, en la utilització dels seus continguts caldrà indicar de forma clara el nom i cognoms de la persona autora i el títol de la tesi doctoral. No s'autoritza la seva reproducció o altres formes d'explotació efectuades amb finalitats de lucre ni la seva comunicació pública des d'un lloc aliè al servei TDX. Tampoc s'autoritza la presentació del seu contingut en una finestra o marc aliè a TDX (framing). Aquesta reserva de drets afecta tant als continguts de la tesi com als seus resums i índexs.
Accession Number: edstdx.10803.404984
Database: TDX
FullText Text:
  Availability: 0
CustomLinks:
  – Url: http://hdl.handle.net/10803/404984#
    Name: EDS - TDX (ns324271)
    Category: fullText
    Text: View record in TDX
Header DbId: edstdx
DbLabel: TDX
An: edstdx.10803.404984
RelevancyScore: 1314
AccessLevel: 3
PubType: Dissertation/ Thesis
PubTypeId: dissertation
PreciseRelevancyScore: 1314.16516113281
IllustrationInfo
Items – Name: Title
  Label: Title
  Group: Ti
  Data: Feature engineering for author profiling and identification: on the relevance of syntax and discourse
– Name: Author
  Label: Authors
  Group: Au
  Data: <searchLink fieldCode="AR" term="%22Soler+Company%2C+Juan%22">Soler Company, Juan</searchLink>
– Name: Author
  Label: Contributors
  Group: Au
  Data: University/Department: Universitat Pompeu Fabra. Departament de Tecnologies de la Informació i les Comunicacions
– Name: Author
  Label: Thesis Advisors
  Group: Au
  Data: Wanner, Leo
– Name: TitleSource
  Label: Source
  Group: Src
  Data: TDX (Tesis Doctorals en Xarxa)
– Name: Publisher
  Label: Publisher Information
  Group: PubInfo
  Data: Universitat Pompeu Fabra, 2017.
– Name: DatePubCY
  Label: Publication Year
  Group: Date
  Data: 2017
– Name: PhysDesc
  Label: Physical Description
  Group: PhysDesc
  Data: 188 p.
– Name: Subject
  Label: Subject Terms
  Group: Su
  Data: <searchLink fieldCode="DE" term="%22Author+profiling%22">Author profiling</searchLink><br /><searchLink fieldCode="DE" term="%22Author+identification%22">Author identification</searchLink><br /><searchLink fieldCode="DE" term="%22Text+classification%22">Text classification</searchLink><br /><searchLink fieldCode="DE" term="%22Stylometry%22">Stylometry</searchLink><br /><searchLink fieldCode="DE" term="%22Gender+identification%22">Gender identification</searchLink><br /><searchLink fieldCode="DE" term="%22Machine+learning%22">Machine learning</searchLink><br /><searchLink fieldCode="DE" term="%22Natural+language+processing%22">Natural language processing</searchLink><br /><searchLink fieldCode="DE" term="%22Syntax%22">Syntax</searchLink><br /><searchLink fieldCode="DE" term="%22Discourse%22">Discourse</searchLink><br /><searchLink fieldCode="DE" term="%22Feature+engineering%22">Feature engineering</searchLink><br /><searchLink fieldCode="DE" term="%22Perfilament+d'autors%22">Perfilament d'autors</searchLink><br /><searchLink fieldCode="DE" term="%22Identificació+d'autors%22">Identificació d'autors</searchLink><br /><searchLink fieldCode="DE" term="%22Classificació+de+text%22">Classificació de text</searchLink><br /><searchLink fieldCode="DE" term="%22Identificació+de+gènere%22">Identificació de gènere</searchLink><br /><searchLink fieldCode="DE" term="%22Estilometría%22">Estilometría</searchLink><br /><searchLink fieldCode="DE" term="%22Aprenentatge+automàtic%22">Aprenentatge automàtic</searchLink><br /><searchLink fieldCode="DE" term="%22Processat+del+llenguatge%22">Processat del llenguatge</searchLink><br /><searchLink fieldCode="DE" term="%22Sintaxis%22">Sintaxis</searchLink><br /><searchLink fieldCode="DE" term="%22Discurs%22">Discurs</searchLink>
– Name: Abstract
  Label: Description
  Group: Ab
  Data: Author profiling and identification are two areas of data-driven computational linguistics that have gained a lot of relevance due to their potential applications in, e.g., forensic linguistic studies, marketing analysis, and historic/literary authorship verification. Author profiling aims to identify demographic traits of the authors, while author identification aims to identify the authors themselves by searching for distinctive linguistic patterns that distinguish them. The majority of approaches in the related work tends to focus on the content of the texts. We argue that focusing on structure rather than content can be more effective. The main focus of the thesis is thus on feature engineering, the development, evaluation and application of the feature set in the context of machine learning techniques to author profiling and identification. We prove the profiling potential of syntactic and iscourse features, which achieve state-of-the-art performance in many different scenarios, especially when combined with other features.
– Name: Abstract
  Label: Description (Translated)
  Group: Ab
  Data: El perfilament i la identificació d’autors són camps de la lingüística computacional que han guanyat rellevància als últims anys gràcies a les seves potencials aplicacions al camp de la lingüística forense o a la verificació d’autoria de textos històrics. El perfilament d’autors té com a objectiu identificar trets demogràfics dels autors; la identificació d’autors tracta d’identificar l’autor del text. Per fer-ho, es busquen automàticament patrons lingüístics per diferenciar entre autors/trets demogràfics. La majoria de treballs anteriors, es centren en el contingut dels texts. Nosaltres argumentem que analitzar l’estructura del text pot ser una alternativa més efectiva. El focus d’aquesta tesi està per tant, al feature engineering: la extracció avaluació i utilització d’un conjunt de característiques lingüístiques amb algoritmes d’aprenentatge automàtic per a perfilar/identificar autors. Demostrem que les característiques sintàctiques i discursives són rellevants i que combinades amb altres, obtenen resultats a l’altura de l’estat de l’art.<br />Programa de doctorat en Tecnologies de la Informació i les Comunicacions
– Name: TypeDocument
  Label: Document Type
  Group: TypDoc
  Data: Dissertation/Thesis
– Name: Format
  Label: File Description
  Group: SrcInfo
  Data: application/pdf
– Name: Language
  Label: Language
  Group: Lang
  Data: English
– Name: URL
  Label: Access URL
  Group: URL
  Data: <link linkTarget="URL" linkTerm="http://hdl.handle.net/10803/404984" linkWindow="_blank">http://hdl.handle.net/10803/404984</link>
– Name: Copyright
  Label: Rights
  Group: Cpyrght
  Data: ADVERTIMENT. L'accés als continguts d'aquesta tesi doctoral i la seva utilització ha de respectar els drets de la persona autora. Pot ser utilitzada per a consulta o estudi personal, així com en activitats o materials d'investigació i docència en els termes establerts a l'art. 32 del Text Refós de la Llei de Propietat Intel·lectual (RDL 1/1996). Per altres utilitzacions es requereix l'autorització prèvia i expressa de la persona autora. En qualsevol cas, en la utilització dels seus continguts caldrà indicar de forma clara el nom i cognoms de la persona autora i el títol de la tesi doctoral. No s'autoritza la seva reproducció o altres formes d'explotació efectuades amb finalitats de lucre ni la seva comunicació pública des d'un lloc aliè al servei TDX. Tampoc s'autoritza la presentació del seu contingut en una finestra o marc aliè a TDX (framing). Aquesta reserva de drets afecta tant als continguts de la tesi com als seus resums i índexs.
– Name: AN
  Label: Accession Number
  Group: ID
  Data: edstdx.10803.404984
PLink https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=edstdx&AN=edstdx.10803.404984
RecordInfo BibRecord:
  BibEntity:
    Languages:
      – Text: English
    PhysicalDescription:
      Pagination:
        PageCount: 188
    Subjects:
      – SubjectFull: Author profiling
        Type: general
      – SubjectFull: Author identification
        Type: general
      – SubjectFull: Text classification
        Type: general
      – SubjectFull: Stylometry
        Type: general
      – SubjectFull: Gender identification
        Type: general
      – SubjectFull: Machine learning
        Type: general
      – SubjectFull: Natural language processing
        Type: general
      – SubjectFull: Syntax
        Type: general
      – SubjectFull: Discourse
        Type: general
      – SubjectFull: Feature engineering
        Type: general
      – SubjectFull: Perfilament d'autors
        Type: general
      – SubjectFull: Identificació d'autors
        Type: general
      – SubjectFull: Classificació de text
        Type: general
      – SubjectFull: Identificació de gènere
        Type: general
      – SubjectFull: Estilometría
        Type: general
      – SubjectFull: Aprenentatge automàtic
        Type: general
      – SubjectFull: Processat del llenguatge
        Type: general
      – SubjectFull: Sintaxis
        Type: general
      – SubjectFull: Discurs
        Type: general
    Titles:
      – TitleFull: Feature engineering for author profiling and identification: on the relevance of syntax and discourse
        Type: main
  BibRelationships:
    HasContributorRelationships:
      – PersonEntity:
          Name:
            NameFull: Soler Company, Juan
    IsPartOfRelationships:
      – BibEntity:
          Dates:
            – D: 06
              M: 07
              Type: published
              Y: 2017
ResultId 1