Dissertation/ Thesis

A Corpus Driven Computational Intelligence Framework for Deception Detection in Financial Text

Bibliographic Details
Title: A Corpus Driven Computational Intelligence Framework for Deception Detection in Financial Text
Authors: Minhas, Saliha Z
Contributors: Hussain, Amir
Publisher Information: University of Stirling
Publication Year: 2016
Collection: University of Stirling: Stirling Digital Research Repository
Subject Terms: Machine Learning, Financial Statement Fraud, Classififcation, Clustering, Language, Readability, Corpus Linguistics, Financial Fraud, Deception Detection, Unstructured Text, Language and computers Data processing, Corpora (Linguistics) Data processing, Misleading financial statements, Fraud
Description: Financial fraud rampages onwards seemingly uncontained. The annual cost of fraud in the UK is estimated to be as high as £193bn a year [1] . From a data science perspective and hitherto less explored this thesis demonstrates how the use of linguistic features to drive data mining algorithms can aid in unravelling fraud. To this end, the spotlight is turned on Financial Statement Fraud (FSF), known to be the costliest type of fraud [2]. A new corpus of 6.3 million words is composed of102 annual reports/10-K (narrative sections) from firms formally indicted for FSF juxtaposed with 306 non-fraud firms of similar size and industrial grouping. Differently from other similar studies, this thesis uniquely takes a wide angled view and extracts a range of features of different categories from the corpus. These linguistic correlates of deception are uncovered using a variety of techniques and tools. Corpus linguistics methodology is applied to extract keywords and to examine linguistic structure. N-grams are extracted to draw out collocations. Readability measurement in financial text is advanced through the extraction of new indices that probe the text at a deeper level. Cognitive and perceptual processes are also picked out. Tone, intention and liquidity are gauged using customised word lists. Linguistic ratios are derived from grammatical constructs and word categories. An attempt is also made to determine ‘what’ was said as opposed to ‘how’. Further a new module is developed to condense synonyms into concepts. Lastly frequency counts from keywords unearthed from a previous content analysis study on financial narrative are also used. These features are then used to drive machine learning based classification and clustering algorithms to determine if they aid in discriminating a fraud from a non-fraud firm. The results derived from the battery of models built typically exceed classification accuracy of 70%. The above process is amalgamated into a framework. The process outlined, driven by empirical data demonstrates in ...
Document Type: doctoral or postdoctoral thesis
File Description: application/pdf
Language: English
Relation: http://hdl.handle.net/1893/25345
Availability: http://hdl.handle.net/1893/25345
http://dspace.stir.ac.uk/bitstream/1893/25345/1/FINAL-%20MAIN.pdf
http://dspace.stir.ac.uk/bitstream/1893/25345/2/FINAL%20-%20APPENDICES.pdf
Accession Number: edsbas.CF49E630
Database: BASE
FullText Text:
  Availability: 0
CustomLinks:
  – Url: http://hdl.handle.net/1893/25345#
    Name: EDS - BASE (ns324271)
    Category: fullText
    Text: View record from BASE
Header DbId: edsbas
DbLabel: BASE
An: edsbas.CF49E630
RelevancyScore: 709
AccessLevel: 3
PubType: Dissertation/ Thesis
PubTypeId: dissertation
PreciseRelevancyScore: 709.17822265625
IllustrationInfo
Items – Name: Title
  Label: Title
  Group: Ti
  Data: A Corpus Driven Computational Intelligence Framework for Deception Detection in Financial Text
– Name: Author
  Label: Authors
  Group: Au
  Data: <searchLink fieldCode="AR" term="%22Minhas%2C+Saliha+Z%22">Minhas, Saliha Z</searchLink>
– Name: Author
  Label: Contributors
  Group: Au
  Data: Hussain, Amir
– Name: Publisher
  Label: Publisher Information
  Group: PubInfo
  Data: University of Stirling
– Name: DatePubCY
  Label: Publication Year
  Group: Date
  Data: 2016
– Name: Subset
  Label: Collection
  Group: HoldingsInfo
  Data: University of Stirling: Stirling Digital Research Repository
– Name: Subject
  Label: Subject Terms
  Group: Su
  Data: <searchLink fieldCode="DE" term="%22Machine+Learning%22">Machine Learning</searchLink><br /><searchLink fieldCode="DE" term="%22Financial+Statement+Fraud%22">Financial Statement Fraud</searchLink><br /><searchLink fieldCode="DE" term="%22Classififcation%22">Classififcation</searchLink><br /><searchLink fieldCode="DE" term="%22Clustering%22">Clustering</searchLink><br /><searchLink fieldCode="DE" term="%22Language%22">Language</searchLink><br /><searchLink fieldCode="DE" term="%22Readability%22">Readability</searchLink><br /><searchLink fieldCode="DE" term="%22Corpus+Linguistics%22">Corpus Linguistics</searchLink><br /><searchLink fieldCode="DE" term="%22Financial+Fraud%22">Financial Fraud</searchLink><br /><searchLink fieldCode="DE" term="%22Deception+Detection%22">Deception Detection</searchLink><br /><searchLink fieldCode="DE" term="%22Unstructured+Text%22">Unstructured Text</searchLink><br /><searchLink fieldCode="DE" term="%22Language+and+computers+Data+processing%22">Language and computers Data processing</searchLink><br /><searchLink fieldCode="DE" term="%22Corpora+%28Linguistics%29+Data+processing%22">Corpora (Linguistics) Data processing</searchLink><br /><searchLink fieldCode="DE" term="%22Misleading+financial+statements%22">Misleading financial statements</searchLink><br /><searchLink fieldCode="DE" term="%22Fraud%22">Fraud</searchLink>
– Name: Abstract
  Label: Description
  Group: Ab
  Data: Financial fraud rampages onwards seemingly uncontained. The annual cost of fraud in the UK is estimated to be as high as £193bn a year [1] . From a data science perspective and hitherto less explored this thesis demonstrates how the use of linguistic features to drive data mining algorithms can aid in unravelling fraud. To this end, the spotlight is turned on Financial Statement Fraud (FSF), known to be the costliest type of fraud [2]. A new corpus of 6.3 million words is composed of102 annual reports/10-K (narrative sections) from firms formally indicted for FSF juxtaposed with 306 non-fraud firms of similar size and industrial grouping. Differently from other similar studies, this thesis uniquely takes a wide angled view and extracts a range of features of different categories from the corpus. These linguistic correlates of deception are uncovered using a variety of techniques and tools. Corpus linguistics methodology is applied to extract keywords and to examine linguistic structure. N-grams are extracted to draw out collocations. Readability measurement in financial text is advanced through the extraction of new indices that probe the text at a deeper level. Cognitive and perceptual processes are also picked out. Tone, intention and liquidity are gauged using customised word lists. Linguistic ratios are derived from grammatical constructs and word categories. An attempt is also made to determine ‘what’ was said as opposed to ‘how’. Further a new module is developed to condense synonyms into concepts. Lastly frequency counts from keywords unearthed from a previous content analysis study on financial narrative are also used. These features are then used to drive machine learning based classification and clustering algorithms to determine if they aid in discriminating a fraud from a non-fraud firm. The results derived from the battery of models built typically exceed classification accuracy of 70%. The above process is amalgamated into a framework. The process outlined, driven by empirical data demonstrates in ...
– Name: TypeDocument
  Label: Document Type
  Group: TypDoc
  Data: doctoral or postdoctoral thesis
– Name: Format
  Label: File Description
  Group: SrcInfo
  Data: application/pdf
– Name: Language
  Label: Language
  Group: Lang
  Data: English
– Name: NoteTitleSource
  Label: Relation
  Group: SrcInfo
  Data: http://hdl.handle.net/1893/25345
– Name: URL
  Label: Availability
  Group: URL
  Data: http://hdl.handle.net/1893/25345<br />http://dspace.stir.ac.uk/bitstream/1893/25345/1/FINAL-%20MAIN.pdf<br />http://dspace.stir.ac.uk/bitstream/1893/25345/2/FINAL%20-%20APPENDICES.pdf
– Name: AN
  Label: Accession Number
  Group: ID
  Data: edsbas.CF49E630
PLink https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=edsbas&AN=edsbas.CF49E630
RecordInfo BibRecord:
  BibEntity:
    Languages:
      – Text: English
    Subjects:
      – SubjectFull: Machine Learning
        Type: general
      – SubjectFull: Financial Statement Fraud
        Type: general
      – SubjectFull: Classififcation
        Type: general
      – SubjectFull: Clustering
        Type: general
      – SubjectFull: Language
        Type: general
      – SubjectFull: Readability
        Type: general
      – SubjectFull: Corpus Linguistics
        Type: general
      – SubjectFull: Financial Fraud
        Type: general
      – SubjectFull: Deception Detection
        Type: general
      – SubjectFull: Unstructured Text
        Type: general
      – SubjectFull: Language and computers Data processing
        Type: general
      – SubjectFull: Corpora (Linguistics) Data processing
        Type: general
      – SubjectFull: Misleading financial statements
        Type: general
      – SubjectFull: Fraud
        Type: general
    Titles:
      – TitleFull: A Corpus Driven Computational Intelligence Framework for Deception Detection in Financial Text
        Type: main
  BibRelationships:
    HasContributorRelationships:
      – PersonEntity:
          Name:
            NameFull: Minhas, Saliha Z
      – PersonEntity:
          Name:
            NameFull: Hussain, Amir
    IsPartOfRelationships:
      – BibEntity:
          Dates:
            – D: 01
              M: 01
              Type: published
              Y: 2016
          Identifiers:
            – Type: issn-locals
              Value: edsbas
ResultId 1