Dissertation/ Thesis
A Corpus Driven Computational Intelligence Framework for Deception Detection in Financial Text
| Title: | A Corpus Driven Computational Intelligence Framework for Deception Detection in Financial Text |
|---|---|
| Authors: | Minhas, Saliha Z |
| Contributors: | Hussain, Amir |
| Publisher Information: | University of Stirling |
| Publication Year: | 2016 |
| Collection: | University of Stirling: Stirling Digital Research Repository |
| Subject Terms: | Machine Learning, Financial Statement Fraud, Classififcation, Clustering, Language, Readability, Corpus Linguistics, Financial Fraud, Deception Detection, Unstructured Text, Language and computers Data processing, Corpora (Linguistics) Data processing, Misleading financial statements, Fraud |
| Description: | Financial fraud rampages onwards seemingly uncontained. The annual cost of fraud in the UK is estimated to be as high as £193bn a year [1] . From a data science perspective and hitherto less explored this thesis demonstrates how the use of linguistic features to drive data mining algorithms can aid in unravelling fraud. To this end, the spotlight is turned on Financial Statement Fraud (FSF), known to be the costliest type of fraud [2]. A new corpus of 6.3 million words is composed of102 annual reports/10-K (narrative sections) from firms formally indicted for FSF juxtaposed with 306 non-fraud firms of similar size and industrial grouping. Differently from other similar studies, this thesis uniquely takes a wide angled view and extracts a range of features of different categories from the corpus. These linguistic correlates of deception are uncovered using a variety of techniques and tools. Corpus linguistics methodology is applied to extract keywords and to examine linguistic structure. N-grams are extracted to draw out collocations. Readability measurement in financial text is advanced through the extraction of new indices that probe the text at a deeper level. Cognitive and perceptual processes are also picked out. Tone, intention and liquidity are gauged using customised word lists. Linguistic ratios are derived from grammatical constructs and word categories. An attempt is also made to determine ‘what’ was said as opposed to ‘how’. Further a new module is developed to condense synonyms into concepts. Lastly frequency counts from keywords unearthed from a previous content analysis study on financial narrative are also used. These features are then used to drive machine learning based classification and clustering algorithms to determine if they aid in discriminating a fraud from a non-fraud firm. The results derived from the battery of models built typically exceed classification accuracy of 70%. The above process is amalgamated into a framework. The process outlined, driven by empirical data demonstrates in ... |
| Document Type: | doctoral or postdoctoral thesis |
| File Description: | application/pdf |
| Language: | English |
| Relation: | http://hdl.handle.net/1893/25345 |
| Availability: | http://hdl.handle.net/1893/25345 http://dspace.stir.ac.uk/bitstream/1893/25345/1/FINAL-%20MAIN.pdf http://dspace.stir.ac.uk/bitstream/1893/25345/2/FINAL%20-%20APPENDICES.pdf |
| Accession Number: | edsbas.CF49E630 |
| Database: | BASE |
| FullText | Text: Availability: 0 CustomLinks: – Url: http://hdl.handle.net/1893/25345# Name: EDS - BASE (ns324271) Category: fullText Text: View record from BASE |
|---|---|
| Header | DbId: edsbas DbLabel: BASE An: edsbas.CF49E630 RelevancyScore: 709 AccessLevel: 3 PubType: Dissertation/ Thesis PubTypeId: dissertation PreciseRelevancyScore: 709.17822265625 |
| IllustrationInfo | |
| Items | – Name: Title Label: Title Group: Ti Data: A Corpus Driven Computational Intelligence Framework for Deception Detection in Financial Text – Name: Author Label: Authors Group: Au Data: <searchLink fieldCode="AR" term="%22Minhas%2C+Saliha+Z%22">Minhas, Saliha Z</searchLink> – Name: Author Label: Contributors Group: Au Data: Hussain, Amir – Name: Publisher Label: Publisher Information Group: PubInfo Data: University of Stirling – Name: DatePubCY Label: Publication Year Group: Date Data: 2016 – Name: Subset Label: Collection Group: HoldingsInfo Data: University of Stirling: Stirling Digital Research Repository – Name: Subject Label: Subject Terms Group: Su Data: <searchLink fieldCode="DE" term="%22Machine+Learning%22">Machine Learning</searchLink><br /><searchLink fieldCode="DE" term="%22Financial+Statement+Fraud%22">Financial Statement Fraud</searchLink><br /><searchLink fieldCode="DE" term="%22Classififcation%22">Classififcation</searchLink><br /><searchLink fieldCode="DE" term="%22Clustering%22">Clustering</searchLink><br /><searchLink fieldCode="DE" term="%22Language%22">Language</searchLink><br /><searchLink fieldCode="DE" term="%22Readability%22">Readability</searchLink><br /><searchLink fieldCode="DE" term="%22Corpus+Linguistics%22">Corpus Linguistics</searchLink><br /><searchLink fieldCode="DE" term="%22Financial+Fraud%22">Financial Fraud</searchLink><br /><searchLink fieldCode="DE" term="%22Deception+Detection%22">Deception Detection</searchLink><br /><searchLink fieldCode="DE" term="%22Unstructured+Text%22">Unstructured Text</searchLink><br /><searchLink fieldCode="DE" term="%22Language+and+computers+Data+processing%22">Language and computers Data processing</searchLink><br /><searchLink fieldCode="DE" term="%22Corpora+%28Linguistics%29+Data+processing%22">Corpora (Linguistics) Data processing</searchLink><br /><searchLink fieldCode="DE" term="%22Misleading+financial+statements%22">Misleading financial statements</searchLink><br /><searchLink fieldCode="DE" term="%22Fraud%22">Fraud</searchLink> – Name: Abstract Label: Description Group: Ab Data: Financial fraud rampages onwards seemingly uncontained. The annual cost of fraud in the UK is estimated to be as high as £193bn a year [1] . From a data science perspective and hitherto less explored this thesis demonstrates how the use of linguistic features to drive data mining algorithms can aid in unravelling fraud. To this end, the spotlight is turned on Financial Statement Fraud (FSF), known to be the costliest type of fraud [2]. A new corpus of 6.3 million words is composed of102 annual reports/10-K (narrative sections) from firms formally indicted for FSF juxtaposed with 306 non-fraud firms of similar size and industrial grouping. Differently from other similar studies, this thesis uniquely takes a wide angled view and extracts a range of features of different categories from the corpus. These linguistic correlates of deception are uncovered using a variety of techniques and tools. Corpus linguistics methodology is applied to extract keywords and to examine linguistic structure. N-grams are extracted to draw out collocations. Readability measurement in financial text is advanced through the extraction of new indices that probe the text at a deeper level. Cognitive and perceptual processes are also picked out. Tone, intention and liquidity are gauged using customised word lists. Linguistic ratios are derived from grammatical constructs and word categories. An attempt is also made to determine ‘what’ was said as opposed to ‘how’. Further a new module is developed to condense synonyms into concepts. Lastly frequency counts from keywords unearthed from a previous content analysis study on financial narrative are also used. These features are then used to drive machine learning based classification and clustering algorithms to determine if they aid in discriminating a fraud from a non-fraud firm. The results derived from the battery of models built typically exceed classification accuracy of 70%. The above process is amalgamated into a framework. The process outlined, driven by empirical data demonstrates in ... – Name: TypeDocument Label: Document Type Group: TypDoc Data: doctoral or postdoctoral thesis – Name: Format Label: File Description Group: SrcInfo Data: application/pdf – Name: Language Label: Language Group: Lang Data: English – Name: NoteTitleSource Label: Relation Group: SrcInfo Data: http://hdl.handle.net/1893/25345 – Name: URL Label: Availability Group: URL Data: http://hdl.handle.net/1893/25345<br />http://dspace.stir.ac.uk/bitstream/1893/25345/1/FINAL-%20MAIN.pdf<br />http://dspace.stir.ac.uk/bitstream/1893/25345/2/FINAL%20-%20APPENDICES.pdf – Name: AN Label: Accession Number Group: ID Data: edsbas.CF49E630 |
| PLink | https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=edsbas&AN=edsbas.CF49E630 |
| RecordInfo | BibRecord: BibEntity: Languages: – Text: English Subjects: – SubjectFull: Machine Learning Type: general – SubjectFull: Financial Statement Fraud Type: general – SubjectFull: Classififcation Type: general – SubjectFull: Clustering Type: general – SubjectFull: Language Type: general – SubjectFull: Readability Type: general – SubjectFull: Corpus Linguistics Type: general – SubjectFull: Financial Fraud Type: general – SubjectFull: Deception Detection Type: general – SubjectFull: Unstructured Text Type: general – SubjectFull: Language and computers Data processing Type: general – SubjectFull: Corpora (Linguistics) Data processing Type: general – SubjectFull: Misleading financial statements Type: general – SubjectFull: Fraud Type: general Titles: – TitleFull: A Corpus Driven Computational Intelligence Framework for Deception Detection in Financial Text Type: main BibRelationships: HasContributorRelationships: – PersonEntity: Name: NameFull: Minhas, Saliha Z – PersonEntity: Name: NameFull: Hussain, Amir IsPartOfRelationships: – BibEntity: Dates: – D: 01 M: 01 Type: published Y: 2016 Identifiers: – Type: issn-locals Value: edsbas |
| ResultId | 1 |