Other/Unknown Material
A High Precision Algorithm for Automatic Extraction of High-frequency Words Based on Statistics
| Τίτλος: | A High Precision Algorithm for Automatic Extraction of High-frequency Words Based on Statistics |
|---|---|
| Συγγραφείς: | XUAN, Zhaoguo, DANG, Yanzhong, JIANG, Shaohua, ZHAO, Mingwei |
| Συνεισφορές: | Jifa, Gu, Gerhard, Chroust |
| Στοιχεία εκδότη: | JAIST Press |
| Θεματικοί όροι: | Chinese-word segmentation, statistics algorithm, high-frequency words, Chinese information processing, psy, info |
| Περιγραφή: | Automatic Chinese Word Segmentation is one of the basic research issues on text categorization, automatic summarization and information retrieval as well as other Chinese Information Processing tasks. In this paper we put forward a high precision algorithm for extracting high-frequency words without thesaurus. It firstly counts the frequencies of co-occurrence patterns of Chinese characters from documents, then eliminates the “bridge-connection” frequencies and therefore obtains the support frequencies of patterns. Afterwards, the words are identified and acquired according to the support frequencies instead of the primary appearing frequencies. The proposed algorithm is tested in the task of extracting words from several sets of scientific document abstracts, and the results show that this algorithm can improve both precision and recall of extracted lexical set to some extent. This algorithm can either be applied to text categorization and automatic summarization. ; The original publication is available at JAIST Press http://www.jaist.ac.jp/library/jaist-press/index.html ; IFSR 2005 : Proceedings of the First World Congress of the International Federation for Systems Research : The New Roles of Systems Sciences For a Knowledge-based Society : Nov. 14-17, 2133, Kobe, Japan ; Symposium 6, Session 4 : Vision of Knowledge Civilization Future Technology |
| Τύπος εγγράφου: | other/unknown material |
| Γλώσσα: | English |
| Relation: | http://hdl.handle.net/10119/3923 |
| Διαθεσιμότητα: | http://hdl.handle.net/10119/3923 |
| Rights: | undefined |
| Αριθμός Καταχώρησης: | edsbas.97A94C21 |
| Βάση Δεδομένων: | BASE |
| FullText | Text: Availability: 0 CustomLinks: – Url: http://hdl.handle.net/10119/3923# Name: EDS - BASE (ns324271) Category: fullText Text: View record from BASE |
|---|---|
| Header | DbId: edsbas DbLabel: BASE An: edsbas.97A94C21 RelevancyScore: 602 AccessLevel: 3 PubType: Other/Unknown Material PubTypeId: unknown PreciseRelevancyScore: 602 |
| IllustrationInfo | |
| Items | – Name: Title Label: Title Group: Ti Data: A High Precision Algorithm for Automatic Extraction of High-frequency Words Based on Statistics – Name: Author Label: Authors Group: Au Data: <searchLink fieldCode="AR" term="%22XUAN%2C+Zhaoguo%22">XUAN, Zhaoguo</searchLink><br /><searchLink fieldCode="AR" term="%22DANG%2C+Yanzhong%22">DANG, Yanzhong</searchLink><br /><searchLink fieldCode="AR" term="%22JIANG%2C+Shaohua%22">JIANG, Shaohua</searchLink><br /><searchLink fieldCode="AR" term="%22ZHAO%2C+Mingwei%22">ZHAO, Mingwei</searchLink> – Name: Author Label: Contributors Group: Au Data: Jifa, Gu<br />Gerhard, Chroust – Name: Publisher Label: Publisher Information Group: PubInfo Data: JAIST Press – Name: Subject Label: Subject Terms Group: Su Data: <searchLink fieldCode="DE" term="%22Chinese-word+segmentation%22">Chinese-word segmentation</searchLink><br /><searchLink fieldCode="DE" term="%22statistics+algorithm%22">statistics algorithm</searchLink><br /><searchLink fieldCode="DE" term="%22high-frequency+words%22">high-frequency words</searchLink><br /><searchLink fieldCode="DE" term="%22Chinese+information+processing%22">Chinese information processing</searchLink><br /><searchLink fieldCode="DE" term="%22psy%22">psy</searchLink><br /><searchLink fieldCode="DE" term="%22info%22">info</searchLink> – Name: Abstract Label: Description Group: Ab Data: Automatic Chinese Word Segmentation is one of the basic research issues on text categorization, automatic summarization and information retrieval as well as other Chinese Information Processing tasks. In this paper we put forward a high precision algorithm for extracting high-frequency words without thesaurus. It firstly counts the frequencies of co-occurrence patterns of Chinese characters from documents, then eliminates the “bridge-connection” frequencies and therefore obtains the support frequencies of patterns. Afterwards, the words are identified and acquired according to the support frequencies instead of the primary appearing frequencies. The proposed algorithm is tested in the task of extracting words from several sets of scientific document abstracts, and the results show that this algorithm can improve both precision and recall of extracted lexical set to some extent. This algorithm can either be applied to text categorization and automatic summarization. ; The original publication is available at JAIST Press http://www.jaist.ac.jp/library/jaist-press/index.html ; IFSR 2005 : Proceedings of the First World Congress of the International Federation for Systems Research : The New Roles of Systems Sciences For a Knowledge-based Society : Nov. 14-17, 2133, Kobe, Japan ; Symposium 6, Session 4 : Vision of Knowledge Civilization Future Technology – Name: TypeDocument Label: Document Type Group: TypDoc Data: other/unknown material – Name: Language Label: Language Group: Lang Data: English – Name: NoteTitleSource Label: Relation Group: SrcInfo Data: http://hdl.handle.net/10119/3923 – Name: URL Label: Availability Group: URL Data: http://hdl.handle.net/10119/3923 – Name: Copyright Label: Rights Group: Cpyrght Data: undefined – Name: AN Label: Accession Number Group: ID Data: edsbas.97A94C21 |
| PLink | https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=edsbas&AN=edsbas.97A94C21 |
| RecordInfo | BibRecord: BibEntity: Languages: – Text: English Subjects: – SubjectFull: Chinese-word segmentation Type: general – SubjectFull: statistics algorithm Type: general – SubjectFull: high-frequency words Type: general – SubjectFull: Chinese information processing Type: general – SubjectFull: psy Type: general – SubjectFull: info Type: general Titles: – TitleFull: A High Precision Algorithm for Automatic Extraction of High-frequency Words Based on Statistics Type: main BibRelationships: HasContributorRelationships: – PersonEntity: Name: NameFull: XUAN, Zhaoguo – PersonEntity: Name: NameFull: DANG, Yanzhong – PersonEntity: Name: NameFull: JIANG, Shaohua – PersonEntity: Name: NameFull: ZHAO, Mingwei – PersonEntity: Name: NameFull: Jifa, Gu – PersonEntity: Name: NameFull: Gerhard, Chroust IsPartOfRelationships: – BibEntity: Identifiers: – Type: issn-locals Value: edsbas |
| ResultId | 1 |