Other/Unknown Material

A High Precision Algorithm for Automatic Extraction of High-frequency Words Based on Statistics

Λεπτομέρειες βιβλιογραφικής εγγραφής
Τίτλος: A High Precision Algorithm for Automatic Extraction of High-frequency Words Based on Statistics
Συγγραφείς: XUAN, Zhaoguo, DANG, Yanzhong, JIANG, Shaohua, ZHAO, Mingwei
Συνεισφορές: Jifa, Gu, Gerhard, Chroust
Στοιχεία εκδότη: JAIST Press
Θεματικοί όροι: Chinese-word segmentation, statistics algorithm, high-frequency words, Chinese information processing, psy, info
Περιγραφή: Automatic Chinese Word Segmentation is one of the basic research issues on text categorization, automatic summarization and information retrieval as well as other Chinese Information Processing tasks. In this paper we put forward a high precision algorithm for extracting high-frequency words without thesaurus. It firstly counts the frequencies of co-occurrence patterns of Chinese characters from documents, then eliminates the “bridge-connection” frequencies and therefore obtains the support frequencies of patterns. Afterwards, the words are identified and acquired according to the support frequencies instead of the primary appearing frequencies. The proposed algorithm is tested in the task of extracting words from several sets of scientific document abstracts, and the results show that this algorithm can improve both precision and recall of extracted lexical set to some extent. This algorithm can either be applied to text categorization and automatic summarization. ; The original publication is available at JAIST Press http://www.jaist.ac.jp/library/jaist-press/index.html ; IFSR 2005 : Proceedings of the First World Congress of the International Federation for Systems Research : The New Roles of Systems Sciences For a Knowledge-based Society : Nov. 14-17, 2133, Kobe, Japan ; Symposium 6, Session 4 : Vision of Knowledge Civilization Future Technology
Τύπος εγγράφου: other/unknown material
Γλώσσα: English
Relation: http://hdl.handle.net/10119/3923
Διαθεσιμότητα: http://hdl.handle.net/10119/3923
Rights: undefined
Αριθμός Καταχώρησης: edsbas.97A94C21
Βάση Δεδομένων: BASE
FullText Text:
  Availability: 0
CustomLinks:
  – Url: http://hdl.handle.net/10119/3923#
    Name: EDS - BASE (ns324271)
    Category: fullText
    Text: View record from BASE
Header DbId: edsbas
DbLabel: BASE
An: edsbas.97A94C21
RelevancyScore: 602
AccessLevel: 3
PubType: Other/Unknown Material
PubTypeId: unknown
PreciseRelevancyScore: 602
IllustrationInfo
Items – Name: Title
  Label: Title
  Group: Ti
  Data: A High Precision Algorithm for Automatic Extraction of High-frequency Words Based on Statistics
– Name: Author
  Label: Authors
  Group: Au
  Data: <searchLink fieldCode="AR" term="%22XUAN%2C+Zhaoguo%22">XUAN, Zhaoguo</searchLink><br /><searchLink fieldCode="AR" term="%22DANG%2C+Yanzhong%22">DANG, Yanzhong</searchLink><br /><searchLink fieldCode="AR" term="%22JIANG%2C+Shaohua%22">JIANG, Shaohua</searchLink><br /><searchLink fieldCode="AR" term="%22ZHAO%2C+Mingwei%22">ZHAO, Mingwei</searchLink>
– Name: Author
  Label: Contributors
  Group: Au
  Data: Jifa, Gu<br />Gerhard, Chroust
– Name: Publisher
  Label: Publisher Information
  Group: PubInfo
  Data: JAIST Press
– Name: Subject
  Label: Subject Terms
  Group: Su
  Data: <searchLink fieldCode="DE" term="%22Chinese-word+segmentation%22">Chinese-word segmentation</searchLink><br /><searchLink fieldCode="DE" term="%22statistics+algorithm%22">statistics algorithm</searchLink><br /><searchLink fieldCode="DE" term="%22high-frequency+words%22">high-frequency words</searchLink><br /><searchLink fieldCode="DE" term="%22Chinese+information+processing%22">Chinese information processing</searchLink><br /><searchLink fieldCode="DE" term="%22psy%22">psy</searchLink><br /><searchLink fieldCode="DE" term="%22info%22">info</searchLink>
– Name: Abstract
  Label: Description
  Group: Ab
  Data: Automatic Chinese Word Segmentation is one of the basic research issues on text categorization, automatic summarization and information retrieval as well as other Chinese Information Processing tasks. In this paper we put forward a high precision algorithm for extracting high-frequency words without thesaurus. It firstly counts the frequencies of co-occurrence patterns of Chinese characters from documents, then eliminates the “bridge-connection” frequencies and therefore obtains the support frequencies of patterns. Afterwards, the words are identified and acquired according to the support frequencies instead of the primary appearing frequencies. The proposed algorithm is tested in the task of extracting words from several sets of scientific document abstracts, and the results show that this algorithm can improve both precision and recall of extracted lexical set to some extent. This algorithm can either be applied to text categorization and automatic summarization. ; The original publication is available at JAIST Press http://www.jaist.ac.jp/library/jaist-press/index.html ; IFSR 2005 : Proceedings of the First World Congress of the International Federation for Systems Research : The New Roles of Systems Sciences For a Knowledge-based Society : Nov. 14-17, 2133, Kobe, Japan ; Symposium 6, Session 4 : Vision of Knowledge Civilization Future Technology
– Name: TypeDocument
  Label: Document Type
  Group: TypDoc
  Data: other/unknown material
– Name: Language
  Label: Language
  Group: Lang
  Data: English
– Name: NoteTitleSource
  Label: Relation
  Group: SrcInfo
  Data: http://hdl.handle.net/10119/3923
– Name: URL
  Label: Availability
  Group: URL
  Data: http://hdl.handle.net/10119/3923
– Name: Copyright
  Label: Rights
  Group: Cpyrght
  Data: undefined
– Name: AN
  Label: Accession Number
  Group: ID
  Data: edsbas.97A94C21
PLink https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=edsbas&AN=edsbas.97A94C21
RecordInfo BibRecord:
  BibEntity:
    Languages:
      – Text: English
    Subjects:
      – SubjectFull: Chinese-word segmentation
        Type: general
      – SubjectFull: statistics algorithm
        Type: general
      – SubjectFull: high-frequency words
        Type: general
      – SubjectFull: Chinese information processing
        Type: general
      – SubjectFull: psy
        Type: general
      – SubjectFull: info
        Type: general
    Titles:
      – TitleFull: A High Precision Algorithm for Automatic Extraction of High-frequency Words Based on Statistics
        Type: main
  BibRelationships:
    HasContributorRelationships:
      – PersonEntity:
          Name:
            NameFull: XUAN, Zhaoguo
      – PersonEntity:
          Name:
            NameFull: DANG, Yanzhong
      – PersonEntity:
          Name:
            NameFull: JIANG, Shaohua
      – PersonEntity:
          Name:
            NameFull: ZHAO, Mingwei
      – PersonEntity:
          Name:
            NameFull: Jifa, Gu
      – PersonEntity:
          Name:
            NameFull: Gerhard, Chroust
    IsPartOfRelationships:
      – BibEntity:
          Identifiers:
            – Type: issn-locals
              Value: edsbas
ResultId 1