Word spotting και clustering εικόνων χειρόγραφων λέξεων

Η συγκεκριμένη διπλωματική εργασία, ασχολείται με την τεχνική του Word Spotting καθώς και την ομαδοποίηση (Clustering) εικόνων χειρόγραφων λέξεων. Σκοπός της εργασίας ήταν η τροποποίηση και βελτίωση κάποιων ήδη υπάρχον μεθόδων, για την ομαδοποίηση χειρόγραφων λέξεων, γραμμένες από πλήθος διαφορετικώ...

Full description

Saved in:
Bibliographic Details
Main Authors: Ευαγόρου, Ανδρέας - Χρίστος, Πολυχρόνης, Μάριος - Γεώργιος
Other Authors: Καβαλλιεράτου, Εργίνα
Language:Greek
Published: 2015
Subjects:
Online Access:https://vsmart.lib.aegean.gr/webopac/List.csp?SearchT1=%CE%95%CF%85%CE%B1%CE%B3%CF%8C%CF%81%CE%BF%CF%85%2C+%CE%91%CE%BD%CE%B4%CF%81%CE%AD%CE%B1%CF%82&Index1=Keywordsbib&Database=1&SearchMethod=Find_1&SearchTerm1=%CE%95%CF%85%CE%B1%CE%B3%CF%8C%CF%81%CE%BF%CF%85%2C+%CE%91%CE%BD%CE%B4%CF%81%CE%AD%CE%B1%CF%82&OpacLanguage=gre&Profile=Default&EncodedRequest=*23*16*FE*09*C0*BE*F5*B0*CB*24*DD*B2*0E*E0*0E*D0&EncodedQuery=*23*16*FE*09*C0*BE*F5*B0*CB*24*DD*B2*0E*E0*0E*D0&Source=SysQR&PageType=Start&PreviousList=RecordListFind&WebPageNr=1&NumberToRetrieve=50&WebAction=NewSearch&StartValue=0&RowRepeat=0&ExtraInfo=&SortIndex=Year&SortDirection=-1&Resource=&SavingIndicator=&RestrType=&RestrTerms=&RestrShowAll=&LinkToIndex=
http://hdl.handle.net/11610/8801
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1828461936895852544
author Ευαγόρου, Ανδρέας - Χρίστος
Πολυχρόνης, Μάριος - Γεώργιος
author2 Καβαλλιεράτου, Εργίνα
author_facet Καβαλλιεράτου, Εργίνα
Ευαγόρου, Ανδρέας - Χρίστος
Πολυχρόνης, Μάριος - Γεώργιος
author_sort Ευαγόρου, Ανδρέας - Χρίστος
collection DSpace
description Η συγκεκριμένη διπλωματική εργασία, ασχολείται με την τεχνική του Word Spotting καθώς και την ομαδοποίηση (Clustering) εικόνων χειρόγραφων λέξεων. Σκοπός της εργασίας ήταν η τροποποίηση και βελτίωση κάποιων ήδη υπάρχον μεθόδων, για την ομαδοποίηση χειρόγραφων λέξεων, γραμμένες από πλήθος διαφορετικών ατόμων. Χρησιμοποιήθηκαν συνολικά πέντε μέθοδοι, εκ των οποίων οι δύο πάρθηκαν από τη βιβλιογραφία ως πρότυπο και με βάση αυτές, δημιουργήθηκαν και υλοποιήθηκαν τρεις νέες μέθοδοι με σκοπό να φέρουν καλύτερα αποτελέσματα από τις δύο ήδη υπάρχουσες. Για το σκοπό της σύγκρισης των δύο υπάρχον και των τριών νέων μεθόδων, δημιουργήθηκαν τρία διαφορετικά σύνολα δεδομένων, τα οποία αξιολογήθηκαν με βάση τέσσερα μέτρα αξιολόγησης (Precision, Recall, F-measure, Purity). Αξίζει να σημειωθεί, ότι μέσα από τη σύγκριση των εν λόγω μεθόδων, προέκυψαν καλύτερα αποτελέσματα ομαδοποίησης στις τρεις νέες μεθόδους που δημιουργήθηκαν από τις δύο ήδη υπάρχουσες, επιτυγχάνοντας έτσι το σκοπό της εργασίας. Συγκεκριμένα στο κεφάλαιο 1 δίνεται μία μικρή εισαγωγή της εργασίας αυτής. Στο κεφάλαιο 2, παρουσιάζονται κάποια γενικά στοιχεία του Word Spotting, αρχίζοντας με τον ορισμό της τεχνικής του, και συνεχίζοντας με τα προβλήματα τα οποία καλείται να αντιμετωπίσει, την βασική του ιδέα, τα στάδια από τα οποία αποτελείται, καθώς και κάποιες εργασίες που έχουν γίνει γύρω από το θέμα αυτό. Επίσης, γίνεται μια μικρή αναφορά σχετικά με την αποτυχία του Optical Character Recognition (OCR), στις περιπτώσεις χειρόγραφων εγγράφων. Στο κεφάλαιο 3, παρουσιάζονται βασικά στάδια υλοποίησης, που περιλαμβάνονται μέσα στο Word Spotting, τα οποία είναι απαραίτητα τόσο για την επιτυχία του Word Spotting, όσο για τη σωστή ομαδοποίηση των χειρόγραφων εικόνων, καθώς επίσης δίνεται και μια περιγραφή δύο ειδών μετασχηματισμού, οι οποίοι είναι απαραίτητοι για την υλοποίηση των μεθόδων της εργασίας. Στο κεφάλαιο 4, αρχικά δίνεται μια μικρή εισαγωγή του κεφαλαίου, την οποία ακολουθεί μια σύντομη περιγραφή των 5 μεθόδων που υλοποιήθηκαν για τους σκοπούς της εργασίας. Στη συνέχεια, παρουσιάζονται τα σύνολα δεδομένων που χρησιμοποιούνται στα πειράματα της εργασίας, καθώς και η μέθοδος αξιολόγησης των πειραμάτων αυτών. Τέλος, δίνεται αναλυτική περιγραφή της υλοποίησης των πέντε σεναρίων, δηλαδή, των δύο ήδη υπάρχον και των τριών νέων προτεινόμενων, τα οποία έχουν στόχο την ομαδοποίηση εικόνων χειρόγραφων λέξεων διάφορων συγγραφέων. Στο κεφάλαιο 5, παρουσιάζονται τα συγκριτικά αποτελέσματα που έδωσαν οι 5 μέθοδοι, και δίνονται τα συμπεράσματα που εξάχθηκαν από τα αποτελέσματα αυτά.
id oai:hellanicus.lib.aegean.gr:11610-8801
institution Hellanicus
language Greek
publishDate 2015
record_format dspace
spelling oai:hellanicus.lib.aegean.gr:11610-88012025-02-07T14:23:09Z Word spotting και clustering εικόνων χειρόγραφων λέξεων Ευαγόρου, Ανδρέας - Χρίστος Πολυχρόνης, Μάριος - Γεώργιος Καβαλλιεράτου, Εργίνα Συσταδοποίηση εικόνων Αναγνώριση προτύπων Χειρόγραφα έγγραφα Word spotting Clustering Pattern recognition Handwritten word images Pattern recognition systems Η συγκεκριμένη διπλωματική εργασία, ασχολείται με την τεχνική του Word Spotting καθώς και την ομαδοποίηση (Clustering) εικόνων χειρόγραφων λέξεων. Σκοπός της εργασίας ήταν η τροποποίηση και βελτίωση κάποιων ήδη υπάρχον μεθόδων, για την ομαδοποίηση χειρόγραφων λέξεων, γραμμένες από πλήθος διαφορετικών ατόμων. Χρησιμοποιήθηκαν συνολικά πέντε μέθοδοι, εκ των οποίων οι δύο πάρθηκαν από τη βιβλιογραφία ως πρότυπο και με βάση αυτές, δημιουργήθηκαν και υλοποιήθηκαν τρεις νέες μέθοδοι με σκοπό να φέρουν καλύτερα αποτελέσματα από τις δύο ήδη υπάρχουσες. Για το σκοπό της σύγκρισης των δύο υπάρχον και των τριών νέων μεθόδων, δημιουργήθηκαν τρία διαφορετικά σύνολα δεδομένων, τα οποία αξιολογήθηκαν με βάση τέσσερα μέτρα αξιολόγησης (Precision, Recall, F-measure, Purity). Αξίζει να σημειωθεί, ότι μέσα από τη σύγκριση των εν λόγω μεθόδων, προέκυψαν καλύτερα αποτελέσματα ομαδοποίησης στις τρεις νέες μεθόδους που δημιουργήθηκαν από τις δύο ήδη υπάρχουσες, επιτυγχάνοντας έτσι το σκοπό της εργασίας. Συγκεκριμένα στο κεφάλαιο 1 δίνεται μία μικρή εισαγωγή της εργασίας αυτής. Στο κεφάλαιο 2, παρουσιάζονται κάποια γενικά στοιχεία του Word Spotting, αρχίζοντας με τον ορισμό της τεχνικής του, και συνεχίζοντας με τα προβλήματα τα οποία καλείται να αντιμετωπίσει, την βασική του ιδέα, τα στάδια από τα οποία αποτελείται, καθώς και κάποιες εργασίες που έχουν γίνει γύρω από το θέμα αυτό. Επίσης, γίνεται μια μικρή αναφορά σχετικά με την αποτυχία του Optical Character Recognition (OCR), στις περιπτώσεις χειρόγραφων εγγράφων. Στο κεφάλαιο 3, παρουσιάζονται βασικά στάδια υλοποίησης, που περιλαμβάνονται μέσα στο Word Spotting, τα οποία είναι απαραίτητα τόσο για την επιτυχία του Word Spotting, όσο για τη σωστή ομαδοποίηση των χειρόγραφων εικόνων, καθώς επίσης δίνεται και μια περιγραφή δύο ειδών μετασχηματισμού, οι οποίοι είναι απαραίτητοι για την υλοποίηση των μεθόδων της εργασίας. Στο κεφάλαιο 4, αρχικά δίνεται μια μικρή εισαγωγή του κεφαλαίου, την οποία ακολουθεί μια σύντομη περιγραφή των 5 μεθόδων που υλοποιήθηκαν για τους σκοπούς της εργασίας. Στη συνέχεια, παρουσιάζονται τα σύνολα δεδομένων που χρησιμοποιούνται στα πειράματα της εργασίας, καθώς και η μέθοδος αξιολόγησης των πειραμάτων αυτών. Τέλος, δίνεται αναλυτική περιγραφή της υλοποίησης των πέντε σεναρίων, δηλαδή, των δύο ήδη υπάρχον και των τριών νέων προτεινόμενων, τα οποία έχουν στόχο την ομαδοποίηση εικόνων χειρόγραφων λέξεων διάφορων συγγραφέων. Στο κεφάλαιο 5, παρουσιάζονται τα συγκριτικά αποτελέσματα που έδωσαν οι 5 μέθοδοι, και δίνονται τα συμπεράσματα που εξάχθηκαν από τα αποτελέσματα αυτά. In this dissertation, some stages of the Word Spotting technique are applied, and more particularly, the stages that are needed, so that the Clustering of handwritten word images can be performed successfully. The main goal of this project was to produce the implementations of 2 methods that have been already proposed in an older paper, and based on those implementations, apply some modifications, so that some new, better performed, methods can be created. A significant difference between the initial methods and the ones proposed here, is that the initial methods were created with the target of applying the Word Spotting technique on handwritten document images of one (or a small group of) writer(s), while, on the other hand, the proposed methods of this dissertation, were created with the target of applying the Clustering process of handwritten word images of many writers. Although not all the stages of the Word Spotting technique were applied in the presented methods, it is our belief that the hard work has been done, and that adding the rest of the stages is not of huge importance. More specifically, the first of the initial methods made use of the Dynamic Time Warping (DTW) algorithm in the Matching stage, while the second one made use of the Discrete Fourier Transform (DFT), to transform the extracted features of the word images. While our implementation of those two methods might not be an exact match of the implementations of the original paper, we believe that it came very close to the originals, as we followed almost all the stages presented in that paper. The method that we implemented and was purporting the original first method was method 1 of this project, while the method that we implemented and was purporting the original second method was method 3. One of the new methods created in this project (method 2), was a modification of method 1, where the modifications that were implemented were mostly in the pre-processing stage, the feature extraction stage and the Clustering stage. In more detail, the pre-processing was chosen so that it would work best for our data sets, some new features were added in the list of extracted features from each word image, and a different Clustering algorithm was used, with a different distance metric. Another one of the new methods created during this dissertation (method 4), was a modification of method 3, where the modifications that were implemented were mostly in the pre-processing stage, and the feature extraction stage. In more detail, the pre-processing again was chosen so that it would work best for our data sets, and some new features were added in the list of extracted features from each word image. The last one of the methods proposed (method 5), is an alternate approach of method 4, where instead of the DFT, the Discrete Cosine Transform (DCT) was used to transform the features extracted from each image. Three data sets were used in our experiments, where two of them were of good quality (one of few images (200), and one of many images (1008)), and one of them (this set was also consisted of 200 images), contained randomly taken images from a larger set that contained both good and bad quality’s word images. For the evaluation of the performance of the Clustering process, four performance indicators were used: precision, recall, F-measure, and purity. For a small description of the dissertation, in chapter 1, a small introduction is given, while in chapter 2, a general approach of the Word Spotting technique is given. Specifically, in chapter 2 a definition of the technique, and some of the problems that it has to face are provided, the reason why Optical Character Recognition (OCR) fails in the case of handwritten documents is given, the basic idea of Word Spotting is presented, followed by a description of its stages. Finally, the chapter ends with a brief discussion of some related work that has been done on this field. In chapter 3, the basic stages of the Word Spotting technique which are implemented in this dissertation are presented, while a small description of the two transformations that are used in some of the implemented methods is also provided. Proceeding to chapter 4, a small introduction of the chapter is found, followed by a brief description of the methods implemented in this dissertation. Then, a presentation of the data sets is given, and the evaluation method of our experiments is explained, an extensive description of the implementations of the methods is also given, and finally, we present a Graphical User Interface (GUI), which was created so that the use of our methods can be performed with ease. Chapter 5 contains the final results of all 5 methods, and the conclusions extracted from those results. By viewing the final results, it is our opinion that we have accomplished our goal, by managing to modify the original methods and got them to give better results, and by making them able to be applied successfully in handwritten word images written by many different writers. 2015-11-17T10:32:28Z 2015-11-17T10:32:28Z 2011 https://vsmart.lib.aegean.gr/webopac/List.csp?SearchT1=%CE%95%CF%85%CE%B1%CE%B3%CF%8C%CF%81%CE%BF%CF%85%2C+%CE%91%CE%BD%CE%B4%CF%81%CE%AD%CE%B1%CF%82&Index1=Keywordsbib&Database=1&SearchMethod=Find_1&SearchTerm1=%CE%95%CF%85%CE%B1%CE%B3%CF%8C%CF%81%CE%BF%CF%85%2C+%CE%91%CE%BD%CE%B4%CF%81%CE%AD%CE%B1%CF%82&OpacLanguage=gre&Profile=Default&EncodedRequest=*23*16*FE*09*C0*BE*F5*B0*CB*24*DD*B2*0E*E0*0E*D0&EncodedQuery=*23*16*FE*09*C0*BE*F5*B0*CB*24*DD*B2*0E*E0*0E*D0&Source=SysQR&PageType=Start&PreviousList=RecordListFind&WebPageNr=1&NumberToRetrieve=50&WebAction=NewSearch&StartValue=0&RowRepeat=0&ExtraInfo=&SortIndex=Year&SortDirection=-1&Resource=&SavingIndicator=&RestrType=&RestrTerms=&RestrShowAll=&LinkToIndex= http://hdl.handle.net/11610/8801 el application/pdf Σάμος
spellingShingle Συσταδοποίηση εικόνων
Αναγνώριση προτύπων
Χειρόγραφα έγγραφα
Word spotting
Clustering
Pattern recognition
Handwritten word images
Pattern recognition systems
Ευαγόρου, Ανδρέας - Χρίστος
Πολυχρόνης, Μάριος - Γεώργιος
Word spotting και clustering εικόνων χειρόγραφων λέξεων
title Word spotting και clustering εικόνων χειρόγραφων λέξεων
title_full Word spotting και clustering εικόνων χειρόγραφων λέξεων
title_fullStr Word spotting και clustering εικόνων χειρόγραφων λέξεων
title_full_unstemmed Word spotting και clustering εικόνων χειρόγραφων λέξεων
title_short Word spotting και clustering εικόνων χειρόγραφων λέξεων
title_sort word spotting και clustering εικόνων χειρόγραφων λέξεων
topic Συσταδοποίηση εικόνων
Αναγνώριση προτύπων
Χειρόγραφα έγγραφα
Word spotting
Clustering
Pattern recognition
Handwritten word images
Pattern recognition systems
url https://vsmart.lib.aegean.gr/webopac/List.csp?SearchT1=%CE%95%CF%85%CE%B1%CE%B3%CF%8C%CF%81%CE%BF%CF%85%2C+%CE%91%CE%BD%CE%B4%CF%81%CE%AD%CE%B1%CF%82&Index1=Keywordsbib&Database=1&SearchMethod=Find_1&SearchTerm1=%CE%95%CF%85%CE%B1%CE%B3%CF%8C%CF%81%CE%BF%CF%85%2C+%CE%91%CE%BD%CE%B4%CF%81%CE%AD%CE%B1%CF%82&OpacLanguage=gre&Profile=Default&EncodedRequest=*23*16*FE*09*C0*BE*F5*B0*CB*24*DD*B2*0E*E0*0E*D0&EncodedQuery=*23*16*FE*09*C0*BE*F5*B0*CB*24*DD*B2*0E*E0*0E*D0&Source=SysQR&PageType=Start&PreviousList=RecordListFind&WebPageNr=1&NumberToRetrieve=50&WebAction=NewSearch&StartValue=0&RowRepeat=0&ExtraInfo=&SortIndex=Year&SortDirection=-1&Resource=&SavingIndicator=&RestrType=&RestrTerms=&RestrShowAll=&LinkToIndex=
http://hdl.handle.net/11610/8801
work_keys_str_mv AT euagorouandreaschristos wordspottingkaiclusteringeikonōncheirographōnlexeōn
AT polychronēsmariosgeōrgios wordspottingkaiclusteringeikonōncheirographōnlexeōn