A character-recognition system for Hangeul

Author

Johan Sageryd

Summary, in English

This work presents a rule-based character-recognition system for the Korean script, Hangeul. An input raster image representing one Korean character (Hangeul syllable) is thinned down to a skeleton, and the individual lines extracted. The lines, along with information on how they are interconnected, are translated into a set of hierarchical graphs, which can be easily traversed and compared with a set of reference structures represented in the same way. Hangeul consists of consonant and vowel graphemes, which are combined into blocks representing syllables. Each reference structure describes one possible variant of such a grapheme. The reference structures that best match the structures found in the input are combined to form a full Hangeul syllable. Testing all of the 11 172 possible characters, each rendered as a 200-pixel-squared raster image using the gothic font AppleGothic Regular, had a recognition accuracy of 80.6 percent. No separation logic exists to be able to handle characters whose graphemes are overlapping or conjoined; with such characters removed from the set, thereby reducing the total number of characters to 9 352, an accuracy of 96.3 percent was reached. Hand-written characters were also recognised, to a certain degree. The work shows that it is possible to create a workable character-recognition system with reasonably simple means.

Department/s

Language Technology Program

Publishing year

2009

Language

English

Full text

Document type

Student publication for Master's degree (one year)

Topic

Technology and Engineering
Languages and Literatures

Keywords

graph matching
graph thinning
Hangeul
Korean
character recognition

Supervisor

Johan Frid

A character-recognition system for Hangeul

Summary, in English

Contact information

Shortcuts

Find us on social media

Collaboration and networks