A Comparison of Distributed, PAM, and Trie Data Structure Dictionaries in Automatic Spelling Correction for Indonesian Formal Text

Mukhlizar Nirwan Samsuri, Arlisa Yuliawati, Ika Alfina

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

Spelling errors can be divided into two groups, non-word errors and word errors. A non-word errors produce words that do not exist in dictionary, while word errors is a real word but not the right word. In this work, we address the non-word errors spelling correction for Indonesian formal text. The objective of our work is to compare the effectiveness of three kinds of dictionary structure for spelling correction, distributed dictionary, PAM (partition around medoids) dictionary, and dictionary using trie data structure, with the baseline of a simple flat dictionary. We conduct experiments with two variations of edit distances, i.e. Levenshtein and Damerau-Levenshtein, and utilized n-grams for ranking suggestion. We also build a gold standard of 200 sentences that consists of 4,323 tokens with 288 of them are non-word errors. Among the various combinations of dictionary type and edit distance, the trie data structure with Damerau-Levenshtein distance gets the best accuracy to produce candidate correction, i.e. 95.89% in 45.31 seconds. Furthermore, the combination of trie data structure with Damerau-Levenshtein distance also gets the best accuracy in choosing the best candidate, i.e. 73.15%.

Original languageEnglish
Title of host publication2022 5th International Seminar on Research of Information Technology and Intelligent Systems, ISRITI 2022
PublisherInstitute of Electrical and Electronics Engineers Inc.
Pages525-530
Number of pages6
ISBN (Electronic)9781665455121
DOIs
Publication statusPublished - 2022
Event5th International Seminar on Research of Information Technology and Intelligent Systems, ISRITI 2022 - Virtual, Online, Indonesia
Duration: 8 Dec 20229 Dec 2022

Publication series

Name2022 5th International Seminar on Research of Information Technology and Intelligent Systems, ISRITI 2022

Conference

Conference5th International Seminar on Research of Information Technology and Intelligent Systems, ISRITI 2022
Country/TerritoryIndonesia
CityVirtual, Online
Period8/12/229/12/22

Keywords

  • automatic spelling correction
  • distributed dictionary
  • non-word error
  • Partition Around Medoids
  • trie data structure

Fingerprint

Dive into the research topics of 'A Comparison of Distributed, PAM, and Trie Data Structure Dictionaries in Automatic Spelling Correction for Indonesian Formal Text'. Together they form a unique fingerprint.

Cite this