MultiLexNorm: A Shared Task on Multilingual Lexical Normalization

Rob van der Goot, Alan Ramponi, Arkaitz Zubiaga, Barbara Plank, Benjamin Muller, Iñaki San Vicente Roncal, Nikola Ljubešić, Özlem Çetinoğlu, Rahmad Mahendra, Talha Çolakoğlu, Timothy Baldwin, Tommaso Caselli, Wladimir Sidorenko

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

26 Citations (Scopus)

Abstract

Lexical normalization is the task of transforming an utterance into its standardized form. This task is beneficial for downstream analysis, as it provides a way to harmonize (often spontaneous) linguistic variation. Such variation is typical for social media on which information is shared in a multitude of ways, including diverse languages and code-switching. Since the seminal work of Han and Baldwin (2011) a decade ago, lexical normalization has attracted attention in English and multiple other languages. However, there exists a lack of a common benchmark for comparison of systems across languages with a homogeneous data and evaluation setup. The MULTILEXNORM shared task sets out to fill this gap. We provide the largest publicly available multilingual lexical normalization benchmark including 12 language variants. We propose a homogenized evaluation setup with both intrinsic and extrinsic evaluation. As extrinsic evaluation, we use dependency parsing and part-of-speech tagging with adapted evaluation metrics (a-LAS, a-UAS, and a-POS) to account for alignment discrepancies. The shared task hosted at W-NUT 2021 attracted 9 participants and 18 submissions. The results show that neural normalization systems outperform the previous state-of-the-art system by a large margin. Downstream parsing and part-of-speech tagging performance is positively affected but to varying degrees, with improvements of up to 1.72 a-LAS, 0.85 a-UAS, and 1.54 a-POS for the winning system.

Original languageEnglish
Title of host publicationW-NUT 2021 - 7th Workshop on Noisy User-Generated Text, Proceedings of the Conference
EditorsWei Xu, Alan Ritter, Tim Baldwin, Afshin Rahimi
PublisherAssociation for Computational Linguistics (ACL)
Pages493-509
Number of pages17
ISBN (Electronic)9781954085909
Publication statusPublished - 2021
Event7th Workshop on Noisy User-Generated Text, W-NUT 2021 - Virtual, Online
Duration: 11 Nov 2021 → …

Publication series

NameW-NUT 2021 - 7th Workshop on Noisy User-Generated Text, Proceedings of the Conference

Conference

Conference7th Workshop on Noisy User-Generated Text, W-NUT 2021
CityVirtual, Online
Period11/11/21 → …

Fingerprint

Dive into the research topics of 'MultiLexNorm: A Shared Task on Multilingual Lexical Normalization'. Together they form a unique fingerprint.

Cite this