Simple But Not Naïve: Fine-Grained Arabic Dialect Identification Using Only N-Grams

Sohaila Eltanbouly; May Bashendy; Tamer Elsayed

2019 ACL ACL 2019

Simple But Not Naïve: Fine-Grained Arabic Dialect Identification Using Only N-Grams

Abstract

AbstractThis paper presents the participation of Qatar University team in MADAR shared task, which addresses the problem of sentence-level fine-grained Arabic Dialect Identification over 25 different Arabic dialects in addition to the Modern Standard Arabic. Arabic Dialect Identification is not a trivial task since different dialects share some features, e.g., utilizing the same character set and some vocabularies. We opted to adopt a very simple approach in terms of extracted features and classification models; we only utilize word and character n-grams as features, and Na ̈ıve Bayes models as classifiers. Surprisingly, the simple approach achieved non-na ̈ıve performance. The official results, reported on a held-out testing set, show that the dialect of a given sentence can be identified at an accuracy of 64.58% by our best submitted run.

🧭 Keyword Pioneer — word n-gram

🐝 Cross-Pollinator — Artificial Intelligence, Computer Science, Computer Vision, Data Science & Analytics, Deep Learning, Healthcare & Medicine, Interdisciplinary, Knowledge & Reasoning, Machine Learning, Mathematics & Optimization, Natural Language Processing, Reinforcement Learning, Security & Privacy, Speech & Audio

Authors

Sohaila Eltanbouly , May Bashendy , Tamer Elsayed

Topics

Machine Learning > Core Methods > Classification Machine Learning > Core Methods > Feature Learning Machine Learning > Learning Types > Classification

Keywords

text classification naive baye character n-gram dialect identification naive bayes classifier arabic dialect identification word n-gram

Download PDF

Related papers

What do phone embeddings learn about Phonology? 2019

Unsupervised Morphological Segmentation for Low-Resource Polysynthetic Languages 2019

Understanding Undesirable Word Embedding Associations 2019

Inferential Machine Comprehension: Answering Questions by Recursively Deducing the Evidence Chain from Text 2019

Domain Adaptation of Neural Machine Translation by Lexicon Induction 2019