Toward Micro-Dialect Identification in Diaglossic and Code-Switched Environments

Muhammad Abdul-Mageed; Chiyu Zhang; AbdelRahim Elmadany; Lyle Ungar

2020 EMNLP EMNLP 2020

Toward Micro-Dialect Identification in Diaglossic and Code-Switched Environments

Abstract

AbstractAlthough prediction of dialects is an important language processing task, with a wide range of applications, existing work is largely limited to coarse-grained varieties. Inspired by geolocation research, we propose the novel task of Micro-Dialect Identification (MDI) and introduce MARBERT, a new language model with striking abilities to predict a fine-grained variety (as small as that of a city) given a single, short message. For modeling, we offer a range of novel spatially and linguistically-motivated multi-task learning models. To showcase the utility of our models, we introduce a new, large-scale dataset of Arabic micro-varieties (low-resource) suited to our tasks. MARBERT predicts micro-dialects with 9.9% F1, 76 better than a majority class baseline. Our new language model also establishes new state-of-the-art on several external tasks.

🌉 Interdisciplinary Bridge — Artificial Intelligence and Deep Learning and Machine Learning and Natural Language Processing

🧭 Keyword Pioneer — arabic micro-variety

🐣 Hot Topic Early Bird — arabic language

🐝 Cross-Pollinator — Artificial Intelligence, Computer Science, Computer Vision, Data Science & Analytics, Deep Learning, Healthcare & Medicine, Interdisciplinary, Knowledge & Reasoning, Machine Learning, Mathematics & Optimization, Natural Language Processing, Reinforcement Learning, Robotics, Security & Privacy, Speech & Audio

Authors

Muhammad Abdul-Mageed , Chiyu Zhang , AbdelRahim Elmadany , Lyle Ungar

Topics

Machine Learning > Core Methods > Classification Machine Learning > Core Methods > Representation Learning Natural Language Processing > Understanding > Syntax Machine Learning > Learning Types > Multi-Task Learning Machine Learning > Learning Types > Classification Deep Learning > Learning Types > Multi-Task Learning Artificial Intelligence > Core AI > Natural Language Processing Deep Learning > Models > Language Models

Keywords

multi-task learning text classification fine-grained classification language model arabic language dialect identification dialect classification arabic micro-variety micro-dialect identification arabic micro-varieties

Download PDF

Related papers

Fast semantic parsing with well-typedness guarantees 2020

Detecting Objectifying Language in Online Professor Reviews 2020

Analogous Process Structure Induction for Sub-event Sequence Prediction 2020

Aspect Sentiment Classification with Aspect-Specific Opinion Spans 2020

Robust and Interpretable Grounding of Spatial References with Relation Networks 2020