Neural Grammatical Error Correction Systems with Unsupervised Pre-training on Synthetic Data

Roman Grundkiewicz; Marcin Junczys-Dowmunt; Kenneth Heafield

2019 ACL ACL 2019

Neural Grammatical Error Correction Systems with Unsupervised Pre-training on Synthetic Data

Abstract

AbstractConsiderable effort has been made to address the data sparsity problem in neural grammatical error correction. In this work, we propose a simple and surprisingly effective unsupervised synthetic error generation method based on confusion sets extracted from a spellchecker to increase the amount of training data. Synthetic data is used to pre-train a Transformer sequence-to-sequence model, which not only improves over a strong baseline trained on authentic error-annotated data, but also enables the development of a practical GEC system in a scenario where little genuine error-annotated data is available. The developed systems placed first in the BEA19 shared task, achieving 69.47 and 64.24 F0.5 in the restricted and low-resource tracks respectively, both on the W&I+LOCNESS test set. On the popular CoNLL 2014 test set, we report state-of-the-art results of 64.16 M² for the submitted system, and 61.30 M² for the constrained system trained on the NUCLE and Lang-8 data.

🌉 Interdisciplinary Bridge — Machine Learning and Natural Language Processing

🧭 Keyword Pioneer — confusion set

🐝 Cross-Pollinator — Artificial Intelligence, Computer Vision, Data Science & Analytics, Deep Learning, Healthcare & Medicine, Interdisciplinary, Machine Learning, Mathematics & Optimization, Natural Language Processing, Reinforcement Learning, Robotics, Speech & Audio

🐣 Hot Topic Early Bird — synthetic data generation

Authors

Roman Grundkiewicz , Marcin Junczys-Dowmunt , Kenneth Heafield

Topics

Machine Learning > Learning Types > Unsupervised Learning Deep Learning > Architectures > Transformers Natural Language Processing > Generation > Text Generation Natural Language Processing > Applications > Text Generation Deep Learning > Learning Types > Self-Supervised Learning Deep Learning > Models > Transformers Natural Language Processing > Applications > Natural Language Understanding

Keywords

grammatical error correction synthetic data generation synthetic datum unsupervised pre-training confusion set transformer model

Download PDF

Related papers

What do phone embeddings learn about Phonology? 2019

Unsupervised Morphological Segmentation for Low-Resource Polysynthetic Languages 2019

Understanding Undesirable Word Embedding Associations 2019

Inferential Machine Comprehension: Answering Questions by Recursively Deducing the Evidence Chain from Text 2019

Domain Adaptation of Neural Machine Translation by Lexicon Induction 2019