Exploring Robustness of Machine Translation Metrics: A Study of Twenty-Two Automatic Metrics in the WMT22 Metric Task

Xiaoyu Chen; Daimeng Wei; Hengchao Shang; Zongyao Li; Zhanglin Wu; Zhengzhe Yu; Ting Zhu; Mengli Zhu; Ning Xie; Lizhi Lei; Shimin Tao; Hao Yang; Ying Qin

2022 EMNLP EMNLP 2022

Exploring Robustness of Machine Translation Metrics: A Study of Twenty-Two Automatic Metrics in the WMT22 Metric Task

Abstract

AbstractContextual word embeddings extracted from pre-trained models have become the basis for many downstream NLP tasks, including machine translation automatic evaluations. Metrics that leverage embeddings claim better capture of synonyms and changes in word orders, and thus better correlation with human ratings than surface-form matching metrics (e.g. BLEU). However, few studies have been done to examine robustness of these metrics. This report uses a challenge set to uncover the brittleness of reference-based and reference-free metrics. Our challenge set1 aims at examining metrics’ capability to correlate synonyms in different areas and to discern catastrophic errors at both word- and sentence-levels. The results show that although embedding-based metrics perform relatively well on discerning sentence-level negation/affirmation errors, their performances on relating synonyms are poor. In addition, we find that some metrics are susceptible to text styles so their generalizability compromised.

🌉 Interdisciplinary Bridge — Machine Learning and Natural Language Processing

🧭 Keyword Pioneer — embedding-based metrics

🐝 Cross-Pollinator — Artificial Intelligence, Computer Science, Computer Vision, Data Science & Analytics, Deep Learning, Healthcare & Medicine, Interdisciplinary, Knowledge & Reasoning, Machine Learning, Mathematics & Optimization, Natural Language Processing, Reinforcement Learning, Robotics, Security & Privacy, Speech & Audio

Authors

Xiaoyu Chen , Daimeng Wei , Hengchao Shang , Zongyao Li , Zhanglin Wu , Zhengzhe Yu , Ting Zhu , Mengli Zhu , Ning Xie , Lizhi Lei , Shimin Tao , Hao Yang , Ying Qin

Topics

Machine Learning > Optimization & Theory > Optimization Natural Language Processing > Applications > Machine Translation Machine Learning > Learning Types > Evaluation

Keywords

machine translation robustness evaluation word embedding contextual embedding automatic evaluation machine translation metric challenge set metric robustness synonym detection embedding-based metrics

Download PDF

Generative Entity Typing with Curriculum Learning 2022

Towards Reinterpreting Neural Topic Models via Composite Activations 2022

Weakly Supervised Headline Dependency Parsing 2022

Cross-modal Transfer Between Vision and Language for Protest Detection 2022

Exploring Robustness of Machine Translation Metrics: A Study of Twenty-Two Automatic Metrics in the WMT22 Metric Task

Abstract

Authors

Topics

Keywords

Related papers