Evaluation Examples are not Equally Informative: How should that change NLP Leaderboards?

Pedro Rodriguez; Joe Barrow; Alexander Hoyle; John P. Lalor; Robin Jia; Jordan Boyd-Graber

2021 ACL ACL 2021

Evaluation Examples are not Equally Informative: How should that change NLP Leaderboards?

Abstract

AbstractLeaderboards are widely used in NLP and push the field forward. While leaderboards are a straightforward ranking of NLP models, this simplicity can mask nuances in evaluation items (examples) and subjects (NLP models). Rather than replace leaderboards, we advocate a re-imagining so that they better highlight if and where progress is made. Building on educational testing, we create a Bayesian leaderboard model where latent subject skill and latent item difficulty predict correct responses. Using this model, we analyze the ranking reliability of leaderboards. Afterwards, we show the model can guide what to annotate, identify annotation errors, detect overfitting, and identify informative examples. We conclude with recommendations for future benchmark tasks.

❓ The Questioner

🌉 Interdisciplinary Bridge — Artificial Intelligence and Machine Learning

🐝 Cross-Pollinator — Artificial Intelligence, Computer Science, Computer Vision, Data Science & Analytics, Deep Learning, Healthcare & Medicine, Interdisciplinary, Knowledge & Reasoning, Machine Learning, Mathematics & Optimization, Natural Language Processing, Reinforcement Learning, Robotics, Security & Privacy, Speech & Audio

Authors

Pedro Rodriguez , Joe Barrow , Alexander Hoyle , John P. Lalor , Robin Jia , Jordan Boyd-Graber

Topics

Machine Learning > Optimization & Theory > Bayesian Inference Machine Learning > Optimization & Theory > Statistical Learning Artificial Intelligence > Bayesian & Probabilistic > Bayesian Inference Machine Learning > Optimization & Theory > Evaluation

Keywords

bayesian inference model evaluation latent variable model ranking annotation error item difficulty

Download PDF

Related papers

Out-of-Scope Intent Detection with Self-Supervision and Discriminative Training 2021

A Non-Autoregressive Edit-Based Approach to Controllable Text Simplification 2021

How Did This Get Funded?! Automatically Identifying Quirky Scientific Achievements 2021

Exploring Discourse Structures for Argument Impact Classification 2021

Language Embeddings for Typology and Cross-lingual Transfer Learning 2021