Evaluating Multiple System Summary Lengths: A Case Study

Ori Shapira; David Gabay; Hadar Ronen; Judit Bar-Ilan; Yael Amsterdamer; Ani Nenkova; Ido Dagan

2018 EMNLP EMNLP 2018

Evaluating Multiple System Summary Lengths: A Case Study

Abstract

AbstractPractical summarization systems are expected to produce summaries of varying lengths, per user needs. While a couple of early summarization benchmarks tested systems across multiple summary lengths, this practice was mostly abandoned due to the assumed cost of producing reference summaries of multiple lengths. In this paper, we raise the research question of whether reference summaries of a single length can be used to reliably evaluate system summaries of multiple lengths. For that, we have analyzed a couple of datasets as a case study, using several variants of the ROUGE metric that are standard in summarization evaluation. Our findings indicate that the evaluation protocol in question is indeed competitive. This result paves the way to practically evaluating varying-length summaries with simple, possibly existing, summarization benchmarks.

🧭 Keyword Pioneer — summary length

🐣 Hot Topic Early Bird — summarization evaluation

🐝 Cross-Pollinator — Artificial Intelligence, Computer Science, Computer Vision, Data Science & Analytics, Deep Learning, Healthcare & Medicine, Interdisciplinary, Knowledge & Reasoning, Machine Learning, Mathematics & Optimization, Natural Language Processing, Reinforcement Learning, Speech & Audio

Authors

Ori Shapira , David Gabay , Hadar Ronen , Judit Bar-Ilan , Yael Amsterdamer , Ani Nenkova , Ido Dagan

Topics

Natural Language Processing > Generation > Summarization

Keywords

summarization evaluation evaluation methodology rouge metric summary length

Download PDF

Related papers

Speeding Up Neural Machine Translation Decoding by Cube Pruning 2018

Limitations in learning an interpreted language with recurrent models 2018

Results of the sixth edition of the BioASQ Challenge 2018

Neural Segmental Hypergraphs for Overlapping Mention Recognition 2018

Hybrid Neural Attention for Agreement/Disagreement Inference in Online Debates 2018