On Leakage of Code Generation Evaluation Datasets

Alexandre Matton; Tom Sherborne; Dennis Aumiller; Elena Tommasone; Milad Alizadeh; Jingyi He; Raymond Ma; Maxime Voisin; Ellen Gilsenan-McMahon; Matthias Gallé

2024 EMNLP EMNLP 2024

On Leakage of Code Generation Evaluation Datasets

Abstract

AbstractIn this paper, we consider contamination by code generation test sets, in particular in their use in modern large language models.We discuss three possible sources of such contamination and show findings supporting each of them: (i) direct data leakage, (ii) indirect data leakage through the use of synthetic data and (iii) overfitting to evaluation sets during model selection.To address this, we release Less Basic Python Problems (LBPP): an uncontaminated new benchmark of 161 prompts with their associated Python solutions. LBPP is released at https://huggingface.co/datasets/CohereForAI/lbpp

🌉 Interdisciplinary Bridge — Artificial Intelligence and Computer Science and Deep Learning and Machine Learning

🧭 Keyword Pioneer — dataset leakage

🐝 Cross-Pollinator — Artificial Intelligence, Computer Science, Computer Vision, Data Science & Analytics, Deep Learning, Healthcare & Medicine, Interdisciplinary, Knowledge & Reasoning, Machine Learning, Mathematics & Optimization, Natural Language Processing, Reinforcement Learning, Robotics, Security & Privacy, Speech & Audio

Authors

Alexandre Matton , Tom Sherborne , Dennis Aumiller , Elena Tommasone , Milad Alizadeh , Jingyi He , Raymond Ma , Maxime Voisin , Ellen Gilsenan-McMahon , Matthias Gallé

Topics

Artificial Intelligence > Core AI > Autonomous Vehicles Machine Learning > Optimization & Theory > Theory Machine Learning > Application Areas > Privacy Machine Learning > Application Areas > Risk Management Computer Science > Applications > Software Engineering Machine Learning > Learning Types > Robustness Deep Learning > Optimization & Theory > Evaluation

Keywords

benchmark evaluation model selection code generation synthetic datum data contamination benchmark contamination dataset leakage

Download PDF

Related papers

EmbodiedBERT: Cognitively Informed Metaphor Detection Incorporating Sensorimotor Information 2024

Mitigating Matthew Effect: Multi-Hypergraph Boosted Multi-Interest Self-Supervised Learning for Conversational Recommendation 2024

Learning to Extract Structured Entities Using Language Models 2024

Towards Understanding Jailbreak Attacks in LLMs: A Representation Space Analysis 2024

CSSL: Contrastive Self-Supervised Learning for Dependency Parsing on Relatively Free Word Ordered and Morphologically Rich Low Resource Languages 2024