Leashing the Inner Demons: Self-Detoxification for Language Models

Canwen Xu; Zexue He; Zhankui He; Julian McAuley

2022 AAAI AAAI 2022

Leashing the Inner Demons: Self-Detoxification for Language Models

Abstract

Abstract Language models (LMs) can reproduce (or amplify) toxic language seen during training, which poses a risk to their practical application. In this paper, we conduct extensive experiments to study this phenomenon. We analyze the impact of prompts, decoding strategies and training corpora on the output toxicity. Based on our findings, we propose a simple yet effective unsupervised method for language models to ``detoxify'' themselves without an additional large corpus or external discriminator. Compared to a supervised baseline, our proposed method shows better toxicity reduction with good generation quality in the generated content under multiple settings. Warning: some examples shown in the paper may contain uncensored offensive content.

🌉 Interdisciplinary Bridge — Artificial Intelligence and Deep Learning and Machine Learning and Natural Language Processing

🐝 Cross-Pollinator — Artificial Intelligence, Computer Science, Computer Vision, Data Science & Analytics, Deep Learning, Healthcare & Medicine, Interdisciplinary, Knowledge & Reasoning, Machine Learning, Mathematics & Optimization, Natural Language Processing, Reinforcement Learning, Robotics, Security & Privacy, Speech & Audio

Authors

Canwen Xu , Zexue He , Zhankui He , Julian McAuley

Topics

Artificial Intelligence > Core AI > AI Safety Natural Language Processing > Generation > Language Modeling Machine Learning > Learning Types > Fairness Deep Learning > Learning Types > Generative Models Artificial Intelligence > Core AI > Safety Deep Learning > Learning Types > Text Generation

Keywords

self-supervised learning text generation toxicity detection language model toxicity reduction

Download PDF

Related papers

Dynamic Spatial Propagation Network for Depth Completion 2022

FedFR: Joint Optimization Federated Framework for Generic and Personalized Face Recognition 2022

Memory-Guided Semantic Learning Network for Temporal Sentence Grounding 2022

AnchorFace: Boosting TAR@FAR for Practical Face Recognition 2022

Parallel and High-Fidelity Text-to-Lip Generation 2022