Draft on the Fly: Adaptive Self-Speculative Decoding using Cosine Similarity

Michael R. Metel; Peng Lu; Boxing Chen; Mehdi Rezagholizadeh; Ivan Kobyzev

2024 EMNLP EMNLP 2024

Draft on the Fly: Adaptive Self-Speculative Decoding using Cosine Similarity

Abstract

AbstractWe present a simple on the fly method for faster inference of large language models. Unlike other (self-)speculative decoding techniques, our method does not require fine-tuning or black-box optimization to generate a fixed draft model, relying instead on simple rules to generate varying draft models adapted to the input context. We show empirically that our light-weight algorithm is competitive with the current SOTA for self-speculative decoding, while being a truly plug-and-play method.

🌉 Interdisciplinary Bridge — Artificial Intelligence and Deep Learning and Machine Learning

🐝 Cross-Pollinator — Artificial Intelligence, Computer Science, Computer Vision, Data Science & Analytics, Deep Learning, Healthcare & Medicine, Interdisciplinary, Knowledge & Reasoning, Machine Learning, Mathematics & Optimization, Natural Language Processing, Reinforcement Learning, Robotics, Security & Privacy, Speech & Audio

Authors

Michael R. Metel , Peng Lu , Boxing Chen , Mehdi Rezagholizadeh , Ivan Kobyzev

Topics

Artificial Intelligence > Core AI > Model Compression Machine Learning > Optimization & Theory > Neural Network Optimization Machine Learning > Application Areas > Efficient Computing Artificial Intelligence > Core AI > Large Language Models Deep Learning > Optimization & Theory > Efficient Computing

Keywords

inference acceleration speculative decoding model efficiency cosine similarity self-speculative decoding large language model

Download PDF

Related papers

EmbodiedBERT: Cognitively Informed Metaphor Detection Incorporating Sensorimotor Information 2024

Mitigating Matthew Effect: Multi-Hypergraph Boosted Multi-Interest Self-Supervised Learning for Conversational Recommendation 2024

Learning to Extract Structured Entities Using Language Models 2024

Towards Understanding Jailbreak Attacks in LLMs: A Representation Space Analysis 2024

CSSL: Contrastive Self-Supervised Learning for Dependency Parsing on Relatively Free Word Ordered and Morphologically Rich Low Resource Languages 2024