When will my ML Job finish? Toward providing Completion Time Estimates through Predictability-Centric Scheduling

Abdullah Bin Faisal; Noah Martin; Hafiz Mohsin Bashir; Swaminathan Lamelas; Fahad R. Dogar

2024 OSDI OSDI 2024

When will my ML Job finish? Toward providing Completion Time Estimates through Predictability-Centric Scheduling

Abstract

In this paper, we make a case for providing job completion time estimates to GPU cluster users, similar to providing the delivery date of a package or arrival time of a booked ride. Our analysis reveals that providing predictability can come at the expense of performance and fairness. Existing GPU schedulers optimize for extreme points in the trade-off space, making them either extremely unpredictable or impractical. To address this challenge, we present PCS, a new scheduling framework that aims to provide predictability while balancing other traditional objectives. The key idea behind PCS is to use Weighted-Fair-Queueing (WFQ) and find a suitable configuration of different WFQ parameters (e.g., queue weights) that meets specific goals for predictability. It uses a simulation-aided search strategy to efficiently discover WFQ configurations that lie around the Pareto front of the trade-off space between these objectives. We implement and evaluate PCS in the context of scheduling ML training workloads on GPUs. Our evaluation, on a small-scale GPU testbed and larger-scale simulations, shows that PCS can provide accurate completion time estimates while marginally compromising on performance and fairness.

❓ The Questioner

🧭 Keyword Pioneer — predictability-centric scheduling

🐝 Cross-Pollinator — Artificial Intelligence, Computer Science, Deep Learning, Machine Learning, Mathematics & Optimization, Reinforcement Learning

🌉 Interdisciplinary Bridge — Computer Science and Machine Learning and Mathematics & Optimization

Authors

Abdullah Bin Faisal , Noah Martin , Hafiz Mohsin Bashir , Swaminathan Lamelas , Fahad R. Dogar

Topics

Machine Learning > Optimization & Theory > Optimization Mathematics & Optimization > Optimization > Online Algorithms Computer Science > Systems > Operating Systems

Keywords

job completion time predictability-centric scheduling gpu cluster weighted fair queueing simulation-aided search ml training workload completion time estimate

Download PDF

Related papers

ServerlessLLM: Low-Latency Serverless Inference for Large Language Models 2024

SquirrelFS: using the Rust compiler to check file-system crash consistency 2024

μSlope: High Compression and Fast Search on Semi-Structured Logs 2024

Using Dynamically Layered Definite Releases for Verifying the RefFS File System 2024

DSig: Breaking the Barrier of Signatures in Data Centers 2024