Small-to-Large Generalization: Training Data Influences Models Consistently Across Scale

Alaa Khaddaj; Logan Engstrom; Aleksander Madry

2025 ICLR ICLR 2025

Small-to-Large Generalization: Training Data Influences Models Consistently Across Scale

Abstract

Choice of training data distribution greatly influences model behavior. Yet, in large-scale settings, precisely characterizing *how* changes in training data affects predictions is often difficult due to model training costs. Current practice is to instead extrapolate from scaled down, inexpensive-to-train proxy models. However, changes in data do not influence smaller and larger models identically. Therefore, understanding how choice of data affects large-scale models raises the question: how does training data distribution influence model behavior across compute scale? We find that small- and large-scale language model predictions (generally) *do* highly correlate across choice of training data. Equipped with these findings, we characterize how proxy scale affects effectiveness in two downstream proxy model applications: data attribution and dataset selection.

Authors

Alaa Khaddaj , Logan Engstrom , Aleksander Madry

Download PDF

Related papers

Gramian Multimodal Representation Learning and Alignment 2025

Separation Power of Equivariant Neural Networks 2025

What should a neuron aim for? Designing local objective functions based on information theory 2025

Regret-Optimal List Replicable Bandit Learning: Matching Upper and Lower Bounds 2025

CHASE-SQL: Multi-Path Reasoning and Preference Optimized Candidate Selection in Text-to-SQL 2025