Beyond a Single Extractor: Re-thinking HTML-to-Text Extraction for LLM Pre-training

Jeffrey Li; Joshua P Gardner; Doug Kang; Fangping Shi; Karanjeet Singh; Chun-Liang Li; Herumb Shandilya; David Leo Wright Hall; Oncel Tuzel; Percy Liang; Ludwig Schmidt; Hadi Pouransari; Fartash Faghri

2026 EACL EACL 2026

Beyond a Single Extractor: Re-thinking HTML-to-Text Extraction for LLM Pre-training

Abstract

AbstractOne of the first pre-processing steps for constructing web-scale LLM pretraining datasets involves extracting text from HTML. Despite the immense diversity of web content, existing open-source datasets predominantly apply a single fixed extractor to all webpages. In this work, we investigate whether this practice leads to suboptimal coverage and utilization of Internet data. We first show that while different extractors may lead to similar model performance on standard language understanding tasks, the pages surviving a fixed filtering pipeline can differ substantially. This suggests a simple intervention: by taking a Union over different extractors, we can increase the token yield of DCLM-Baseline by up to 71% while maintaining benchmark performance. We further show that for structured content such as tables and code blocks, extractor choice can significantly impact downstream task performance, with differences of up to 10 percentage points (p.p.) on WikiTQ and 3 p.p. on HumanEval.

🌉 Interdisciplinary Bridge — Machine Learning and Natural Language Processing

🧭 Keyword Pioneer — html extraction

🐝 Cross-Pollinator — Artificial Intelligence, Computer Science, Computer Vision, Data Science & Analytics, Deep Learning, Healthcare & Medicine, Interdisciplinary, Knowledge & Reasoning, Machine Learning, Mathematics & Optimization, Natural Language Processing, Reinforcement Learning, Robotics, Security & Privacy, Speech & Audio

Authors

Jeffrey Li , Joshua P Gardner , Doug Kang , Fangping Shi , Karanjeet Singh , Chun-Liang Li , Herumb Shandilya , David Leo Wright Hall , Oncel Tuzel , Percy Liang , Ludwig Schmidt , Hadi Pouransari , Fartash Faghri

Topics

Machine Learning > Application Areas > Efficient Computing Natural Language Processing > Resources & Methods > Large Language Models

Keywords

text extraction large language model pretraining datum html extraction token yield

Download PDF

Related papers

Investigating Gender Stereotypes in Large Language Models via Social Determinants of Health 2026

A Benchmark for Audio Reasoning Capabilities of Multimodal Large Language Models 2026

InfiGUIAgent: A Multimodal Generalist GUI Agent with Native Reasoning and Reflection 2026

Generative Personality Simulation via Theory-Informed Structured Interview 2026

Word Surprisal Correlates with Sentential Contradiction in LLMs 2026