Declarative Techniques for NL Queries over Heterogeneous Data

Elham Khabiri; Jeffrey O. Kephart; Fenno F. Heath Iii; Srideepika Jayaraman; Yingjie Li; Fateh A. Tipu; Dhruv Shah; Achille Fokoue; Anu Bhamidipaty

2025 EMNLP EMNLP 2025

Declarative Techniques for NL Queries over Heterogeneous Data

Abstract

AbstractIn many industrial settings, users wish to ask questions in natural language, the answers to which require assembling information from diverse structured data sources. With the advent of Large Language Models (LLMs), applications can now translate natural language questions into a set of API calls or database calls, execute them, and combine the results into an appropriate natural language response. However, these applications remain impractical in realistic industrial settings because they do not cope with the data source heterogeneity that typifies such environments. In this work, we simulate the heterogeneity of real industry settings by introducing two extensions of the popular Spider benchmark dataset that require a combination of database and API calls. Then, we introduce a declarative approach to handling such data heterogeneityand demonstrate that it copes with data source heterogeneity significantly better than state-of-the-art LLM-based agentic or imperative code generation systems. Our augmented benchmarks are available to the research community.

🌉 Interdisciplinary Bridge — Artificial Intelligence and Computer Science and Natural Language Processing

🧭 Keyword Pioneer — database api integration

🐝 Cross-Pollinator — Artificial Intelligence, Computer Science, Computer Vision, Data Science & Analytics, Deep Learning, Healthcare & Medicine, Interdisciplinary, Knowledge & Reasoning, Machine Learning, Mathematics & Optimization, Natural Language Processing, Reinforcement Learning, Robotics, Security & Privacy, Speech & Audio

Authors

Elham Khabiri , Jeffrey O. Kephart , Fenno F. Heath Iii , Srideepika Jayaraman , Yingjie Li , Fateh A. Tipu , Dhruv Shah , Achille Fokoue , Anu Bhamidipaty

Topics

Artificial Intelligence > Core AI > Agent Systems Artificial Intelligence > Core AI > Multi-Agent Systems Artificial Intelligence > Core AI > Planning Natural Language Processing > Applications > Information Retrieval Natural Language Processing > Applications > Machine Translation Natural Language Processing > Applications > Question Answering Computer Science > Systems > Databases Artificial Intelligence > Core AI > Large Language Models Computer Science > Applications > Databases Artificial Intelligence > Core AI > Natural Language Processing

Keywords

data integration semantic parsing heterogeneous datum declarative programming query planning llm-based agent natural language query large language model api call api integration declarative approach database api integration

Download PDF

Related papers

Bit-Flip Error Resilience in LLMs: A Comprehensive Analysis and Defense Framework 2025

VoiceCraft-X: Unifying Multilingual, Voice-Cloning Speech Synthesis and Speech Editing 2025

Model-based Large Language Model Customization as Service 2025

ZoomEye: Enhancing Multimodal LLMs with Human-Like Zooming Capabilities through Tree-Based Image Exploration 2025

SlideCoder: Layout-aware RAG-enhanced Hierarchical Slide Generation from Design 2025