Lune

ICDE2026Top-tier venue

Accurate Table Question Answering with Accessible LLMs

Yangfan Jiang, Fei Wei, Ergute Bao, Yaliang Li, Bolin Ding, Yin Yang, Xiaokui Xiao

2026Year
1Citations

Abstract

Given a table T in a database and a question Q expressed in natural language, the table question answering (TQA) task aims to return an accurate answer to Q based on the content of T . The current state-of-the-art solutions leverage large language models (LLMs) to obtain high-quality answers. Most of these solutions, however, rely on proprietary, large-scale LLMs that require costly API access, which can be a significant financial barrier for many users. This paper focuses on TQA with smaller, open-weight LLMs that can run on a desktop or even a laptop. This is challenging since such LLMs typically have weaker capabilities compared to large proprietary ones, e.g., in terms of context understanding and instruction following. As a result, existing solutions suffer from a substantial performance drop when paired with such accessible LLMs.

We observe that one main reason for the poor performance of existing solutions with small open-weight LLMs is that these methods tend to ask the LLM to perform a highly sophisticated task with a long, complex prompt, which is often beyond the capabilities of such LLMs. Motivated by this, we present Orchestra, a multi-agent approach designed to unlock the potential of accessible LLMs to enable high-quality, cost-effective TQA for a broader audience. The main idea of Orchestra is to carefully coordinate a group of LLM agents, each performing a relatively simple task, through a structured, layered workflow to solve complicated TQA tasks -akin to the coordination of an orchestra. This approach effectively reduces the complexity of prompts faced by each LLM agent, significantly improving the reliability of their outputs.

We have implemented Orchestra on top of AgentScope, an open-source multi-agent framework, and evaluated it across multiple TQA benchmarks using a wide range of open-weight LLMs. Experimental results demonstrate that Orchestra achieves strong performance even with small-to medium-sized LLMs. For instance, when paired with Qwen2.5-14B, Orchestra attains a test accuracy of 72.1% on the WikiTQ benchmark, which approaches the best prior result 75.3% achieved with GPT-4; meanwhile, when paired with a larger Qwen / Llama / DeepSeek model, Orchestra beats all previous solutions and establishes new state-of-the-art results on all benchmarks in our experiments.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext d02dd90c-f4c2-48f8-a878-3081e9f79c7c

Builds on40

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines