Long-Document QA with Chain-of-Structured-Thought and Fine-Tuned SLMs
Zhuowen Liang, Xiaotian Lin, Zhengxuan Zhang, Yuyu Luo, Haixun Wang, Nan Tang
摘要
Large language models (LLMs) are widely applied to data analytics over documents, yet direct reasoning over long, noisy documents remains brittle and error-prone. Hence, we study document question answering (QA) that consolidates dispersed evidence into a structured output (e.g., a table, graph, or chunks) to support reliable, verifiable QA. We propose a two-pillar framework, LiteCoST, to achieve both high accuracy and low latency with small language models (SLMs). Pillar 1: Chain-of-Structured-Thought (CoST). We introduce a CoST template, a schema-aware instruction that guides a strong LLM to produce both a step-wise CoST trace and the corresponding structured output. The process induces a minimal structure, normalizes entities/units, aligns records, serializes the output, and verifies/refines it, yielding auditable supervision. Pillar 2: SLM fine-tuning. The compact models are trained on LLM-generated CoST data in two stages: Supervised Fine-Tuning for structural alignment, followed by Group Relative Policy Optimization (GRPO) incorporating triple rewards for answer/format quality and process consistency. By distilling structure-first behavior into SLMs, this approach achieves LLM-comparable quality on multi-domain long-document QA using 3B/7B SLMs, while delivering 2–4× lower latency than GPT-4o and DeepSeek-R1 (671B). The code is available at https://github.com/HKUSTDial/LiteCoST.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- VisJudge-Bench: Aesthetics and Quality Assessment of VisualizationsYupeng Xie, Zhiyang Zhang, Yifan Wu, Sirong Lu 等ICLR 2026 · 被引用 23 次
- Document-to-Database: Extraction Meets Relational SemanticsZhengxuan Zhang, Zhuowen Liang, Jiazhuo Chen, Haixun Wang 等VLDB 2026 · 被引用 3 次
它引用的顶会 Paper12
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni 等NeurIPS 2020 · 被引用 19,162 次
- The Dawn of Natural Language to SQL: Are We Fully Ready? [Experiment, Analysis & Benchmark ]Boyan Li, Yuyu Luo, Chengliang Chai, Guoliang Li 等VLDB 2024 · 被引用 137 次
- LongBench: A Bilingual, Multitask Benchmark for Long Context UnderstandingYushi Bai, Xin Lv, Jiajie Zhang, Hongchang Lyu 等ACL 2024 · 被引用 94 次
- DeepEye-SQL: A Software-Engineering-Inspired Text-to-SQL FrameworkBoyan Li, Chong Chen, Zhujun Xue, Yinan Mei 等SIGMOD 2026 · 被引用 40 次
- LEAD: Iterative Data Selection for Efficient LLM Instruction TuningXiaotian Lin, Yanlin Qi, Yizhang Zhu, Themis Palpanas 等VLDB 2026 · 被引用 16 次
相关 Paper
- Replacing Multi-Step Assembly of Data Preparation Pipelines with One-Step LLM Pipeline Generation for Table QAFengyu Li, Junhao Zhu, Kaishi Song, Lu Chen 等VLDB 2026 · 被引用 1 次
- Making Slow Thinking Faster: Compressing LLM Chain-of-Thought via Step EntropyZeju Li, Jianyuan Zhong, Ziyang Zheng, Xiangyu Wen 等ICLR 2026 · 被引用 35 次
- Boosting Small Language Models for Text-to-SQL with Fine-Grained Execution Feedback and Cost-Efficient RewardsThanh Dat Hoang, Thanh Trung Huynh, Matthias Weidlich, Thanh Tam Nguyen 等ICDE 2026 · 被引用 3 次
- Thinking in Latent Space: Progressive Multimodal Simplification for Visual ReasoningYuesen Tang, Yiming Yang, Tengfei Bao, Yu TongICML 2026
- Rethinking LLM Reasoning: From Explicit Trajectories to Latent RepresentationsCong Jiang, Xiaofeng Zhang, Fangzhi Zhu, XiaoWei Chen 等ICLR 2026
