ESTJ: Enhancing Structured Tendency Judgment in Hybrid-Modal Table Understanding
Shu-Xun Yang, Xian-Ling Mao, Heyan Huang
Abstract
Hybrid-modal table understanding (HMTU), which targets leveraging multi-modal table evidence for multi-hop reasoning, has garnered widespread attention. Existing models primarily focus on effectively integrating multi-modal table evidence to enhance the table understanding capabilities of multi-modal large language models (MLLMs). However, these models ignore the fact that different types of table understanding questions lean toward different modalities of table evidence. Consequently, these models suffer from low utilization efficiency and poor interpretability. To address these issues, in this paper, we propose a modality preference alignment model, called ESTJ, which Enhances Structured Tendency Judgment in HMTU. Specifically, ESTJ first samples modality preference data from the responses generated by MLLMs. Then, it alleviates modality preference imbalance by adhering to the principle of least modality priority. Finally, ESTJ performs direct preference optimization (DPO) training based on structured tendency judgment to align modality preference effectively. Experimental results on TableQA and TableFV tasks demonstrate that our proposed model outperforms state-of-the-art baselines. Additionally, these results present fascinating phenomena and unveil profound insights into modality preference for table understanding.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 6550a4da-7e98-4907-9a4f-9094530dc0abRelated papers
- Table as a Modality for Large Language ModelsLiyao Li, Chao Ye, Wentao Ye, Yifei Sun et al.NeurIPS 2025 · 5 citations
- Multimodal Table UnderstandingMingyu Zheng, Xinwei Feng, Qingyi Si, Qiaoqiao She et al.ACL 2024
- MoD-DPO: Towards Mitigating Cross-modal Hallucinations in Omni LLMs using Modality Decoupled Preference OptimizationAshutosh Chaubey, Jiacheng Pang, Mohammad SoleymaniCVPR 2026 · 7 citations
- Beyond Fixed Biases: Decoding the Role of Reasoning Uncertainty in MLLM Modality ConflictsZhuoran Zhang, Tengyue Wang, Xilin Gong, Yang Shi et al.ICML 2026
- Rethinking Pragmatics in Large Language Models: Towards Open-Ended Evaluation and Preference TuningShengguang Wu, Shusheng Yang, Zhenglun Chen, Qi SuEMNLP 2024 · 2 citations
