CSPO: Alleviating Reward Ambiguity for Structured Table-to-LaTeX Generation
Yunfan Yang, Cuiling Lan, Jitao Sang, Yan Lu
摘要
Tables contain rich structured information, yet when stored as images their contents remain "locked" within pixels. Converting table images into LaTeX code enables faithful digitization and reuse, but current multimodal large language models (MLLMs) often fail to preserve structural, style, or content fidelity. Conventional post-training with reinforcement learning (RL) typically relies on a single aggregated reward, leading to reward ambiguity that conflates multiple behavioral aspects and hinders effective optimization. We propose Component-Specific Policy Optimization (CSPO), an RL framework that disentangles optimization across LaTeX tables components-structure, style, and content. In particular, CSPO assigns component-specific rewards and backpropagates each signal only through the tokens relevant to its component, alleviating reward ambiguity and enabling targeted component-wise optimization. To comprehensively assess performance, we introduce a set of hierarchical evaluation metrics. Extensive experiments demonstrate the effectiveness of CSPO, underscoring the importance of component-specific optimization for reliable structured generation.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper3
- Pix2Struct: Screenshot Parsing as Pretraining for Visual Language UnderstandingKenton Lee, Mandar Joshi, Iulia Raluca Turc, Hexiang Hu 等ICML 2023 · 被引用 426 次
- Nougat: Neural Optical Understanding for Academic DocumentsLukas Blecher, Guillem Cucurull, Thomas Scialom, Robert StojnicICLR 2024 · 被引用 243 次
- LATTE: Improving Latex Recognition for Tables and Formulae with Iterative RefinementNan Jiang, Shanchao Liang, Chengxiao Wang, Jiannan Wang 等AAAI 2025 · 被引用 8 次
相关 Paper
- Table2LaTeX-RL: High-Fidelity LaTeX Code Generation from Table Images via Reinforced Multimodal Language ModelsJun Ling, Yao Qi, Tao Huang, Shibo Zhou 等NeurIPS 2025 · 被引用 9 次
- COPO: Causal-Oriented Policy Optimization for Hallucinations of MLLMsPeizheng Guo, Jingyao Wang, Wenwen Qiang, Jiahuan Zhou 等CVPR 2026 · 被引用 1 次
- Dissecting Post-Training: Uncovering the Complementary Roles of SFT and RL for Document ParsingJun-Peng Jiang, An-Yang Ji, Shiyin Lu, Guodong Zheng 等ICML 2026
- Teaching Models to Improve on TapeLiat Bezalel, Eyal Orgad, Amir GlobersonAAAI 2025
- GRACE: Generative Representation Learning via Contrastive Policy OptimizationJiashuo Sun, Shixuan Liu, Zhaochen Su, Xianrui Zhong 等ICLR 2026 · 被引用 7 次
