SPOR: A Comprehensive and Practical Evaluation Method for Compositional Generalization in Data-to-Text Generation
Ziyao Xu, Houfeng Wang
摘要
Compositional generalization is an important ability of language models and has many different manifestations. For data-to-text generation, previous research on this ability is limited to a single manifestation called Systematicity and lacks consideration of large language models (LLMs), which cannot fully cover practical application scenarios. In this work, we propose SPOR, a comprehensive and practical evaluation method for compositional generalization in data-to-text generation. SPOR includes four aspects of manifestations (Systematicity, Productivity, Order invariance, and Rule learnability) and allows high-quality evaluation without additional manual annotations based on existing datasets. We demonstrate SPOR on two different datasets and evaluate some existing language models including LLMs. We find that the models are deficient in various aspects of the evaluation and need further improvement. Our work shows the necessity for comprehensive research on different manifestations of compositional generalization in data-to-text generation and provides a framework for evaluation. The dataset and code are available at https://github.com/xzy-xzy/SPOR .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Can Large Language Models Always Solve Easy Problems if They Can Solve Harder Ones?Zhe Yang, Yichang Zhang, Tianyu Liu, Jian Yang 等EMNLP 2024 · 被引用 2 次
- Paraphrasing as Zero-shot Translation with Feature-guided Diversity EnhancementZiyue Yan, Hongying Zan, Xinglin Lyu, Hongfei XuACL 2026
- Investigating More Explainable and Partition-Free Compositionality Estimation for LLMs: A Rule-Generation PerspectiveZiyao Xu, Cong Wang, Houfeng WangACL 2026
- From A and B to A+B: Can Large Language Models Solve Compositional Math Problems?Xisheng Xiao, Hanlin ZhaoEMNLP 2025
它引用的顶会 Paper6
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad 等ACL 2020 · 被引用 1,224 次
- COGS: A Compositional Generalization Challenge Based on Semantic InterpretationNajoung Kim, Tal LinzenEMNLP 2020 · 被引用 149 次
- Making Transformers Solve Compositional TasksSantiago Ontañón, Joshua Ainslie, Zachary Fisher, Vaclav CvicekACL 2022 · 被引用 87 次
- Improving Compositional Generalization with Self-Training for Data-to-Text GenerationSanket Vaibhav Mehta, Jinfeng Rao, Yi Tay, Mihir Kale 等ACL 2022 · 被引用 34 次
相关 Paper
- Benchmarking and Improving Compositional Generalization of Multi-aspect Controllable Text GenerationTianqi Zhong, Zhaoyi Li, Quan Wang, Linqi Song 等ACL 2024
- @ CREPE: Can Vision-Language Foundation Models Reason Compositionally?Zixian Ma, Jerry Hong, Mustafa Omer Gul, Mona Gandhi 等CVPR 2023
- The Paradox of the Compositionality of Natural Language: A Neural Machine Translation Case StudyVerna Dankers, Elia Bruni, Dieuwke HupkesACL 2022
- Compositional Generalization for Multi-Label Text Classification: A Data-Augmentation ApproachYuyang Chai, Zhuang Li, Jiahui Liu, Lei Chen 等AAAI 2024 · 被引用 18 次
- Can Models Learn Skill Composition from Examples?Haoyu Zhao, Simran Kaur, Dingli Yu, Anirudh Goyal 等NeurIPS 2024 · 被引用 20 次
