Do Vision and Text Cues Exhibit Evidential Coupling? UFO: A Benchmark for Compositional Multimodal Reasoning in Unified Models
Zhongyu Yang, Dannong Xu, Yonghan Zhang, Kefan Chen, Xinyi Wang, Yang Xu, Wei Pang, Yingfang Yuan
摘要
Unified Foundation Models (UFMs), which support interleaved multimodal generation and understanding, have been proposed as a promising paradigm for reasoning about dynamic world states, yet it remains unclear whether the visual content they generate functions as grounded evidence for subsequent reasoning or merely as auxiliary output. Existing benchmarks largely evaluate generation and understanding as separate capabilities and do not test their functional dependence during reasoning. We introduce UFO, a benchmark designed to evaluate whether UFMs generate and use image and text cues as evidence for compositional multimodal reasoning. UFO spans three state-transition regimes: state determination, state reconstruction, and state augmentation, which correspond to progressively smaller transformations of the underlying world state. Our analysis reveals a significant modality gap, as models often achieve high prediction accuracy even when the generated visual cues exert limited influence on their decisions, indicating weakened evidential coupling and a reliance on textual shortcuts rather than robust cross-modal grounding. Code and dataset can be found at https://01yzzyu.github.io/UFO/
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper24
- MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual ContextsPan Lu, Hritik Bansal, Tony Xia, Jiacheng Liu 等ICLR 2024 · 被引用 1,472 次
- OmniGen2: Towards Instruction-Aligned Multimodal GenerationChenyuan Wu, Jiahao Wang, Pengfei Zheng, Ruiran Yan 等CVPR 2026 · 被引用 231 次
- MMMU: A Massive Multi-Discipline Multimodal Understanding and Reasoning Benchmark for Expert AGIXiang Yue, Yuansheng Ni, Tianyu Zheng, Kai Zhang 等CVPR 2024 · 被引用 213 次
- LongVT: Incentivizing "Thinking with Long Videos" via Native Tool CallingZuhao Yang, Sudong Wang, Kaichen Zhang, Keming Wu 等CVPR 2026 · 被引用 63 次
- Uni-CoT: Towards Unified Chain-of-Thought Reasoning Across Text and VisionLuozheng Qin, Jia Gong, Yuqing Sun, Tianjiao Li 等ICLR 2026 · 被引用 55 次
相关 Paper
- Uni-MMMU: A Massive Multi-discipline Multimodal Unified BenchmarkKai Zou, Ziqi Huang, Yuhao Dong, Shulin Tian 等ACL 2026 · 被引用 19 次
- ROVER: Benchmarking Reciprocal Cross-Modal Reasoning for Omnimodal GenerationYongyuan Liang, Wei Chow, Feng Li, Ziqiao Ma 等ICLR 2026 · 被引用 13 次
- GGBench: A Geometric Generative Reasoning Benchmark for Unified Multimodal ModelsJingxuan Wei, Caijun Jia, Xi Bai, Xinglong Xu 等CVPR 2026 · 被引用 7 次
- GIR-Bench: Versatile Benchmark for Generating Images with ReasoningHongxiang Li, Yaowei Li, Bin Lin, Yuwei Niu 等ICLR 2026 · 被引用 15 次
- CRIT: Graph-Based Automatic Data Synthesis to Enhance Cross-Modal Multi-Hop ReasoningJunyoung Sung, Seungwoo Lyu, Minjun Kim, Sumin An 等CVPR 2026 · 被引用 2 次
