Conformal Prediction and MLLM aided Uncertainty Quantification in Scene Graph Generation
Sayak Nag, Udita Ghosh, Calvin-Khang Ta, Sarosij Bose, Jiachen Li, Amit K. Roy-Chowdhury
摘要
Scene Graph Generation (SGG) aims to represent visual scenes by identifying objects and their pairwise relationships, providing a structured understanding of image content. However, inherent challenges like long-tailed class distributions and prediction variability necessitate uncertainty quantification in SGG for its practical viability. In this paper, we introduce a novel Conformal Prediction based framework, adaptive to any existing SGG method, for quantifying their predictive uncertainty by constructing well-calibrated prediction sets over their generated scene graphs. These scene graph prediction sets are designed to achieve statistically rigorous coverage guarantees under exchangeability assumptions. Additionally, to ensure the prediction sets contain the most practically interpretable scene graphs, we propose an effective MLLM-based postprocessing strategy for selecting the most visually and semantically plausible scene graphs within each set. We show that our proposed approach can produce diverse possible scene graphs from an image, assess the reliability of SGG methods, and improve overall SGG performance.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Drive My Way: Preference Alignment of Vision-Language-Action Model for Personalized DrivingZehao Wang, Huaide Jiang, Shuaiwu Dong, Yuping Wang 等CVPR 2026 · 被引用 7 次
- PSG-Nav: Probabilistic Scene Graph Navigation via Multiverse Decision MakingRufeng Chen, Yue Chang, Xiaqiang Tang, Hechang Chen 等ICML 2026 · 被引用 3 次
- CURVE: Learning Causality-Inspired Invariant Representations for Robust Scene Understanding via Uncertainty-Guided RegularizationYue Liang, JIATONG DU, Ziyi Yang, Yanjun Huang 等ICML 2026 · 被引用 2 次
它引用的顶会 Paper18
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
- Classification with Valid and Adaptive CoverageYaniv Romano, Matteo Sesia, Emmanuel J. CandèsNeurIPS 2020 · 被引用 586 次
- Conformal Time-series ForecastingKamile Stankeviciute, Ahmed M. Alaa, Mihaela van der SchaarNeurIPS 2021 · 被引用 233 次
相关 Paper
- Conformal Structured PredictionBotong Zhang, Shuo Li, Osbert BastaniICLR 2025
- Consistent Scene Graph Generation by Constraint OptimizationBoqi Chen, Kristóf Marussy, Sebastian Pilarski, Oszkár Semeráth 等ASE 2022 · 被引用 5 次
- Inference-Time Conformal Reasoning with Valid Factuality Control for Large Language ModelsTing Wang, Yuanjie Shi, Yan Yan, Huan ZhangICML 2026
- Iterative Scene Graph GenerationSiddhesh Khandelwal, Leonid SigalNeurIPS 2022 · 被引用 47 次
- Conformal Prediction Meets Long-tail ClassificationShuqi Liu, Jianguo Huang, Luke OngAAAI 2026
