3D-Aware Multi-Task Learning with Cross-View Correlations for Dense Scene Understanding
Xiaoye Wang, Chen Tang, Xiangyu Yue, Wei-Hong Li
摘要
This paper addresses the challenge of training a single network to jointly perform multiple dense prediction tasks, such as segmentation and depth estimation, i.e., multi-task learning (MTL). Current approaches mainly capture cross-task relations in the 2D image space, often leading to unstructured features lacking 3D-awareness. We argue that 3D-awareness is vital for modeling cross-task correlations essential for comprehensive scene understanding. We propose to address this problem by integrating correlations across views, i.e., cost volume, as geometric consistency in the MTL network. Specifically, we introduce a lightweight Cross-view Module (CvM), shared across tasks, to exchange information across views and capture cross-view correlations, integrated with a feature from MTL encoder for multi-task predictions. This module is architecture-agnostic and can be applied to both single and multi-view data. Extensive results on NYUv2 and PASCAL-Context demonstrate that our method effectively injects geometric consistency into existing MTL methods to improve performance.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- PRISM: Synergizing Vision Foundation Models via Self-organized Expert SpecializationYing Tang, Dong Li, Youjia Zhang, Zikai Song 等ICML 2026
- MTPano: Multi-Task Panoramic Scene Understanding via Label-Free Integration of Dense Prediction PriorsJingdong Zhang, Xiaohang Zhan, Lingzhi Zhang, Yizhou Wang 等SIGGRAPH 2026
它引用的顶会 Paper34
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- 3D Gaussian Splatting for Real-Time Radiance Field RenderingBernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, George DrettakisSIGGRAPH 2023 · 被引用 5,687 次
- Gradient Surgery for Multi-Task LearningTianhe Yu, Saurabh Kumar, Abhishek Gupta, Sergey Levine 等NeurIPS 2020 · 被引用 2,261 次
- Conflict-Averse Gradient Descent for Multi-task learningBo Liu, Xingchao Liu, Xiaojie Jin, Peter Stone 等NeurIPS 2021 · 被引用 686 次
相关 Paper
- Multi-task Learning with 3D-Aware RegularizationWei-Hong Li, Steven McDonagh, Ales Leonardis, Hakan BilenICLR 2024 · 被引用 10 次
- Learning to Fuse Monocular and Multi-view Cues for Multi-frame Depth Estimation in Dynamic ScenesRui Li, Dong Gong, Wei Yin, Hao Chen 等CVPR 2023
- Point-Based Multi-View Stereo NetworkRui Chen, Songfang Han, Jing Xu, Hao SuICCV 2019 · 被引用 403 次
- Contrastive Multi-Task Dense PredictionSiwei Yang, Hanrong Ye, Dan XuAAAI 2023 · 被引用 13 次
- Going Beyond Multi-Task Dense Prediction with Synergy Embedding ModelsHuimin Huang, Yawen Huang, Lanfen Lin, Ruofeng Tong 等CVPR 2024
