Segment Any 3D Object with Language
Seungjun Lee, Yuyang Zhao, Gim Hee Lee
2025Year
7Top-tier citations
Abstract
edu.sg https://cvrp-sole.github.io sink (a) "Can I wash my hands?" (b) "Brown Furnitures." laptop (c) "Device to play game." Figure 1: Qualitative results of SOLE with various language instructions. SOLE is highly generalizable and can segment corresponding instances with various language instructions, including but not limited to (a) visual questions, (b) attributes description, and (c) functional description.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers7
- A Unified Framework for 3D Scene UnderstandingWei Xu, Chunsheng Shi, Sifan Tu, Xin Zhou et al.NeurIPS 2024 · 25 citations
- D3D-VLP: Dynamic 3D Vision-Language-Planning Model for Embodied Grounding and NavigationZihan Wang, Seungjun Lee, Guangzhao Dai, Gim Hee LeeCVPR 2026 · 9 citations
- Segment Any Events with LanguageSeungjun Lee, Gim Hee LeeICLR 2026 · 3 citations
- EmbodiedSplat: Online Feed-Forward Semantic 3DGS for Open-Vocabulary 3D Scene UnderstandingSeungjun Lee, Zihan Wang, Yunsong Wang, Gim Hee LeeCVPR 2026 · 2 citations
- MV3DIS: Multi-View Mask Matching via 3D Guides for Zero-Shot 3D Instance SegmentationYibo Zhao, Yigong Zhang, Jin XieCVPR 2026 · 1 citation
Builds on34
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech et al.NeurIPS 2022 · 6,707 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
- Deep Hough Voting for 3D Object Detection in Point CloudsCharles R. Qi, Or Litany, Kaiming He, Leonidas J. GuibasICCV 2019 · 1,467 citations
- Open-vocabulary Object Detection via Vision and Language Knowledge DistillationXiuye Gu, Tsung-Yi Lin, Weicheng Kuo, Yin CuiICLR 2022 · 1,274 citations
Related papers
- SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual GroundingRong Li, Shijie Li, Lingdong Kong, Xulei Yang et al.CVPR 2025
- Intermediate Connectors and Geometric Priors for Language-Guided Affordance Segmentation on Unseen Object CategoriesYicong Li, Yiyang Chen, Zhenyuan Ma, Junbin Xiao et al.ICCV 2025 · 3 citations
- PanoGS: Gaussian-based Panoptic Segmentation for 3D Open Vocabulary Scene UnderstandingHongjia Zhai, Hai Li, Zhenzhe Li, Xiaokun Pan et al.CVPR 2025
- Conversational Image Segmentation: Grounding Abstract Concepts with Scalable SupervisionAadarsh Sahoo, Georgia GkioxariCVPR 2026 · 3 citations
- GLIGEN: Open-Set Grounded Text-to-Image GenerationYuheng Li, Haotian Liu, Qingyang Wu, Fangzhou Mu et al.CVPR 2023
