Modeling Long-Tail Relations in the Operating Room via In-Context Multimodal Learning
Boqiang Xu, Wei Zhang, Ding Ma, Jian Liang, Zhenan Sun, Zhen Lei
Abstract
Operating room (OR) scene graph generation (SGG) enables holistic modeling of OR domains by encoding interactions among medical staff, tools, and equipment as triplet-based structured scene graphs. Although existing OR SGG methods demonstrate satisfactory overall performance, they exhibit substantially lower accuracy on long-tail categories compared to head categories in OR data. We introduce SGG-ICL, a novel framework that represents the first attempt to address the long-tail problem in OR SGG by leveraging in-context learning (ICL). SGG-ICL first identifies long-tail samples via an Adaptive Router module and selectively applies ICL only to these samples. This selective routing strategy enhances performance on long-tail categories without degrading head-category accuracy. Subsequently, SGG-ICL constructs a candidate pool through multimodal retrieval and then employs a trained MLLM Reranker to re-rank the candidates, selecting the most similar examples to the test sample for ICL. The reranker is supervised by IoU scores derived from annotated SGG triplets and exploits rich multimodal information to estimate pairwise sample similarity. Experimental results show that SGG-ICL improves accuracy on long-tail categories by 6.9%, while also achieving a 2.6% improvement in overall accuracy.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext defe34ab-e647-48f6-9c3d-4e657ffe2d15Builds on11
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- Robust Speech Recognition via Large-Scale Weak SupervisionAlec Radford, Jong Wook Kim, Tao Xu, Greg Brockman et al.ICML 2023 · 6,966 citations
Related papers
- Scene Graph Generation Strategy with Co-occurrence Knowledge and Learnable Term FrequencyHyeongjin Kim, Sangwon Kim, Dasom Ahn, Jong Taek Lee et al.ICML 2024 · 8 citations
- Learning to Generate Structured Meshes with In-Context: Toward Generalization in Mesh GenerationJing Xiao, Xinhai Chen, Jiaming Peng, Jie LiuAAAI 2026
- VideoStir: Understanding Long Videos via Spatio-Temporally Structured and Intent-Aware RAGHonghao Fu, Miao Xu, Yiwei Wang, Dailing Zhang et al.ACL 2026 · 2 citations
- Adaptive Self-training Framework for Fine-grained Scene Graph GenerationKibum Kim, Kanghoon Yoon, Yeonjun In, Jinyoung Moon et al.ICLR 2024 · 14 citations
- Fast Contextual Scene Graph Generation with Unbiased Context AugmentationTianlei Jin, Fangtai Guo, Qiwei Meng, Shiqiang Zhu et al.CVPR 2023
