Live or Lie: Action-Aware Capsule Multiple Instance Learning for Risk Assessment in Live Streaming Platforms
Yiran Qiao, Jing Chen, Xiang Ao, Qiwei Zhong, Yang Liu, Qing He
Abstract
Live streaming has become a cornerstone of today's internet, enabling massive real-time social interactions. However, it faces severe risks arising from sparse, coordinated malicious behaviors among multiple participants, which are often concealed within normal activities and challenging to detect timely and accurately. In this work, we provide a pioneering study on risk assessment in live streaming rooms, characterized by weak supervision where only room-level labels are available. We formulate the task as a Multiple Instance Learning (MIL) problem, treating each room as a bag and defining structured user-timeslot capsules as instances. These capsules represent subsequences of user actions within specific time windows, encapsulating localized behavioral patterns. Based on this formulation, we propose AC-MIL, an Action-aware Capsule MIL framework that models both individual behaviors and group-level coordination patterns. AC-MIL captures multi-granular semantics and behavioral cues through a serial and parallel architecture that jointly encodes temporal dynamics and cross-user dependencies. These signals are integrated for robust room-level risk prediction, while also offering interpretable evidence at the behavior segment level. Extensive experiments on large-scale industrial datasets from Douyin demonstrate that AC-MIL significantly outperforms MIL and sequential baselines, establishing new state-of-the-art performance in room-level risk assessment for live streaming. Moreover, AC-MIL provides capsule-level interpretability, enabling identification of risky behavior segments as actionable evidence for intervention. The project page is available at: https://qiaoyran.github.io/AC-MIL/.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a0e5e5eb-c3e8-41e5-aaa9-7a038b55e564Cited by top-tier papers2
- Deja Vu in Plots: Leveraging Cross-Session Evidence with Retrieval-Augmented LLMs for Live Streaming Risk AssessmentYiran Qiao, Xiang Ao, Jing Chen, Yang Liu et al.SIGIR 2026 · 2 citations
- Outsmarting the Chameleon: Counterfactual Decoupling for Tactical OOD Shifts in Live Streaming Risk AssessmentYiran Qiao, Jing Chen, Jiaqi Xu, Yang Liu et al.KDD 2026 · 1 citation
Builds on11
- Informer: Beyond Efficient Transformer for Long Sequence Time-Series ForecastingHaoyi Zhou, Shanghang Zhang, Jieqi Peng, Shuai Zhang et al.AAAI 2021 · 7,289 citations
- Reformer: The Efficient TransformerNikita Kitaev, Lukasz Kaiser, Anselm LevskayaICLR 2020 · 2,878 citations
- Pick and Choose: A GNN-based Imbalanced Learning Approach for Fraud DetectionYang Liu, Xiang Ao, Zidi Qin, Jianfeng Chi et al.WWW 2021 · 527 citations
- H2-FDetector: A GNN-based Fraud Detector with Homophilic and Heterophilic ConnectionsFengzhao Shi, Yanan Cao, Yanmin Shang, Yuchen Zhou et al.WWW 2022 · 149 citations
- Additive MIL: Intrinsically Interpretable Multiple Instance Learning for PathologySyed Ashar Javed, Dinkar Juyal, Harshith Padigela, Amaro Taylor-Weiner et al.NeurIPS 2022 · 124 citations
Related papers
- TAIL-MIL: Time-Aware and Instance-Learnable Multiple Instance Learning for Multivariate Time Series Anomaly DetectionJaeseok Jang, Hyuk-Yoon KwonAAAI 2025 · 5 citations
- Room Matters: Dynamic Room-level Collaboration Information Modeling for Live Streaming RecommendationKe Guo, Changle Qu, Xiao Zhang, Liqin Zhao et al.WWW 2026
- VideoModerator: A Risk-aware Framework for Multimodal Video Moderation in E-CommerceTan Tang, Yanhong Wu, Yingcai Wu, Lingyun Yu et al.IEEE VIS 2021 · 31 citations
- ASM-Loc: Action-aware Segment Modeling for Weakly-Supervised Temporal Action LocalizationBo He, Xitong Yang, Le Kang, Zhiyu Cheng et al.CVPR 2022 · 104 citations
- Learning from Noisy Supervision: A Denoising-Debiasing Framework for Weakly Supervised Video Anomaly DetectionYaxin Zhao, Yang Wang, Wenya Guo, Sihan Xu et al.CVPR 2026
