Scaling Choice Models of Relational Social Data
Jan Overgoor, George Pakapol Supaniratisai, Johan Ugander
Abstract
Many prediction problems on social networks, from recommendations to anomaly detection, can be approached by modeling network data as a sequence of relational events and then leveraging the resulting model for prediction. Conditional logit models of discrete choice are a natural approach to modeling relational events as "choices" in a framework that envelops and extends many long-studied models of network formation. The conditional logit model is simplistic, but it is particularly attractive because it allows for efficient consistent likelihood maximization via negative sampling, something that isn't true for mixed logit and many other richer models. The value of negative sampling is particularly pronounced because choice sets in relational data are often enormous. Given the importance of negative sampling, in this work we introduce a model simplification technique for mixed logit models that we call "de-mixing", whereby standard mixture models of network formation---particularly models that mix local and global link formation---are reformulated to operate their modes over disjoint choice sets. This reformulation reduces mixed logit models to conditional logit models, opening the door to negative sampling while also circumventing other standard challenges with maximizing mixture model likelihoods. To further improve scalability, we also study importance sampling for more efficiently selecting negative samples, finding that it can greatly speed up inference in both standard and de-mixed models. Together, these steps make it possible to much more realistically model network formation in very large graphs. We illustrate the relative gains of our improvements on synthetic datasets with known ground truth as well as a large-scale dataset of public transactions on the Venmo platform.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itRelated papers
- Domain-Informed Negative Sampling Strategies for Dynamic Graph Embedding in Meme Stock-Related Social NetworksYunming Hui, Inez Maria Zwetsloot, Simon Trimborn, Stevan RudinacWWW 2025 · 2 citations
- Reconciling Competing Sampling Strategies of Network EmbeddingYuchen Yan, Baoyu Jing, Lihui Liu, Ruijie Wang et al.NeurIPS 2023 · 34 citations
- Diffusion-based Negative Sampling on Graphs for Link PredictionTrung-Kien Nguyen, Yuan FangWWW 2024 · 27 citations
- SCE: Scalable Network Embedding from Sparsest CutShengzhong Zhang, Zengfeng Huang, Haicang Zhou, Ziang ZhouKDD 2020 · 9 citations
- HiLoMix: Robust High- and Low-Frequency Graph Learning Framework for Mixing Address AssociationXiaofan Tu, Tiantian Duan, Shuyi Miao, Hanwen Zhang et al.AAAI 2026 · 1 citation
