UniNet: Scalable Network Representation Learning with Metropolis-Hastings Sampling
Xingyu Yao, Yingxia Shao, Bin Cui, Lei Chen
摘要
Network representation learning (NRL) technique has been successfully adopted in various data mining and machine learning applications. Random walk based NRL is one popular paradigm, which uses a set of random walks to capture the network structural information, and then employs word2vec models to learn the low-dimensional representations. However, until now there is lack of a framework, which unifies existing random walk based NRL models and supports to efficiently learn from large networks. The main obstacle comes from the diverse random walk models and the inefficient sampling method for the random walk generation. In this paper, we first introduce a new and efficient edge sampler based on Metropolis-Hastings sampling technique, and theoretically show the convergence property of the edge sampler to arbitrary discrete probability distributions. Then we propose a random walk model abstraction, in which users can easily define different transition probability by specifying dynamic edge weights and random walk states. The abstraction is efficiently supported by our edge sampler, since our sampler can draw samples from unnormalized probability distribution in constant time complexity. Finally, with the new edge sampler and random walk model abstraction, we carefully implement a scalable NRL framework called UniNet. We conduct comprehensive experiments with five random walk based NRL models over eleven real-world datasets, and the results clearly demonstrate the efficiency of UniNet over billion-edge networks. The code of UniNet is released at: https://github.com/shaoyx/UniNet.
1 For the embedding learning step, in NLP or machine learning community, researchers have proposed optimization techniques to improve the efficiency of training word embedding models [13], [27]
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- xGCN: An Extreme Graph Convolutional Network for Large-scale Social Link PredictionXiran Song, Jianxun Lian, Hong Huang, Zihan Luo 等WWW 2023 · 被引用 33 次
- An I/O-Efficient Disk-based Graph System for Scalable Second-Order Random Walk of Large GraphsHongzheng Li, Yingxia Shao, Junping Du, Bin Cui 等VLDB 2022 · 被引用 19 次
- Space-Efficient Random Walks on Streaming GraphsSerafeim Papadias, Zoi Kaoudi, Jorge-Arnulfo Quiané-Ruiz, Volker MarklVLDB 2023 · 被引用 8 次
它引用的顶会 Paper1
相关 Paper
- Accelerating graph sampling for graph machine learning using GPUsAbhinav Jangda, Sandeep Polisetty, Arjun Guha, Marco SerafiniEuroSys 2021 · 被引用 79 次
- SCE: Scalable Network Embedding from Sparsest CutShengzhong Zhang, Zengfeng Huang, Haicang Zhou, Ziang ZhouKDD 2020 · 被引用 9 次
- Fast Unsupervised Graph Embedding via Graph Zoom LearningZiyang Liu, Chaokun Wang, Yunkai Lou, Hao FengICDE 2023 · 被引用 6 次
- TGL: A General Framework for Temporal GNN Training onBillion-Scale GraphsHongkuan Zhou, Da Zheng, Israt Nisa, Vassilis N. Ioannidis 等VLDB 2022 · 被引用 109 次
- Distributed Graph Embedding with Information-Oriented Random WalksPeng Fang, Arijit Khan, Siqiang Luo, Fang Wang 等VLDB 2023 · 被引用 18 次
