Distributed Graph Embedding with Information-Oriented Random Walks
Peng Fang, Arijit Khan, Siqiang Luo, Fang Wang, Dan Feng, Zhenli Li, Wei Yin, Yuchao Cao
摘要
Graph embedding maps graph nodes to low-dimensional vectors, and is widely adopted in machine learning tasks. The increasing availability of billion-edge graphs underscores the importance of learning efficient and effective embeddings on large graphs, such as link prediction on Twitter with over one billion edges. Most existing graph embedding methods fall short of reaching high data scalability. In this paper, we present a general-purpose, distributed, information-centric random walk-based graph embedding framework, DistGER, which can scale to embed billion-edge graphs. DistGER incrementally computes information-centric random walks. It further leverages a multi-proximity-aware, streaming, parallel graph partitioning strategy, simultaneously achieving high local partition quality and excellent workload balancing across machines. DistGER also improves the distributed Skip-Gram learning model to generate node embeddings by optimizing the access locality, CPU throughput, and synchronization efficiency. Experiments on real-world graphs demonstrate that compared to state-of-the-art distributed graph embedding frameworks, including KnightKing, DistDGL, and Pytorch-BigGraph, DistGER exhibits 2.33×--129× acceleration, 45% reduction in cross-machines communication, and >10% effectiveness improvement in downstream tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- LD2: Scalable Heterophilous Graph Neural Network with Decoupled EmbeddingsNingyi Liao, Siqiang Luo, Xiang Li, Jieming ShiNeurIPS 2023 · 被引用 23 次
- GENTI: GPU-powered Walk-based Subgraph Extraction for Scalable Representation Learning on Dynamic GraphsZihao Yu, Ningyi Liao, Siqiang LuoVLDB 2024 · 被引用 8 次
- TIGER: Training Inductive Graph Neural Network for Large-scale Knowledge Graph ReasoningKai Wang, Yuwei Xu, Siqiang LuoVLDB 2024 · 被引用 3 次
- OMeGa: Boosting Large-scale Graph Embeddings with Heterogeneous Memory ProcessingPeng Fang, Siqiang Luo, Fang Wang, Bolong Zheng 等ICDE 2025 · 被引用 1 次
- FeLoG: Scalable and Efficient Distributed Graph Embedding with Feedback Loop MechanismPeng Fang, Arijit Khan, Ziqiang Wu, Zhenli Li 等VLDB 2026
它引用的顶会 Paper13
- ByteGNN: Efficient Graph Neural Network Training at Large ScaleChenguang Zheng, Hongzhi Chen, Yuxuan Cheng, Zhezheng Song 等VLDB 2022 · 被引用 107 次
- Accelerating graph sampling for graph machine learning using GPUsAbhinav Jangda, Sandeep Polisetty, Arjun Guha, Marco SerafiniEuroSys 2021 · 被引用 79 次
- Homogeneous Network Embedding for Massive Graphs via Reweighted Personalized PageRankRenchi Yang, Jieming Shi, Xiaokui Xiao, Yin Yang 等VLDB 2020 · 被引用 77 次
- Marius: Learning Massive Graph Embeddings on a Single MachineJason Mohoney, Roger Waleffe, Henry Xu, Theodoros Rekatsinas 等OSDI 2021 · 被引用 75 次
- GraphWalker: An I/O-Efficient and Resource-Friendly Graph Analytic System for Fast and Scalable Random WalksRui Wang, Yongkun Li, Hong Xie, Yinlong Xu 等USENIX ATC 2020 · 被引用 64 次
相关 Paper
- Parallel Training of Knowledge Graph Embedding Models: A Comparison of TechniquesAdrian Kochsiek, Rainer GemullaVLDB 2022 · 被引用 33 次
- Scalable Robust Graph Embedding with SparkChi Thang Duong, Dung Hoang, Hongzhi Yin, Matthias Weidlich 等VLDB 2022 · 被引用 3 次
- Billion-Scale Bipartite Graph Embedding: A Global-Local Induced ApproachXueyi Wu, Yuanyuan Xu, Wenjie Zhang, Ying ZhangVLDB 2024 · 被引用 21 次
- DGL-KE: Training Knowledge Graph Embeddings at ScaleDa Zheng, Xiang Song, Chao Ma, Zeyuan Tan 等SIGIR 2020 · 被引用 132 次
- FlowWalker: A Memory-efficient and High-performance GPU-based Dynamic Graph Random Walk FrameworkJunyi Mei, Shixuan Sun, Chao Li, Cheng Xu 等VLDB 2024 · 被引用 10 次
