An Efficient Algorithm for Distance-based Structural Graph Clustering
Kaixin Liu, Sibo Wang, Yong Zhang, Chunxiao Xing
Abstract
Structural graph clustering (SCAN) is a classic graph clustering algorithm. In SCAN, a key step is to compute the structural similarity between vertices according to the overlap ratio of one-hop neighborhoods. Given two vertices u and v, existing studies only consider the case when u and v are neighbors. However, the structural similarity between non-neighboring vertices in SCAN is always zero, and using only one-hop neighbors on weighted graphs discards the weights on each edge. Both may not reflect the true closeness of two vertices and may fail to return high-quality clustering results. To tackle this issue, we define and study the distance-based structural graph clustering problem. Given a distance threshold d and two vertices u and v, the structural similarity between u and v is defined as the ratio of their respective neighbors within a distance of no more than d. We show that the newly defined distance-based SCAN achieves better clustering results compared to the vanilla version of SCAN. However, the new definition brings challenges in the computation of final clustering results. To tackle this efficiency issue, we propose DistanceSCAN, an efficient approximate algorithm for solving the distance-based SCAN problem. The main idea of DistanceSCAN is to use all-distances bottom-k sketches (ADS) to speed up the computation of similarities. Given the ADS, we can derive the similarity between two vertices with a bounded cost of O(k). However, to ensure that the estimated similarity has an approximation guarantee, the value of k still needs to be set to as large as thousands. This brings high computational costs when computing the similarities between neighboring vertices. To tackle this issue, we further construct histograms to prune the structural similarity computations of vertices pairs. Extensive experiments on real datasets validate the effectiveness and efficiency of DistanceSCAN.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 208c40cb-8eff-43f9-a781-d930a2adb7c6Cited by top-tier papers3
- Breaking the Entanglement of Homophily and Heterophily in Semi-supervised Node ClassificationHenan Sun, Xunkai Li, Zhengyu Wu, Daohan Su et al.ICDE 2024 · 9 citations
- Searching and Detecting Structurally Similar Communities in Large Heterogeneous Information NetworksShu Wang, Yixiang Fang, Wensheng LuoVLDB 2025 · 3 citations
- Efficient Structural Clustering Over HypergraphsDong Pan, Xu Zhou, Lingwei Li, Quanqing Xu et al.ICDE 2025
Related papers
- Effective Indexing for Dynamic Structural Graph ClusteringFangyuan Zhang, Sibo WangVLDB 2022 · 18 citations
- Index-based Structural Clustering on Directed GraphsLingkai Meng, Long Yuan, Zi Chen, Xuemin Lin et al.ICDE 2022 · 21 citations
- HINSCAN: Efficient Structural Graph Clustering Over Heterogeneous Information NetworksLong Yuan, Xiaotong Sun, Zi Chen, Peng Cheng et al.ICDE 2025 · 4 citations
- Efficient and Accurate SimRank-based Similarity Joins: Experiments, Analysis, and ImprovementQian Ge, Yu Liu, Yinghao Zhao, Yuetian Sun et al.VLDB 2024 · 4 citations
- S^3AND: Efficient Subgraph Similarity Search Under Aggregated Neighbor Difference SemanticsQi Wen, Yutong Ye, Xiang Lian, Mingsong ChenVLDB 2025 · 3 citations
