Edge sampling and graph parameter estimation via vertex neighborhood accesses
Jakub Tetek, Mikkel Thorup
摘要
In this paper, we consider the problems from the area of sublineartime algorithms of edge sampling, edge counting, and triangle counting. Part of our contribution is that we consider three different settings, differing in the way in which one may access the neighborhood of a given vertex. In previous work, people have considered indexed neighbor access, with a query returning the 𝑖-th neighbor of a given vertex. Full neighborhood access model, which has a query that returns the entire neighborhood at a unit cost, has recently been considered in the applied community. Between these, we propose hash-ordered neighbor access, inspired by coordinated sampling, where we have a global fully random hash function, and can access neighbors in order of their hash values, paying a constant for each accessed neighbor. For edge sampling and counting, our new lower bounds are in the most powerful full neighborhood access model. We provide matching upper bounds in the weaker hash-ordered neighbor access model. Our new faster algorithms can be provably implemented efficiently on massive graphs in external memory and with the current APIs for, e.g., Twitter or Wikipedia. For triangle counting, we provide a separation: a better upper bound with full neighborhood access than the known lower bounds with indexed neighbor access. The technical core of our paper is our edge-sampling algorithm on which the other results depend. We now describe our results on the classic problems of edge and triangle counting. We give an algorithm that uses hash-ordered neighbor access to approximately count edges in time Õ to the state of the art without hash-ordered neighbor access of Õ ( 𝑛 𝜀 2 √ 𝑚 ) by Eden, Ron, and Seshadhri [ICALP 2017]). We present an Ω( 𝑛 𝜀 √ 𝑚 ) lower bound for 𝜀 ≥ √ 𝑚/𝑛 in the full neighborhood access model. This improves the lower bound of Ω( 𝑛 √ 𝜀𝑚 ) by Goldreich and Ron [Rand. Struct. Alg. 2008]) and it matches our new upper bound for 𝜀 ≥ √ 𝑚/𝑛. We also show an algorithm that uses the more standard assumption of pair queries ("are the vertices 𝑢 and 𝑣 adjacent?"),
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Faster Estimation of the Average Degree of a Graph Using Random Edges and Structural QueriesLorenzo Beretta, Deeparnab Chakrabarty, C. SeshadhriSODA 2026
- Approximately Counting and Sampling Hamiltonian Motifs in Sublinear TimeTalya Eden, Reut Levi, Dana Ron, Ronitt RubinfeldSTOC 2025
它引用的顶会 Paper1
相关 Paper
- Equivalences between triangle and range query problemsLech Duraj, Krzysztof Kleiner, Adam Polak, Virginia Vassilevska WilliamsSODA 2020 · 被引用 5 次
- Sublinear time approximation of the cost of a metric k-nearest neighbor graphArtur Czumaj, Christian SohlerSODA 2020 · 被引用 3 次
- Estimating the Number of Induced Subgraphs from Incomplete Data and Neighborhood QueriesDimitris Fotakis, Thanasis Pittas, Stratis SkoulakisAAAI 2021
- Computing and Testing Small Connectivity in Near-Linear Time and Queries via Fast Local Cut AlgorithmsSebastian Forster, Danupon Nanongkai, Liu Yang, Thatchaphol Saranurak 等SODA 2020 · 被引用 29 次
- Efficiently Counting Triangles in Large Temporal GraphsYuyang Xia, Yixiang Fang, Wensheng LuoSIGMOD 2025 · 被引用 3 次
