One-Pass Diversified Sampling with Application to Terabyte-Scale Genomic Sequence Streams
Benjamin Coleman, Benito Geordie, Li Chou, Ryan A. Leo Elworth, Todd J. Treangen, Anshumali Shrivastava
摘要
A popular approach to reduce the size of a massive dataset is to apply efficient online sampling to the stream of data as it is read or generated. Online sampling routines are currently restricted to variations of reservoir sampling, where each sample is selected uniformly and independently of other samples. This renders them unsuitable for large-scale applications in computational biology, such as metagenomic community profiling and protein function annotation, which suffer from severe class imbalance. To maintain a representative and diverse sample, we must identify and preferentially select data that are likely to belong to rare classes. We argue that existing schemes for diversity sampling have prohibitive overhead for large-scale problems and high-throughput streams. We propose an efficient sampling routine that uses an online representation of the data distribution as a prefilter to retain elements from rare groups. We apply this method to several genomic data analysis tasks and demonstrate significant speedup in downstream analysis without sacrificing the quality of the results. Because our algorithm is 2x faster and uses 1000x less memory than coreset, reservoir and sketch-based alternatives, we anticipate that it will become a useful preprocessing step for applications with large-scale streaming data. Whole-genome shotgun sequencing (WGS) has inspired nu-* Equal contribution
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- How to train data-efficient LLMsNoveen Sachdeva, Benjamin Coleman, Wang-Cheng Kang, Jianmo Ni 等ICLR 2026 · 被引用 106 次
- One-Pass Distribution Sketch for Measuring Data Heterogeneity in Federated LearningZichang Liu, Zhaozhuo Xu, Benjamin Coleman, Anshumali ShrivastavaNeurIPS 2023 · 被引用 19 次
它引用的顶会 Paper2
相关 Paper
- Finer Metagenomic Reconstruction via Biodiversity OptimizationSimon Foucart, David KoslickiNeurIPS 2020 · 被引用 1 次
- Stabilizing Linear Passive-Aggressive Online Learning with Weighted Reservoir SamplingSkyler Wu, Fred Lu, Edward Raff, James HoltNeurIPS 2024 · 被引用 5 次
- Fast Processing and Querying of 170TB of Genomics Data via a Repeated And Merged BloOm Filter (RAMBO)Gaurav Gupta, Minghao Yan, Benjamin Coleman, Bryce Kille 等SIGMOD 2021 · 被引用 19 次
- HistSketch: A Compact Data Structure for Accurate Per-Key Distribution MonitoringJintao He, Jiaqi Zhu, Qun HuangICDE 2023 · 被引用 21 次
- GREAT: Generalized Reservoir Sampling based Triangle Counting Estimation over Streaming GraphsSiyue Wu, Dingming Wu, Sinhong Cheuk, Tsz Nam Chan 等VLDB 2025
