Exploring 3D Dataset Pruning
Xiaohan Zhao, Xinyi Shang, Jiacheng Liu, Zhiqiang Shen
Abstract
Dataset pruning remains underexplored for 3D modalities, where inherent class imbalance persists across both training and test sets. This creates a divergence in evaluation: overall accuracy favors natural frequency, reflecting practical usage; while mean accuracy demands balanced generalization. Instead of forcing a premature trade-off, we advocate for base principles that remain universally robust and beneficial across diverse priors. We cast pruning as a quadrature approximation on population risk and decompose the error bound into representation error (fidelity to the underlying manifold) and prior-mismatch bias (distribution shift), clarifying what can be improved jointly across priors. To address prior-mismatch bias, we decouple likelihood from prior in the posterior and transfer the structural likelihood via distillation with a calibrated teacher and geometry-preserving constraints. Simultaneously, to reduce representation error, we audit common pruning signals and choose geometric embedding, which exhibits greater robustness given the high inductive bias of 3D models. We also prioritize a safety floor before selection, capturing high-reward regions beneficial across priors. Finally, acknowledging that no single subset optimally satisfies divergent evaluation priors, we augment these principles with a steering wrapper that interpolates between stratified seeding and global selection. Empirical results demonstrate that our framework elevates the performance floor while offering flexibility for different prior preferences.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on13
- Decoupling Representation and Classifier for Long-Tailed RecognitionBingyi Kang, Saining Xie, Marcus Rohrbach, Zhicheng Yan et al.ICLR 2020 · 1,496 citations
- Revisiting Point Cloud Classification: A New Benchmark Dataset and Classification Model on Real-World DataMikaela Angelina Uy, Quang-Hieu Pham, Binh-Son Hua, Duc Thanh Nguyen et al.ICCV 2019 · 1,003 citations
- Deep Learning on a Data Diet: Finding Important Examples Early in TrainingMansheej Paul, Surya Ganguli, Gintare Karolina DziugaiteNeurIPS 2021 · 806 citations
- Variational Adversarial Active LearningSamarth Sinha, Sayna Ebrahimi, Trevor DarrellICCV 2019 · 662 citations
- Coresets for Data-efficient Training of Machine Learning ModelsBaharan Mirzasoleiman, Jeff A. Bilmes, Jure LeskovecICML 2020 · 494 citations
Related papers
- FUSION: Dataset Pruning via Fusing Uncertainty with Structural Information for Optimal Neural Training in Crystal Property PredictionXiean Wang, Pin Chen, Liqin Tan, Yutong Lu et al.AAAI 2026
- MIND: Decoupling Model-Induced Label Noise via Latent Manifold DisentanglementDayong RenICML 2026
- From Extrinsic to Intrinsic: Geodesic-Guided Representation Learning for 3D Geometric DataYuming ZHAO, Junhui Hou, Qijian Zhang, Jia Qin et al.ICML 2026
- Uncovering the Latent Potential of Deep Intermediate RepresentationsArnesh Batra, Arush Gumber, Aniket Khandelwal, Jashn Khemani et al.ICML 2026 · 1 citation
- GeoDM: Geometry-aware Distribution Matching for Dataset DistillationXuhui Li, Zhengquan Luo, Zihui Cui, Kai Zhao et al.ICML 2026 · 2 citations
