Learning Structured Representations by Embedding Class Hierarchy
Siqi Zeng, Remi Tachet des Combes, Han Zhao
Abstract
To embed structured knowledge within labels into feature representations, prior work (Zeng et al., 2022) proposed to use the Cophenetic Correlation Coefficient (CPCC) as a regularizer during supervised learning. This regularizer calculates pairwise Euclidean distances of class means and aligns them with the corresponding shortest path distances derived from the label hierarchy tree. However, class means may not be good representatives of the class conditional distributions, especially when they are multi-mode in nature. To address this limitation, under the CPCC framework, we propose to use the Earth Mover's Distance (EMD) to measure the pairwise distances among classes in the feature space. We show that our exact EMD method generalizes previous work, and recovers the existing algorithm when class-conditional distributions are Gaussian. To further improve the computational efficiency of our method, we introduce the Optimal Transport-CPCC family by exploring four EMD approximation variants. Our most efficient OT-CPCC variant, the proposed Fast FlowTree algorithm, runs in linear time in the size of the dataset, while maintaining competitive performance across datasets and tasks. The code is available at https://github.com/uiuctml/OTCPCC .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers7
- Learning Structured Representations with Hyperbolic EmbeddingsAditya Sinha, Siqi Zeng, Makoto Yamada, Han ZhaoNeurIPS 2024 · 24 citations
- The LLM Bottleneck: Why Open-Source Vision LLMs Struggle with Hierarchical Visual RecognitionYuwen Tan, Yuan Qing, Boqing GongCVPR 2026 · 6 citations
- Free-Grained Hierarchical Visual RecognitionSeulki Park, Zilin Wang, Stella X. YuCVPR 2026 · 3 citations
- Taxonomy-Aware Representation Alignment for Hierarchical Visual Recognition with Large Multimodal ModelsHulingxiao He, Zhi Tan, Yuxin PengCVPR 2026 · 3 citations
- Learning Hierarchical Knowledge in Text-Rich Networks with Taxonomy-Informed Representation LearningYunhui Liu, Yongchao Liu, Yinfeng Chen, Chuntao Hong et al.KDD 2026 · 1 citation
Builds on6
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- BREEDS: Benchmarks for Subpopulation ShiftShibani Santurkar, Dimitris Tsipras, Aleksander MadryICLR 2021 · 193 citations
- Scalable Nearest Neighbor Search for Optimal TransportArturs Backurs, Yihe Dong, Piotr Indyk, Ilya P. Razenshteyn et al.ICML 2020 · 60 citations
- Supervised Tree-Wasserstein DistanceYuki Takezawa, Ryoma Sato, Makoto YamadaICML 2021 · 14 citations
- HIER: Metric Learning Beyond Class Labels via Hierarchical RegularizationSungyeon Kim, Boseung Jeong, Suha KwakCVPR 2023
Related papers
- Learning Structured Representations by Embedding Class Hierarchy with Fast Optimal TransportSiqi Zeng, Sixian Du, Makoto Yamada, Han ZhaoICLR 2025
- Efficient Discrete Multi Marginal Optimal Transport RegularizationRonak Mehta, Jeffery Kline, Vishnu Suresh Lokhande, Glenn Fung et al.ICLR 2023
- A linear time approximation of Wasserstein distance with word embedding selectionSho Otao, Makoto YamadaEMNLP 2023 · 2 citations
- Fast Regularized Discrete Optimal Transport with Group-Sparse RegularizersYasutoshi Ida, Sekitoshi Kanai, Kazuki Adachi, Atsutoshi Kumagai et al.AAAI 2023 · 3 citations
- Improving Semi-Supervised Semantic Segmentation with Sliced-Wasserstein Feature Alignment and UniformityChen-Yi Lu, Kasra Derakhshandeh, Somali ChaterjiCVPR 2025
