Retriever: Learning Content-Style Representation as a Token-Level Bipartite Graph
Dacheng Yin, Xuanchi Ren, Chong Luo, Yuwang Wang, Zhiwei Xiong, Wenjun Zeng
摘要
This paper addresses the unsupervised learning of content-style decomposed representation. We first give a definition of style and then model the content-style representation as a token-level bipartite graph. An unsupervised framework, named Retriever, is proposed to learn such representations. First, a cross-attention module is employed to retrieve permutation invariant (P.I.) information, defined as style, from the input data. Second, a vector quantization (VQ) module is used, together with man-induced constraints, to produce interpretable content tokens. Last, an innovative link attention module serves as the decoder to reconstruct data from the decomposed content and style, with the help of the linking keys. Being modal-agnostic, the proposed Retriever is evaluated in both speech and image domains. The state-of-the-art zero-shot voice conversion performance confirms the disentangling ability of our framework. Top performance is also achieved in the part discovery task for images, verifying the interpretability of our representation. In addition, the vivid part-based style transfer quality demonstrates the potential of Retriever to support various fascinating generative tasks. Project page at https://ydcustc.github.io/retriever-demo/ .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Visual Concepts TokenizationTao Yang, Yuwang Wang, Yan Lu, Nanning ZhengNeurIPS 2022 · 被引用 19 次
- NANSY++: Unified Voice Synthesis with Neural Analysis and SynthesisHyeong-Seok Choi, Jinhyeok Yang, Juheon Lee, Hyeongju KimICLR 2023 · 被引用 8 次
- Unsupervised Disentanglement of Content and Style via Variance-Invariance ConstraintsYuxuan Wu, Ziyu Wang, Bhiksha Raj, Gus XiaICLR 2025
- StreamVoice: Streamable Context-Aware Language Modeling for Real-time Zero-Shot Voice ConversionZhichao Wang, Yuanzhe Chen, Xinsheng Wang, Lei Xie 等ACL 2024
它引用的顶会 Paper14
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 被引用 9,451 次
- Perceiver: General Perception with Iterative AttentionAndrew Jaegle, Felix Gimeno, Andy Brock, Oriol Vinyals 等ICML 2021 · 被引用 1,399 次
- Early Convolutions Help Transformers See BetterTete Xiao, Mannat Singh, Eric Mintun, Trevor Darrell 等NeurIPS 2021 · 被引用 974 次
- Perceiver IO: A General Architecture for Structured Inputs & OutputsAndrew Jaegle, Sebastian Borgeaud, Jean-Baptiste Alayrac, Carl Doersch 等ICLR 2022 · 被引用 797 次
相关 Paper
- Improving Zero-Shot Voice Style Transfer via Disentangled Representation LearningSiyang Yuan, Pengyu Cheng, Ruiyi Zhang, Weituo Hao 等ICLR 2021 · 被引用 64 次
- Transductive Learning for Unsupervised Text Style TransferFei Xiao, Liang Pang, Yanyan Lan, Yan Wang 等EMNLP 2021 · 被引用 20 次
- VoiceMixer: Adversarial Voice Style MixupSang-Hoon Lee, Ji-Hoon Kim, Hyunseung Chung, Seong-Whan LeeNeurIPS 2021 · 被引用 46 次
- Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised DisentanglementXueyao Zhang, Xiaohui Zhang, Kainan Peng, Zhenyu Tang 等ICLR 2025
- StableVC: Style Controllable Zero-Shot Voice Conversion with Conditional Flow MatchingJixun Yao, Yuguang Yang, Yu Pan, Ziqian Ning 等AAAI 2025 · 被引用 13 次
