Learning Signal-Agnostic Manifolds of Neural Fields
Yilun Du, Katie Collins, Josh Tenenbaum, Vincent Sitzmann
Abstract
Deep neural networks have been used widely to learn the latent structure of datasets, across modalities such as images, shapes, and audio signals. However, existing models are generally modality-dependent, requiring custom architectures and objectives to process different classes of signals. We leverage neural fields to capture the underlying structure in image, shape, audio and cross-modal audiovisual domains in a modality-independent manner. We cast our task as one of learning a manifold, where we aim to infer a low-dimensional, locally linear subspace in which our data resides. By enforcing coverage of the manifold, local linearity, and local isometry, our model -- dubbed GEM -- learns to capture the underlying structure of datasets across modalities. We can then travel along linear regions of our manifold to obtain perceptually consistent interpolations between samples, and can further use GEM to recover points on our manifold and glean not only diverse completions of input images, but cross-modal hallucinations of audio or image signals. Finally, we show that by walking across the underlying manifold of GEM, we may generate new samples in our signal domains. Code and additional results are available at https://yilundu.github.io/gem/.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 60c2a9ab-ab2b-4b63-aac8-4bfc241b00f5Cited by top-tier papers34
- Diffusion Probabilistic FieldsPeiye Zhuang, Samira Abnar, Jiatao Gu, Alexander G. Schwing et al.ICLR 2023 · 3,587 citations
- From data to functa: Your data point is a function and you can treat it like oneEmilien Dupont, Hyunjik Kim, S. M. Ali Eslami, Danilo Jimenez Rezende et al.ICML 2022 · 209 citations
- HyperDiffusion: Generating Implicit Neural Fields with Weight-Space DiffusionZiya Erkoç, Fangchang Ma, Qi Shan, Matthias Nießner et al.ICCV 2023 · 174 citations
- Learning Neural Acoustic FieldsAndrew F. Luo, Yilun Du, Michael J. Tarr, Josh Tenenbaum et al.NeurIPS 2022 · 153 citations
- AV-NeRF: Learning Neural Fields for Real-World Audio-Visual Scene SynthesisSusan Liang, Chao Huang, Yapeng Tian, Anurag Kumar et al.NeurIPS 2023 · 77 citations
Builds on13
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray et al.ICML 2021 · 6,356 citations
- Implicit Neural Representations with Periodic Activation FunctionsVincent Sitzmann, Julien N. P. Martel, Alexander W. Bergman, David B. Lindell et al.NeurIPS 2020 · 4,008 citations
- Implicit Geometric Regularization for Learning ShapesAmos Gropp, Lior Yariv, Niv Haim, Matan Atzmon et al.ICML 2020 · 1,001 citations
- GRAF: Generative Radiance Fields for 3D-Aware Image SynthesisKatja Schwarz, Yiyi Liao, Michael Niemeyer, Andreas GeigerNeurIPS 2020 · 1,001 citations
- Learning Shape Templates With Structured Implicit FunctionsKyle Genova, Forrester Cole, Daniel Vlasic, Aaron Sarna et al.ICCV 2019 · 427 citations
Related papers
- MM-LDM: Multi-Modal Latent Diffusion Model for Sounding Video GenerationMingzhen Sun, Weining Wang, Yanyuan Qiao, Jiahui Sun et al.ACM MM 2024 · 4 citations
- Neural Fields as Distributions: Signal Processing Beyond Euclidean SpaceDaniel Rebain, Soroosh Yazdani, Kwang Moo Yi, Andrea TagliasacchiCVPR 2024 · 1 citation
- 3DShape2VecSet: A 3D Shape Representation for Neural Fields and Generative Diffusion ModelsBiao Zhang, Jiapeng Tang, Matthias Nießner, Peter WonkaSIGGRAPH 2023 · 172 citations
- Autoencoder Image Interpolation by Shaping the Latent SpaceAlon Oring, Zohar Yakhini, Yacov Hel-OrICML 2021 · 41 citations
- Latent Graph Inference using Product ManifoldsHaitz Sáez de Ocáriz Borde, Anees Kazi, Federico Barbero, Pietro LiòICLR 2023 · 1 citation
