One-shot recognition of any material anywhere using contrastive learning with physics-based rendering
Manuel S. Drehwald, Sagi Eppel, Jolina Li, Han Hao, Alán Aspuru-Guzik
摘要
Visual recognition of materials and their states is essential for understanding the world, from determining whether food is cooked, metal is rusted, or a chemical reaction has occurred. However, current image recognition methods are limited to specific classes and properties and can’t handle the vast number of material states in the world. To address this, we present MatSim: the first dataset and benchmark for computer vision-based recognition of similarities and transitions between materials and textures, focusing on identifying any material under any conditions using one or a few examples. The dataset contains synthetic and natural images. Synthetic images were rendered using giant collections of textures, objects, and environments generated by computer graphics artists. We use mixtures and gradual transitions between materials to allow the system to learn cases with smooth transitions between states (like gradually cooked food). We also render images with materials inside transparent containers to support beverage and chemistry lab use cases. We use this dataset to train a Siamese net that identifies the same material in different objects, mixtures, and environments. The descriptor generated by this net can be used to identify the states of materials and their subclasses using a single image. We also present the first few-shot material recognition benchmark with natural images from a wide range of fields, including the state of foods and beverages, types of grounds, and many other use cases. We show that a net trained on the MatSim synthetic dataset outperforms state-of-the-art models like Clip on the benchmark and also achieves good results on other unsupervised material classification tasks. Dataset, generation code and trained models have been made available at: https://github.com/ZuseZ4/MatSim-Dataset-Generator-Scripts-And-Neural-net
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Make-it-Real: Unleashing Large Multimodal Model for Painting 3D Objects with Realistic MaterialsYe Fang, Zeyi Sun, Tong Wu, Jiaqi Wang 等NeurIPS 2024 · 被引用 21 次
- LUWA Dataset: Learning Lithic Use-Wear Analysis on Microscopic ImagesJing Zhang, Irving Fang, Hao Wu, Akshat Kaushik 等CVPR 2024 · 被引用 5 次
- Harnessing the Power of Foundation Models for Accurate Material ClassificationQINGRAN LIN, Fengwei Yang, Chaolun ZhuCVPR 2026 · 被引用 3 次
- RF-MatID: Dataset and Benchmark for Radio Frequency Material IdentificationXinyan Chen, Qinchun Li, Ruiqin Ma, Jiaqi Bai 等ICLR 2026 · 被引用 1 次
它引用的顶会 Paper5
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- A ConvNet for the 2020sZhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer 等CVPR 2022 · 被引用 6,782 次
- The Many Faces of Robustness: A Critical Analysis of Out-of-Distribution GeneralizationDan Hendrycks, Steven Basart, Norman Mu, Saurav Kadavath 等ICCV 2021 · 被引用 2,294 次
- Correspondence-Free Material Reconstruction using Sparse Surface ConstraintsSebastian Weiss, Robert Maier, Daniel Cremers, Rüdiger Westermann 等CVPR 2020
相关 Paper
- Hierarchical Material Recognition from Local AppearanceMatthew Beveridge, Shree K. NayarICCV 2025 · 被引用 5 次
- STATUS Bench: A Rigorous Benchmark for Evaluating Object State Understanding in Vision-Language ModelsMahiro Ukai, Shuhei Kurita, Nakamasa InoueACM MM 2025
- MatCLIP: Light- and Shape-Insensitive Assignment of PBR Material ModelsMichael Birsak, John Femiani, Biao Zhang, Peter WonkaSIGGRAPH 2025 · 被引用 2 次
- PhotoMat: A Material Generator Learned from Single Flash PhotosXilong Zhou, Milos Hasan, Valentin Deschaintre, Paul Guerrero 等SIGGRAPH 2023 · 被引用 31 次
- ABO: Dataset and Benchmarks for Real-World 3D Object UnderstandingJasmine Collins, Shubham Goel, Kenan Deng, Achleshwar Luthra 等CVPR 2022 · 被引用 117 次
