Chop & Learn: Recognizing and Generating Object-State Compositions
Nirat Saini, Hanyu Wang, Archana Swaminathan, Vinoj Jayasundara, Bo He, Kamal Gupta, Abhinav Shrivastava
Abstract
Recognizing and generating object-state compositions has been a challenging task, especially when generalizing to unseen compositions. In this paper, we study the task of cutting objects in different styles and the resulting object state changes. We propose a new benchmark suite Chop & Learn, to accommodate the needs of learning objects and different cut styles using multiple viewpoints. We also propose a new task of Compositional Image Generation, which can transfer learned cut styles to different objects, by generating novel object-state images. Moreover, we also use the videos for Compositional Action Recognition, and show valuable uses of this dataset for multiple video tasks. Project website: https://chopnlearn.github.io .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4cbd7f53-af60-4fb2-b589-70c8a60eb4ddCited by top-tier papers7
- SCHEMA: State CHangEs MAtter for Procedure Planning in Instructional VideosYulei Niu, Wenliang Guo, Long Chen, Xudong Lin et al.ICLR 2024 · 26 citations
- PICTURE: PhotorealistIC Virtual Try-on from UnconstRained dEsignsShuliang Ning, Duomin Wang, Yipeng Qin, Zirong Jin et al.CVPR 2024 · 10 citations
- Active Object Detection with Knowledge Aggregation and Distillation from Large ModelsDejie Yang, Yang LiuCVPR 2024 · 9 citations
- MOSCATO: Predicting Multiple Object State Change through ActionsParnian Zameni, Yuhan Shen, Ehsan ElhamifarICCV 2025 · 4 citations
- Beyond Seen Primitive Concepts and Attribute-Object Compositional LearningNirat Saini, Khoi Pham, Abhinav ShrivastavaCVPR 2024 · 4 citations
Builds on24
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray et al.ICML 2021 · 6,356 citations
Related papers
- OSCBench: Benchmarking Object State Change in Text-to-Video GenerationXianjing Han, Bin Zhu, Shiqi Hu, Franklin Mingzhe Li et al.ACL 2026 · 2 citations
- Video State-Changing Object SegmentationJiangwei Yu, Xiang Li, Xinran Zhao, Hongming Zhang et al.ICCV 2023 · 16 citations
- A Dynamic Learning Method towards Realistic Compositional Zero-Shot LearningXiaoming Hu, Zilei WangAAAI 2024 · 10 citations
- Watch Less, Do More: Implicit Skill Discovery for Video-Conditioned PolicyJiangxing Wang, Zongqing LuICLR 2025
- Look Less Think More: Rethinking Compositional Action RecognitionRui Yan, Peng Huang, Xiangbo Shu, Junhao Zhang et al.ACM MM 2022 · 17 citations
