PEANUT: A Human-AI Collaborative Tool for Annotating Audio-Visual Data
Zheng Zhang, Zheng Ning, Chenliang Xu, Yapeng Tian, Toby Jia-Jun Li
Abstract
Audio-visual learning seeks to enhance the computer’s multi-modal perception leveraging the correlation between the auditory and visual modalities. Despite their many useful downstream tasks, such as video retrieval, AR/VR, and accessibility, the performance and adoption of existing audio-visual models have been impeded by the availability of high-quality datasets. Annotating audio-visual datasets is laborious, expensive, and time-consuming. To address this challenge, we designed and developed an efficient audio-visual annotation tool called Peanut. Peanut’s human-AI collaborative pipeline separates the multi-modal task into two single-modal tasks, and utilizes state-of-the-art object detection and sound-tagging models to reduce the annotators’ effort to process each frame and the number of manually-annotated frames needed. A within-subject user study with 20 participants found that Peanut can significantly accelerate the audio-visual data annotation process while maintaining high annotation accuracy.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 107e1971-8adc-4e72-801d-381e9f9515b5Cited by top-tier papers4
- SPICA: Interactive Video Content Exploration through Augmented Audio Descriptions for Blind or Low-Vision ViewersZheng Ning, Brianna L. Wimer, Kaiwen Jiang, Keyi Chen et al.CHI 2024 · 27 citations
- VideoDiff: Human-AI Video Co-Creation with AlternativesMina Huh, Ding Li, Kim Pimmel, Hijung Valentina Shin et al.CHI 2025 · 26 citations
- Supporting Co-Adaptive Machine Teaching through Human Concept Learning and Cognitive TheoriesSimret Araya Gebreegziabher, Yukun Yang, Elena L. Glassman, Toby Jia-Jun LiCHI 2025 · 5 citations
- ReFEree: Reference-Free and Fine-Grained Method for Evaluating Factual Consistency in Real-World Code SummarizationSuyoung Bae, CheolWon Na, Jaehoon Lee, Yumin Lee et al.ACL 2026 · 1 citation
Builds on26
- "Everyone wants to do the model work, not the data work": Data Cascades in High-Stakes AINithya Sambasivan, Shivani Kapania, Hannah Highfill, Diana Akrong et al.CHI 2021 · 725 citations
- Novice-AI Music Co-Creation via AI-Steering Tools for Deep Generative ModelsRyan Louie, Andy Coenen, Cheng Zhi Huang, Michael Terry et al.CHI 2020 · 265 citations
- Dual Attention Matching for Audio-Visual Event LocalizationYu Wu, Linchao Zhu, Yan Yan, Yi YangICCV 2019 · 233 citations
- Co-Separating Sounds of Visual ObjectsRuohan Gao, Kristen GraumanICCV 2019 · 224 citations
- Discriminative Sounding Objects Localization via Self-supervised Audiovisual MatchingDi Hu, Rui Qian, Minyue Jiang, Xiao Tan et al.NeurIPS 2020 · 156 citations
Related papers
- AVoCaDO: An Audiovisual Video Captioner Driven by Temporal OrchestrationXinlong Chen, Yue Ding, Weihong Lin, Jingyun Hua et al.ICLR 2026 · 27 citations
- Bio-Inspired Audiovisual Multi-Representation Integration via Self-Supervised LearningZhaojian Li, Bin Zhao, Yuan YuanACM MM 2023 · 3 citations
- Few-Shot Audio-Visual Class-Incremental Learning with Temporal Prompting and RegularizationYawen Cui, Li Liu, Zitong Yu, Guanjie Huang et al.AAAI 2025 · 2 citations
- Aligning Audio-Visual Joint Representations with an Agentic WorkflowShentong Mo, Yibing SongNeurIPS 2024 · 7 citations
- Learnable Irrelevant Modality Dropout for Multimodal Action Recognition on Modality-Specific Annotated VideosSaghir Alfasly, Jian Lu, Chen Xu, Yuru ZouCVPR 2022 · 28 citations
