VADER: Video Alignment Differencing and Retrieval
Alexander Black, Simon Jenni, Tu Bui, Md. Mehrab Tanjim, Stefano Petrangeli, Ritwik Sinha, Viswanathan Swaminathan, John P. Collomosse
Abstract
We propose VADER, a spatio- temporal matching, alignment, and change summarization method to help fight misinformation spread via manipulated videos. VADER matches and coarsely aligns partial video fragments to candidate videos using a robust visual descriptor and scalable search over adaptively chunked video content. A transformer- based alignment module then refines the temporal localization of the query fragment within the matched video. A space- time comparator module identifies regions of manipulation between aligned content, invariant to any changes due to any residual temporal misalignments or artifacts arising from non- editorial changes of the content. Robustly matching video to a trusted source enables conclusions to be drawn on video provenance, enabling informed trust decisions on content encountered. Code and data are available at https://github.com/AlexBlck/vader
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on11
- Leveraging Frequency Analysis for Deep Fake Image RecognitionJoel Frank, Thorsten Eisenhofer, Lea Schönherr, Asja Fischer et al.ICML 2020 · 848 citations
- WildDeepfake: A Challenging Real-World Dataset for Deepfake DetectionBojia Zi, Minghao Chang, Jingjing Chen, Xingjun Ma et al.ACM MM 2020 · 443 citations
- COTR: Correspondence Transformer for Matching Across ImagesWei Jiang, Eduard Trulls, Jan Hosang, Andrea Tagliasacchi et al.ICCV 2021 · 318 citations
- Detecting Photoshopped Faces by Scripting PhotoshopSheng-Yu Wang, Oliver Wang, Richard Zhang, Andrew Owens et al.ICCV 2019 · 147 citations
- Responsible Disclosure of Generative Models Using Scalable FingerprintingNing Yu, Vladislav Skripniuk, Dingfan Chen, Larry S. Davis et al.ICLR 2022 · 118 citations
Related papers
- RelayFormer: A Unified Local-Global Attention Framework for Scalable Image and Video Manipulation LocalizationWen Huang, Jiarui Yang, Tao Dai, Jiawei Li et al.ICLR 2026
- Video Similarity and Alignment Learning on Partial Video Copy DetectionZhen Han, Xiangteng He, Mingqian Tang, Yiliang LvACM MM 2021 · 34 citations
- Consistent and Invariant Generalization Learning for Short-video Misinformation DetectionHanghui Guo, Weijie Shi, Mengze Li, Juncheng Li et al.ACM MM 2025 · 1 citation
- Explainable Forensics of Manipulated Segments in Untrimmed Long VideosYue Feng, Jingjing Li, Qijia Lu, Wei Ji et al.ICML 2026
- Deepfake Video Detection with Spatiotemporal Dropout TransformerDaichi Zhang, Fanzhao Lin, Yingying Hua, Pengju Wang et al.ACM MM 2022 · 46 citations
