TV-TREES: Multimodal Entailment Trees for Neuro-Symbolic Video Reasoning
Kate Sanders, Nathaniel Weir, Benjamin Van Durme
Abstract
It is challenging for models to understand complex, multimodal content such as television clips, and this is in part because video-language models often rely on single-modality reasoning and lack interpretability. To combat these issues we propose TV-TREES, the first multimodal entailment tree generator. TV-TREES serves as an approach to video understanding that promotes interpretable joint-modality reasoning by searching for trees of entailment relationships between simple text-video evidence and higher-level conclusions that prove question-answer pairs. We also introduce the task of multimodal entailment tree generation to evaluate reasoning quality. Our method’s performance on the challenging TVQA benchmark demonstrates interpretable, state-of-the-art zero-shot performance on full clips, illustrating that multimodal entailment tree generation can be a best-of-both-worlds alternative to black-box systems.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ecaee8ca-64d2-4bd7-b2c4-ebed9449d17dCited by top-tier papers6
- VideoForest: Person-Anchored Hierarchical Reasoning for Cross-Video Question AnsweringYiran Meng, Junhong Ye, Wei Zhou, Guanghui Yue et al.ACM MM 2025 · 1 citation
- Bonsai: Interpretable Tree-Adaptive Grounded ReasoningKate Sanders, Benjamin Van DurmeAAAI 2026 · 1 citation
- Enhancing Systematic Decompositional Natural Language Inference Using Informal LogicNathaniel Weir, Kate Sanders, Orion Weller, Shreya Sharma et al.EMNLP 2024
- Natural Language Inference Improves Compositionality in Vision-Language ModelsPaola Cascante-Bonilla, Yu Hou, Yang Trista Cao, Hal Daumé III et al.ICLR 2025
- VideoTree: Adaptive Tree-based Video Representation for LLM Reasoning on Long VideosZiyang Wang, Shoubin Yu, Elias Stengel-Eskin, Jaehong Yoon et al.CVPR 2025
Builds on16
- Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question AnsweringPan Lu, Swaroop Mishra, Tanglin Xia, Liang Qiu et al.NeurIPS 2022 · 2,727 citations
- ViLT: Vision-and-Language Transformer Without Convolution or Region SupervisionWonjae Kim, Bokyung Son, Ildoo KimICML 2021 · 2,258 citations
- Unified Vision-Language Pre-Training for Image Captioning and VQALuowei Zhou, Hamid Palangi, Lei Zhang, Houdong Hu et al.AAAI 2020 · 1,047 citations
- HERO: Hierarchical Encoder for Video+Language Omni-representation Pre-trainingLinjie Li, Yen-Chun Chen, Yu Cheng, Zhe Gan et al.EMNLP 2020 · 387 citations
- Zero-Shot Video Question Answering via Frozen Bidirectional Language ModelsAntoine Yang, Antoine Miech, Josef Sivic, Ivan Laptev et al.NeurIPS 2022 · 305 citations
Related papers
- Commonsense Video Question Answering through Video-Grounded Entailment Tree ReasoningHuabin Liu, Filip Ilievski, Cees G. M. SnoekCVPR 2025
- Divide and Conquer: Exploring Language-centric Tree Reasoning for Video Question-AnsweringZhaohe Liao, Jiangtong Li, Siyu Sun, Qingyang Liu et al.ICML 2025
- Explaining Answers with Entailment TreesBhavana Dalvi, Peter Jansen, Oyvind Tafjord, Zhengnan Xie et al.EMNLP 2021 · 6 citations
- Clover: Towards A Unified Video-Language Alignment and Fusion ModelJingjia Huang, Yinan Li, Jiashi Feng, Xinglong Wu et al.CVPR 2023
- Inferential Knowledge-Enhanced Integrated Reasoning for Video Question AnsweringJianguo Mao, Wenbin Jiang, Hong Liu, Xiangdong Wang et al.AAAI 2023 · 1 citation
