TriEx: A Game-based Tri-View Framework for Explaining Internal Reasoning in Multi-Agent LLMs
Ziyi Wang, Chen Zhang, Wenjun Peng, Qi Wu, Xinyu Wang
Abstract
Explainability for Large Language Model (LLM) agents is especially challenging in interactive, partially observable settings, where decisions depend on evolving beliefs and other agents. We present TriEx, a tri-view explainability framework that instruments sequential decision making with aligned artifacts: (i) structured first-person self-reasoning bound to an action, (ii) explicit second-person belief states about opponents updated over time, and (iii) third-person oracle audits grounded in environment-derived reference signals. This design turns explanations from free-form narratives into evidence-anchored objects that can be compared and checked across time and perspectives. Using imperfect-information strategic games as a controlled testbed, we show that TriEx enables scalable analysis of explanation faithfulness, belief dynamics, and evaluator reliability, revealing systematic mismatches between what agents say, what they believe, and what they do. Our results highlight explainability as an interaction-dependent property and motivate multi-view, evidence-grounded evaluation for LLM agents. Code is available at https://github.com/Einsam1819/TriEx.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fa885ba8-7998-4e32-b726-c7f8b424d1b2Builds on12
- Locating and Editing Factual Associations in GPTKevin Meng, David Bau, Alex Andonian, Yonatan BelinkovNeurIPS 2022 · 3,415 citations
- Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought PromptingMiles Turpin, Julian Michael, Ethan Perez, Samuel R. BowmanNeurIPS 2023 · 1,792 citations
- Towards Best Practices of Activation Patching in Language Models: Metrics and MethodsFred Zhang, Neel NandaICLR 2024 · 233 citations
- SmartPlay : A Benchmark for LLMs as Intelligent AgentsYue Wu, Xuan Tang, Tom M. Mitchell, Yuanzhi LiICLR 2024 · 121 citations
- Overthinking the Truth: Understanding how Language Models Process False DemonstrationsDanny Halawi, Jean-Stanislas Denain, Jacob SteinhardtICLR 2024 · 83 citations
Related papers
- Theory of Mind for Multi-Agent Collaboration via Large Language ModelsHuao Li, Yu Quan Chong, Simon Stepputtis, Joseph Campbell et al.EMNLP 2023 · 57 citations
- DiLLS: Interactive Diagnosis of LLM-based Multi-agent Systems via Layered Summary of Agent BehaviorsRui Sheng, Yukun Yang, Chuhan Shi, Yanna Lin et al.CHI 2026 · 2 citations
- Sensemaking in Multi-Agent LLM Interfaces: How Users Interpret Transparency and Trustworthiness CuesSaumya Pareek, Jarod Govers, Naja Kathrine Kollerup, Emily Wong et al.CHI 2026 · 1 citation
- Noise, Adaptation, and Strategy: Assessing LLM Fidelity in Decision-MakingYuanjun Feng, Vivek Choudhary, Yash Raj ShresthaEMNLP 2025 · 2 citations
- InEx: Hallucination Mitigation via Introspection and Cross-Modal Multi-Agent CollaborationZhongyu Yang, Yingfang Yuan, Xuanming Jiang, Baoyi An et al.AAAI 2026 · 5 citations
