Tracing Logit Trajectories Across Layer Depth: Dataset-Level Explainability for Language Models
Jeesu Jung, Sangkeun Jung
Abstract
Sentence-level explanations can miss the bigger picture of how a black-box model behaves across data, which matters most for complex criteria like safety that cannot be defined by a single rule. We trace Logit-Trajectory, which tracks adjacent-layer logit updates as vectors and aggregates them into a reproducible datasetlevel trajectory pattern, enabling depth-wise explainability through signals such as coherence and angular rotation. Across 6 languages and 5 NLP tasks, we show these trajectory summaries reveal consistent depth-wise patterns that divergence-and similarity-based baselines often wash out due to scalarization. As a case study where dataset-level intermediate decision structure matters, we evaluate safety classification, reporting both trajectory-level visual separability and classification performance.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on6
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Jailbroken: How Does LLM Safety Training Fail?Alexander Wei, Nika Haghtalab, Jacob SteinhardtNeurIPS 2023 · 2,230 citations
- WildTeaming at Scale: From In-the-Wild Jailbreaks to (Adversarially) Safer Language ModelsLiwei Jiang, Kavel Rao, Seungju Han, Allyson Ettinger et al.NeurIPS 2024 · 247 citations
- Explaining How Transformers Use Context to Build PredictionsJavier Ferrando, Gerard I. Gállego, Ioannis Tsiamas, Marta R. Costa-jussàACL 2023 · 9 citations
Related papers
- Truth as a Trajectory: What Internal Representations Reveal About Large Language Model ReasoningHamed Damirchi, Ignacio Meza De La Jara, Ehsan Abbasnejad, Afshar Shamsi et al.ACL 2026 · 8 citations
- Internal Chain-of-Thought: Empirical Evidence for Layer-wise Subtask Scheduling in LLMsZhipeng Yang, Junzhuo Li, Siyu Xia, Xuming HuEMNLP 2025
- Probing the Safety Robustness of LLMs in Latent SpaceTianle Gu, Kexin Huang, Zongqi Wang, Yixu Wang et al.ACL 2026
- State-Dependent Safety Failures in Multi-Turn Language Model Interactionpengcheng li, Jie Zhang, Tianwei Zhang, Han Qiu et al.ICML 2026 · 3 citations
- Back Attention: Understanding and Enhancing Multi-Hop Reasoning in Large Language ModelsZeping Yu, Yonatan Belinkov, Sophia AnaniadouEMNLP 2025 · 2 citations
