Tracing Logit Trajectories Across Layer Depth: Dataset-Level Explainability for Language Models
Jeesu Jung, Sangkeun Jung
摘要
Sentence-level explanations can miss the bigger picture of how a black-box model behaves across data, which matters most for complex criteria like safety that cannot be defined by a single rule. We trace Logit-Trajectory, which tracks adjacent-layer logit updates as vectors and aggregates them into a reproducible datasetlevel trajectory pattern, enabling depth-wise explainability through signals such as coherence and angular rotation. Across 6 languages and 5 NLP tasks, we show these trajectory summaries reveal consistent depth-wise patterns that divergence-and similarity-based baselines often wash out due to scalarization. As a case study where dataset-level intermediate decision structure matters, we evaluate safety classification, reporting both trajectory-level visual separability and classification performance.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper6
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Jailbroken: How Does LLM Safety Training Fail?Alexander Wei, Nika Haghtalab, Jacob SteinhardtNeurIPS 2023 · 被引用 2,230 次
- WildTeaming at Scale: From In-the-Wild Jailbreaks to (Adversarially) Safer Language ModelsLiwei Jiang, Kavel Rao, Seungju Han, Allyson Ettinger 等NeurIPS 2024 · 被引用 247 次
- Explaining How Transformers Use Context to Build PredictionsJavier Ferrando, Gerard I. Gállego, Ioannis Tsiamas, Marta R. Costa-jussàACL 2023 · 被引用 9 次
相关 Paper
- Truth as a Trajectory: What Internal Representations Reveal About Large Language Model ReasoningHamed Damirchi, Ignacio Meza De La Jara, Ehsan Abbasnejad, Afshar Shamsi 等ACL 2026 · 被引用 8 次
- Internal Chain-of-Thought: Empirical Evidence for Layer-wise Subtask Scheduling in LLMsZhipeng Yang, Junzhuo Li, Siyu Xia, Xuming HuEMNLP 2025
- Probing the Safety Robustness of LLMs in Latent SpaceTianle Gu, Kexin Huang, Zongqi Wang, Yixu Wang 等ACL 2026
- State-Dependent Safety Failures in Multi-Turn Language Model Interactionpengcheng li, Jie Zhang, Tianwei Zhang, Han Qiu 等ICML 2026 · 被引用 3 次
- Back Attention: Understanding and Enhancing Multi-Hop Reasoning in Large Language ModelsZeping Yu, Yonatan Belinkov, Sophia AnaniadouEMNLP 2025 · 被引用 2 次
