Abstraction Alignment: Comparing Model-Learned and Human-Encoded Conceptual Relationships
Angie W. Boggust, Hyemin Bang, Hendrik Strobelt, Arvind Satyanarayan
2025年份
4被引次数
2顶会引用
摘要
A model's confidence distribution is a reflection of its underlying knowledge.
Human abstractions represent the concepts and relationships we expect models to learn.
Abstraction alignment measures how much of a model's uncertainty can be explained by the human abstractions. SUBGRAPH PREFERENCE: Confidence in different regions of the abstraction. ABSTRACTION MATCH: Uncertainty reduced by a level of abstraction. CONCEPT CO-CONFUSION: Concepts the model regularly confuses.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Semantic Regexes: Auto-Interpreting LLM Features with a Structured LanguageAngie W. Boggust, Donghao Ren, Yannick Assogba, Dominik Moritz 等ICLR 2026 · 被引用 3 次
- Editable XAI: Toward Bidirectional Human-AI Alignment with Co-Editable Explanations of Interpretable AttributesHaoyang Chen, Jingwen Bai, Fang Tian, Brian Y. LimCHI 2026 · 被引用 2 次
它引用的顶会 Paper31
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Contrastive Representation DistillationYonglong Tian, Dilip Krishnan, Phillip IsolaICLR 2020 · 被引用 1,305 次
- How Language Model Hallucinations Can SnowballMuru Zhang, Ofir Press, William Merrill, Alisa Liu 等ICML 2024 · 被引用 406 次
- XRAI: Better Attributions Through RegionsAndrei Kapishnikov, Tolga Bolukbasi, Fernanda B. Viégas, Michael TerryICCV 2019 · 被引用 251 次
- How does GPT-2 compute greater-than?: Interpreting mathematical abilities in a pre-trained language modelMichael Hanna, Ollie Liu, Alexandre VariengienNeurIPS 2023 · 被引用 251 次
相关 Paper
- From Sampling to Cognition: Modeling Internal Cognitive Confidence in Language Models for Robust Uncertainty CalibrationHao Li, Tao He, Jiafeng Liang, Zheng Chu 等AAAI 2026
- Confidence Under the Hood: An Investigation into the Confidence-Probability Alignment in Large Language ModelsAbhishek Kumar, Robert Morabito, Sanzhar Umbet, Jad Kabbara 等ACL 2024 · 被引用 9 次
- Neural Causal AbstractionsKevin Xia, Elias BareinboimAAAI 2024 · 被引用 17 次
- Shared Interest: Measuring Human-AI Alignment to Identify Recurring Patterns in Model BehaviorAngie W. Boggust, Benjamin Hoover, Arvind Satyanarayan, Hendrik StrobeltCHI 2022 · 被引用 51 次
- High Accuracy and Hidden Disparities: Investigating Foundation Model Performance in Clinical Cognitive AssessmentAbhay Sheel Anand, Deepak Ganesan, Ravi KarkarCHI 2026 · 被引用 1 次
