Abstraction Alignment: Comparing Model-Learned and Human-Encoded Conceptual Relationships
Angie W. Boggust, Hyemin Bang, Hendrik Strobelt, Arvind Satyanarayan
Abstract
A model's confidence distribution is a reflection of its underlying knowledge.
Human abstractions represent the concepts and relationships we expect models to learn.
Abstraction alignment measures how much of a model's uncertainty can be explained by the human abstractions. SUBGRAPH PREFERENCE: Confidence in different regions of the abstraction. ABSTRACTION MATCH: Uncertainty reduced by a level of abstraction. CONCEPT CO-CONFUSION: Concepts the model regularly confuses.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext eaa63be1-5bf5-482d-a645-1f0c5b64e7a0Cited by top-tier papers2
- Semantic Regexes: Auto-Interpreting LLM Features with a Structured LanguageAngie W. Boggust, Donghao Ren, Yannick Assogba, Dominik Moritz et al.ICLR 2026 · 3 citations
- Editable XAI: Toward Bidirectional Human-AI Alignment with Co-Editable Explanations of Interpretable AttributesHaoyang Chen, Jingwen Bai, Fang Tian, Brian Y. LimCHI 2026 · 2 citations
Builds on31
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Contrastive Representation DistillationYonglong Tian, Dilip Krishnan, Phillip IsolaICLR 2020 · 1,305 citations
- How Language Model Hallucinations Can SnowballMuru Zhang, Ofir Press, William Merrill, Alisa Liu et al.ICML 2024 · 406 citations
- XRAI: Better Attributions Through RegionsAndrei Kapishnikov, Tolga Bolukbasi, Fernanda B. Viégas, Michael TerryICCV 2019 · 251 citations
- How does GPT-2 compute greater-than?: Interpreting mathematical abilities in a pre-trained language modelMichael Hanna, Ollie Liu, Alexandre VariengienNeurIPS 2023 · 251 citations
Related papers
- From Sampling to Cognition: Modeling Internal Cognitive Confidence in Language Models for Robust Uncertainty CalibrationHao Li, Tao He, Jiafeng Liang, Zheng Chu et al.AAAI 2026
- Confidence Under the Hood: An Investigation into the Confidence-Probability Alignment in Large Language ModelsAbhishek Kumar, Robert Morabito, Sanzhar Umbet, Jad Kabbara et al.ACL 2024 · 9 citations
- Neural Causal AbstractionsKevin Xia, Elias BareinboimAAAI 2024 · 17 citations
- Shared Interest: Measuring Human-AI Alignment to Identify Recurring Patterns in Model BehaviorAngie W. Boggust, Benjamin Hoover, Arvind Satyanarayan, Hendrik StrobeltCHI 2022 · 51 citations
- High Accuracy and Hidden Disparities: Investigating Foundation Model Performance in Clinical Cognitive AssessmentAbhay Sheel Anand, Deepak Ganesan, Ravi KarkarCHI 2026 · 1 citation
