Policy Maps: Tools for Guiding the Unbounded Space of LLM Behaviors
Michelle S. Lam, Fred Hohman, Dominik Moritz, Jeffrey P. Bigham, Kenneth Holstein, Mary Beth Kery
2025年份
4被引次数
5顶会引用
摘要
Figure 1: Policy maps chart LLM policy coverage over an unbounded space of model behaviors.Here, an AI practitioner is designing a policy for how an LLM should summarize violent text.Policy map abstractions (right) allow the policy designer to interactively author and test policies that govern a model's behavior using if-then rules over concepts.The designer can create any desired concept by providing a simple text definition to capture cases of model behavior.Our Policy Projector tool (center) renders cases, concepts, and policies as visual map layers to aid iterative policy design.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- PolicyPad: Collaborative Prototyping of LLM PoliciesK. J. Kevin Feng, Tzu-Sheng Kuo, Quan Ze Jim Chen, Inyoung Cheong 等CHI 2026 · 被引用 2 次
- Evalet: Evaluating Large Language Models through Functional FragmentationTae Soo Kim, Heechan Lee, Yoonjoo Lee, Joseph Seering 等CHI 2026 · 被引用 1 次
- PASTA: A Scalable Framework for Multi-Policy AI Compliance EvaluationYu Yang, Ig-Jae Kim, Dongwook YoonCHI 2026 · 被引用 1 次
- Data-Prompt Co-Evolution: Growing Test Sets to Refine LLM BehaviorMinjae Lee, Minsuk KahngCHI 2026 · 被引用 1 次
- Agile Deliberation: Concept Deliberation for Subjective Visual ClassificationLeijie Wang, Otilia Stretcu, Wei Qiao, Thomas Denby 等CVPR 2026
它引用的顶会 Paper34
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- Inference-Time Intervention: Eliciting Truthful Answers from a Language ModelKenneth Li, Oam Patel, Fernanda B. Viégas, Hanspeter Pfister 等NeurIPS 2023 · 被引用 1,549 次
- Chatbot Arena: An Open Platform for Evaluating LLMs by Human PreferenceWei-Lin Chiang, Lianmin Zheng, Ying Sheng, Anastasios Nikolas Angelopoulos 等ICML 2024 · 被引用 1,212 次
相关 Paper
- MeetMap: Real-Time Collaborative Dialogue Mapping with LLMs in Online MeetingsXinyue Chen, Nathan Yap, Xinyi Lu, Aylin Gunal 等CSCW 2025 · 被引用 12 次
- Inverting the Shield: Systematically Generating Safety Tests from Policy SpecificationsXiaoyue Lu, Xianglin Yang, Haijun Liu, Jiahao Liu 等ACL 2026
- Learning with Language-Guided State AbstractionsAndi Peng, Ilia Sucholutsky, Belinda Z. Li, Theodore R. Sumers 等ICLR 2024 · 被引用 20 次
- PoliGraph: Automated Privacy Policy Analysis using Knowledge GraphsHao Cui, Rahmadi Trimananda, Athina Markopoulou, Scott JordanUSENIX Security 2023
- MapStory: Prototyping Editable Map Animations with LLM AgentsAditya Gunturu, Ben Pearman, Keiichi Ihara, Morteza Faraji 等UIST 2025
