Policy Maps: Tools for Guiding the Unbounded Space of LLM Behaviors
Michelle S. Lam, Fred Hohman, Dominik Moritz, Jeffrey P. Bigham, Kenneth Holstein, Mary Beth Kery
Abstract
Figure 1: Policy maps chart LLM policy coverage over an unbounded space of model behaviors.Here, an AI practitioner is designing a policy for how an LLM should summarize violent text.Policy map abstractions (right) allow the policy designer to interactively author and test policies that govern a model's behavior using if-then rules over concepts.The designer can create any desired concept by providing a simple text definition to capture cases of model behavior.Our Policy Projector tool (center) renders cases, concepts, and policies as visual map layers to aid iterative policy design.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers5
- PolicyPad: Collaborative Prototyping of LLM PoliciesK. J. Kevin Feng, Tzu-Sheng Kuo, Quan Ze Jim Chen, Inyoung Cheong et al.CHI 2026 · 2 citations
- Evalet: Evaluating Large Language Models through Functional FragmentationTae Soo Kim, Heechan Lee, Yoonjoo Lee, Joseph Seering et al.CHI 2026 · 1 citation
- PASTA: A Scalable Framework for Multi-Policy AI Compliance EvaluationYu Yang, Ig-Jae Kim, Dongwook YoonCHI 2026 · 1 citation
- Data-Prompt Co-Evolution: Growing Test Sets to Refine LLM BehaviorMinjae Lee, Minsuk KahngCHI 2026 · 1 citation
- Agile Deliberation: Concept Deliberation for Subjective Visual ClassificationLeijie Wang, Otilia Stretcu, Wei Qiao, Thomas Denby et al.CVPR 2026
Builds on34
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- Inference-Time Intervention: Eliciting Truthful Answers from a Language ModelKenneth Li, Oam Patel, Fernanda B. Viégas, Hanspeter Pfister et al.NeurIPS 2023 · 1,549 citations
- Chatbot Arena: An Open Platform for Evaluating LLMs by Human PreferenceWei-Lin Chiang, Lianmin Zheng, Ying Sheng, Anastasios Nikolas Angelopoulos et al.ICML 2024 · 1,212 citations
Related papers
- MeetMap: Real-Time Collaborative Dialogue Mapping with LLMs in Online MeetingsXinyue Chen, Nathan Yap, Xinyi Lu, Aylin Gunal et al.CSCW 2025 · 12 citations
- Inverting the Shield: Systematically Generating Safety Tests from Policy SpecificationsXiaoyue Lu, Xianglin Yang, Haijun Liu, Jiahao Liu et al.ACL 2026
- Learning with Language-Guided State AbstractionsAndi Peng, Ilia Sucholutsky, Belinda Z. Li, Theodore R. Sumers et al.ICLR 2024 · 20 citations
- PoliGraph: Automated Privacy Policy Analysis using Knowledge GraphsHao Cui, Rahmadi Trimananda, Athina Markopoulou, Scott JordanUSENIX Security 2023
- MapStory: Prototyping Editable Map Animations with LLM AgentsAditya Gunturu, Ben Pearman, Keiichi Ihara, Morteza Faraji et al.UIST 2025
