Standardize: Aligning Language Models with Expert-Defined Standards for Content Generation
Joseph Marvin Imperial, Gail Forey, Harish Tayyar Madabushi
Abstract
Domain experts across engineering, healthcare, and education follow strict standards for producing quality content such as technical manuals, medication instructions, and children's reading materials. However, current works in controllable text generation have yet to explore using these standards as references for control. Towards this end, we introduce STANDARD-IZE, a retrieval-style in-context learning-based framework to guide large language models to align with expert-defined standards. Focusing on English language standards in the education domain as a use case, we consider the Common European Framework of Reference for Languages (CEFR) and Common Core Standards (CCS) for the task of open-ended content generation. Our findings show that models can gain 45% to 100% increase in precise accuracy across open and commercial LLMs evaluated, demonstrating that the use of knowledge artifacts extracted from standards and integrating them in the generation process can effectively guide models to produce better standardaligned content. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9e4328bd-59f1-4c2a-bb84-726285b359b9Cited by top-tier papers1
Ask how each one uses itBuilds on7
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu et al.ICLR 2022 · 4,966 citations
- Multitask Prompted Training Enables Zero-Shot Task GeneralizationVictor Sanh, Albert Webson, Colin Raffel, Stephen H. Bach et al.ICLR 2022 · 1,976 citations
- Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP TasksYizhong Wang, Swaroop Mishra, Pegah Alipoormolabashi, Yeganeh Kordi et al.EMNLP 2022 · 238 citations
- Controlled Text Generation with Natural Language InstructionsWangchunshu Zhou, Yuchen Eleanor Jiang, Ethan Wilcox, Ryan Cotterell et al.ICML 2023 · 121 citations
Related papers
- Systematic Task Exploration with LLMs: A Study in Citation Text GenerationFurkan Sahinuç, Ilia Kuznetsov, Yufang Hou, Iryna GurevychACL 2024
- UniversalCEFR: Enabling Open Multilingual Research on Language Proficiency AssessmentJoseph Marvin Imperial, Abdullah Barayan, Regina Stodden, Rodrigo Wilkens et al.EMNLP 2025 · 2 citations
- DocCGen: Document-based Controlled Code GenerationSameer Pimparkhede, Mehant Kammakomati, Srikanth Tamilselvam, Prince Kumar et al.EMNLP 2024 · 4 citations
- CertiCoder: Towards MISRA-Compliant C Code Generation with LLMsMin Gou, Zhiyu Yao, Hualong Ma, Ende Zhang et al.FSE 2026
- Unlocking Anticipatory Text Generation: A Constrained Approach for Large Language Models DecodingLifu Tu, Semih Yavuz, Jin Qu, Jiacheng Xu et al.EMNLP 2024 · 2 citations
