Tailor: Generating and Perturbing Text with Semantic Controls
Alexis Ross, Tongshuang Wu, Hao Peng, Matthew E. Peters, Matt Gardner
Abstract
Controlled text perturbation is useful for evaluating and improving model generalizability. However, current techniques rely on training a model for every target perturbation, which is expensive and hard to generalize. We present Tailor, a semantically-controlled text generation system. Tailor builds on a pretrained seq2seq model and produces textual outputs conditioned on control codes derived from semantic representations. We craft a set of operations to modify the control codes, which in turn steer generation towards targeted attributes. These operations can be further composed into higher-level ones, allowing for flexible perturbation strategies. We demonstrate the effectiveness of these perturbations in multiple applications. First, we use Tailor to automatically create high-quality contrast sets for four distinct natural language processing (NLP) tasks. These contrast sets contain fewer spurious artifacts and are complementary to manually annotated ones in their lexical diversity. Second, we show that Tailor perturbations can improve model generalization through data augmentation. Perturbing just ∼2% of training data leads to a 5.8-point gain on an NLI challenge set measuring reliance on syntactic heuristics. * denotes equal contribution.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4223ed43-2ddb-4edf-9b30-afc6ebe4a29dCited by top-tier papers21
- Red Teaming Language Models with Language ModelsEthan Perez, Saffron Huang, H. Francis Song, Trevor Cai et al.EMNLP 2022 · 239 citations
- Adversarial training for high-stakes reliabilityDaniel M. Ziegler, Seraphina Nix, Lawrence Chan, Tim Bauman et al.NeurIPS 2022 · 79 citations
- Generating Data to Mitigate Spurious Correlations in Natural Language Inference DatasetsYuxiang Wu, Matt Gardner, Pontus Stenetorp, Pradeep DasigiACL 2022 · 74 citations
- Feature-Level Debiased Natural Language UnderstandingYougang Lyu, Piji Li, Yechang Yang, Maarten de Rijke et al.AAAI 2023 · 12 citations
- All That Glisters Is Not Gold: A Benchmark for Reference-Free Counterfactual Financial Misinformation DetectionYuechen Jiang, Zhiwei Liu, Yupeng Cao, Yueru He et al.ACL 2026 · 9 citations
Builds on11
- Plug and Play Language Models: A Simple Approach to Controlled Text GenerationSumanth Dathathri, Andrea Madotto, Janice Lan, Jane Hung et al.ICLR 2020 · 1,166 citations
- Neural Text Generation With Unlikelihood TrainingSean Welleck, Ilia Kulikov, Stephen Roller, Emily Dinan et al.ICLR 2020 · 683 citations
- Learning The Difference That Makes A Difference With Counterfactually-Augmented DataDivyansh Kaushik, Eduard H. Hovy, Zachary Chase LiptonICLR 2020 · 625 citations
- Competency Problems: On Finding and Removing Artifacts in Language DataMatt Gardner, William Merrill, Jesse Dodge, Matthew E. Peters et al.EMNLP 2021 · 72 citations
- Beyond Accuracy: Behavioral Testing of NLP Models with CheckListMarco Túlio Ribeiro, Tongshuang Wu, Carlos Guestrin, Sameer SinghACL 2020 · 51 citations
Related papers
- Tailor: A Soft-Prompt-Based Approach to Attribute-Based Controlled Text GenerationKexin Yang, Dayiheng Liu, Wenqiang Lei, Baosong Yang et al.ACL 2023 · 25 citations
- Contrastive Learning with Adversarial Perturbations for Conditional Text GenerationSeanie Lee, Dong Bok Lee, Sung Ju HwangICLR 2021 · 117 citations
- Let the CAT out of the bag: Contrastive Attributed explanations for TextSaneem A. Chemmengath, Amar Prakash Azad, Ronny Luss, Amit DhurandharEMNLP 2022 · 6 citations
- Controllable Meaning Representation to Text Generation: Linearization and Data Augmentation StrategiesChris Kedzie, Kathleen R. McKeownEMNLP 2020 · 16 citations
- Attention Biasing and Context Augmentation for Zero-Shot Control of Encoder-Decoder Transformers for Natural Language GenerationDevamanyu Hazarika, Mahdi Namazifar, Dilek Hakkani-TürAAAI 2022 · 6 citations
