Generating Data for Symbolic Language with Large Language Models
Jiacheng Ye, Chengzu Li, Lingpeng Kong, Tao Yu
摘要
While large language models (LLMs) bring not only performance but also complexity, recent work has started to turn LLMs into data generators rather than task inferencers, where another affordable task model is trained for efficient deployment and inference. However, such an approach has primarily been applied to natural language tasks, and has not yet been explored for symbolic language tasks with complex structured outputs (e.g., semantic parsing and code generation). In this paper, we propose SYMGEN which utilizes LLMs for generating various annotationexpensive symbolic language data. SYMGEN consists of an informative prompt to steer generation and an agreement-based verifier to improve data correctness. We conduct extensive experiments on six symbolic language tasks across various settings. Compared with the LLMs, we demonstrate the 1%-sized task model can achieve comparable or better performance, largely cutting inference and deployment costs. We also show that generated data with only a few human demonstrations can be as effective as over 10 times the amount of human-annotated data when training the task model, saving a considerable amount of annotation effort. SYMGEN takes a step toward data generation for annotation-expensive complex tasks, and we release the code at https://github.com/HKUNLP/SymGen .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- Agentic Context Engineering: Evolving Contexts for Self-Improving Language ModelsQizheng Zhang, Changran Hu, Shubhangi Upasani, Boyuan Ma 等ICLR 2026 · 被引用 374 次
- Compositional Exemplars for In-context LearningJiacheng Ye, Zhiyong Wu, Jiangtao Feng, Tao Yu 等ICML 2023 · 被引用 188 次
- FlexKBQA: A Flexible LLM-Powered Framework for Few-Shot Knowledge Base Question AnsweringZhenyu Li, Sunqi Fan, Yu Gu, Xiuxing Li 等AAAI 2024 · 被引用 143 次
- DetGPT: Detect What You Need via ReasoningRenjie Pi, Jiahui Gao, Shizhe Diao, Rui Pan 等EMNLP 2023 · 被引用 57 次
- Natural Language Dataset Generation Framework for Visualizations Powered by Large Language ModelsHyung-Kwon Ko, Hyeon Jeon, Gwanmo Park, Dae Hyun Kim 等CHI 2024 · 被引用 20 次
它引用的顶会 Paper24
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- Transformers are RNNs: Fast Autoregressive Transformers with Linear AttentionAngelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, François FleuretICML 2020 · 被引用 2,665 次
- Calibrate Before Use: Improving Few-shot Performance of Language ModelsZihao Zhao, Eric Wallace, Shi Feng, Dan Klein 等ICML 2021 · 被引用 1,843 次
- CodeT5: Identifier-aware Unified Pre-trained Encoder-Decoder Models for Code Understanding and GenerationYue Wang, Weishi Wang, Shafiq R. Joty, Steven C. H. HoiEMNLP 2021 · 被引用 1,224 次
- Train Short, Test Long: Attention with Linear Biases Enables Input Length ExtrapolationOfir Press, Noah A. Smith, Mike LewisICLR 2022 · 被引用 1,168 次
相关 Paper
- Making Large Language Models Better Data CreatorsDong-Ho Lee, Jay Pujara, Mohit Sewak, Ryen White 等EMNLP 2023 · 被引用 13 次
- Human-LLM Collaborative Annotation Through Effective Verification of LLM LabelsXinru Wang, Hannah Kim, Sajjadur Rahman, Kushan Mitra 等CHI 2024 · 被引用 127 次
- Can Large Language Models Understand Symbolic Graphics Programs?Zeju Qiu, Weiyang Liu, Haiwen Feng, Zhen Liu 等ICLR 2025
- Symbol-LLM: Towards Foundational Symbol-centric Interface For Large Language ModelsFangzhi Xu, Zhiyong Wu, Qiushi Sun, Siyu Ren 等ACL 2024
- LLM-Aided Automatic Modeling for Security Protocol VerificationZiyu Mao, Jingyi Wang, Jun Sun, Shengchao Qin 等ICSE 2025 · 被引用 3 次
