Lune

EMNLP2023Top-tier venue

Generating Data for Symbolic Language with Large Language Models

Jiacheng Ye, Chengzu Li, Lingpeng Kong, Tao Yu

2023Year
8Citations
9Top-tier citations

Abstract

While large language models (LLMs) bring not only performance but also complexity, recent work has started to turn LLMs into data generators rather than task inferencers, where another affordable task model is trained for efficient deployment and inference. However, such an approach has primarily been applied to natural language tasks, and has not yet been explored for symbolic language tasks with complex structured outputs (e.g., semantic parsing and code generation). In this paper, we propose SYMGEN which utilizes LLMs for generating various annotationexpensive symbolic language data. SYMGEN consists of an informative prompt to steer generation and an agreement-based verifier to improve data correctness. We conduct extensive experiments on six symbolic language tasks across various settings. Compared with the LLMs, we demonstrate the 1%-sized task model can achieve comparable or better performance, largely cutting inference and deployment costs. We also show that generated data with only a few human demonstrations can be as effective as over 10 times the amount of human-annotated data when training the task model, saving a considerable amount of annotation effort. SYMGEN takes a step toward data generation for annotation-expensive complex tasks, and we release the code at https://github.com/HKUNLP/SymGen .

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 7e80860c-a7b9-4593-9df9-97c2ba8b4a77

Cited by top-tier papers9

Ask how each one uses it

Builds on24

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines