Lune

EMNLP2023顶会

ByteSized32: A Corpus and Challenge Task for Generating Task-Specific World Models Expressed as Text Games

Ruoyao Wang, Graham Todd, Xingdi Yuan, Ziang Xiao, Marc-Alexandre Côté, Peter A. Jansen

2023年份
1被引次数
3顶会引用

摘要

In this work we investigate the capacity of language models to generate explicit, inter pretable, and interactive world models of sci entific and common-sense reasoning tasks. We operationalize this as a task of generating text games, expressed as hundreds of lines of PYTHON code. To facilitate this task, we introduce BYTESIZED321, a corpus of 32 reasoning-focused text games totalling 20k lines of PYTHON code. We empirically demon strate that GPT-4 can use these games as tem plates for single-shot in-context learning, suc cessfully producing runnable games on unseen topics in 28% of cases. When allowed to self reflect on program errors, game runnability substantially increases to 57%. While evalu ating simulation fidelity is labor intensive, we introduce a suite of automated metrics to assess game fidelity, technical validity, adherence to task specifications, and winnability, showing a high-degree of agreement with expert human ratings. We pose this as a challenge task to spur further development at the juncture of world modeling and code generation. ©2023 Association for Computational Linguistics.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

引用它的顶会 Paper3

问问它们各自怎么用它

它引用的顶会 Paper13

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖