Evaluating Spatiotemporal Consistency in Automatically Generated Sewing Instructions
Luisa Geiger, Mareike Hartmann, Michael Sullivan, Alexander Koller
Abstract
In this paper, we propose a novel, automatic tree-based evaluation metric for LLMgenerated step-by-step assembly instructions, that more accurately reflects spatiotemporal aspects of construction than traditional metrics such as BLEU and BERT similarity scores. We apply our proposed metric to the domain of sewing instructions, and show that our metric better correlates with manually-annotated error counts as well as human quality ratings, demonstrating our metric's superiority for evaluating the spatiotemporal soundness of sewing instructions. Further experiments show that our metric is more robust than traditional approaches against artificially-constructed counterfactual examples that are specifically constructed to confound metrics that rely on textual similarity.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on6
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger et al.ICLR 2020 · 8,443 citations
- ViLT: Vision-and-Language Transformer Without Convolution or Region SupervisionWonjae Kim, Bokyung Son, Ildoo KimICML 2021 · 2,258 citations
- Abstract Visual Reasoning with Tangram ShapesAnya Ji, Noriyuki Kojima, Noah Rush, Alane Suhr et al.EMNLP 2022 · 18 citations
Related papers
- Identifying Reliable Evaluation Metrics for Scientific Text RevisionLéane Jourdan, Nicolas Hernandez, Florian Boudin, Richard DufourACL 2025
- LexInstructEval: Lexical Instruction Following Evaluation for Large Language ModelsHuimin Ren, Yan Liang, Baiqiao Su, Chaobo Sun et al.AAAI 2026
- SimLLM: Calculating Semantic Similarity in Code Summaries using a Large Language Model-Based ApproachXin Jin, Zhiqiang LinFSE 2024 · 8 citations
- Global Explainability of BERT-Based Evaluation Metrics by Disentangling along Linguistic FactorsMarvin Kaster, Wei Zhao, Steffen EgerEMNLP 2021 · 13 citations
- Metric Calculating Benchmark: Code-Verifiable Complicate Instruction Following Benchmark for Large Language ModelsHyeonseok Moon, Seongtae Hong, Jaehyung Seo, Heuiseok LimEMNLP 2025
