Self-Consistency for LLM-Based Motion Trajectory Generation and Verification
Jiaju Ma, R. Kenny Jones, Jiajun Wu, Maneesh Agrawala
摘要
Self-consistency has proven to be an effective technique for improving LLM performance on natural language reasoning tasks in a lightweight, unsupervised manner. In this work, we study how to adapt self-consistency to visual domains. Specifically, we consider the generation and verification of LLM-produced motion graphics trajectories. Given a prompt (e.g., "Move the circle in a spiral path"), we first sample diverse motion trajectories from an LLM, and then identify groups of consistent trajectories via clustering.
Our key insight is to model the family of shapes associated with a prompt as a prototype trajectory paired with a group of geometric transformations (e.g., rigid, similarity, and affine). Two trajectories can then be considered consistent if one can be transformed into the other under the warps allowable by the transformation group. We propose an algorithm that automatically recovers a shape family, using hierarchical relationships between a set of candidate transformation groups. Our approach improves the accuracy of LLM-based trajectory generation by 4-6%. We further extend our method to support verification, observing 11% precision gains over VLM baselines. Our code and dataset are available at https://majiaju.io/trajectory- self-consistency.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper14
- Self-Consistency Improves Chain of Thought Reasoning in Language ModelsXuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V. Le 等ICLR 2023 · 被引用 681 次
- TIFA: Accurate and Interpretable Text-to-Image Faithfulness Evaluation with Question AnsweringYushi Hu, Benlin Liu, Jungo Kasai, Yizhong Wang 等ICCV 2023 · 被引用 400 次
- Davidsonian Scene Graph: Improving Reliability in Fine-grained Evaluation for Text-to-Image GenerationJaemin Cho, Yushi Hu, Jason M. Baldridge, Roopal Garg 等ICLR 2024 · 被引用 139 次
- LLMR: Real-time Prompting of Interactive Worlds using Large Language ModelsFernanda De La Torre, Cathy Mengying Fang, Han Huang, Andrzej Banburski-Fahey 等CHI 2024 · 被引用 124 次
- SceneCraft: An LLM Agent for Synthesizing 3D Scenes as Blender CodeZiniu Hu, Ahmet Iscen, Aashi Jain, Thomas Kipf 等ICML 2024 · 被引用 105 次
相关 Paper
- MoVer: Motion Verification for Motion Graphics AnimationsJiaju Ma, Maneesh AgrawalaSIGGRAPH 2025 · 被引用 8 次
- Enhancing the Outcome Reward-based RL Training of MLLMs with Self-Consistency SamplingJiahao Wang, Weiye Xu, Aijun Yang, Wengang Zhou 等NeurIPS 2025 · 被引用 4 次
- Towards Self-Refinement of Vision-Language Models with Triangular ConsistencyYunlong Deng, Guangyi Chen, Tianpei Gu, Lingjing Kong 等NeurIPS 2025 · 被引用 3 次
- Improving Retrieval Augmented Language Model with Self-ReasoningYuan Xia, Jingbo Zhou, Zhenhui Shi, Jun Chen 等AAAI 2025 · 被引用 42 次
- The Art of Interrogation: Consistency Amplifies Factuality in Spatial ReasoningThéo Uscidda, Marta Gazulla, Maks Ovsjanikov, Federico Tombari 等ICML 2026
