SOSControl: Enhancing Human Motion Generation Through Saliency-Aware Symbolic Orientation and Timing Control
Ho Yin Au, Junkun Jiang, Jie Chen
Abstract
Traditional text-to-motion frameworks often lack precise control, and existing approaches based on joint keyframe locations provide only positional guidance, making it challenging and unintuitive to specify body part orientations and motion timing. To address these limitations, we introduce the Salient Orientation Symbolic (SOS) script, a programmable symbolic framework for specifying body part orientations and motion timing at keyframes. We further propose an automatic SOS extraction pipeline that employs temporally-constrained agglomerative clustering for frame saliency detection and a Saliency-based Masking Scheme (SMS) to generate sparse, interpretable SOS scripts directly from motion data. Moreover, we present the SOSControl framework, which treats the available orientation symbols in the sparse SOS script as salient and prioritizes satisfying these constraints during motion generation. By incorporating SMS-based data augmentation and gradient-based iterative optimization, the framework enhances alignment with user-specified constraints. Additionally, it employs a ControlNet-based ACTOR-PAE Decoder to ensure smooth and natural motion outputs. Extensive experiments demonstrate that the SOS extraction pipeline generates human-interpretable scripts with symbolic annotations at salient keyframes, while the SOSControl framework outperforms existing baselines in motion quality, controllability, and generalizability with respect to motion timing and body part orientation control.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on14
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
- Action-Conditioned 3D Human Motion Synthesis with Transformer VAEMathis Petrovich, Michael J. Black, Gül VarolICCV 2021 · 672 citations
- Generating Diverse and Natural 3D Human Motions from TextChuan Guo, Shihao Zou, Xinxin Zuo, Sen Wang et al.CVPR 2022 · 462 citations
- Human Motion Diffusion as a Generative PriorYoni Shafir, Guy Tevet, Roy Kapon, Amit Haim BermanoICLR 2024 · 371 citations
- Guided Motion Diffusion for Controllable Human Motion SynthesisKorrawe Karunratanakul, Konpat Preechakul, Supasorn Suwajanakorn, Siyu TangICCV 2023 · 240 citations
Related papers
- Motion Synthesis with Sparse and Flexible Keyjoint ControlInwoo Hwang, Jinseok Bae, Donggeun Lim, Young Min KimICCV 2025 · 2 citations
- FineXtrol: Controllable Motion Generation via Fine-Grained TextKeming Shen, Bizhu Wu, Junliang Chen, Xiaoqin Wang et al.AAAI 2026 · 3 citations
- PMG: Progressive Motion Generation via Sparse Anchor Postures Curriculum LearningYingjie Xi, Jian Jun Zhang, Xiaosong YangACM MM 2025 · 1 citation
- ParTY: Part-Guidance for Expressive Text-to-Motion SynthesisKunHo Heo, SuYeon Kim, Yonghyun Gwon, Youngbin Kim et al.CVPR 2026
- AutoKeyframe: Autoregressive Keyframe Generation for Human Motion Synthesis and EditingBowen Zheng, Ke Chen, Yuxin Yao, Zijiao Zeng et al.SIGGRAPH 2025 · 3 citations
