BABEL: Bodies, Action and Behavior With English Labels
Abhinanda R. Punnakkal, Arjun Chandrasekaran, Nikos Athanasiou, Alejandra Quiros-Ramirez, Michael J. Black
Abstract
Understanding the semantics of human movement -the what, how and why of the movement -is an important problem that requires datasets of human actions with semantic labels. Existing datasets take one of two approaches. Large-scale video datasets contain many action labels but do not contain ground-truth 3D human motion. Alternatively, motion-capture (mocap) datasets have precise body motions but are limited to a small number of actions. To address this, we present BABEL, a large dataset with language labels describing the actions being performed in mocap sequences. BABEL consists of language labels for over 43 hours of mocap sequences from AMASS, containing over 250 unique actions. Each action label in BABEL is precisely aligned with the duration of the corresponding action in the mocap sequence. BABELalso allows overlap of multiple actions, that may each span different durations. This results in a total of over 66000 action segments. The dense annotations can be leveraged for tasks like action recognition, temporal localization, motion synthesis, etc. To demonstrate the value of BABEL as a benchmark, we evaluate the performance of models on 3D action recognition. We demonstrate that BABEL poses interesting learning challenges that are applicable to real-world scenarios, and can serve as a useful benchmark for progress in 3D action recognition. The dataset, baseline methods, and evaluation code are available and supported for academic research purposes at https://babel.is.tue.mpg.de/ . bob head, head bang, nod head, nod, move head side ways, tilt head back, move head around, twist head, tilt head to the right, move head in a circle pace quickly, walk away, runway walk, power walk, wobble walk, slow walk back and forth, speed walk, sidle, strolls, catwalk, walking backwards, walking in circles, zigzag, stride, saunter, plod, shamble, strut, wander, hobble, tread, swagger dance with partner, ballet dancing, interpretive dancing, ginga dance, doing a nutty dance, expression dance, slow dancing with imaginary partner, teacup dance, waltz, fish flop, plie, shimmy, sway scratch hands, scratch chin, scratch arm, scratch face, scratch head, scratch nose, scratching side, scratch waist with right hand, itch butt, itch leg lean backwards to the left, lean to left, lean right, lean left knee forward, lean against object, tilt right place object on upper shelf, put object on table, place ice in glass, put book on shelf, place drink, put box on shelf, place object on ground, place item in waist band, loading stuff into a truck twist wrists in a circle, twist body to the left, twist body to the right, jump twist, whirl body, twirl leg, twist left leg, twist head, twist cork, twirl, barrel roll
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 336416ed-5619-4490-9576-65a7e02b3fa8Cited by top-tier papers112
- Generating Diverse and Natural 3D Human Motions from TextChuan Guo, Shihao Zou, Xinxin Zuo, Sen Wang et al.CVPR 2022 · 462 citations
- Human Motion Diffusion as a Generative PriorYoni Shafir, Guy Tevet, Roy Kapon, Amit Haim BermanoICLR 2024 · 371 citations
- FLAME: Free-Form Language-Based Motion Synthesis & EditingJihoon Kim, Jiseob Kim, Sungjoon ChoiAAAI 2023 · 276 citations
- HUMANISE: Language-conditioned Human Motion Generation in 3D ScenesZan Wang, Yixin Chen, Tengyu Liu, Yixin Zhu et al.NeurIPS 2022 · 207 citations
- TMR: Text-to-Motion Retrieval Using Contrastive 3D Human Motion SynthesisMathis Petrovich, Michael J. Black, Gül VarolICCV 2023 · 192 citations
Builds on5
- AMASS: Archive of Motion Capture As Surface ShapesNaureen Mahmood, Nima Ghorbani, Nikolaus F. Troje, Gerard Pons-Moll et al.ICCV 2019 · 1,784 citations
- Action2Motion: Conditioned Generation of 3D Human MotionsChuan Guo, Xinxin Zuo, Sen Wang, Shihao Zou et al.ACM MM 2020 · 394 citations
- HACS: Human Action Clips and Segments Dataset for Recognition and Temporal LocalizationHang Zhao, Antonio Torralba, Lorenzo Torresani, Zhicheng YanICCV 2019 · 298 citations
- Robust motion in-betweeningFélix G. Harvey, Mike Yurick, Derek Nowrouzezahrai, Christopher J. PalSIGGRAPH 2020 · 269 citations
- Skeleton-Based Action Recognition With Shift Graph Convolutional NetworkKe Cheng, Yifan Zhang, Xiangyu He, Weihan Chen et al.CVPR 2020
Related papers
- Breaking The Limits of Text-conditioned 3D Motion Synthesis with Elaborative DescriptionsYijun Qian, Jack Urbanek, Alexander G. Hauptmann, Jungdam WonICCV 2023 · 15 citations
- Humoto: A 4D Dataset of Mocap Human Object InteractionsJiaxin Lu, Chun-Hao Paul Huang, Uttaran Bhattacharya, Qixing Huang et al.ICCV 2025 · 4 citations
- InterAct: Advancing Large-Scale Versatile 3D Human-Object Interaction GenerationSirui Xu, Dongting Li, Yucheng Zhang, Xiyan Xu et al.CVPR 2025
- Further Understanding Videos through Adverbs: A New Video TaskBo Pang, Kaiwen Zha, Yifan Zhang, Cewu LuAAAI 2020 · 18 citations
- Perceiving and Acting in First-Person: A Dataset and Benchmark for Egocentric Human-Object-Human InteractionsLiang Xu, Chengqun Yang, Zili Lin, Fei Xu et al.ICCV 2025 · 2 citations
