Bifold and Semantic Reasoning for Pedestrian Behavior Prediction
Amir Rasouli, Mohsen Rohani, Jun Luo
Abstract
Pedestrian behavior prediction is one of the major challenges for intelligent driving systems. Pedestrians often exhibit complex behaviors influenced by various contextual elements. To address this problem, we propose BiPed, a multitask learning framework that simultaneously predicts trajectories and actions of pedestrians by relying on multi-modal data. Our method benefits from 1) a bifold encoding approach where different data modalities are processed in-dependently allowing them to develop their own representations, and jointly to produce a representation for all modalities using shared parameters; 2) a novel interaction modeling technique that relies on categorical semantic parsing of the scenes to capture interactions between target pedestrians and their surroundings; and 3) a bifold prediction mechanism that uses both independent and shared decoding of multimodal representations. Using public pedestrian behavior benchmark datasets for driving, PIE and JAAD, we highlight the benefits of the proposed method for behavior prediction and show that our model achieves state-of-the-art performance and improves trajectory and action prediction by up to 22% and 9% respectively. We further investigate the contributions of the proposed reasoning techniques via extensive ablation studies.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9e04253b-2e39-494e-9296-427467f0d997Cited by top-tier papers3
- Trajectory Unified Transformer for Pedestrian Trajectory PredictionLiushuai Shi, Le Wang, Sanping Zhou, Gang HuaICCV 2023 · 100 citations
- TrEP: Transformer-Based Evidential Prediction for Pedestrian Intention with UncertaintyZhengming Zhang, Renran Tian, Zhengming DingAAAI 2023 · 84 citations
- Understanding Interaction as You Need: Intention-Driven Pedestrian Behavior PredictionHang Yu, Yansen Yu, Jiayan QiuAAAI 2026
Builds on19
- PIE: A Large-Scale Dataset and Models for Pedestrian Intention Estimation and Trajectory PredictionAmir Rasouli, Iuliia Kotseruba, Toni Kunic, John K. TsotsosICCV 2019 · 411 citations
- EvolveGraph: Multi-Agent Trajectory Prediction with Dynamic Relational ReasoningJiachen Li, Fan Yang, Masayoshi Tomizuka, Chiho ChoiNeurIPS 2020 · 258 citations
- Unsupervised Multi-Task Feature Learning on Point CloudsKaveh Hassani, Mike HaleyICCV 2019 · 205 citations
- PAMTRI: Pose-Aware Multi-Task Learning for Vehicle Re-Identification Using Highly Randomized Synthetic DataZheng Tang, Milind Naphade, Stan Birchfield, Jonathan Tremblay et al.ICCV 2019 · 146 citations
- Deep Floor Plan Recognition Using a Multi-Task Network With Room-Boundary-Guided AttentionZhiliang Zeng, Xianzhi Li, Ying Kin Yu, Chi-Wing FuICCV 2019 · 125 citations
Related papers
- MMTL-UniAD: A Unified Framework for Multimodal and Multi-Task Learning in Assistive Driving PerceptionWenzhuo Liu, Wenshuo Wang, Yicheng Qiao, Qiannan Guo et al.CVPR 2025
- Three Steps to Multimodal Trajectory Prediction: Modality Clustering, Classification and SynthesisJianhua Sun, Yuxuan Li, Haoshu Fang, Cewu LuICCV 2021 · 91 citations
- Euro-PVI: Pedestrian Vehicle Interactions in Dense Urban CentersApratim Bhattacharyya, Daniel Olmeda Reino, Mario Fritz, Bernt SchieleCVPR 2021
- Bi-Modal Learning for Networked Time SeriesYoungeun Nam, Jihye Na, Susik Yoon, Hwanjun Song et al.KDD 2025
- Large Scale Interactive Motion Forecasting for Autonomous Driving : The Waymo Open Motion DatasetScott Ettinger, Shuyang Cheng, Benjamin Caine, Chenxi Liu et al.ICCV 2021 · 817 citations
