CrossLight: Offline-to-Online Reinforcement Learning for Cross-City Traffic Signal Control
Qian Sun, Rui Zha, Le Zhang, Jingbo Zhou, Yu Mei, Zhiling Li, Hui Xiong
Abstract
The recent advancements in Traffic Signal Control (TSC) have highlighted the potential of Reinforcement Learning (RL) as a promising solution to alleviate traffic congestion. Current research in this area primarily concentrates on either online or offline learning strategies, aiming to create optimized policies for specific cities. Nevertheless, the transferability of these policies to new cities is impeded by constraints such as the limited availability of high-quality data and the expensive and risky exploration process. To this end, in this paper, we present an innovative cross-city Traffic Signal Control (TSC) paradigm called CrossLight. Our approach involves meta training using offline data from source cities and adaptively fine-tuning in the target city. This novel methodology aims to address the challenges of transferring TSC policies across different cities effectively. In our proposed approach, we start by acquiring meta-decision pattern knowledge through trajectory dynamics reconstruction via pre-training in source cities. To address disparities in road network topologies between cities, we dynamically construct city topological structures based on the extracted meta-knowledge during the offline meta-training phase. These structures are then used to distill pattern-structure aware representations of decision trajectories from the source cities. To identify effective initial parameters for the learnable components, we employ the Model-Agnostic Meta-Learning (MAML) framework, a popular meta-learning approach. During adaptive fine-tuning in the target city, we introduce a replay buffer that is iteratively updated using online interactions with a rank and filter mechanism. This mechanism, along with a carefully designed exploration strategy, ensures a balance between exploitation and exploration, thereby fostering both the diversity and quality of the trajectories for fine-tuning. Finally, extensive experiments across four cities validate that CrossLight achieves comparable performance in new cities with minimal fine-tuning iterations, surpassing both existing online and offline methods. This success underscores that our CrossLight framework emerges as a groundbreaking and potent paradigm, offering a feasible and effective solution to the intelligent transportation community.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b91d9478-eead-4c81-a9c4-3b8e16565332Cited by top-tier papers1
Ask how each one uses itBuilds on15
- Decision Transformer: Reinforcement Learning via Sequence ModelingLili Chen, Kevin Lu, Aravind Rajeswaran, Kimin Lee et al.NeurIPS 2021 · 2,557 citations
- Offline Reinforcement Learning with Implicit Q-LearningIlya Kostrikov, Ashvin Nair, Sergey LevineICLR 2022 · 1,402 citations
- Toward A Thousand Lights: Decentralized Deep Reinforcement Learning for Large-Scale Traffic Signal ControlChacha Chen, Hua Wei, Nan Xu, Guanjie Zheng et al.AAAI 2020 · 450 citations
- Efficient Online Reinforcement Learning with Offline DataPhilip J. Ball, Laura Smith, Ilya Kostrikov, Sergey LevineICML 2023 · 326 citations
- Online Decision TransformerQinqing Zheng, Amy Zhang, Aditya GroverICML 2022 · 256 citations
Related papers
- MetaLight: Value-Based Meta-Reinforcement Learning for Traffic Signal ControlXinshi Zang, Huaxiu Yao, Guanjie Zheng, Nan Xu et al.AAAI 2020 · 185 citations
- Self-Supervised Cross-City Trajectory Representation Learning Based on Meta-LearningYanwei Yu, Hong Xia, Shaoxuan Gu, Xingyu Zhao et al.AAAI 2026
- TransformerLight: A Novel Sequence Modeling Based Traffic Signaling Mechanism via Gated TransformerQiang Wu, Mingyuan Li, Jun Shen, Linyuan Lü et al.KDD 2023 · 17 citations
- Optimizing Traffic Control with Model-Based Learning: A Pessimistic Approach to Data-Efficient Policy InferenceMayuresh Kunjir, Sanjay Chawla, Siddarth Chandrasekar, Devika Jay et al.KDD 2023 · 3 citations
- Selective Cross-City Transfer Learning for Traffic Prediction via Source City Region Re-WeightingYilun Jin, Kai Chen, Qiang YangKDD 2022 · 76 citations
