Urban flow prediction with masked and contrastive spatio-temporal pre-training
Urban flow prediction forecasts future flows across city regions from historical flow data and supports services such as trip planning, congestion control, and public safety. Three papers (Zhang et al., CIKM 2023; Zhang et al., Knowledge-Based Systems 2023; Zhang et al., Neural Networks 2025) develop self-supervised spatio-temporal pre-training for this task.
MC-STL (CIKM 2023; code repository MCSTL) pre-trains two encoders to capture regional correlations that change over time: one learns to reconstruct masked regions, and its attention weights also form the adjacency matrix of a GCN that extracts important regions; the other learns a contrastive task using global-local cross-attention. ST-FCL (Knowledge-Based Systems 2023; code repository ST-CSL) applies contrastive learning separately in a temporal view and a spatial view to capture global periodicity and the hidden relations between functionally similar regions. MR-UFP (Neural Networks 2025; code repository MR-UPF) pairs spatial-temporal random masking and contrastive pre-training with a multi-scale region-classification auxiliary task, and reports robust results when training data are reduced.
Papers compared
| Method | Published | Core idea | Evaluation | Code |
|---|---|---|---|---|
| MR-UFP | Neural Networks, vol. 182, article 106900 (2025) | Masked and contrastive spatial-temporal pre-training plus a hand-labeled multi-scale region-classification task trained jointly with flow prediction. | TaxiBJ and BikeNYC, 13 baselines, RMSE/MAE; full-data RMSE 14.32 and 4.45; tested on 20–100% data subsets. | GitHub |
| ST-FCL | Knowledge-Based Systems, vol. 282, article 111104 (2023) | Contrastively pretrained temporal and spatial view encoders per time span, fused with external factors, for grid urban flow prediction. | TaxiBJ and BikeNYC, full data and 20–100% subsets; RMSE, MAE, R²; TaxiBJ RMSE 14.71 vs ATFM 15.41. | GitHub |
| MC-STL | CIKM 2023, pp. 3298–3307 | Two pre-trained encoders: spatio-temporal masked ViT whose attention builds a GCN graph, plus temporal-order contrastive cross-attention encoder. | TaxiBJ (P1-P4, full) and BikeNYC, RMSE/MAE vs nine baselines; full TaxiBJ RMSE 14.53 (ST-GSP 14.72). | GitHub |
MR-UFP: Enhancing urban flow prediction via mutual reinforcement with multi-scale regional information
Xu Zhang, Mengxin Cao, Yongshun Gong, Xiaoming Wu, Xiangjun Dong, Ying Guo, Long Zhao, Chengqi Zhang · Neural Networks, vol. 182, article 106900 (2025)
MR-UFP predicts the inflow and outflow of city grid regions, even with limited training data, by pre-training encoders with spatial-temporal masking and contrastive learning and training flow prediction jointly with a multi-scale region-classification task. On full TaxiBJ and BikeNYC it beats all 13 baselines (RMSE 14.32 and 4.45).
- Two encoders are pre-trained on historical flows: EncM, a ViT encoder that learns to reconstruct spatio-temporally masked regions (75% masking worked best), and EncC, a global–local cross-attention encoder trained with a triplet loss.
- Regions are hand-labeled from real maps at three scales (road/non-road, coarse, fine-grained; 2/5/142 classes on TaxiBJ), and classifying them is an auxiliary task trained jointly with flow prediction.
- A joint loss, 2·L_P·L_rcf/(L_P + L_rcf), keeps either loss from dominating; on TaxiBJ P1 it gives RMSE 15.36 versus 15.41 for a plain L_P + L_rcf sum.
- On full TaxiBJ and BikeNYC, MR-UFP beats all 13 baselines (RMSE 14.32 vs 14.53 for MC-STL; 4.45 vs 4.63 for DeepMeshCity) and is best on most 20–100% data subsets.
ST-FCL: Spatio-temporal fusion and contrastive learning for urban flow prediction
Xu Zhang, Yongshun Gong, Chengqi Zhang, Xiaoming Wu, Ying Guo, Wenpeng Lu, Long Zhao, Xiangjun Dong · Knowledge-Based Systems, vol. 282, article 111104 (2023)
ST-FCL predicts grid-level urban inflow and outflow by fusing temporal and spatial views, learned through contrastive pretraining, with an external-factor view. On the full TaxiBJ dataset it reaches RMSE 14.71, against 15.41 for the best baseline, ATFM.
- A temporal view extraction module captures global flow periodicity through contrastive pretraining with a triplet loss. Two proposed schemes, closest sampling and aggregated (Top-K) sampling, build its positive and negative pairs.
- A spatial view extraction module mines implicit associations between similar functional zones. Its InfoNCE-style loss uses multiple positives: regions whose distance to the anchor region is below a threshold λ.
- Mix Layers fuse each span's samples before six pretrained encoders (temporal and spatial for closeness, period and trend). Concatenated outputs go to a convolutional decoder; external-factor embeddings are added to input maps.
- Full TaxiBJ: RMSE 14.71, MAE 8.85 (ATFM 15.41, 9.18). Full BikeNYC: RMSE 4.71, MAE 2.50 (best baselines 4.84 PDFormer, 2.70 STGCN). Table 2 reports p < 0.01 for every metric.
MC-STL: Mask- and Contrast-Enhanced Spatio-Temporal Learning for Urban Flow Prediction
Xu Zhang, Yongshun Gong, Xinxin Zhang, Xiaoming Wu, Chengqi Zhang, Xiangjun Dong · CIKM 2023, pp. 3298–3307
MC-STL pre-trains two encoders for urban flow prediction: a ViT encoder learns to reconstruct regions masked at different timestamps, and its attention weights also build a GCN adjacency matrix; a global-local cross-attention encoder learns a temporal-order contrastive task. It reaches RMSE 14.53 on full TaxiBJ.
- Spatio-temporal random masking hides different regions at different timestamps. On TaxiBJ it gives RMSE 14.53 / MAE 8.70, against 14.91 / 8.86 with conventional image random masking.
- A graph processing module (GPM) sums the self-attention weights of the mask-pre-trained ViT encoder over heads and batch and uses them as the adjacency matrix of a two-layer GCN.
- The contrastive pretext task treats each temporally ordered batch as a unit: a temporally shuffled anchor is the negative, a unit from another period is the positive, and triplet loss trains the global-local cross-attention encoder.
- Against nine baselines, MC-STL has the lowest RMSE and MAE on all six evaluation sets (TaxiBJ P1-P4, full TaxiBJ, BikeNYC); on full TaxiBJ it reaches RMSE 14.53 versus 14.72 for ST-GSP.
Key concepts
- Urban (crowd) flow prediction
- Forecasting how many people or vehicles enter (inflow) and leave (outflow) each cell of a city grid in the next time interval. Papers: MC-STL, ST-FCL, MR-UFP
- TaxiBJ
- Beijing taxi inflow/outflow on a 32×32 grid at 30-minute intervals, with holiday and weather data; a standard crowd-flow benchmark. Papers: MC-STL, ST-FCL, MR-UFP
- BikeNYC
- New York City bike inflow/outflow on a 16×8 grid at 1-hour intervals. Papers: MC-STL, ST-FCL, MR-UFP
- Closeness, period and trend
- The three temporal views (recent, daily and weekly history) used by crowd-flow models since ST-ResNet. Papers: ST-FCL
- Spatio-temporal pre-training
- Self-supervised tasks, such as masked reconstruction or contrastive learning, that train encoders on flow data before prediction. Papers: MC-STL, MR-UFP
Frequently asked questions
Which papers use masked pre-training for urban flow prediction?
MC-STL (CIKM 2023) introduces a mask-enhanced pre-training task across the spatial and temporal dimensions, and MR-UFP (Neural Networks 2025) uses spatial-temporal random masking together with spatial-temporal contrastive pre-training.
How do MC-STL, ST-FCL and MR-UFP differ?
MC-STL pre-trains two encoders, one by reconstructing masked regions and one with a contrastive task using global-local cross-attention, and uses the masked encoder's attention weights as the adjacency matrix of a GCN that extracts important regions. ST-FCL builds separate temporal-view and spatial-view contrastive modules. MR-UFP adds a multi-scale region-classification auxiliary task, trained jointly with flow prediction so that the two tasks reinforce each other, and balances their losses with a joint loss.
Which datasets are used for urban flow prediction in these papers?
All three papers (MC-STL, ST-FCL, MR-UFP) evaluate on TaxiBJ and BikeNYC with RMSE and MAE.
Which baselines do they compare with?
They compare with crowd-flow and spatio-temporal models such as ST-ResNet, ATFM, AGCRN, STGCN and ST-GSP; ST-FCL and MR-UFP also include PDFormer and STAEformer, and MR-UFP includes MC-STL and ST-FCL.
Fields, tasks, datasets and related search terms
Task
urban flow prediction crowd flow prediction inflow and outflow prediction traffic prediction with limited data spatio-temporal forecasting traffic flow forecasting citywide crowd flow forecasting grid-based traffic flow forecasting
Method
multi-task learning auxiliary region classification task multi-scale region information masked pre-training spatial-temporal contrastive pre-training environment feature fusion attention-derived graph GCN multi-view contrastive learning temporal-view contrastive learning spatial-view contrastive learning closeness, period and trend periodicity modeling functional region similarity multi-view fusion masked spatio-temporal pre-training self-supervised pre-training for traffic prediction spatio-temporal contrastive learning Vision Transformer (ViT) encoder global-local cross-attention
Datasets
TaxiBJ BikeNYC
Compared with
ST-ResNet ATFM AGCRN STGCN PDFormer STAEformer DeepMeshCity MC-STL ST-FCL STTN ST-GSP ST-SSL Multi-STGCnet
中文
城市流量预测 人群流量预测 交通流量预测 多任务学习 多尺度区域 掩码预训练 对比预训练 时空预测 多视图对比学习 周期性建模 功能区相似性 流入流出预测 时空对比学习 自监督预训练
Fields (broad to narrow)
artificial intelligence machine learning data mining spatio-temporal data mining time series forecasting intelligent transportation systems smart cities urban computing traffic prediction crowd flow prediction urban flow prediction multi-task spatio-temporal learning multi-view contrastive spatio-temporal learning self-supervised spatio-temporal pre-training
Related areas
traffic forecasting multi-task learning auxiliary learning urban region profiling region representation learning text-enhanced spatio-temporal modeling self-supervised learning masked pre-training contrastive learning spatio-temporal graph neural networks data-efficient learning external factors (weather, holidays) contrastive representation learning multi-view learning periodic pattern mining functional zones urban region representation mobility modeling resource-limited learning deep spatio-temporal residual networks graph neural networks Transformers for time series masked autoencoders representation learning human mobility spatio-temporal heterogeneity citywide prediction grid-based prediction
Applications
traffic management resource allocation route optimization urban planning bike-sharing operations public safety event crowd management bike-sharing rebalancing congestion control trip planning ride-hailing dispatch
Related benchmarks (not used in this paper)
TaxiNYC BikeDC PEMS-BAY METR-LA PeMS04 PeMS08 BikeChicago
中文 (扩展)
交通预测 时空数据挖掘 智能交通 智慧城市 城市计算 多任务学习 辅助任务学习 区域画像 区域表征学习 数据高效学习 天气与节假日外部因素 路径优化 资源分配 时间序列预测 对比表征学习 多视图学习 周期模式挖掘 城市功能区 区域表征 小样本/低资源时空预测 城市规划 时空图神经网络 图神经网络 掩码自编码器 自监督学习 表征学习 人类移动性 网格流量预测 交通管理 共享单车调度
中文概述
基于掩码与对比预训练的城市流量预测
城市流量预测根据历史流量预测城市各区域未来的流量,服务于出行规划、拥堵控制与公共安全。三篇论文围绕自监督时空预训练展开:MC-STL(CIKM 2023)结合掩码重建预训练与基于全局-局部交叉注意力的对比预训练;ST-FCL(Knowledge-Based Systems 2023)在时间视图与空间视图上分别做对比学习;MR-UFP(Neural Networks 2025)将时空随机掩码、对比预训练与多尺度区域分类辅助任务结合,在训练数据减少时依然表现稳健。
相关检索词:城市流量预测;人群流量预测;交通流量预测;流入流出预测;时空预测;时空预训练;TaxiBJ;BikeNYC
Cite these papers
@article{zhang2025mrufp,
title = {Enhancing urban flow prediction via mutual reinforcement with multi-scale regional information},
author = {Zhang, Xu and Cao, Mengxin and Gong, Yongshun and Wu, Xiaoming and Dong, Xiangjun and Guo, Ying and Zhao, Long and Zhang, Chengqi},
journal = {Neural Networks},
year = {2025},
volume = {182},
pages = {106900},
doi = {10.1016/j.neunet.2024.106900},
issn = {0893-6080},
url = {https://doi.org/10.1016/j.neunet.2024.106900}
}
@article{zhang2023stcsl,
title = {Spatio-temporal fusion and contrastive learning for urban flow prediction},
author = {Zhang, Xu and Gong, Yongshun and Zhang, Chengqi and Wu, Xiaoming and Guo, Ying and Lu, Wenpeng and Zhao, Long and Dong, Xiangjun},
journal = {Knowledge-Based Systems},
year = {2023},
volume = {282},
pages = {111104},
doi = {10.1016/j.knosys.2023.111104},
issn = {0950-7051},
url = {https://doi.org/10.1016/j.knosys.2023.111104}
}
@inproceedings{zhang2023mcstl,
title = {Mask- and Contrast-Enhanced Spatio-Temporal Learning for Urban Flow Prediction},
author = {Zhang, Xu and Gong, Yongshun and Zhang, Xinxin and Wu, Xiaoming and Zhang, Chengqi and Dong, Xiangjun},
booktitle = {Proceedings of the 32nd ACM International Conference on Information and Knowledge Management},
year = {2023},
series = {CIKM '23},
pages = {3298--3307},
publisher = {ACM},
doi = {10.1145/3583780.3614958},
url = {https://doi.org/10.1145/3583780.3614958}
}