ST-FCL: Spatio-temporal fusion and contrastive learning for urban flow prediction

Xu Zhang, Yongshun Gong, Chengqi Zhang, Xiaoming Wu, Ying Guo, Wenpeng Lu, Long Zhao, Xiangjun Dong

Knowledge-Based Systems, vol. 282, article 111104 (2023)

Also indexed in: Semantic Scholar · OpenAlex · dblp · Publisher page

TL;DR. ST-FCL predicts grid-level urban inflow and outflow by fusing temporal and spatial views, learned through contrastive pretraining, with an external-factor view. On the full TaxiBJ dataset it reaches RMSE 14.71, against 15.41 for the best baseline, ATFM.

Quick facts

Method nameST-FCL; ST-CSL (name of the authors' code repository, github.com/CodeZx6/ST-CSL)
AuthorsXu Zhang, Yongshun Gong, Chengqi Zhang, Xiaoming Wu, Ying Guo, Wenpeng Lu, Long Zhao, Xiangjun Dong
DOI10.1016/j.knosys.2023.111104
TaskNext-step urban flow prediction: inflow and outflow for every cell of a city grid
Contrastive pretrainingTemporal view: triplet loss on Gaussian-blur-augmented data, with closest or aggregated (Top-K) sampling. Spatial view: InfoNCE-style loss with multiple positives chosen by a distance threshold λ
ArchitectureMix Layers feed six pretrained encoders (temporal and spatial for the closeness, period and trend spans); their outputs are concatenated and decoded by convolutional layers
External factorsHour of the timestamp, weekend or holiday, temperature, wind speed and weather conditions, embedded and added to the input flow maps
DatasetsTaxiBJ (Beijing taxis, 30-min intervals, 32×32 grid, four periods P1–P4) and BikeNYC (New York bikes, 1-h intervals, 16×8 grid)
MetricsRMSE, MAE and R² (coefficient of determination), reported with standard errors over five random seeds
BaselinesHA, ARIMA, AutoEncoder, RNN, Seq2Seq, FNN, ST-ResNet, STGCN, ATFM, AGCRN, STTN, ST-GSP, STAEformer and PDFormer
Headline resultFull TaxiBJ: RMSE 14.71, MAE 8.85, R² 0.9806 (ATFM: RMSE 15.41, MAE 9.18). Full BikeNYC: RMSE 4.71, MAE 2.50, R² 0.9286
Training cost17.34 s per training epoch; 2564 MiB GPU memory at batch size 32; 200 epochs to beat the baselines (ATFM needs about 1000, STGSP about 1500)
ImplementationPyTorch on an NVIDIA A100 GPU; Adam optimizer, batch size 64, max–min normalization to [0, 1]
AuthorshipXu Zhang and Yongshun Gong contributed equally; Xiangjun Dong is the corresponding author
PublicationKnowledge-Based Systems, vol. 282 (December 2023), article 111104; available online 21 October 2023. Not open access: © 2023 Elsevier B.V., all rights reserved
Codehttps://github.com/CodeZx6/ST-CSL
BibTeX keyzhang2023stcsl

Main results

SettingResultCompared with
TaxiBJ, full dataset (test set: last 28 days)RMSE = 14.71 ± 0.03; MAE = 8.85 ± 0.01; R² = 0.9806 ± 0.0003
“As shown in Table 2, the RMSE was reduced to 14.71, the MAE was reduced to 8.85”
Best baseline RMSE and MAE: ATFM 15.41 ± 0.04 and 9.18 ± 0.03; best baseline R²: STGSP 0.9786 ± 0.0004; p < 0.01 for all three metrics
BikeNYC, full dataset (test set: last 10 days)RMSE = 4.71 ± 0.01; MAE = 2.50 ± 0.00; R² = 0.9286 ± 0.0001
“As shown in Table 2, the RMSE of our method was 4.71, the MAE was 2.50”
Best baselines: PDFormer RMSE 4.84 ± 0.03; STGCN MAE 2.70 ± 0.03; STTN R² 0.9163 ± 0.0006; p < 0.01 for all three metrics
TaxiBJ ablation, full dataset (Table 7)Full model with external view: RMSE = 14.62, MAE = 8.83
“considering the environment during training resulted in a 6.3% decrease in the RMSE”
Same model without the external view: RMSE 15.60, MAE 9.20 (6.3% RMSE reduction)
TaxiBJ P1, 20/40/60/80/100% data subsets (test set: last 10% of each)RMSE gain over best baseline = +0.78%, +2.54%, +5.04%, +2.42%, +7.53%
“For example, for the RMSE of P1, ST-FCL outperformed the best baseline by 0.78%, 2.54%, 5.04%, 2.42%, and 7.53% on the respective sub-datasets.”
RMSE p < 0.05 on the 40–100% subsets; the 20% gain is not significant (p = 0.18) (Table 3)
BikeNYC, 20/40/60/80/100% data subsets (test set: last 10% of each)RMSE gain over best baseline = +7.41%, +1.43%, +10.85%, +17.47%, +2.45%
“our proposed method achieved better RMSE results of 7.41%, 1.43%, 10.85%, 17.47% and 2.45% compared to the existing methods”
ST-FCL also has the lowest MAE on every subset; RMSE p < 0.05 on all five subsets, but the 100% MAE gain is not significant (p = 0.21) (Table 9)
Training cost (Table 8)17.34 s per epoch; 2564 MiB GPU memory at batch size 32; 200 epochs to beat the baselines
“In contrast, our proposed ST-FCL required only 200 training epochs to achieve better results than the baselines.”
ATFM: 25.08 s, 1719 MiB, about 1000 epochs; STGSP: 14.48 s, 2711 MiB, about 1500 epochs

Quoted text is verbatim from the paper.

Key points

Abstract

Urban flow prediction is critical for urban planning, management, and safety. However, owing to the inherent instability of urban flows, prediction accuracy requires the fusion of multi-view influencing factors. Current prediction methods are insensitive to periodic changes in urban flows, and rarely consider the implied spatial and temporal correlations between similar functional areas. Thus, we propose a method based on spatiotemporal fusion and contrastive learning. We construct an extraction module of spatial and temporal views based on contrastive learning in the spatial and temporal dimensions. With the temporal view extraction, we can obtain the distribution and change in the global urban flow periodicity. Using the spatial view extraction, we can obtain the implied flow variation relationship between similar regions. Therefore, our model can capture the high-level semantic features of urban flow changes from multiple views. We also design a multi-view fusion and prediction network to combine multi-view representations that impact urban flow. We experimentally evaluated our proposed approach using two real-world datasets and demonstrated state-of-the-art forecasting performance, as well as effectiveness in resource-limited environments.

Abstract shown up to the publicly indexed portion; see the publisher page for the full text.

Frequently asked questions

Is ST-CSL the same as ST-FCL?

Yes. The Knowledge-Based Systems paper calls the model ST-FCL. ST-CSL is the name of the authors' code repository for the paper (github.com/CodeZx6/ST-CSL).

What limitation of earlier methods does ST-FCL target?

According to the abstract, current methods are insensitive to periodic changes in urban flows and rarely consider the implied spatial and temporal correlations between similar functional areas.

How does ST-FCL build positive and negative samples for contrastive learning?

For the temporal view, closest sampling takes the most and least similar samples to the anchor (by Euclidean distance) as positive and negative, and aggregated sampling builds them from distance-weighted Top-K samples; the temporal encoder is trained with a triplet loss on Gaussian-blur-augmented data. For the spatial view, regions whose distance to the anchor region is below a threshold λ are positives and the rest are negatives, and the loss is an InfoNCE-style loss with multiple positives.

How accurate is ST-FCL?

On the full TaxiBJ dataset ST-FCL reaches RMSE 14.71, MAE 8.85 and R² 0.9806; the strongest baseline on RMSE and MAE, ATFM, reaches 15.41 and 9.18. On the full BikeNYC dataset it reaches RMSE 4.71, MAE 2.50 and R² 0.9286, against best baseline values of 4.84 (PDFormer) and 2.70 (STGCN). Table 2 reports p < 0.01 against the best baseline for every metric.

Does ST-FCL work with limited historical data?

The paper tests 20%, 40%, 60%, 80% and 100% subsets of each TaxiBJ period (P1–P4) and of BikeNYC, with the last 10% of each subset as the test set. ST-FCL has the lowest RMSE and MAE on every subset, and its gain over the best baseline is significant (p < 0.05) in most but not all comparisons. On TaxiBJ P1, closest sampling works better with 20% and 40% of the data, and aggregated sampling with 60%, 80% and 100%.

How efficient is ST-FCL to train?

In Table 8, one ST-FCL training epoch takes 17.34 s, and ST-FCL uses 2564 MiB of GPU memory at batch size 32. ST-FCL needs only 200 epochs to beat the baselines, while ATFM needs about 1000 and STGSP about 1500 epochs to reach their results. ATFM uses less GPU memory (1719 MiB) and STGSP is faster per epoch (14.48 s).

How do I cite ST-FCL?

Cite it as: Zhang, X., Gong, Y., Zhang, C., Wu, X., Guo, Y., Lu, W., Zhao, L., & Dong, X. (2023). Spatio-temporal fusion and contrastive learning for urban flow prediction. Knowledge-Based Systems, 282, Article 111104. https://doi.org/10.1016/j.knosys.2023.111104 BibTeX is available at https://codezx6.github.io/papers/st-csl.html#cite and in https://codezx6.github.io/publications.bib.

Is the code for ST-FCL available?

Yes, at https://github.com/CodeZx6/ST-CSL.

中文摘要

基于时空融合与对比学习的城市流量预测(ST-FCL)

城市流量预测对城市规划、管理与安全至关重要。现有方法对周期变化不敏感,也很少考虑相似功能区之间隐含的时空关联。ST-FCL(代码库名 ST-CSL)基于对比学习构建时间与空间视图提取模块:时间视图获取全局流量周期性的分布与变化,空间视图获取相似区域间隐含的流量变化关系;多视图融合与预测网络融合各视图表示及天气、节假日等外部因素。该方法在 TaxiBJ 与 BikeNYC 上取得最优预测性能,并在资源受限环境中有效。

关键词:城市流量预测;人群流量预测;对比学习;时空融合;多视图表示融合;多视图影响因素

Keywords

urban flow prediction crowd flow prediction contrastive learning spatio-temporal fusion multi-view representation fusion multi-view influencing factors spatio-temporal systems traffic prediction

Also referred to as: ST-CSL; spatio-temporal fusion and contrastive learning; temporal view extraction module; spatial view extraction module; external view extraction module; multi-view fusion and prediction network.

Research area map and related search terms

Field path (broad to narrow): artificial intelligence › machine learning › data mining › spatio-temporal data mining › time series forecasting › intelligent transportation systems › smart cities › urban computing › traffic prediction › crowd flow prediction › urban flow prediction › multi-view contrastive spatio-temporal learning › ST-FCL

Task

urban flow prediction crowd flow prediction inflow and outflow prediction traffic flow forecasting spatio-temporal forecasting

Method

multi-view contrastive learning temporal-view contrastive learning spatial-view contrastive learning closeness, period and trend periodicity modeling functional region similarity multi-view fusion

Datasets

TaxiBJ BikeNYC

Compared with

ST-ResNet STGCN ATFM AGCRN STTN ST-GSP STAEformer PDFormer

中文

城市流量预测 人群流量预测 交通流量预测 时空预测 多视图对比学习 周期性建模 功能区相似性

Related areas

traffic forecasting spatio-temporal graph neural networks contrastive representation learning self-supervised learning multi-view learning periodic pattern mining functional zones urban region representation mobility modeling resource-limited learning deep spatio-temporal residual networks

Applications

urban planning traffic management public safety event crowd management bike-sharing rebalancing

Related benchmarks (not used in this paper)

TaxiNYC BikeDC PEMS-BAY METR-LA PeMS04 PeMS08

中文 (扩展)

交通预测 时空数据挖掘 时间序列预测 智能交通 智慧城市 城市计算 对比表征学习 多视图学习 周期模式挖掘 城市功能区 区域表征 小样本/低资源时空预测 城市规划

Terms are grouped by role. "Related areas", "Applications" and "Related benchmarks (not used in this paper)" describe the surrounding field, not results of this paper. See the site-wide research area map.

Related papers

Urban flow prediction

Other topics

Cite this paper

@article{zhang2023stcsl,
  title        = {Spatio-temporal fusion and contrastive learning for urban flow prediction},
  author       = {Zhang, Xu and Gong, Yongshun and Zhang, Chengqi and Wu, Xiaoming and Guo, Ying and Lu, Wenpeng and Zhao, Long and Dong, Xiangjun},
  journal      = {Knowledge-Based Systems},
  year         = {2023},
  volume       = {282},
  pages        = {111104},
  doi          = {10.1016/j.knosys.2023.111104},
  issn         = {0950-7051},
  url          = {https://doi.org/10.1016/j.knosys.2023.111104}
}
Zhang, X., Gong, Y., Zhang, C., Wu, X., Guo, Y., Lu, W., Zhao, L., & Dong, X. (2023). Spatio-temporal fusion and contrastive learning for urban flow prediction. Knowledge-Based Systems, 282, Article 111104. https://doi.org/10.1016/j.knosys.2023.111104